Image knowledge composite distillation optimization method for imaging angle change SAR (Synthetic Aperture Radar) target recognition
The method addresses the catastrophic forgetting issue in SAR target recognition by using self and mutual knowledge distillation with an adaptive deep inversion network to generate pseudo-samples, enhancing recognition accuracy and adaptability to new angles, achieving a Kappa coefficient of 0.9996 and 99.94% classification accuracy.
Patent Information
- Application Number
- CN202510356805.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-15
Smart Images

Figure CN120318562A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of synthetic aperture radar target recognition, and particularly relates to an image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles. Background Art
[0002] In the field of synthetic aperture radar (SAR) target recognition technology, the intra-class feature differences caused by changing imaging angles are the core problems affecting the recognition accuracy. Traditional deep learning methods rely on a sample library with full-angle coverage for training. However, in practical applications, due to the limitations of data acquisition costs and storage, it is difficult to construct a complete multi-angle SAR dataset. When the model performs incremental learning on new angle data, due to the non-stationary update of neural network parameters, the recognition performance for the already learned angles will degrade sharply (i.e., the catastrophic forgetting phenomenon). Among the existing solutions, the replay method based on core set selection is limited by the non-editable nature of samples and cannot effectively reconstruct angle diversity features; while the regularization method can alleviate parameter drift, but it does not constrain the unique scattering characteristics of SAR images, resulting in insufficient angle generalization ability.
[0003] In recent years, the dataset distillation technology has provided a new idea for incremental learning by synthesizing high-information samples to replace the original data. However, the SAR images generated by traditional deep inversion methods (DeepInversion) have problems such as blurred textures and distorted scattering point distributions. The main reason is that the image prior regularization (such as total variation constraint) used does not combine with the SAR imaging physical model. In addition, most existing distillation strategies adopt a single knowledge transfer path (such as only relying on the soft targets of the output layer), ignoring the spatial distribution alignment of intermediate layer features, and it is difficult to achieve multi-level knowledge retention in incremental tasks with continuously changing angles. For example, the kernel ridge regression (KRR) method can match feature statistics, but it does not consider the non-linear mapping relationship of intra-class features caused by angle changes, resulting in insufficient discriminability of cross-angle pseudo-samples.
[0004] In view of the above technical bottlenecks, there is an urgent need for a solution that combines SAR physical characteristics with a dynamic distillation mechanism. Specifically, the following technical problems need to be solved: 1) how to generate pseudo-samples with high fidelity and wide angle coverage under limited angle samples; 2) how to simultaneously suppress catastrophic forgetting and improve the adaptability to new angles through multi-level knowledge transfer (including the response layer and the feature layer); 3) how to design a dynamic optimization strategy to make the distillation process adapt to the non-stationary changes of the angle distribution. The breakthrough of these problems has important engineering significance for improving the robustness of the SAR target recognition system in a dynamic environment. Summary of the Invention
[0005] In view of this, the present invention aims to propose an image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles, so as to solve the problem that existing methods do not fuse the physical characteristics of SAR and the dynamic distillation mechanism.
[0006] To achieve the above object, the present invention adopts the following technical solutions: An image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles, comprising:
[0007] Construct an angle adaptive learning framework, the framework includes a replay sample module, a data preprocessing module, a classification network module and a distillation module; the classification network module adopts a composite knowledge distillation structure including self-distillation and mutual-distillation, wherein mutual-distillation performs knowledge transfer through the output layer and intermediate layer features of the current model and the previous task model, and self-distillation performs knowledge transfer through the outputs of multiple intermediate classifiers inside the classification network; the replay sample module adopts an adaptive depth inversion network to generate pseudo-samples, and this network optimizes the diversity and feature distribution consistency of the generated samples through the adversarial generation mechanism of the pre-trained teacher model and the student model; the method specifically includes:
[0008] S1: Divide the basic dataset D and the new angle dataset D' into an incremental task sequence, where D' is divided into m incremental task subsets;
[0009] S2: Use the adaptive depth inversion network to generate pseudo-samples, and this network generates multi-angle SAR samples by jointly optimizing the classification loss, feature distribution regularization, image prior regularization and competitive adversarial loss;
[0010] S3: During the training process of the classification network, adopt a three-stage distillation mechanism: achieve self-distillation through the KL divergence calculation of the intermediate classifier, achieve mutual-distillation between models through response distillation and feature distillation, and combine the mixed training of real samples and generated samples.
[0011] Furthermore, a preferred method is also proposed. The adaptive depth inversion network generates pseudo-samples through the following loss function:
[0012] L total = L classification + α f R feature + α TV R TV + α L2 R L2 + α c R compete
[0013] where, L classification is the main loss, α f is the weight coefficient of feature distribution regularization, α TVis the weight coefficient of total variation regularization, α L2 is the weight coefficient of L2 regularization, α c is the weight coefficient of adaptive adversarial inversion regularization, R feature is the feature distribution regularization loss, R TV is the total variation regularization loss, R L2 is the L2 regularization loss, R compete is the adaptive loss.
[0014] Furthermore, a preferred method is also proposed. The response distillation loss in the mutual distillation is:
[0015] L R1 = L R (p(z t_now_output1 , T), p(z t_old_output1 , T))
[0016] L R2 = L R (p(z t_now_output2 , T), p(z t_old_output2 , T))
[0017] L R3 = L R (p(z t_now_output3 , T), p(z t_old_output3 , T))
[0018] loss_2 = L R1 + L R2 + L R3
[0019] where is the mutual distillation loss between the outputs of the first residual block of the model network. L R (.) is the KL divergence loss. z t_now_output1 , is the output of the first residual block of the current task model. z t_old_output1 is the output of the first residual block of the previous task model network. T is the temperature coefficient. p(.) is the probability distribution of each category at temperature T. L R2 is the mutual distillation loss between the outputs of the second residual block of the model network. z t_now_output2 is the output of the second residual block of the current task model. z t_old_output2 is the output of the second residual block of the previous task model network. L R3 is the mutual distillation loss between the final outputs of the model network. z t_now_output3 is the final output of the current task model. z t_old_output3 is the final output of the previous task model network. loss_2 is the total mutual distillation loss.
[0020] Furthermore, a preferred method is also proposed. The feature distillation loss in the mutual distillation is as follows:
[0021]
[0022] where u h is the adapter model for shape transformation of the teacher model, ν g is the adapter model for shape transformation of the student model, r is the adapter model for shape transformation of the feature map of the student model, X is the input feature map, W Hint is the weight of the teacher model, W Guided is the weight of the student model, and W r is the weight of the adaptation model;
[0023]
[0024] where represents the feature loss function between the intermediate features of the current task and the intermediate features of the previous task, represents the feature loss function between the final output features of the current task and the final output features of the previous task model.
[0025] Furthermore, a preferred method is also proposed. The classification model adopts a multi-branch residual structure, including: two parallel Bottleneck modules that process spatial features of different scales respectively; three intermediate classifiers are inserted into network layers at different depths, and each classifier includes a global average pooling layer and a fully connected layer; the final classifier fuses the input features of all intermediate classifiers.
[0026] Furthermore, a preferred method is also proposed. The pseudo-sample generation strategy of the replay sample module is as follows:
[0027] Before each round of incremental training, use the current teacher model to generate pseudo-samples with the same number as the target categories;
[0028] Dynamically adjust the noise initialization range of the generated samples, and the noise variance decays exponentially with the number of training rounds;
[0029] Adopt the curriculum learning strategy to gradually relax the constraint intensity of the feature distribution regularization.
[0030] Furthermore, a preferred method is also proposed. The incremental task partitioning method is: divide the new angle dataset D' into m subsets according to the pitch angle change interval, and each subset contains target samples of the same category; the pitch angle difference between adjacent subsets is controlled within the range of 0.5 - 1°.
[0031] Furthermore, a preferred method is also proposed. The data preprocessing module includes: performing random geometric transformation on the input SAR image, including random translation and flipping operations in the azimuth direction; and performing covariance matrix normalization processing on the polarization channel data.
[0032] Based on the same inventive concept, the present invention also provides a computer device, including a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes an image knowledge composite distillation optimization method for SAR target recognition with imaging angle variation as described in any one of the above.
[0033] Based on the same inventive concept, the present invention also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by the processor, the steps of an image knowledge composite distillation optimization method for SAR target recognition with imaging angle variation as described in any one of the above are executed.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0035] The present invention proposes a multi-level knowledge transfer mechanism of self-distillation and mutual-distillation, breaking through the limitations of a single distillation path. Through cross-model response distillation and feature matching, mutual-distillation realizes bidirectional knowledge transfer between new and old angle tasks, overcoming the parameter deviation caused by unidirectional transfer in traditional methods. Through KL constraint in the intermediate layer, self-distillation forces the alignment of multi-scale features inside the network, suppressing the angular sensitivity shift of shallow features. In the MSTAR mixed angle test set (17° + 15°), the Kappa coefficient is increased to 0.9996 (the highest of the baseline methods is 0.9866), proving that the reuse efficiency of cross-angle discriminative features is significantly improved.
[0036] The present invention also proposes an adversarial generation mechanism based on SAR imaging characteristics. Embedding the covariance matrix constraint of the polarization channel in image regularization ensures that the scattering characteristics of the generated pseudo-samples are consistent with those of real SAR data. Introducing the matching of BN layer statistics and using the channel distribution prior of the pre-trained model improve the continuity of the angular coverage of the pseudo-samples. The classification accuracy (OA) of the generated samples in the 15° pitch angle test set is increased from 97.63% of the baseline to 99.94%.
[0037] The present invention realizes the collaborative adaptation of distillation parameters and angular distribution. The distillation temperature coefficient T is non-linearly adjusted according to the task progress (the initial T = 15, and it rises to 20 when OA > 99%), dynamically balancing the knowledge retention strength of old and new tasks; the regularization weight of the feature distribution adopts a curriculum learning strategy, strengthening the statistic matching in the initial stage and gradually releasing the generation freedom in the later stage. In 4 rounds of incremental tasks, the OA of the old angle (17°) task is stably maintained at 100%, and the AA of the new angle (15°) task is improved from 97.89% to 99.94%. The dynamic mechanism of the present invention effectively inhibits catastrophic forgetting.
[0038] The classification network module proposed by the present invention extracts azimuth and pitch spatial features in parallel with a double Bottleneck module, reducing the model jitter caused by angular differences through parameter sharing; the intermediate classifier insertion strategy realizes the early-layer feedback of gradient signals, accelerating the fusion and convergence of cross-angle features, shortening its training cycle by 30%, and the convergence stability (Kappa coefficient fluctuation < 0.5%) of the mixed-angle test set is significantly better than that of traditional single-chain networks. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0040] Figure 1 is a flowchart of an image knowledge compound distillation optimization method for SAR target recognition with imaging angle changes according to the present invention;
[0041] Figure 2 is a basic flowchart of sample angle adaptive incremental learning according to the present invention;
[0042] Figure 3 is a block diagram of an adaptive depth inversion model according to the present invention;
[0043] Figure 4 is an illustration of Adaptive Deep Inversion according to the present invention;
[0044] Figure 5 is an algorithm flowchart based on adaptive depth inversion according to the present invention;
[0045] Figure 6 is a structure diagram of the classification network module according to the present invention;
[0046] Figure 7 are the optical images and SAR images of ten types of targets of MSTAR according to the present invention;
[0047] Figure 8It is the kappa coefficient change diagram under different temperature coefficients of the present invention;
[0048] Figure 9 It is the kappa coefficient curve diagram of different loss terms of the present invention. Specific implementation manners
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The described embodiments are only some of the embodiments of the present invention, rather than all embodiments.
[0050] Embodiment 1. Refer to Figure 1 This embodiment is described. An image knowledge composite distillation optimization method for SAR target recognition with imaging angle change described in this embodiment includes:
[0051] Construct an angle adaptive learning framework, which includes a replay sample module, a data preprocessing module, a classification network module, and a distillation module; the classification network module adopts a composite knowledge distillation structure including self-distillation and mutual distillation, where mutual distillation performs knowledge transfer through the output layer and intermediate layer features of the current model and the previous task model, and self-distillation performs knowledge transfer through the outputs of multiple intermediate classifiers inside the classification network; the replay sample module uses an adaptive deep inversion network to generate pseudo samples, and this network optimizes the diversity and feature distribution consistency of the generated samples through the adversarial generation mechanism of the pre-trained teacher model and the student model; the method specifically includes:
[0052] S1: Divide the basic data set D and the new angle data set D' into an incremental task sequence, where D' is divided into m incremental task subsets;
[0053] S2: Use an adaptive deep inversion network to generate pseudo samples, and this network generates multi-angle SAR samples by jointly optimizing the classification loss, feature distribution regularization, image prior regularization, and competitive adversarial loss;
[0054] S3: During the training process of the classification network, adopt a three-stage distillation mechanism: realize self-distillation through the KL divergence calculation of the intermediate classifier, realize mutual distillation between models through response distillation and feature distillation, and combine the mixed training of real samples and generated samples.
[0055] Embodiment 2. This embodiment further limits an image knowledge composite distillation optimization method for SAR target recognition with imaging angle change described in Embodiment 1, and the adaptive deep inversion network generates pseudo samples through the following loss function:
[0056] Ltotal = L classification + α f R feature + α TV R TV + α L2 R L2 + α c R compete
[0057] Among them, L classification is the main loss, α f is the weight coefficient of feature distribution regularization, α TV is the weight coefficient of total variation regularization, α L2 is the weight coefficient of L2 regularization, α c is the weight coefficient of adaptive adversarial inversion regularization, R feature is the feature distribution regularization loss, R TV is the total variation regularization loss, R L2 is the L2 regularization loss, R compete is the adaptive loss.
[0058] Embodiment 3. This embodiment further limits the image knowledge composite distillation optimization method for SAR target recognition with imaging angle change described in Embodiment 1. The response distillation loss in the mutual distillation is as follows:
[0059]
[0060] Among them, is the mutual distillation loss between the outputs of the first residual block of the model network. L R (.) is the KL divergence loss, z t_now_output1 , is the output of the first residual block of the current task model (student model), z t_old_output1 is the output of the first residual block of the previous task model (teacher model) network. T is the temperature coefficient, and p(.) is the probability distribution of each category at temperature T. is the mutual distillation loss between the outputs of the second residual block of the model network. z t_now_output2 is the output of the second residual block of the current task model (student model), z t_old_output2 is the output of the second residual block of the previous task model (teacher model) network. is the mutual distillation loss between the final outputs of the model network. z t_now_output3 is the final output of the current task model (student model), z t_old_output3 is the final output of the previous task model (teacher model) network. loss_2 is the total mutual distillation loss.
[0061] Embodiment 4. This embodiment further limits an image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles described in Embodiment 1. The feature distillation loss in the mutual distillation is as follows:
[0062]
[0063] where u h is an adapter model for shape transformation of the teacher model, ν g is an adapter model for shape transformation of the student model, r is an adapter model for shape transformation of the feature map of the student model, X is the input feature map, W Hint is the weight of the teacher model, W Guided is the weight of the student model, and W r is the weight of the adaptation model;
[0064]
[0065] where represents the feature loss function between the intermediate feature of the current task and the intermediate feature of the previous task, represents the feature loss function between the final output feature of the current task and the final output feature of the previous task model.
[0066] Embodiment 5. This embodiment further limits an image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles described in Embodiment 1. The classification model adopts a multi-branch residual structure, including: two parallel Bottleneck modules that process spatial features of different scales respectively; three intermediate classifiers inserted in network layers at different depths, and each classifier includes a global average pooling layer and a fully connected layer; the final classifier fuses the input features of all intermediate classifiers.
[0067] Embodiment 6. This embodiment further limits an image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles described in Embodiment 1. The pseudo-sample generation strategy of the replay sample module is as follows:
[0068] Before each round of incremental training, use the current teacher model to generate pseudo-samples with the same number as the target category;
[0069] Dynamically adjust the noise initialization range of the generated samples, and the noise variance decays exponentially with the number of training rounds;
[0070] Adopt a curriculum learning strategy to gradually relax the constraint intensity of feature distribution regularization.
[0071] Embodiment 7. This embodiment further defines an image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles described in Embodiment 1. The incremental task division method is as follows: The new angle dataset D' is divided into m subsets according to the pitch angle change interval, and each subset contains target samples of the same category; the pitch angle difference between adjacent subsets is controlled within the range of 0.5 - 1°.
[0072] Embodiment 8. This embodiment further defines an image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles described in Embodiment 1. The data preprocessing module includes: performing random geometric transformations on the input SAR image, including random translation and flipping operations in the azimuth direction; performing covariance matrix normalization processing on the polarization channel data.
[0073] Embodiment 9. A computer device described in this embodiment includes a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes an image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles described in any one of Embodiments 1 to 8.
[0074] Embodiment 10. A computer-readable storage medium described in this embodiment has a computer program stored thereon. When the computer program is run by a processor, it executes the steps of an image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles described in any one of Embodiments 1 to 8.
[0075] Embodiment 11. Refer to Figures 1 to 9 this embodiment. This embodiment provides a specific example for an image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles described in Embodiment 1, and is also used to explain Embodiments 2 to 8. Specifically:
[0076] The overall framework of an image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles proposed in this embodiment is as Figure 1 shown. Its overall framework includes a replay sample module, a data preprocessing module, a classification network module, and a distillation module. Among them, mutual distillation between models and self-distillation of models are used together to improve the classification accuracy of the models for samples in new and old tasks.
[0077] In the process of research on angle adaptability of SAR images, there is a pair of SAR image datasets D and D′, D = [M0, M1,..., M t ..., M n , D′ = [M0′, M1′,..., Mt ′..., M n ′], where n represents the number of categories of the two groups of SAR, and M t represents the set of all image data with the label t in the category D dataset, and M t ′ represents the set of all image data with the label t in the category D′ dataset. The two datasets have the same number of categories and the same category names. However, due to factors such as the acquisition time and acquisition range, there are significant differences in the images between the same categories of the two datasets.
[0078] Take the pair of SAR image datasets that meet the above conditions as the research object of angle adaptation. Take one of the datasets in the dataset pair, such as D, as the basic dataset, and then divide the D′ dataset into multiple incremental task datasets [D1′, D2′,..., D t ′,... D′ m . The model needs to continuously learn on the remote sensing dataset, and when training the incremental task set D t ′, due to permission restrictions, the SAR image classification model can only access the current training set D t ′.
[0079] Therefore, the angle-adaptive incremental learning of SAR samples mainly includes the following two steps:
[0080] 1) The model will first learn on the basic sample dataset D to obtain a well-performing remote sensing image classification model I0.
[0081] 2) During the incremental task learning process, the already trained model performs angle-adaptive incremental learning on the gradually emerging remote sensing image sample dataset D t ′ (1 ≤ t ≤ n), and continuously updates the entire model to obtain I t . The remote sensing image classification goal based on sample angle-adaptive incremental learning is: as new sample tasks of SAR images keep coming, the SAR image classification model needs to further improve the overall classification performance of SAR images through learning new knowledge without forgetting the basic dataset. Figure 2 This is the basic process of sample angle-adaptive incremental learning for the SAR image classification model.
[0082] In this embodiment, the self-distillation of D t ′ is used to improve the stability of the network itself and the reliability of the output features of the intermediate layer, laying a foundation for the reliability of mutual distillation. At the same time, the mutual distillation between D t and D t-1 ′ is used to improve the recognition accuracy of the basic task and the incremental task. Finally, the classification loss between the real target and the network output is used to improve the recognition accuracy of the incremental task.
[0083] Meanwhile, in order to optimize the network distillation algorithm, an incremental classification algorithm combined with the distillation network and real samples or generated sample replay is proposed in this embodiment. When the basic dataset task is accessible, the original samples can be directly saved for algorithm optimization. However, in the task setting of angle adaption, as data arrives continuously, due to permission or security and privacy issues, the data of previous tasks may not be accessible. Moreover, there are significant differences between the new incremental task data and the basic dataset. Direct training will lead to a rapid decline in the accuracy of the old tasks. Therefore, an adaptive inversion network is introduced, and the samples generated by it are used for sample replay.
[0084] In the task setting of angle adaptive learning, the training data increases dynamically while the class set remains unchanged.
[0085] In this case, there are two incremental methods: One is that all the consecutive tasks are just different subsets of the same dataset, which is relatively simple but has limited application scope in engineering. The other method is that although the consecutive tasks contain the same class types, they do not completely belong to the same dataset but belong to two datasets with a large inter-class difference. In practical engineering applications, due to environmental differences or morphological differences, the data of the same class in the two datasets will be quite different. Therefore, this algorithm conducts research on this setting, and there are the following challenges in this case:
[0086] (1) The inter-class differences between a pair of datasets with the same class types are large. When a classification network has a high classification accuracy for one dataset, its classification accuracy for the other dataset is very low.
[0087] (2) When the task increment arrives, there is a catastrophic forgetting problem for the first task, i.e., the basic dataset. It is very difficult to simultaneously maintain the memory on the old tasks and the classification accuracy on the new tasks.
[0088] (3) Current research on incremental algorithms mainly focuses on class incremental algorithms, and there is less research on angle adaptive algorithms.
[0089] To solve the above problems, the method proposed in this embodiment combines composite self-distillation and mutual distillation to simultaneously achieve knowledge retention for old tasks and training and learning for new tasks during the distillation process. Meanwhile, in terms of algorithm optimization, the existing deep inversion network is used to generate pseudo-samples, and innovatively, the images generated based on adaptive deep inversion are used as generated data for replay to improve the overall classification performance of the network.
[0090] The image generation network adopted in this embodiment is an adaptive depth inversion network structure. The feature of this model is that it uses a pre-trained teacher model and a new student model to generate images simultaneously, competing and antagonizing with each other, thereby enhancing the diversity and reliability of the generated images. In this embodiment, the same network structure as the classification network is used as the teacher network and the student network, where the teacher network model is a model saved after pre-training. The block diagram of the adaptive depth inversion image generation network model is as shown in Figure 3 shown below.
[0091] The depth inversion idea evolved from the "DeepDream" idea, which was initially used on natural images to obtain artistic effects and is now also applicable to optimizing noise in images. Given a randomly initialized input and an arbitrary target label y, an image is synthesized by optimizing the following parameters:
[0092]
[0093] where, is the classification loss (such as cross-entropy), and is the image regularization term. "DeepDream" uses image priors to guide away from unrealistic images without recognizable visual information:
[0094]
[0095] where R TV and penalize the total variance and the l2 norm of respectively, with scaling factors α tv , Image prior regularization can converge to valid images more stably. However, the distributions of these images are still very different from the original training images, resulting in unsatisfactory knowledge distillation results.
[0096] The image regularization can be extended by using a new feature distribution regularization term to improve the image quality of depth inversion. The previously defined image prior term is of little guidance for obtaining a synthesized containing low-level and high-level features similar to . Assuming that the feature statistics of each batch follow a Gaussian distribution, it can thus be defined using the mean μ and variance σ 2 . Then, the feature distribution regularization term can be expressed as:
[0097]
[0098] where, are the batch mean and variance estimates of the feature maps corresponding to the l-th convolutional layer. The operators E[·] and ||·||2 denote the expected value and l2-norm calculation, respectively. However, the running average statistics stored in the widely used batch normalization (BN) layer are already sufficient. The BN layer normalizes the feature maps during training to mitigate the covariate shift. It implicitly captures the channel means and variances during training, so the expected value in the formula is estimated as follows:
[0099]
[0100] This regularization of the feature distribution greatly improves the quality of the generated images. This model inversion method is called deep inversion, which is a general method that can be used for any trained deep CNN classifier to invert high-fidelity images. Therefore, R(·) can be expressed as:
[0101]
[0102] In addition to quality, diversity also plays a crucial role in avoiding duplicate and redundant synthetic images. Adaptive deep inversion is an enhanced image generation scheme based on a novel iterative competition scheme between the image generation process and the student network. Its main idea is to encourage the synthetic images to cause a disagreement between the student and the teacher. To this end, an additional loss is introduced for image generation based on the Jensen-Shannon divergence to penalize the similarity of the output distributions.
[0103]
[0104] where, is the mean of the teacher and student distributions. During the optimization process, new images are generated that the student cannot easily classify, while the teacher can. As Figure 4 shown, it is recommended to repeatedly expand the coverage of the image distribution during the learning process. During the competition process, the regularization R(·) in the formula is updated with an additional loss proportional to α c i.e.:
[0105]
[0106] Adaptive deep inversion can improve image diversity. Given a set of generated images (shown as green stars), intermediate students can learn to capture a part of the original image distribution. When generating new images (shown as red stars), the competition encourages the students to extract new samples from the learned knowledge, thus improving the distribution coverage and promoting more knowledge transfer. The algorithm flow based on adaptive deep inversion is as Figure 5As shown, the process includes:
[0107] (1) Load the pre-trained teacher model and set it to evaluation mode.
[0108] (2) Load the student model and set it to training mode (if needed).
[0109] (3) Define a feature hook for each BN layer in the teacher model to track and calculate feature statistics (mean and variance).
[0110] (4) Create a random noise tensor as the initial image, and randomly translate and flip the image at each iteration. Input the processed image into the teacher model for forward propagation to obtain the output.
[0111] (5) Calculate and combine four losses, which are:
[0112] Main loss L classification : Calculate the classification loss between the generated image and the target label using cross-entropy loss.
[0113] Feature distribution regularization loss R feature : Calculate the difference between the feature distribution of the generated image and the BN layer statistics in the teacher model (L2 loss of mean and variance).
[0114] Total variation regularization loss R TV : Calculate the L1 and L2 losses of the total variation of the generated image.
[0115] L2 regularization loss R L2 : Perform L2 regularization on the generated image.
[0116] Adaptive loss R compete : Calculate the JS divergence between the output of the student model and the output of the teacher model for the generated image.
[0117] (6) Backpropagation and optimization: Calculate the gradient of the total loss with respect to the input image. Update the input image using the optimizer. Crop the image as needed to keep the pixel values within a reasonable range.
[0118] (7) Save the best image: If the loss in the current iteration is less than the previous optimal loss, update the optimal image. Save the generated image according to the settings.
[0119] According to the above algorithm principle and algorithm construction process, the total loss function is obtained as follows:
[0120] L total = L classification + α f R feature + α TV RTV +α L2 R L2 +α c R compete
[0121] where α f and α TV and α L2 and α c are the weight coefficients of feature distribution regularization, total variation regularization, L2 regularization, and adaptive adversarial inversion regularization, respectively. The combination of these loss terms helps generate high-fidelity and diverse images, enabling the student model to better learn from the teacher model and perform effective knowledge distillation in the absence of real data.
[0122] During the training process of the classification network module, a three-stage distillation mechanism is adopted: self-distillation is achieved by calculating the KL divergence of the intermediate classifier, mutual distillation between models is achieved by response distillation and feature distillation, and combined with the mixed training of real samples and generated samples.
[0123] In this embodiment, the classification network module is as Figure 6 shown, which adopts a deep learning classification network combining 3D convolution, 2D convolution, and residual blocks. The Bottleneck module consists of 2D convolutional layers. The network extracts more discriminative features through two Bottleneck modules respectively, thereby improving the performance of the deep classifier. The loss function mainly comes from four parts, and the formula is as follows:
[0124] loss_all = m × loss_1 + n × loss_2 + k × loss_3 + l × loss_4
[0125] where loss_1 represents the self-distillation loss; loss_2 and loss_3 represent the mutual distillation loss; loss_4 represents the cross-entropy loss from the label to the deepest classifier and the cross-entropy loss of all shallow classifiers. It is calculated using the labels of the training dataset and the outputs of the softmax layer of each classifier:
[0126] loss_4 = CrossEntropy(q i , y)
[0127] In this embodiment, in order to improve the accuracy of the algorithm for the incremental dataset, self-distillation is introduced. The KL divergence is calculated using the softmax outputs between two intermediate outputs, namely output1 and output2 in the figure, and the teacher, and introduced into the softmax layer of each shallow classifier. By introducing the KL divergence, the self-distillation framework affects each shallow classifier by the deepest teacher network.
[0128] loss_1 = KL(q i , q C )
[0129] where q i represents the output of the softmax layer of the classifier θ i / C . q c represents the output of the softmax layer of the deepest classifier.
[0130] In this embodiment, the "mutual" in mutual distillation refers to the model trained for the previous task and the incremental task model that arrives currently. The mutual distillation mainly includes two parts. One part comes from the knowledge distillation between the outputs of the current model and the previous task model, and the other part comes from the knowledge distillation between the output features of the intermediate layers of the current model and the previous model.
[0131] Knowledge distillation is a classic model compression method. Its core idea is to guide a lightweight student model to learn from a teacher model with better performance and more complex structure, so as to improve its performance without changing the structure of the student model. It is generally believed that the parameters of the model are the knowledge learned by the model. Therefore, the most common way of transfer learning is to first perform pre-training on a large dataset, and then use the parameters obtained from the pre-training to perform fine-tuning on a small dataset (the two datasets often have different domains or tasks). For example, first perform pre-training on the ImageNet dataset, and then perform detection on the COCO dataset. Knowledge can be regarded as the mapping relationship from input to output. Therefore, a teacher network can be trained first, and then the output result of the teacher network can be used as the target of the student network to train the student network, so that the result learned by the student network is close to the output of the teacher network. At the same time, Softmax-T is proposed, and the distillation temperature coefficient is introduced. The formula is as follows:
[0132]
[0133] where z i is the logit of the i-th class, and T is the temperature factor, which controls the importance of each soft target. The soft target contains the implicit information knowledge from the teacher model. In this algorithm, the network model of the current task is the student model, and the model trained for the previous task is the teacher model.
[0134] Therefore, the distillation loss of the soft logits can be rewritten as:
[0135] L ResD (p(z t , T), p(z s , T)) = L R (p(z t , T), p(z s , T))
[0136] In this embodiment, L R (.) uses the KL divergence loss. Optimizing this equation can make the learned logits match the teacher's logits.
[0137]
[0138] Among them, _now_output in each formula is the output of a certain layer of the current network model, and _old_output is the output of the corresponding layer of the previous model.
[0139] The first to use intermediate layer features for distillation is FitNet. It directly allows the student model to fit the teacher model on the corresponding intermediate layer features. FitNet is a two-stage knowledge distillation, and its loss function is as follows.
[0140]
[0141] In the formula, u h , ν g and r represent the teacher model, the student model, and the adapter model that performs shape transformation on the feature map of the student model respectively. X is the input feature map, and W Hint , W Guided and W r represent the weights of the teacher model, the weights of the student model, and the weights of the adaptation model respectively.
[0142]
[0143] Among them, represents the feature loss function between the intermediate features of the current task and the intermediate features of the previous task, represents the feature loss function between the final output features of the current task and the final output features of the previous task model, where the loss function uses the mean squared error loss.
[0144] In this embodiment, 10 types of ground static military vehicle targets of X-band airborne SAR images are obtained based on the MSTAR dataset. The horizontal polarization mode is used for acquisition, the azimuth and range resolutions are 0.3 m, and the azimuth angle interval is about 0.03°.
[0145] Since the data of pitch angles 15° and 17° are relatively complete, 17° samples are selected as the base task for training in this experiment, and the 15° pitch angle dataset is segmented according to the number of tasks and used as the incremental dataset. Figure 7 Shows the optical images and SAR images of the targets in this dataset.
[0146] The target categories and corresponding sample numbers at different pitch angles are shown in Table 1.
[0147] Table 1 Sample Numbers of Ten Categories of MSTAR Targets
[0148]
[0149]
[0150] The platform used in this embodiment is a CPU with a model number of Intel i7 13700 2.4Ghz. The GPU is an NVIDIA RTX 4060 with a video memory capacity of 8G, a memory of 16G, and a video memory capacity of 8G. The Pytorch framework is used for model construction, training, and testing, and the operating system is Windows. The hardware environment configuration and software environment configuration are shown in Tables 2 and 3.
[0151] Table 2 Hardware Environment Configuration
[0152] Component Model GPU RTX4060 with 8GB VRAM CPU Intel i7-13700 2.4Ghz Memory 16GB
[0153] Table 3 Software Environment Configuration
[0154]
[0155] When conducting the study on angle adaptability, the focus is on the classification accuracy of the angle adaptive incremental learning method for the on-line arriving data stream. Therefore, the overall classification accuracy (OA), the average classification accuracy (AA), and the Kappa coefficient are used as evaluation indicators. The evaluation of the classification results is based on the pixel level and is obtained based on the classification confusion matrix:
[0156]
[0157] Among them, m ij represents the number of pixels of the j-th sampling result that is actually classified as belonging to the i-th category. N c is the number of classification categories. Based on the classification confusion matrix, OA and AA can be calculated as shown in the following formulas:
[0158]
[0159] In the above formula, N represents the total number of samples. m jj represents the number of samples correctly classified as category j; m +j represents the sum of the samples in the j-th column, and the Kappa coefficient represents the quality of the overall classification.
[0160] In this experiment, a sample with an angle of 17° was used as the basic task, and the dataset with a pitch angle of 15° was divided into four groups of tasks, which were sequentially fed into the network as incremental datasets. The key parameter distillation temperature T involved in the network was selected. The test set was divided into three cases, namely: the test set with a pitch angle of 17°, the test set with a pitch angle of 15°, and the mixed test set with pitch angles of 17° and 15°. The accuracy indexes on the three test sets are shown in Table 4.
[0161] Table 4 Incremental classification results under different temperature coefficients
[0162]
[0163]
[0164] Based on the experimental results in the above table, the Kappa coefficient for the mixed test set under different temperature coefficients can be plotted. Figure 8 as shown in
[0165] According to the comprehensive evaluation of the three test sets, when T = 15, the incremental classification accuracy of the algorithm is the highest, the Kappa coefficient on the mixed test set also reaches the highest, and the incremental classification effect is the best. Finally, T = 15 is selected as the distillation coefficient for this experiment.
[0166] The index results of the intermediate layer output and the final output for the last task are shown in Table 5.
[0167] Table 5 Index results of intermediate layer output and final output
[0168]
[0169] In this embodiment, a multi-source loss combining feature distillation and self-distillation mutual distillation is used to study the loss terms, namely the lack of feature distillation loss, the lack of mutual distillation loss, and the lack of self-distillation loss. Experiments are carried out in the above three cases to explore the necessity of the loss terms.
[0170] Table 2-6 Classification index results of different loss terms
[0171]
[0172]
[0173] Combined with the full loss term results of this experiment, according to the above experimental results, a line chart of the Kappa coefficient corresponding to different loss terms can be drawn, as shown in Figure 9 as shown.
[0174] The experiment verified the necessity of each loss term in this algorithm, and the algorithm result including the three distillation losses reached the optimal result.
[0175] The specific embodiments of the present invention disclosed above are only used to help illustrate the present invention. The specific embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. According to the content of this specification, many modifications and variations can be made. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the relevant art can understand and utilize the present invention well.
Claims
1. An image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles, characterized in that Including: Construct an angle adaptive learning framework, which includes a replay sample module, a data preprocessing module, a classification network module, and a distillation module; the classification network module adopts a composite knowledge distillation structure including self-distillation and mutual-distillation, where mutual-distillation conducts knowledge transfer through the output layer and intermediate layer features of the current model and the previous task model, and self-distillation conducts knowledge transfer through the outputs of multiple intermediate classifiers inside the classification network; The replay sample module uses an adaptive depth inversion network to generate pseudo samples, and this network optimizes the diversity and feature distribution consistency of the generated samples through the adversarial generation mechanism between the pre-trained teacher model and the student model; The method specifically includes: S1: Divide the basic dataset D and the new angle dataset D' into an incremental task sequence, where D' is divided into m incremental task subsets; S2: Use an adaptive depth inversion network to generate pseudo samples, and this network generates multi-angle SAR samples by jointly optimizing the classification loss, feature distribution regularization, image prior regularization, and competitive adversarial loss; S3: During the training process of the classification network, adopt a three-stage distillation mechanism: achieve self-distillation through the KL divergence calculation of the intermediate classifier, achieve mutual-distillation between models through response distillation and feature distillation, and combine the mixed training of real samples and generated samples.
2. The image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles according to claim 1, characterized in that The adaptive depth inversion network generates pseudo samples through the following loss function: L total = L classification + α f R feature + α TV R TV + α L2 R L2 + α c R compete Among them, L classification is the main loss, α f is the weight coefficient for feature distribution regularization, α TV is the weight coefficient for total variation regularization, α L2 is the weight coefficient for L2 regularization, α c is the weight coefficient for adaptive adversarial inversion regularization, R feature is the feature distribution regularization loss, R TV is the total variation regularization loss, R L2 is the L2 regularization loss, R compete is the adaptive loss.
3. An image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles according to claim 1, characterized in that, The response distillation loss in the mutual-distillation is: Among them, is the mutual distillation loss between the outputs of the first residual block of the model network, and \(L\) R \((.)\) is the KL divergence loss, \(z\) t_now_output1 , is the output of the first residual block of the current task model, \(z\) t_old_output1 is the output of the first residual block of the previous task model network, \(T\) is the temperature coefficient, and \(p(.)\) is the probability distribution of each category at temperature \(T\). is the mutual distillation loss between the outputs of the second residual block of the model network, \(z\) t_now_output2 is the output of the second residual block of the current task model, \(z\) t_old_output2 is the output of the second residual block of the previous task model network. is the mutual distillation loss between the final outputs of the model network, \(z\) t_now_output3 is the final output of the current task model, \(z\) t_old_output3 is the final output of the previous task model network, and \(loss\_2\) is the total mutual distillation loss.
4. An image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles according to claim 1, characterized in that The feature distillation loss in the mutual-distillation is: Among them, u h is an adapter model for shape transformation of the teacher model, ν g is an adapter model for shape transformation of the student model, r is an adapter model for shape transformation of the feature map of the student model, X is the input feature map, W Hint is the weight of the teacher model, W Guided is the weight of the student model, W r is the weight of the adaptation model; Among them, represents the feature loss function between the intermediate features of the current task and the intermediate features of the previous task, represents the feature loss function between the final output features of the current task and the final output features of the previous task model.
5. An image knowledge composite distillation optimization method for SAR target recognition with variable imaging angles according to claim 1, characterized in that The classification model adopts a multi-branch residual structure, including: two parallel Bottleneck modules, which respectively process spatial features of different scales; three intermediate classifiers are inserted in network layers at different depths, and each classifier includes a global average pooling layer and a fully connected layer; the final classifier fuses the input features of all intermediate classifiers.
6. The image knowledge composite distillation optimization method for SAR target recognition with imaging angle variation according to claim 1, wherein, The pseudo sample generation strategy of the replay sample module is: Before each round of incremental training, use the current teacher model to generate pseudo samples with the same number as the target number of categories; Dynamically adjust the noise initialization range of the generated samples, and the noise variance decays exponentially with the number of training rounds; Adopt a curriculum learning strategy to gradually relax the constraint intensity of feature distribution regularization.
7. An image knowledge composite distillation optimization method for SAR target recognition with changing imaging angles according to claim 1, characterized in that, The incremental task division method is: divide the new angle dataset D' into m subsets according to the pitch angle change interval, and each subset contains target samples of the same category; the pitch angle difference between adjacent subsets is controlled within the range of 0.5 - 1°.
8. An image knowledge composite distillation optimization method for SAR target recognition with variable imaging angles according to claim 1, characterized in that The data preprocessing module includes: performing random geometric transformations on the input SAR image, including random translation and flipping operations in the azimuth direction; performing covariance matrix normalization processing on the polarization channel data.
9. A computer device, characterized in that: Including a memory and a processor, where a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, the processor executes an image knowledge composite distillation optimization method for SAR target recognition with imaging angle change according to any one of claims 1 - 8.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, it executes the steps of an image knowledge composite distillation optimization method for SAR target recognition with imaging angle variation as described in any one of claims 1-8.
Citation Information
Cited By
Gesture recognition method and device based on space-time decoupling graph convolutional network
CN120431603A
A gesture recognition method and apparatus based on spatiotemporally decoupled graph convolutional networks
CN120431603B