Incremental Recognition Method for Radar Working Mode Classes Based on Attention Distillation

By introducing an incremental learning method based on attention distillation in multifunction radar working mode recognition, combining knowledge distillation of old training samples and new training samples, the problem of difficult to adapt to dynamic incremental observation scenarios and catastrophic forgetting in traditional methods is solved, and the continuous learning ability and efficient behavioral perception efficiency of incremental recognition of radar working mode is achieved.

CN118395239BActive Publication Date: 2025-06-17XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410604527.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-15
Publication Date
2025-06-17
Estimated Expiration
2044-05-15

AI Technical Summary

Technical Problem

Traditional multifunctional radar working mode recognition methods are difficult to adapt to dynamic incremental observation scenarios, and the inability to continuously learn complex and changeable unknown radar working modes, and directly using new data to update the model will cause catastrophic forgetting and lose the ability to identify old categories.

Method used

A radar working mode class increment recognition method based on attention distillation is proposed. By selecting the example set from the old training set in the previous stage and combining it with the new training samples, the training set of the current stage is constructed, and a class increment learning algorithm based on attention distillation is introduced during the training process to distillate the knowledge of feature vectors and intermediate layer feature maps.

Benefits of technology

This method can continuously learn the new category on the basis of not forgetting the old category, improve the behavioral perception efficiency of reconnaissance aircraft, alleviate the catastrophic forgetting problem, and improve the information control ability of complex and changeable battlefield environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118395239B_ABST
    Figure CN118395239B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for incremental recognition of radar working mode classes based on attention distillation, including: selecting a number of old training samples from the old training set of the previous stage to form an exemplar set, and forming the training set of the current stage together with the new training samples; inputting the training set into a pre-constructed incremental recognition network for radar working mode classes, and training the incremental recognition network for radar working mode classes based on attention distillation-based class incremental learning, so as to use the trained network for radar working mode recognition in the current stage; wherein, the incremental recognition network for radar working mode classes includes a parallel attention-temporal feature-aware prototype neural network and a distance classifier. This method can continuously update the existing old model with new samples. At the same time, the introduced class incremental learning algorithm based on attention distillation can alleviate the catastrophic forgetting of the new model for old classes, enabling the model to continuously learn new classes without forgetting old classes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of radar operating mode recognition, and particularly relates to a method for radar operating mode class incremental recognition based on attention distillation. Background Art

[0002] With the rapid development of technologies such as electronic information, microelectronics, and game theory, radars have gradually evolved from single-function to multi-functional and multi-task radar systems. Multifunction radars (MFRs) have characteristics such as agile beam changes and complex waveform and timing task scheduling. They can adaptively schedule strategies according to task requirements and external environmental factors to obtain better detection results. However, the above characteristics of multifunction radars make it particularly difficult for reconnaissance aircraft to invert their operating modes and mission intentions from reconnaissance pulse parameters, which seriously affects the threat assessment of the battlefield situation.

[0003] Regarding the topic of multifunction radar operating mode recognition, relevant scholars have carried out a large amount of work. Current multifunction radar operating mode recognition task models mostly target static closed-set mode data samples and do not have the ability to update for new modes and new samples. They cannot continuously learn complex and changing unknown radar operating modes and are difficult to adapt to dynamic incremental observation scenarios, resulting in a weakened information control ability for the battlefield environment. The open-set based incremental learning recognition algorithm provides a new idea for solving this problem. It refers to enabling the model to have the ability of lifelong learning through data stream learning under the limitation of limited memory resources.

[0004] However, directly updating the model with new data will result in catastrophic forgetting - that is, the model learns new classes while forgetting old classes and loses the discriminative ability for old classes, thus leading to a decline in the classification accuracy of the model. The catastrophic forgetting problem is mainly caused by the deep learning network structure and cannot be completely solved. Currently, it can only be alleviated through various algorithms. A deep neural network is determined by two parts: network parameters and network structure, and both aspects can cause catastrophic forgetting. First, in terms of network parameters, when training a new task, the update of network parameters may cause the knowledge learned previously to be forgotten. On the other hand, when the network structure is too complex, it may cause the model to overfit the new data when learning a new task, thus forgetting the knowledge learned previously. Some traditional neural network structures may lack a memory mechanism and cannot effectively store and retain the knowledge learned previously. If the network structure cannot adapt to changes in different tasks or data distributions, it may cause the knowledge learned previously to be forgotten when learning a new task.

[0005] In summary, most of the traditional methods for recognizing the working modes of multifunctional radars are targeted at static observation scenarios. The task models do not have the ability to continuously learn unknown modes, and it is difficult to effectively adapt to dynamic incremental observation scenarios, making it impossible to better adapt to the complex and ever-changing battlefield environment. Directly updating the model with new data will result in catastrophic forgetting and loss of the ability to recognize old categories, thus leading to a decline in the classification accuracy of the model. Therefore, how to enable the model to learn new categories while reducing catastrophic forgetting of old categories is the research difficulty of the current class incremental learning problem. Summary of the Invention

[0006] To solve the above problems existing in the prior art, the present invention provides a method for class incremental recognition of radar working modes based on attention distillation. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0007] The present invention proposes a method for class incremental recognition of radar working modes based on attention distillation, including:

[0008] Select a number of old training samples from the old training set of the previous stage to form an exemplar set, and form the training set of the current stage by combining the exemplar set with the new training samples;

[0009] Input the training set into a pre-constructed network for class incremental recognition of radar working modes, and train the network for class incremental recognition of radar working modes based on attention distillation, so as to use the trained network to recognize the radar working modes in the current stage;

[0010] Among them, the network for class incremental recognition of radar working modes includes a parallel attention-temporal feature-aware prototype neural network and a distance classifier. The parallel attention-temporal feature-aware prototype neural network is used to extract features from the input data to obtain feature vectors; the distance classifier is used to classify the feature vectors to obtain the recognition result of the radar working mode.

[0011] Advantages of the present invention:

[0012] The method for incremental recognition of radar working mode classes based on attention distillation provided by the present invention, on the one hand, draws on the idea of hard example mining, selects an example set from the old training samples of the previous stage, and together with the new training samples, constitutes the training set of the current stage to train the network; on the other hand, in the training process, an attention distillation-based class incremental learning algorithm is introduced. By performing knowledge distillation on the feature vectors and the feature maps of the intermediate layers, the catastrophic forgetting of the new model for the old categories is greatly alleviated. This method can continuously update the existing model with new samples, so that the new model can not only continuously learn new categories without forgetting the old categories, but also enable the radar working mode class incremental recognition network to have the ability to continuously learn new categories without forgetting the old categories, thereby improving the behavior perception efficiency of the reconnaissance aircraft and having strong practical significance for improving our reconnaissance and countermeasure capabilities.

[0013] The following will further elaborate on the present invention in conjunction with the accompanying drawings and embodiments. Description of the Drawings

[0014] Figure 1 is a schematic flowchart of a method for incremental recognition of radar working mode classes based on attention distillation provided by an embodiment of the present invention;

[0015] Figure 2 is a schematic structural diagram of a parallel attention-temporal feature perception prototype neural network provided by an embodiment of the present invention;

[0016] Figure 3 is a schematic diagram of the principle of the training and testing process of a radar working mode class incremental recognition network provided by an embodiment of the present invention. Detailed Embodiment

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] In response to the challenge that the traditional multi-functional radar working mode recognition task model does not have the ability to continuously learn unknown modes, the present invention introduces an incremental learning algorithm and proposes a method for incremental recognition of radar working mode classes based on attention distillation. In the implementation process of this method, when new samples (i.e., new working modes) appear in a certain stage, the training set is jointly constructed by combining the old samples and the new samples, and the network is trained with class incremental learning based on attention distillation, so that the network can not only maintain the memory ability of the known working modes in the incremental recognition of multi-functional radar working mode classes, but also continuously learn unknown working modes. Finally, the performance of the UI network is evaluated through the test set.

[0019] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a radar working mode class incremental recognition method based on attention distillation provided by an embodiment of the present invention. The method mainly includes:

[0020] Step 1: Select several old training samples from the old training set of the previous stage to form an exemplar set, and form the training set of the current stage together with the new training samples.

[0021] Step 2: Input the training set into a pre-constructed radar working mode class incremental recognition network, and train the radar working mode class incremental recognition network based on attention distillation-based class incremental learning, so as to use the trained network to perform radar working mode recognition in the current stage.

[0022] The radar working mode class incremental recognition method based on attention distillation provided by this embodiment will be introduced in detail below.

[0023] First, the pre-constructed radar working mode class incremental recognition network will be introduced.

[0024] The radar working mode class incremental recognition network constructed in this embodiment mainly includes a parallel attention-temporal feature perception prototype neural network and a distance classifier. The parallel attention-temporal feature perception prototype neural network is used to extract features from the input data to obtain feature vectors; the distance classifier is used to classify the feature vectors to obtain the radar working mode recognition result.

[0025] As an optional implementation manner, the parallel attention-temporal feature perception prototype neural network constructed in this embodiment is denoted as M, and mainly includes multiple encoders, a DANet layer, a position encoding layer, multiple linear layers, and multiple ReLU activation layers. The specific structure of the parallel attention-temporal feature perception prototype neural network is as Figure 2 shown, including a first linear layer, a first ReLU activation layer, a position encoding layer, multiple encoders, a DANet layer, and a second linear layer in sequence. Among them, the first linear layer serves as the input port of the parallel attention-temporal feature perception prototype neural network, and its input data is the PDW (Pulse Description Word) parameter of the radar, which mainly includes parameters such as the carrier frequency, PRI (pulse repetition time), bandwidth, amplitude, and pulse width of the radar; the second linear layer serves as the output port of the parallel attention-temporal feature perception prototype neural network and outputs feature vectors.

[0026] Specifically, in this embodiment, the first linear layer contains 5 nodes and is linearly connected to 16 nodes; the second linear layer contains 16 nodes and is linearly connected to 5 nodes.

[0027] The position encoding layer adds a learnable special character x cls and a learnable position encoding E pos , that is

[0028] z0 = [x cls ; x] + E pos ;

[0029] In the formula, z0 represents the output of the position encoding layer, and x represents the input of the position encoding layer.

[0030] Optionally, in this embodiment, 6 encoder modules are provided, and the structure of each encoder module is the same, and each includes: a first normalization layer, a multi-head attention layer, a first Dropout layer, a second normalization layer, a first dilated convolution layer, a second ReLU activation layer, a second Dropout layer, a second dilated convolution layer, a third ReLU activation layer, a third Dropout layer, a third dilated convolution layer, a fourth ReLU activation layer, and a fourth Dropout layer.

[0031] Among them, the first normalization layer, the multi-head attention layer, and the first Dropout layer are connected in sequence; the output of the first Dropout layer is added to the input of the first normalization layer and then used as the input of the second normalization layer; the second normalization layer, the first dilated convolution layer, the second ReLU activation layer, the second Dropout layer, the second dilated convolution layer, the third ReLU activation layer, the third Dropout layer, the third dilated convolution layer, the fourth ReLU activation layer, and the fourth Dropout layer are connected in sequence; the input of the second normalization layer is added to the output of the fourth Dropout layer and then used as the output of the encoder module.

[0032] Furthermore, the normalization layer in the encoder is Layer Normalization (LN) layer normalization, the multi-head attention layer in the encoder has 5 heads, dropout is used in the encoder to avoid overfitting, the convolution kernel size of the first dilated convolution layer in the encoder is 3, the stride is 1, the padding is 1, and the dilation factor is 1, the convolution kernel size of the second dilated convolution layer in the encoder is 3, the stride is 1, the padding is 2, and the dilation factor is 2, the convolution kernel size of the third dilated convolution layer in the encoder is 3, the stride is 1, the padding is 4, and the dilation factor is 4.

[0033] As an alternative implementation, the DANet layer mainly includes a position attention module and a channel attention module. Among them, the position attention module is used to extract position features from the input feature matrix to obtain corresponding position features. It mainly includes 3 Reshape layers, a Reshape-Transpose layer, and a Softmax layer. The specific connection relationship is as Figure 2 shown.

[0034] The channel attention module is used to extract channel features from the input feature matrix to obtain corresponding channel features. It mainly includes 3 convolutional layers, a Reshape layer, a Reshape-Transpose layer, and a Softmax layer. Among them, the structural models of the 3 convolutional layers are exactly the same, and the kernel size of each convolutional layer is 1.

[0035] The channel-position features are obtained by adding the channel features and the position features.

[0036] It can be understood that in addition to Figure 2 the parallel attention-temporal feature perception prototype neural network structure shown, in the process of implementing the present invention, other network structures can also be used to perform feature extraction.

[0037] Furthermore, for the distance classifier, it mainly calculates the Euclidean distance from the feature vector of each training sample to each class prototype, and takes the class of the class prototype corresponding to the shortest Euclidean distance as the classification result, that is, the radar working mode recognition result.

[0038] Next, the training set construction process in the radar working mode class incremental recognition method provided by the present invention is introduced, that is, the specific implementation process of step 1.

[0039] It should be noted that in the initialization stage, the training set of all classes in the initial stage is directly used to perform open-set learning to obtain the original model M1. At this time, M1 can recognize known classes and reject unknown classes.

[0040] When new samples appear in a certain stage, it is necessary to reconstruct the training set to perform class incremental learning training on the model trained in the previous stage, so that the model can learn new working modes without forgetting the old classes.

[0041] Optionally, in step 1, several old training samples are selected from the old training set of the previous stage to form an exemplar set, specifically including:

[0042] 11) First, use the new training samples of the current stage to update the old model M t-1 to obtain the model Among them, the old model Mt-1 It is a parallel attention-temporal feature-aware prototype neural network trained in the previous stage.

[0043] 12) Then, the old training samples from the previous stage are respectively input into the old model M t-1 and the model to correspondingly obtain the feature vectors z and z * .

[0044] 13) Finally, in each old category, K old training samples with the largest Euclidean distance between the feature vectors z and z * are selected to form an exemplar set, denoted as where the Euclidean distance calculation formula for the feature vectors z and z * is:

[0045] ΔL = ||z - z * ||2;

[0046] In the formula, ||·||2 represents the Euclidean norm.

[0047] After obtaining the exemplar set , it is merged with the new samples in the current stage to obtain the training set in this stage

[0048] It can be understood that here, the old test samples in stage T t-1 and the new test samples in the current stage can also be merged to obtain the test set in the current stage for performance testing after the network training is completed.

[0049] Figure 3 Furthermore, in combination with Figure 3 the schematic diagram of the principle of the training and testing process of the radar working mode class incremental recognition network shown, the class incremental learning based on attention distillation in step 2 is used to train the radar working mode class incremental recognition network, specifically including:

[0050] Step a: Take the old model M t-1 as the original model of the new model M t and copy the weights and parameters of the old model M t-1 to the new model M t ; The new model M t is the parallel attention-temporal feature-aware prototype neural network to be trained in the current stage.

[0051] Step b: Input the training set in the current stage into the old model M t-1 and the new model M tFeature extraction is performed, and the MDCE (margin Distance based Cross Entropy) loss, MPL (margin prototype loss) loss, knowledge distillation loss, and attention distillation loss are calculated.

[0052] First, the training set of the current stage is input into the old model M t-1 , and the feature vector g t-1 (x) can be obtained; correspondingly, the training set of the current stage is input into the new model M t , and the feature vector g t (x) can be obtained.

[0053] Incidentally, the class prototype A = {a t | i = 1, 2, … N} corresponding to the training set of the current stage can also be calculated according to the feature vector g i , where

[0054]

[0055] In the formula, n i represents the number of samples in the i-th class training set.

[0056] Then, various losses are calculated; among them, the calculation formula for the MDCE loss is:

[0057]

[0058] In the formula, L MDCE represents the MDCE loss function, K represents the number of training samples in a batch, N represents the number of sample categories, q(y) represents the distribution of sample labels, p(y|x) represents the probability that the training sample x belongs to the category y, and margin represents a hyperparameter.

[0059] The calculation formula for the MPL loss is:

[0060]

[0061] In the formula, L MPL represents the MPL loss function, g t (x) represents the feature vector obtained by inputting the training set of the current stage into the new model M t , a y represents the prototype of the correct category to which the sample belongs, ||·||2 represents the Euclidean norm, and margin represents a hyperparameter.

[0062] The calculation formula for the knowledge distillation loss is:

[0063] L D = ρ ||f t (x) - f t-1 (x)||₂;

[0064] In the formula, L D represents the knowledge distillation loss function, f t (x) represents the feature vector obtained by inputting the old training samples into the new model M t , and f t-1 (x) represents the feature vector obtained by inputting the old training samples into the old model M t-1 . ||·||₂ represents the Euclidean norm.

[0065] ρ represents the weight parameter reflecting the quantitative relationship between the new training samples and the old training samples, and its expression is:

[0066]

[0067] In the formula, N all represents the total number of samples in the training set, N new represents the total number of samples in the new class, and N old represents the total number of samples in the exemplar set.

[0068] The calculation formula of the attention distillation loss is:

[0069]

[0070] In the formula, L AD represents the attention distillation loss function, k represents the number of temporal feature perception modules, and A t ∈ R 1×H×W represents the feature map output by each temporal feature perception module of the new model M t ;

[0071] A t-1 ∈ R 1×H×W represents the feature map output by each temporal feature perception module of the old model M t-1 , ||·||₂ represents the Euclidean norm, and ||·|| p represents the p-norm.

[0072] Step c, calculate the combined loss function according to the MDCE loss, MPL loss, distillation loss, and attention distillation loss.

[0073] The calculation formula of the combined loss function is:

[0074]

[0075] In the formula, Loss represents the combined loss function, and N all is the total number of samples in the training set, is a training set sample, L D represents the knowledge distillation loss function, L AD represents the attention distillation loss function, L MDCE represents the MDCE loss function, L MPL represents the MPL loss function, and λ represents a coefficient with a value between 0 and 1, which can be taken as 0.25 here.

[0076] Step d, backpropagate the joint loss function to update the network parameters of the new model M t of.

[0077] Step e, repeat steps b - d to iteratively train the new model M t until the joint loss function converges, the network training ends, and the trained new model M t is obtained, thus completing the training of the radar working mode class incremental recognition network.

[0078] It can be understood that after the training of the new model M t is completed, it also includes:

[0079] Use the test set of the current stage to test the radar working mode class incremental recognition network to evaluate the network performance, such as Figure 3 shown.

[0080] Among them, the test set of the current stage includes the new test samples of the current stage and the old test samples of the previous stage that is, the test set of the current stage constructed above

[0081] Specifically, when testing the radar working mode class incremental recognition network, the prediction result is obtained through the distance between the distribution of the test set of the current stage in the feature space and the class prototype corresponding to the training set of the current stage. The formula for the predicted label is:

[0082]

[0083] In the formula, y pred represents the predicted label of the test sample data, i represents the number of classes, N represents the number of sample categories, and V i (x) represents the discriminant function of the i-th class, and its expression is:

[0084] V i (x) = ||g t (x) - α i ||2;

[0085] In the formula, g t (x) represents inputting the training set of the current stage into the new model M tThe obtained feature vector, α i represents the class prototype of the i-th class corresponding to the training set in the current stage, and ||·||2 represents the Euclidean norm.

[0086] After the test is completed, the newly trained network can be used for incremental recognition of radar working mode classes. The PDW parameters of the recognition radar are input into the incremental recognition network of radar working mode classes. First, the parallel attention-temporal feature-aware prototype neural network is used to extract features from the input data to obtain the feature vector; then, the distance classifier is used to classify the feature vector, and finally, the recognition result of the radar working mode is obtained.

[0087] The method for incremental recognition of radar working mode classes based on attention distillation provided by the present invention, on the one hand, draws on the idea of hard example mining, selects an example set from the old training samples in the previous stage, and together with the new training samples, constitutes the training set in the current stage to train the network; on the other hand, in the training process, an attention distillation-based class incremental learning algorithm is introduced, and knowledge distillation is carried out on the feature vector and the feature map of the intermediate layer, which greatly alleviates the catastrophic forgetting of the new model for the old categories. This method can continuously update the existing model with new samples, so that the new model can continuously learn new categories without forgetting the old categories, enabling the incremental recognition network of radar working mode classes to have the ability to continuously learn new categories without forgetting the old categories, thereby improving the behavior perception efficiency of the reconnaissance aircraft and having strong practical significance for improving our reconnaissance and countermeasure capabilities.

[0088] The effectiveness of the method proposed by the present invention will be verified and illustrated through experiments below.

[0089] Taking the signal of a non-cooperative multi-functional radar intercepted by a certain receiver as an example in this experiment, the parameter selection for its different working modes is shown in Table 1.

[0090] Table 1 PDW parameter table of eight multi-functional radar working modes

[0091]

[0092] To verify the recognition effect of the method proposed by the present invention, the above-mentioned measured data set is used for three rounds of training, and its recognition accuracy is calculated. The results are shown in Table 2.

[0093] Table 2 Recognition accuracy table of measured data

[0094] K Initial network First-round increment Second-round increment Third-round increment Average recognition accuracy 100 96.96% 92% 89.71% 87.95% 89.89% 200 96.96% 92.53% 90.51% 89.1% 90.71% 300 96.96% 94.13% 92.69% 92.15% 92.99%

[0095] As can be seen from Table 2, with the increase in the radar working modes of network learning, the recognition accuracy of the algorithm continuously decreases, and catastrophic forgetting is inevitable. However, the experiment also shows that after three rounds of class incremental learning, the recognition accuracy in the worst-case scenario of the algorithm only decreases by about 9%, and the final recognition accuracy is 87.95%, greatly alleviating the problem of catastrophic forgetting. It meets the requirements of engineering practice and has high engineering application value.

[0096] The present invention introduces incremental learning into the recognition of MFR working modes, realizes the leap of the network model from open-set recognition to class incremental recognition, and at the same time performs knowledge distillation on the feature vectors and feature maps of the new and old models, greatly alleviating the catastrophic forgetting of the model in the MFR working mode recognition task.

[0097] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A radar working mode incremental recognition method based on attention distillation, characterized in that: include: Selecting a number of old training samples from the old training set of the previous stage to form a sample set, and combining the sample set with the new training samples to form a training set of the current stage; Inputting the training set into a pre-built radar working mode class incremental recognition network, training the radar working mode class incremental recognition network based on class incremental learning of attention distillation, so as to use the trained network to perform radar working mode recognition at the current stage; The radar working mode class incremental recognition network includes a parallel attention-series feature perception prototype neural network and a distance classifier. The parallel attention-series feature perception prototype neural network is used to extract features from input data to obtain a feature vector; the distance classifier is used to classify the feature vector to obtain a radar working mode recognition result; The parallel attention-temporal feature perception prototype neural network includes a first linear layer, a first ReLU activation layer, a position encoding layer, multiple encoders, a DANet layer and a second linear layer in sequence; wherein the first linear layer serves as an input port of the parallel attention-temporal feature perception prototype neural network, and its input data is the PDW parameter of the radar; the second linear layer serves as an output port of the parallel attention-temporal feature perception prototype neural network, and outputs a feature vector.

2. The radar working mode class incremental recognition method based on attention distillation according to claim 1 is characterized in that: Select several old training samples from the old training set in the previous stage to form a sample set, including: Use new training samples in the current stage For old model M t-1 Update and get the model Among them, the old model M t-1 The parallel attention-temporal feature perception prototype neural network trained in the previous stage; The old training samples from the previous stage Input to the old model M t-1 and the model In the equation, we get the corresponding eigenvectors z and z * ; In each old category, select the feature vectors z and z respectively * The K old training samples with the largest Euclidean distance constitute the example set.

3. The radar working mode class incremental recognition method based on attention distillation according to claim 2 is characterized in that: Based on the class incremental learning of attention distillation, the class incremental recognition network of the radar working mode is trained, specifically including: Step a: the old model M t-1 As the new model M t The original model and the old model M t-1 The weights and parameters are copied to the new model M t Above: The new model M t It is the parallel attention-temporal feature perception prototype neural network to be trained in the current stage; Step b: input the training set of the current stage into the old model M t-1 And the new model M t Perform feature extraction and calculate MDCE loss, MPL loss, knowledge distillation loss, and attention distillation loss; Step c, calculating a joint loss function according to the MDCE loss, the MPL loss, the distillation loss and the attention distillation loss; Step d: Back propagate the joint loss function to the new model M t Update the network parameters of Step e: repeat steps b to d to modify the new model M. t Perform iterative training until the joint loss function converges to obtain a trained new model M t , thereby completing the training of the radar working mode class incremental recognition network.

4. The radar working mode class incremental recognition method based on attention distillation according to claim 3 is characterized in that: The calculation formula of the MDCE loss is: Where, L MDCE represents the MDCE loss function, K represents the number of training samples in a batch, N represents the number of sample categories, q(y) represents the distribution of sample labels, p(y|x) represents the probability that training sample x belongs to category y, and margin represents a hyperparameter.

5. The radar working mode class incremental recognition method based on attention distillation according to claim 3 is characterized in that: The calculation formula of the MPL loss is: Where, L MPL represents the MPL loss function, g t (x) indicates that the training set of the current stage is input into the new model M t The obtained feature vector, a y represents the prototype of the correct category to which the sample belongs, ||·||2 represents the Euclidean norm, and margin represents a hyperparameter.

6. The radar working mode class incremental recognition method based on attention distillation according to claim 3 is characterized in that: The calculation formula of the knowledge distillation loss is: L D =ρ||f t (x)-f t-1 (x)||2; Where, L D represents the knowledge distillation loss function, ρ represents the weight parameter reflecting the relationship between the number of new training samples and old training samples, and f t (x) represents the input of old training samples into the new model M t The obtained eigenvector, f t-1 (x) indicates that the old training samples are input into the old model M t-1 The obtained eigenvector, ||·||2 represents the Euclidean norm.

7. The radar working mode class incremental recognition method based on attention distillation according to claim 3 is characterized in that: The calculation formula of the attention distillation loss is: Where, L AD represents the attention distillation loss function, ρ represents the weight parameter reflecting the relationship between the number of new training samples and the number of old training samples, k represents the number of temporal feature perception modules, and A t ∈R 1×H×W Represents the new model M t The feature map output by each temporal feature perception module; A t-1 ∈R 1×H×W Indicates the old model M t-1 Each time series feature perception module outputs a feature map,||·||2 represents the Euclidean norm,||·|| p represents the p-norm.

8. The radar working mode incremental recognition method based on attention distillation according to claim 3 is characterized in that: The calculation formula of the joint loss function is: In the formula, Loss represents the joint loss function, N all is the total number of samples in the training set, is the training set sample, L D represents the knowledge distillation loss function, L AD represents the attention distillation loss function, L MDCE represents the MDCE loss function, L MPL represents the MPL loss function, and λ represents a coefficient with a value between 0 and 1.

9. The radar working mode incremental recognition method based on attention distillation according to claim 1 is characterized in that: After completing the training of the radar working mode incremental recognition network, it also includes: Using the test set of the current stage to test the radar working mode class incremental recognition network to evaluate the network performance; The test set of the current stage includes the new test samples of the current stage. and the old test samples from the previous stage 10. The radar working mode class incremental recognition method based on attention distillation according to claim 9 is characterized in that: When testing the radar working mode class incremental recognition network, the prediction result is obtained by the distance between the distribution of the current stage test set in the feature space and the class prototype corresponding to the current stage training set. The formula for predicting the label is: In the formula, y pred represents the predicted label of the test sample data, i represents the number of classes, N represents the number of sample categories, V i (x) represents the discriminant function of the i-th category, and its expression is: V i (x)=||g t (x)-a i ||2; In the formula, g t (x) indicates that the training set of the current stage is input into the new model M t The obtained feature vector, α i represents the class prototype of the i-th class corresponding to the training set in the current stage, and ||·||2 represents the Euclidean norm.

Citation Information

Patent Citations

  • SAR target class increment identification method based on knowledge robust-heavy balance network

    CN116129219A

  • Deep network incremental learning method for multi-view children tumor pathological image classification

    CN116363461A