SAR Target Incremental Recognition Method Based on Dynamic Structure and Multi-Level Distillation
Through dynamic structure and multi-level distillation training mode, the problem of catastrophic forgetting in SAR target recognition is solved, the recognition accuracy and adaptability of incremental learning are improved, and efficient SAR target recognition is achieved.
Patent Information
- Application Number
- CN202310774347.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-06-28
AI Technical Summary
The existing SAR target recognition methods have catastrophic forgetting when facing incremental learning tasks, making it difficult to adapt to new goals and new data, and the recognition rate is low.
Using a training model based on dynamic structure and multi-level distillation, the backbone network structure is expanded through training auxiliary networks, and the knowledge distillation loss function is used to migrate old model knowledge to slow catastrophic forgetting and improve feature separability.
It significantly improves the accuracy of incremental learning recognition of SAR targets, reduces feature confusion among tasks, and enhances the model's adaptability to new tasks and new data.
Smart Images

Figure CN116977870B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of radar remote sensing, and further relates to a method for SAR target incremental recognition, which can be used for battlefield reconnaissance and situation awareness. Background Art
[0002] Synthetic Aperture Radar (SAR) is an active microwave imaging sensor that obtains two-dimensional high-resolution images by transmitting signals with a large time-bandwidth product and using aperture synthesis. Compared with optical and infrared sensors, SAR has unique advantages of all-weather operation, long operating range, and strong penetration ability, making it an important means for earth observation and widely used in military and civilian fields. With the continuous improvement of SAR systems and the continuous improvement of SAR imaging levels, SAR image interpretation technology has gradually attracted the attention of scholars and researchers in related fields. As a difficult and key step among them, the accurate recognition of key targets has important significance and research value.
[0003] Traditional SAR target recognition methods mainly design manual features and construct classifiers based on the statistical information and physical characteristics of images. However, this requires strong professional knowledge and expert experience, and the accuracy and flexibility of the algorithms are poor, making it difficult to achieve ideal results in practical applications.
[0004] In recent years, with the continuous development of deep learning technology, target recognition methods based on deep neural networks have made significant breakthroughs in the field of computer vision.
[0005] Although the above-mentioned deep learning-based methods provide a feasible approach for SAR target recognition, compared with optical images, the SAR image scene is more complex, the similarity of different category targets is higher, and affected by speckle noise, the edges of the targets are not clear. Therefore, there are still problems of being not robust in complex environments and difficult to distinguish similar categories, resulting in more difficulties for SAR image recognition in the face of incremental learning tasks and more serious catastrophic forgetting phenomena.
[0006] The patent document with the application number 201711257577.0 discloses a "SAR vehicle target recognition method based on an improved convolutional neural network". First, it removes the background clutter in each image of the training samples and crops each SAR image. Then, it constructs an improved convolutional neural network structure based on the caffe architecture, that is, sets the classifier in the target recognition part of the convolutional neural network to a hybrid maximum margin softmax. Finally, it inputs the cropped training samples into the improved convolutional neural network for training to obtain a trained network model, and after removing the background clutter and cropping the test samples, inputs them into the trained improved convolutional neural network model for testing to obtain its recognition rate. However, since this method does not perform relevant optimizations for SAR target incremental learning, it is difficult to adapt to new targets and new data, and there is serious catastrophic forgetting in the incremental learning tasks of multiple scenarios. The targets in different scenarios affect each other, resulting in a low recognition rate. Summary of the Invention
[0007] The purpose of the present invention is to propose a SAR target incremental learning method based on dynamic structure and multi-level distillation for the deficiencies of the above-mentioned existing technologies, so as to alleviate catastrophic forgetting in multi-scenario incremental learning tasks, enhance the separability of intra-task features, reduce inter-task feature confusion, and significantly improve the recognition accuracy of SAR target incremental learning.
[0008] The technical idea of the present invention is to improve the ability of the model to adapt to new tasks and new data in multi-scenario tasks through a training mode of expanding the dynamic structure of the backbone network and multi-level distillation by a training auxiliary network. The implementation steps are as follows:
[0009] (1) Obtain multiple SAR images of multiple types of targets from the initial task data, and randomly divide the images and labels to obtain an initial task training set and an initial task test set;
[0010] (2) Construct an initial backbone network:
[0011] (2a) Establish a feature extractor A of the initial backbone network composed of cascaded multiple residual blocks init , which is used to output multi-scale feature maps;
[0012] (2b) Establish a fully connected layer FC that can match the feature dimension and output category init ;
[0013] (2c) Cascade the feature extractor and the fully connected layer to form an initial backbone network, and select cross-entropy loss as its classification loss;
[0014] (3) Randomly sample a group of SAR images from the training set and input them into the initial backbone network to calculate the loss, and update the network parameters through the stochastic gradient descent algorithm until the network converges to obtain a trained initial backbone network.
[0015] (4) Input the SAR images in the initial task test set into the trained initial network to obtain the recognition results;
[0016] (5) Set up an empty example set, and select less than 10% of the samples from the initial task data according to the herding algorithm and put them into this example set;
[0017] (6) Obtain multiple SAR images of multiple types of targets from the incremental task data, and randomly divide their images and labels to obtain an auxiliary training set and an auxiliary test set;
[0018] (7) Construct an auxiliary network:
[0019] (7a) Establish a feature extractor A of the auxiliary network composed of cascaded multiple residual blocks aux , which is used to output multi-scale feature maps;
[0020] (7b) Establish a fully connected layer FC that can match the feature dimension and output categories of the new task aux ;
[0021] (7c) Cascade the feature extractor and the fully connected layer to form an auxiliary network, and select the cross-entropy loss as its classification loss;
[0022] (8) Randomly sample a group of SAR images from the auxiliary training set and input them into the auxiliary network to calculate the loss. Based on this loss, update the network parameters through the stochastic gradient descent algorithm until the network converges to obtain the trained auxiliary network;
[0023] (9) Use the auxiliary network trained in (8) and the initial network in step (3) to construct a backbone network:
[0024] (9a) Parallelly connect the feature extractor A in the auxiliary network aux with the feature extractor A in the initial network init to obtain the backbone network feature extractor A new and the multi-scale feature map Y output by it new ;
[0025] (9b) Construct a new fully connected layer FC according to the output dimension and classification number of the parallel feature extractor new ;
[0026] (9c) Initialize some parameters of the fully connected layer in (9b) with the parameters of the fully connected layer of the initial network;
[0027] (9d) Initialize the remaining parameters of the fully connected layer in (9b) with the parameters of the fully connected layer of the auxiliary network;
[0028] (9e) Cascading the feature extractor after parallel connection and the new fully-connected layer to form a new backbone network, and using cross-entropy loss Knowledge distillation loss between the new and old models And the knowledge distillation loss in the feature space These three losses are used as its classification loss
[0029] (10) Obtaining multiple SAR images of multi-class targets from the incremental task data and the example set data, and randomly partitioning their images and labels to obtain an incremental training set and an incremental test set;
[0030] (11) Randomly sampling a group of SAR images from the incremental training set and inputting them into the backbone network to calculate its loss, and iteratively updating the network parameters based on this loss through the stochastic gradient descent algorithm until the network converges to obtain a trained backbone network;
[0031] (12) Inputting the SAR images in the incremental test set into the trained backbone network to obtain recognition results.
[0032] Compared with the prior art, the present invention has the following advantages:
[0033] First, since the present invention designs a dynamic structure for expanding the backbone network by training an auxiliary network, it avoids the coverage of new knowledge over old knowledge and effectively improves the adaptability of the model to new data and new tasks;
[0034] Second, since the present invention designs a knowledge distillation loss function for the new and old models Transferring the knowledge in the old model and the auxiliary model to the new backbone network effectively alleviates catastrophic forgetting;
[0035] Third, since the present invention designs a knowledge distillation loss function based on the feature space Using the feature extractor of the old model to guide the new feature extractor to learn the feature differences between tasks effectively improves the separability of different task features in the SAR target incremental learning process.
[0036] The simulation results show that the accuracy of the present invention is improved by 74.76% compared with the existing SAR target recognition method. Description of the Drawings
[0037] Figure 1 It is a flowchart for the implementation of the present invention;
[0038] Figure 2 It is a diagram of the initial model constructed in the present invention;
[0039] Figure 3 It is a diagram of the auxiliary model constructed in the present invention;
[0040] Figure 4 It is the backbone network model diagram constructed in the present invention;
[0041] Figure 5 It is the simulation experiment result diagram of the present invention and the existing method when the initial task is 2 classes;
[0042] Figure 6 It is the simulation experiment result diagram of the present invention and the existing method when the initial task is 8 classes. Specific implementation manners
[0043] The following further elaborates on the examples and effects of the present invention in conjunction with the accompanying drawings.
[0044] Refer to Figure 1 , the implementation steps of this example are as follows:
[0045] Step 1, construct an initial task training set and an initial task test set.
[0046] Obtain multiple SAR images of multiple classes of targets from the known initial task data, and randomly divide the images and labels according to a ratio of 7:3 to obtain an initial task training set and an initial task test set.
[0047] Step 2, construct an initial backbone network.
[0048] Refer to Figure 2 , the implementation of this step is as follows:
[0049] 2.1) Establish an initial backbone network feature extractor A containing 5 cascaded convolutional modules init , where:
[0050] The first convolutional module is cascaded by a 3×3 standard convolutional layer, a batch normalization layer, a ReLU activation layer, and a 3×3 max-pooling downsampling layer;
[0051] The second convolutional module is cascaded by 2 residual blocks with an input dimension of 64 and an output dimension of 64;
[0052] The third convolutional module is cascaded by 2 residual blocks with an input dimension of 128 and an output dimension of 128;
[0053] The fourth convolutional module is cascaded by 2 residual blocks with an input dimension of 256 and an output dimension of 256;
[0054] The fifth convolutional module is cascaded by 2 residual blocks with an input dimension of 512 and an output dimension of 512 and a 1×1 adaptive average pooling layer;
[0055] In this example, the initial drying network feature extractor A init is used to extract multi-scale feature maps. For an input SAR image with width and height being W and H respectively A init the output multi-scale feature maps are expressed as:
[0056]
[0057] where Y i init is the output feature map of the i-th convolutional module ;
[0058] 2.2) Construct the fully-connected layer FC of the initial drying network init :
[0059] The size of this fully-connected layer FC of the initial drying network init is: H init ×N init , where H init is the length of the features extracted by the initial drying network feature extractor A init , and N init is the total number of categories to be learned for the initial task;
[0060] In this example, FC init is used to output the recognition result, expressed as:
[0061] O init = FC init (A init (X)),
[0062] where A init is the initial drying network feature extractor, and X is the SAR image with width and height being W and H respectively.
[0063] 2.3) Concatenate A init and FC init to form the initial backbone network.
[0064] Step three, train the initial backbone network.
[0065] 3.1) Randomly sample a group of SAR images from the initial task training set and input them into the initial drying network to calculate its loss. Based on this loss, update the parameters of the initial drying network through the stochastic gradient descent algorithm:
[0066] 3.1.1) Calculate the loss of the initial drying network
[0067]
[0068] Among them, is a SAR image with width and height being W and H respectively, and y X is the label corresponding to the SAR image X;
[0069] 3.1.2) Solve the initial dry network loss in 3.1.1) The gradient of the initial dry network parameter θ init is expressed as
[0070] 3.1.3) Update the initial dry network parameter according to the gradient solved in 3.1.2) :
[0071]
[0072] where θ i ′ nit is the current updated network parameter, θ init is the network parameter before update; lr is the learning rate, which is set according to the batch size of the input image. In this example, let lr = 0.005.
[0073] 3.2) Repeat step 3.1) until the network converges to obtain the trained initial dry network.
[0074] Step four, identify the initial task and construct an example set.
[0075] 4.1) Input the SAR images in the initial task test set into the trained initial dry network to obtain the initial task recognition result;
[0076] 4.2) Set an empty example set, and select less than 10% of the samples from the initial task data according to the herding algorithm and put them into this example set.
[0077] Step five, construct an incremental task training set and an incremental task test set.
[0078] Obtain multiple SAR images of multiple types of targets from the known incremental task data, and randomly divide its images and labels according to the ratio of 7:3 to obtain an incremental task training set and an incremental task test set.
[0079] Step six, construct an auxiliary network.
[0080] Refer to Figure 3 , the implementation of this step is as follows:
[0081] 6.1) Establish an auxiliary network feature extractor A containing 5 cascaded convolutional modules aux , where:
[0082] The first convolutional module It is composed of a 3×3 standard convolutional layer, a batch normalization layer, a ReLU activation layer, and a 3×3 max-pooling downsampling layer cascaded together;
[0083] The second convolutional module It is composed of 2 residual blocks with an input dimension of 64 and an output dimension of 64 cascaded together;
[0084] The third convolutional module It is composed of 2 residual blocks with an input dimension of 128 and an output dimension of 128 cascaded together;
[0085] The fourth convolutional module It is composed of 2 residual blocks with an input dimension of 256 and an output dimension of 256 cascaded together;
[0086] The fifth convolutional module It is composed of 2 residual blocks with an input dimension of 512 and an output dimension of 512 and a 1×1 adaptive average pooling layer cascaded together;
[0087] In this example, the auxiliary network feature extractor A aux is used to extract multi-scale feature maps. For an input SAR image with width and height of W and H respectively The feature extractor A aux The output multi-scale feature maps are expressed as:
[0088]
[0089] where Y i aux is the output feature map of the i-th convolutional module ;
[0090] 6.2) Construct the auxiliary network fully connected layer FC aux :
[0091] The size of this auxiliary network fully connected layer FC aux is: H aux ×N aux , where H aux is the length of the features extracted by the initial network feature extractor A aux , and N aux is the total number of categories to be learned for the incremental task;
[0092] In this example, FC aux is used to output the recognition result, expressed as:
[0093] O aux = FC aux (A aux (X)),
[0094] where A auxis an auxiliary network feature extractor, is a SAR image with width and height being W and H respectively.
[0095] 6.3) Combine A aux and FC aux in cascade to form an auxiliary network.
[0096] Step Seven: Train the auxiliary network.
[0097] 7.1) Randomly sample a group of SAR images from the auxiliary training set and input them into the auxiliary network to calculate its loss. Based on this loss, update the auxiliary network parameters through the stochastic gradient descent algorithm:
[0098] 7.1.1) Calculate the loss of the auxiliary network
[0099]
[0100] where, is a SAR image with width and height being W and H respectively, and y X is the label corresponding to the SAR image X;
[0101] 7.1.2) Solve the loss of the auxiliary network in 8.1.1) The gradient of the auxiliary network parameter θ aux is expressed as
[0102] 7.1.3) Update the auxiliary network parameters according to the gradient solved in 8.1.2) Update the auxiliary network parameters:
[0103]
[0104] where θ a ′ ux is the updated network parameter, θ aux is the network parameter before update; lr is the learning rate, which is set according to the batch size of the input images. In the example, let lr = 0.005;
[0105] 7.2) Repeat Step 8.1) until the network converges to obtain the trained auxiliary network.
[0106] Step Eight: Construct the backbone network.
[0107] Refer to Figure 4 and the implementation of this step is as follows:
[0108] 8.1) Connect the auxiliary network feature extractor A aux in parallel with the initial backbone network feature extractor A init to obtain the feature extractor A newand its output multi-scale feature map Y new , are respectively represented as follows:
[0109] A new ={A init ,A aux},
[0110]
[0111] where and are respectively the i-th convolutional modules of A init and A aux , is a SAR image with width and height being W and H respectively, and concat means concatenating the vectors in the list in dimension 2 in sequence.
[0112] 8.2) Construct the fully-connected layer FC of the backbone network new :
[0113] 8.2.1) According to the length H init of the features extracted by the initial backbone network feature extractor A init , the total number of categories N init of the initial task, the length H aux of the features extracted by the auxiliary network feature extractor A aux and the total number of categories N aux of the incremental task, establish the fully-connected layer FC init of the backbone network with parameters (H aux +H init )×(N aux +N new ), and its output is O new which is represented as:
[0114] O new =FC new (A new (X))
[0115] where A new is the backbone network feature extractor, is a SAR image with width and height being W and H respectively;
[0116] 8.2.2) Initialize the first H init rows and N new columns of the parameters of the fully-connected layer FC init of the backbone network with all the parameters of the fully-connected layer FC init of the initial network, and initialize the last H aux rows of the parameters of the fully-connected layer FC new of the backbone network with all the parameters of the fully-connected layer FCaux Parameter of row N aux and column;
[0117] 8.3) Combine the backbone network feature extractor A new and the fully connected layer FC of the backbone network new in cascade to form the backbone network.
[0118] Step Nine: Train the backbone network.
[0119] 9.1) Construct an incremental training set and an incremental test set;
[0120] Obtain multiple SAR images of multiple classes of targets from the incremental task data and the example set data, and randomly divide their images and labels according to the ratio of 7:3 to obtain an incremental training set and an incremental test set;
[0121] 9.2) Randomly sample a group of SAR images from the incremental training set and input them into the backbone network to calculate their losses. Based on these losses, update the backbone network parameters through the stochastic gradient descent algorithm:
[0122] 9.2.1) Calculate the classification loss of the backbone network
[0123]
[0124] where, is the cross-entropy loss of the backbone network represents a SAR image with width and height of W and H respectively.
[0125] is the knowledge distillation loss between the old and new models of the backbone
[0126] network X ∈ Task init means the image comes from the initial task data, X ∈ Task increment means the image comes from the incremental task data, T is the temperature coefficient, and its range is selected between 1 and 5;
[0127] is the knowledge distillation loss of the feature space of the backbone network Avgpool1×1 represents 1×1 adaptive average pooling, Y i new is the output feature map of the backbone network feature extractor A new and X ∈ Task increment means the image comes from the incremental task data;
[0128] 9.2.2) Solve the backbone network loss in 9.2.1) For the backbone network parameter θnew The gradient, expressed as:
[0129]
[0130] 9.2.3) According to the gradient solved in 9.2.2) Update the backbone network parameters, expressed as:
[0131]
[0132] where θ n ′ ew is the currently updated network parameter, θ new is the network parameter before update; lr is the learning rate, which is set according to the batch size of the input images. In this example, let lr = 0.005;
[0133] 9.3) Repeat step 9.2) until the network converges to obtain the trained backbone network.
[0134] Step ten, input the SAR images in the incremental test set into the trained backbone network to obtain the recognition results of the SAR images.
[0135] The effects of the present invention can be further illustrated by the following simulation experiments:
[0136] I. Simulation experiment conditions:
[0137] The software platform for the simulation experiment of the present invention is: Windows 11 operating system and Pytorch 1.11.0, and the hardware configuration is: Core i7-11800H CPU and NVIDIA GeForce RTX 3080 Laptop GPU.
[0138] The SAR images used in the simulation experiment of the present invention are respectively from the measured SAR ground stationary target data MSTAR dataset announced by the MSTAR program supported by the Defense Advanced Research Projects Agency (DARPA) of the United States, the high-resolution large aircraft dataset obtained by the spaceborne radar on the Gaofen-3 satellite, and the OpenSARShip dataset released by Shanghai Jiao Tong University. It includes a total of ten types of ground vehicle targets, five types of aircraft targets, and three types of ship targets. The total number of SAR images is 9280, the number of training set images is 4977, and the number of test set images is 4303.
[0139] II. Simulation content and result analysis:
[0140] Simulation 1. Under the above simulation conditions, an incremental recognition experiment of 18 types of targets was carried out using the present invention and the existing "SAR vehicle target recognition method based on improved convolutional neural network". For the initial task, 2 types of targets were learned, and for each subsequent incremental task, 2 types of targets were learned until all 18 types of targets were learned. The experimental results are as Figure 5 shown.
[0141] Simulation 2. Under the above simulation conditions, an incremental recognition experiment of 18 types of targets was carried out using the present invention and the existing "SAR vehicle target recognition method based on improved convolutional neural network". For the initial task, 8 types of targets were learned, and for each subsequent incremental task, 2 types of targets were learned until all 18 types of targets were learned. The experimental results are as Figure 6 shown.
[0142] From Figure 5 and Figure 6 it can be seen that there is a serious catastrophic forgetting phenomenon in the existing technology when performing incremental recognition of SAR targets. After learning the incremental task, the recognition rate of the model drops significantly, while the present invention can ensure that the model still maintains a high recognition rate after learning the incremental task.
[0143] According to the experimental results of Simulation 1 and Simulation 2, the incremental recognition indexes of SAR target recognition were statistically analyzed, and the average accuracy and the final accuracy of all tasks were calculated. The results are shown in Table 1:
[0144] Table 1 Comparison of recognition accuracies of the present invention and the existing method
[0145]
[0146] It can be seen from Table 1 that both the average accuracy and the final accuracy of the present invention are significantly higher than those of the existing technology, indicating that the incremental recognition performance of the present invention is significantly better than that of the existing technology.
Claims
1. An incremental learning method for SAR target recognition based on dynamic structure and multi-level distillation, characterized in that Including the following steps: (1) Obtain multiple SAR images of multiple types of targets from the initial task data, and randomly partition the images and labels thereof to obtain an initial task training set and an initial task test set; (2) Construct an initial backbone network: (2a) Establish the feature extractor A of the initial trunk network composed of cascaded multiple residual blocks init , for outputting multi-scale feature maps; (2b) Establish a fully connected layer FC that can match the feature dimensions and output categories init ; (2c) Cascade a feature extractor and a fully connected layer to form an initial backbone network, and select cross-entropy loss as its classification loss; (3) Randomly sample a group of SAR images from the training set and input them into the initial backbone network to calculate the loss, and update the network parameters through the stochastic gradient descent algorithm until the network converges to obtain a trained initial backbone network; (4) Input the SAR images in the initial task test set into the trained initial backbone network to obtain recognition results; (5) Set an empty example set, and select less than 10% of the samples from the initial task data according to the herding algorithm and put them into this example set; (6) Obtain multiple SAR images of multiple types of targets from the incremental task data, and randomly partition the images and labels thereof to obtain an auxiliary training set and an auxiliary test set; (7) Construct an auxiliary network: (7a) Establish the feature extractor A of the auxiliary network composed of cascaded multiple residual blocks aux , for outputting multi-scale feature maps; (7b) Establish a fully connected layer FC that can match the new task feature dimensions and output categories aux ; (7c) Cascade a feature extractor and a fully connected layer to form an auxiliary network, and select cross-entropy loss as its classification loss; (8) Randomly sample a group of SAR images from the auxiliary training set and input them into the auxiliary network to calculate the loss, and update the network parameters based on this loss through the stochastic gradient descent algorithm until the network converges to obtain a trained auxiliary network; (9) Use the trained auxiliary network in (8) and the initial backbone network in step (3) to construct a backbone network: (9a) The feature extractor A in the auxiliary network aux is connected in parallel with the feature extractor A in the initial backbone network init to obtain the backbone network feature extractor A new and its output multi-scale feature map Y new ; (9b) Construct a new fully connected layer FC based on the output dimension and the number of classifications of the feature extractor after parallel connection. new ; (9c) Initialize some of the parameters of the fully connected layer in (9b) with the parameters of the fully connected layer of the initial backbone network; (9d) Initialize the remaining parameters of the fully connected layer in (9b) with the parameters of the fully connected layer of the auxiliary network; (9e) The feature extractor after parallel connection and the new fully connected layer are cascaded to form a new backbone network, and the cross-entropy loss is used The knowledge distillation loss between the new and old models And the knowledge distillation loss in the feature space These three losses are used as its classification loss (10) Obtain multiple SAR images of multiple types of targets from the incremental task data and the example set data, and randomly partition the images and labels thereof to obtain an incremental training set and an incremental test set; (11) Randomly sample a group of SAR images from the incremental training set and input them into the backbone network to calculate its loss, and iteratively update the network parameters based on this loss through the stochastic gradient descent algorithm until the network converges to obtain a trained backbone network; (12) Input the SAR images in the incremental test set into the trained backbone network to obtain recognition results.
2. The method according to claim 1, characterized in that Step (2a) constitutes the initial drying network feature extractor A init The structural parameters of each convolutional module in The feature extractor A init comprises 5 cascaded convolutional modules The first convolutional module described above is composed of a cascaded structure including a 3×3 standard convolutional layer, a batch normalization layer, a ReLU activation layer, and a 3×3 max-pooling downsampling layer; The second convolutional module described is formed by cascading 2 residual blocks with an input dimension of 64 and an output dimension of 64; The third convolutional module described is formed by cascading 2 residual blocks with an input dimension of 128 and an output dimension of 128; The fourth convolutional module described above is formed by cascading two residual blocks with an input dimension of 256 and an output dimension of 256; The fifth convolutional module described above is formed by cascading two residual blocks with an input dimension of 512 and an output dimension of 512 and a 1×1 adaptive average pooling layer.
3. The method according to claim 1, wherein The initial drying network feature extractor A in step (2a) init The output multi-scale feature maps are represented as follows: where Y i init is the output feature map of the i-th convolutional module , is an SAR image with width and height being W and H respectively is the i-th convolutional module of the feature extractor A init .
4. The method according to claim 1, wherein The parameters and outputs of the initial drying network fully connected layer FC in step (2b) are as follows: init The size parameter of the initial drying network's fully connected layer FC init is: H init ×N init , where H init is the length of the feature extracted by the feature extractor, and N init is the total number of categories to be learned for the initial task; The output of the initial drying network's fully connected layer FC init is: O init = FC init (A init (X)), where A init is the initial drying network feature extractor, and X is a SAR image with width and height W and H respectively.
5. The method according to claim 1, wherein The auxiliary network feature extractor A formed in step (7a) aux Each convolutional module includes five cascaded convolutional modules The first convolutional module described above is composed of a cascaded structure including a 3×3 standard convolutional layer, a batch normalization layer, a ReLU activation layer, and a 3×3 max-pooling downsampling layer; The second convolutional module described is formed by cascading 2 residual blocks with an input dimension of 64 and an output dimension of 64; The third convolutional module described above is formed by cascading two residual blocks with an input dimension of 128 and an output dimension of 128; The fourth convolutional module described above is formed by cascading two residual blocks with an input dimension of 256 and an output dimension of 256; The fifth convolutional module described above is composed of two residual blocks with an input dimension of 512 and an output dimension of 512, and a 1×1 adaptive average pooling layer cascaded together.
6. The method according to claim 1, wherein The auxiliary network feature extractor A in step (7a) aux The output multi-scale feature maps are represented as follows: where Y i aux is the output feature map of the i-th convolutional module , is a SAR image with width and height being W and H respectively, is the i-th convolutional module of the feature extractor A aux .
7. The method according to claim 1, wherein The parameters and output of the auxiliary network fully connected layer FC in step (7b) aux are as follows: The size parameter of the fully connected layer is: H aux × N aux , where H aux is the length of the feature extracted by the feature extractor, and N aux is the total number of categories to be learned for the incremental task; The output of the fully connected layer is: O aux = FC aux (A aux (X)), where A aux is the auxiliary network feature extractor, and X is the SAR image with width and height being W and H respectively.
8. The method according to claim 1, wherein The new backbone network feature extractor A obtained in step (9a) new and the multi-scale feature map Y output thereby new are respectively represented as follows: A new = {A init , A aux}, Among them and are the i-th convolutional modules of A init and A aux respectively. is a SAR image with width and height being W and H respectively. Concat means concatenating the vectors in the list in sequence along dimension 2.
9. The method according to claim 1, wherein The fully connected layer FC constructed in step (9b) new has the following parameters and output: Fully connected layer FC new The size is: (H init + H aux ) × (N init + N aux ), where H init is the length of the features extracted by the feature extractor A init The total number of categories to be learned for the initial task is N init is the length of the features extracted by the feature extractor A aux is the length of the features extracted by the feature extractor A aux The total number of categories to be learned for the incremental task is N aux ; Fully connected layer FC new The output is: new =FC new (A new (X)), where A new is the backbone network feature extractor, is a SAR image with width W and height H respectively.
10. The method according to claim 1, characterized in that, The classification loss in step (9e) is expressed as follows: Among them, wherein, represents a SAR image with a width of W and a height of H, and X ∈ Task init indicates that the image is from the initial task data, and X ∈ Task increment indicates that the image is from the incremental task data; T is the temperature coefficient, which is selected within the range of 1 to 5; Avgpool1×1 represents 1×1 adaptive average pooling.
Citation Information
Patent Citations
SAR vehicle target recognition method based on improved convolutional neural network
CN108280460B
SAR (Synthetic Aperture Radar) target detection and identification method based on multistage enhancement network
CN115909086A
SAR target class increment identification method based on knowledge robust-heavy balance network
CN116129219A