Task-related incremental learning method based on multi-feature fusion

Through task-related incremental learning methods based on multi-feature fusion, catastrophic forgetting and inter-task confusion problems in SAR automatic target recognition are solved, and the effect of maintaining old task knowledge and improving the accuracy of new task recognition in the incremental learning process is achieved.

CN120297364APending Publication Date: 2025-07-11XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510394868.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art has problems of catastrophic forgetting and inter-task confusion in SAR automatic target recognition, especially in the absence of sufficient annotation samples in real-time in military environments, making it difficult to effectively deal with the degradation of model performance.

Method used

Using a task-related incremental learning method based on multi-feature fusion, feature extraction, multi-scale feature fusion and knowledge distillation are performed through the TCIL-FA network, the feature extraction module is dynamically expanded to overcome catastrophic forgetting and inter-task confusion, parameters are passed using the multi-level knowledge distillation method, and classification results are optimized through classifier recalibration and auxiliary classifiers.

Benefits of technology

It effectively overcomes catastrophic forgetting and confusion between tasks, maintains knowledge of old tasks, improves the recognition accuracy and robustness of the model under new tasks, and especially maintains the recognition ability of old categories during incremental learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297364A_ABST
    Figure CN120297364A_ABST
Patent Text Reader

Abstract

The invention discloses a task-related incremental learning method based on multi-feature fusion, and solves the problems of disastrous forgetting and task confusion in the prior art. The method comprises the following steps: acquiring a to-be-classified SAR image; inputting the SAR image to be classified into a pre-trained TCIL-FA network to obtain a classification result; according to the TCIL-FA network, feature extraction is carried out on SAR images to be classified according to an extensible feature extraction module to obtain feature vectors corresponding to feature extractors, and meanwhile, the feature vectors corresponding to the feature extractors are spliced to obtain spliced feature vectors; splicing the feature vectors according to a multi-scale feature fusion module, and carrying out multi-scale feature fusion to obtain fine feature vectors; and finally, performing feature classification on the fine feature vector according to a classification module based on multistage knowledge distillation to obtain a classification result. According to the method, parameters are transmitted through a knowledge distillation method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of synthetic aperture radar, and in particular to a task-related incremental learning method based on multi-feature fusion. Background Art

[0002] SAR Automatic Target Recognition (ATR) is a technology for automatically identifying and classifying targets through Synthetic Aperture Radar (SAR) images. In recent years, with the development of machine learning theory and the improvement of computing power, the effect of automatic target recognition has been significantly improved in many application scenarios. Machine learning promotes the rapid development of target recognition technology by automatically learning features and extracting effective information from a large amount of data.

[0003] However, in the actual combat application environment, due to the dynamic and open characteristics of the military environment, the acquisition of training data is a gradually incremental process. This requires the recognition algorithm to not only be able to train a recognition model when the number of samples and categories is sufficient, but also be able to learn new training samples and new target types during the data increment process to complete more complex and diverse recognition tasks.

[0004] Large deep learning models have achieved remarkable results in many fields such as computer vision, natural language processing, and speech recognition. These models rely on a large amount of data for training and can extract useful features from the data through complex non-linear mappings. However, when faced with the situation of unable to obtain data in real time or the change of data distribution, the performance of the model is often significantly affected. Especially in the military background of SAR automatic target recognition, it may not always be possible to ensure sufficient labeled samples. Therefore, it is necessary to study the application of class incremental learning in SAR automatic target recognition. Class incremental learning is a deep learning method aimed at solving the problem of the model learning new tasks under the condition of a small number of labeled samples while retaining the knowledge of old tasks. Many studies have been proposed to combat catastrophic forgetting. Dynamic Expansion Architecture (DEA), as a promising paradigm, has been gradually widely applied. As the number of tasks increases, it dynamically expands the network, where each task is associated with a dedicated sub-network and its weights are frozen when learning new tasks. It has the advantage of well maintaining the knowledge of old tasks. However, due to the scalable characteristics of the dynamic expansion architecture itself, it is difficult to effectively handle the problems of catastrophic forgetting and task intermixing. Summary of the Invention

[0005] The present invention provides a task-related incremental learning method based on multi-feature fusion, which solves the problems of catastrophic forgetting and task confusion easily occurring in the prior art, and realizes parameter transfer through the method of knowledge distillation.

[0006] The present invention provides a task-related incremental learning method based on multi-feature fusion, and the method includes:

[0007] Input the SAR image to be classified into a pre-trained TCIL-FA network to obtain a classification result; wherein, for the TCIL-FA network, first, the SAR image to be classified is subjected to feature extraction according to an extensible feature extraction module to obtain feature vectors corresponding to each feature extractor, and at the same time, the feature vectors corresponding to each feature extractor are spliced to obtain a spliced feature vector; then, multi-scale feature fusion is performed on the spliced feature vector according to a multi-scale feature fusion module to obtain a refined feature vector; finally, feature classification is performed on the refined feature vector according to a classification module based on multi-level knowledge distillation to obtain a classification result.

[0008] One or more technical solutions provided in the present invention have at least the following technical effects or advantages:

[0009] The present invention performs feature extraction on the SAR image to be classified through multiple feature extractors.

[0010] The number of multiple feature extractors in the extensible feature extraction module is added sequentially with the new tasks, and the model parameters increase with the increase of the incremental steps, which can overcome the problems of catastrophic forgetting and task confusion as much as possible; the extensible feature extraction module can not only extract local features, but also the multi-scale feature fusion after splicing can capture local details and global context information; when performing feature fusion, the correlation between all positions in the feature map is calculated globally, so as to aggregate similar information at a long distance, breaking through the limitation that the traditional convolutional neural network can only process local neighborhoods; the multi-scale feature fusion calculates the correlation between all positions in the feature map globally, so as to aggregate similar information at a long distance, breaking through the limitation that the traditional convolutional neural network can only process local neighborhoods. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a flowchart of the steps of the task-related incremental learning method based on multi-feature fusion provided by an embodiment of the present invention;

[0012] Figure 2 It is a schematic diagram of the TCIL-FA network structure provided by an embodiment of the present invention;

[0013] Figure 3 It is a schematic diagram when the multi-scale feature fuser performs feature fusion provided by an embodiment of the present invention;

[0014] Figure 4 Schematic diagram of the Euclidean distance calculation result provided by the embodiment of the present invention;

[0015] Figure 5 Specific vehicle target optical image and SAR image provided by the embodiment of the present invention;

[0016] Figure 6 Schematic diagram of the recognition accuracy of each method on the MSTAR dataset provided by the embodiment of the present invention;

[0017] Figure 7 Schematic diagram of the network structure with an auxiliary classifier added provided by the embodiment of the present invention. Detailed implementation manners

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] A task-related incremental learning method based on multi-feature fusion, see Figure 1 , and this method includes the following steps S101 to S102.

[0020] S101, obtain the SAR image to be classified.

[0021] S102, input the SAR image to be classified into the pre-trained TCIL-FA network to obtain the classification result.

[0022] Exemplarily, the present invention is an incremental learning method. During the incremental learning process, the training data is gradually obtained as the task progresses. Assume there are T batches of data {D1,…,D T}, where represents the training data at the t-th step (i.e., task t), where is the i-th input image, is the label in the label set C t , and n t represents the number of samples in the dataset D t . In the t-th incremental step, the t-th batch of training data D t will be added to the training set. The label space of the model is the union of all seen classes and the model needs to be able to accurately predict all classes in .

[0023] For the TCIL-FA network, first, the SAR image to be classified is subjected to feature extraction according to the extensible feature extraction module to obtain the feature vectors corresponding to each feature extractor. At the same time, the feature vectors corresponding to each feature extractor are concatenated to obtain a concatenated feature vector. Then, multi-scale feature fusion is performed on the concatenated feature vector according to the multi-scale feature fusion module to obtain a refined feature vector. Finally, feature classification is performed on the refined feature vector according to the classification module based on multi-level knowledge distillation to obtain a classification result.

[0024] The training process of the TCIL-FA network includes:

[0025] (1) Obtain the training set data and initialize the TCIL-FA network structure. Among them, the training set data includes data in T batches {D1, …, D T}; The TCIL-FA network structure includes: T parallel feature extraction and classification units. Each feature extraction and classification unit includes: a feature extractor, a multi-scale feature fuser, and a classifier connected in sequence;

[0026] (2) Determine the number of t feature extraction and classification units according to the data batch;

[0027] (3) Use the multi-level knowledge distillation method to update the parameters of the feature extractor and the classifier in the t-th feature extraction and classification unit according to the 1st to (t - 1)-th feature extraction and classification units.

[0028] Here, using the multi-level knowledge distillation method to update the parameters of the feature extractor and the classifier in the t-th feature extraction and classification unit according to the 1st to (t - 1)-th feature extraction and classification units includes:

[0029] (3.1) Adopt the feature-level knowledge distillation method to calculate the feature-level knowledge distillation loss of the feature extractor in the t-th feature extraction and classification unit according to the 1st to (t - 1)-th feature extraction and classification units, and then update the parameters of the feature extractor according to the feature-level knowledge distillation loss;

[0030] (3.2) Obtain any image in the training set as the input image, calculate the first Logits value output by the fully connected layer in the t-th feature extraction and classification unit corresponding to the input image, and at the same time calculate the second Logits value output by the fully connected layer in the (t - 1)-th feature extraction and classification unit for the input image;

[0031] (3.3) Update the parameters of the feature extractor and the classifier in the t feature extraction and classification units according to the weighted sum of the first Logits value and the second Logits value.

[0032] (4) Until the training of data for T batches is completed, a pre-trained TCIL-FA network is obtained according to T parallel feature extraction and classification units; wherein, referring to Figure 2 , the extensible feature extraction module includes: a plurality of parallel feature extractors; the multi-scale feature fusion module includes the t-th multi-scale feature fusion unit; the classification module includes the t-th classifier.

[0033] Exemplarily, in the training process of the present invention, the feature extractor F1 and the classifier G1 in the first step (i.e., t = 1) are directly constructed using ResNet18. The feature extractor F1 is composed of the first 17 layers of ResNet-18, including a 7×7 convolutional kernel, a 3×3 max pooling layer, and four groups of feature processing layers containing basic residual blocks.

[0034] The classifier G1 is composed of a global average pooling layer and a fully connected layer.

[0035] Then, in each step t ∈ {2, …, T}, a new feature extractor F t is added, while keeping the previous feature extractors F1, …, F t-1 and the previous classifier G t-1 frozen.

[0036] Meanwhile, the parameters of the classifier G t are initialized to the parameters of the classifier G t-1 .

[0037] Taking any training data in the training set as the input image x, the feature vector extracted by the feature extractor F t is F t (x). All the extracted feature vectors are concatenated to obtain the concatenated feature vector u t , and the concatenated feature vector u t is expressed as u t = [F1(x), …, F t (x)].

[0038] In the t-th feature extraction and classification unit, the feature extractor F t learns through the knowledge distillation method from the t-th data batch D t and the feature vector u t-1 obtained from the feature extractor in the (t - 1)-th feature extraction and classification unit.

[0039] The feature-level knowledge distillation method is used to help the feature extractor F t in the t-th feature extraction and classification unit to learn.

[0040] For an input image x ∈ D, using the i-th feature extractor F iAs a teacher to guide the training, the feature-level knowledge distillation loss can be expressed as:

[0041] L f (x) = ‖‖F t (x) - F i (x)‖‖2;

[0042] Where, F t (x) represents the feature vector obtained by the feature extractor of the input image x passing through the feature extraction and classification unit at the t-th layer; F i (x) represents the feature vector obtained by the feature extractor F of the input image x passing through the feature extraction and classification unit corresponding to its data batch i ; L f (x) represents the feature-level knowledge distillation loss; x represents the input image; ‖·‖2 represents the second norm.

[0043] Then, a fine-grained feature vector f is obtained through an attention-based multi-scale feature fusion network t , where the fine-grained feature vector f t is expressed as:

[0044] f t = A t (u t );

[0045] Where, A t (·) represents the feature fusion operation.

[0046] During the incremental learning process, the set of old feature extractors {F1(x), …, F t-1 (x)} are trained in different training steps and cannot effectively extract the features of the data in the new steps {t, …, T}, resulting in a decline in feature quality. Therefore, the present invention uses a multi-scale feature fusion network to optimize the features and combine the transformed features for classification.

[0047] See Figure 3 , the schematic diagram of the multi-scale feature fusion network during feature fusion.

[0048] Channel attention focuses on which feature channels in the concatenated feature vector u t are more important. Average pooling is used to capture the global information of the concatenated feature vector u t , and max pooling is used to capture the salient features, generating a channel attention map.

[0049] The average pooling result and the max pooling result are processed and added through a shared MLP, and after Sigmoid activation, the channel weights are obtained.

[0050] Spatial attention focuses on the concatenated feature vector u tWhich spatial regions are more important. By averaging and max-pooling the concatenated feature vector u along the channel axis t perform average pooling and max pooling on the concatenated feature vector u t and then use a 7×7 convolution to generate a spatial attention map, followed by Sigmoid activation to obtain the spatial weight.

[0051] Input the concatenated feature vector u t u t ∈R C×H×W into the input feature fusion module, and apply the following formula to obtain a 1D channel attention map A c ∈R C×1×1 and a 2D spatial attention map A s ∈R 1×H×W .

[0052] A c =σ(MLP(AvgPool(u t )) + MLP(MaxPool(u t )));

[0053] A s =σ(f 7×7 ([AvgPool(u t ) ; MaxPool(u t )]));

[0054] where σ represents the Sigmoid function; f 7×7 represents the convolution operation with a convolution kernel size of 7×7; AvgPool(·) represents average pooling, MaxPool(·) represents max pooling, and MLP(·) represents the multi-layer perceptron mechanism.

[0055] Fuse the channel attention map A c and the spatial attention map A s to obtain the refined feature vector f t . The feature fusion process is as follows:

[0056]

[0057] where A c and A s represent channel attention and spatial attention respectively, represents element-wise multiplication of vectors.

[0058] Next, input the refined feature vector f t into the classifier G t in the t-th feature extraction and classification unit, and obtain the output of the last fully connected layer of the classifier G t i.e., the output logits value o of the logitst (x). During the training of the TCIL-FA network, the logits-level knowledge distillation method is applied to guide the classifier G t Retain the old knowledge, specifically using the formula:

[0059] o t (x) = G t (f t );

[0060] Logits is the raw output of the fully connected layer of the classifier, without being processed by the activation function, representing the "raw scores" of each category. The output of Softmax is the probability of the classification task, and its input is the logits value. The logits value is usually a value in (-∞, +∞), and the Softmax layer converts it into a value between 0 and 1.

[0061] For each input image x, we can obtain o t (x) and o t-1 (x), which are the outputs of the classifier G t and the classifier G t-1 respectively. Then, the KL divergence is used to calculate the distance between them, and the logits-level knowledge distillation loss is obtained:

[0062]

[0063] where, represents the probability that the prediction result of the input image x passing through the classifier in the (t - 1)-th feature extraction and classification unit is the c-th category; q c (x) represents the probability that the prediction result of the input image x passing through the classifier in the t-th feature extraction and classification unit is the c-th category; represents all categories of the labels in the training data set from the 1st batch to the t-th batch; L l (x) represents the logits-level knowledge distillation loss; x represents the input image.

[0064] Specifically, is an element in o t-1 (x), representing the logits value of the classifier G t-1 , and T is the temperature hyperparameter, taking the constant 1;

[0065] o c (x) is an element in o t (x), representing the logits value of the classifier G t .

[0066] Then, the weighted sum of the loss of the logits-level knowledge distillation method and the loss of the feature-level knowledge distillation method is used as the total knowledge distillation loss, and the total knowledge distillation loss is used to guide the parameter update of the feature extractor F t and the classifier G t The parameter update of is as follows. The total knowledge distillation loss can be written as:

[0067]

[0068] where λ and μ are hyperparameters, and the empirical values are set to 0.5. If λ = 0, it represents a non-replay setting, that is, all losses come from L l . If λ is not 0, it represents a replay setting, and in the replay setting, R is the set of samples of the retained old tasks.

[0069] In a specific embodiment provided by the present invention, as the incremental process progresses and the training data set is continuously expanded, the phenomenon of confusion between new and old categories will occur because the classifier G t will be more biased towards new categories, resulting in a poorer discrimination ability for old categories. In a classification task, each category corresponds to a weight vector, and the norm of the weight vector reflects the influence of the category in the model. If the weight norm of the new category is large, the classifier will be more biased towards predicting the new category, resulting in a decrease in the ability to predict the old category. To solve this problem, the present invention proposes a classifier recalibration method

[0070] Specifically, after using the multi-level knowledge distillation method to update the parameters of the feature extractor and the classifier in the t-th feature extraction and classification unit according to the 1st to the t-1st feature extraction and classification units, it further includes:

[0071] (1) Calculate the first weight norm vector of all categories of the classifier in the t-1st feature extraction and classification unit, and calculate the second weight norm vector of all categories of the classifier in the t-th feature extraction and classification unit;

[0072] (2) Calculate the coefficient for classifier recalibration based on the first weight norm vector and the second weight norm vector;

[0073] (3) Correct the parameters of the classifier in the t-th feature extraction and classification unit according to the coefficient for classifier recalibration to obtain the corrected parameters of the classifier in the t-th feature extraction and classification unit.

[0074] Exemplarily, first, at the end of each batch of data training, calculate the weight norm vectors of the old and new categories in the last fully connected layer of the classifier G t :

[0075]

[0076] where nold Denote the first weight norm vector of all categories of the classifier in the (t - 1)-th feature extraction and classification unit; n new Denote the second weight norm vector of all categories of the classifier in the t-th feature extraction and classification unit; Denote the class index in the t - 1-th batch of data D t-1 ; Denote the class index in the t-th batch of data D t ; w i is the weight vector corresponding to the i-th class in the last layer of the classifier; ||w i || is the L2 norm of w i , that is, the square root of the sum of the squares of the elements of the vector.

[0077] For example: w i =(w i1 , w i2 , …, w in ) has the L2 norm of:

[0078]

[0079] Based on the above norm vectors, calculate the recalibration coefficient γ of the classifier:

[0080] γ = Mean(n old ) / Mean(n new );

[0081] where Mean(·) represents calculating the mean of the elements in the vector. If the average norm of all categories of the classifier in the (t - 1)-th feature extraction and classification unit is greater than the average norm of all categories of the classifier in the t-th feature extraction and classification unit, then γ < 1, and subsequently γ will be used to shrink the new class logits; otherwise, it will be amplified.

[0082] Use the recalibration coefficient γ to correct the logits value o t output by the classifier G rt (x), where

[0083] o rt (x) = W t (o(x)) = (o old (x), γ·o new (x));

[0084] In the formula, o rt (x) is the calibrated logits value, o(x) is the logits value before calibration, W t (·) represents the calibration operation, o old(x) represents the set of logits for all classes of the classifier in the (t-1)-th feature extraction and classification unit, maintaining the original values; o new (x) represents the set of logits for all classes of the classifier in the t-th feature extraction and classification unit, scaled by γ.

[0085] The logits output of the classifier, the logits value o rt (x) passes through the Softmax layer to obtain the class probability distribution Calculating the class label corresponding to the maximum probability is the classification result The specific process is expressed as:

[0086]

[0087] Through classifier recalibration, the mean of the norms of the old and new class logits (all classes of the classifier in the (t-1)-th feature extraction and classification unit and all classes of the classifier in the t-th feature extraction and classification unit) is forced to remain relatively stable, avoiding the model from being biased towards new classes.

[0088] Here, each feature extraction and classification unit further includes: an auxiliary classifier;

[0089] The auxiliary classifier is used to reconstruct the class label space corresponding to the training data of the t-th batch; among them, the parameters of the auxiliary classifier are randomly initialized parameters;

[0090] The auxiliary classifier is used to classify the feature vectors output by the feature extractor in the t-th feature extraction and classification unit to obtain an auxiliary classification result.

[0091] Exemplarily, refer to Figure 7 , in order to force the network to learn the diversity and discriminative features of new classes, the present invention designs an auxiliary classifier A new auxiliary classifier is created in each incremental step Its parameters are randomly initialized and updated through the currently input training data. The existence of the auxiliary classifier is temporary and dynamic, only serving the new feature learning of the current step and not being reused across steps. The auxiliary classifier and the classifier G t The difference is that Before classification, it is necessary to reconstruct the class label space: the old data as a whole is regarded as one class, and the new classes added in the current task remain independent. If the number of new classes added in the current task is Y t , the label space of the auxiliary classifier is expanded to Y t +1, including the new class set Y tand the old category aggregation classes. For example: If there are old categories 1 to 100, and new categories 101 to 110 are added, the reconstructed label space is [old category: 1, new category: 2 to 11]. Send the output of the feature extractor F t into to obtain the probability distribution Calculate the cross-entropy loss using this probability distribution to obtain the auxiliary loss L div :

[0092]

[0093] where D t represents the training data of the t-th batch; represents the first input data x i The second predicted classification result obtained by passing the first input data x through the auxiliary classifier in the t-th feature extraction and classification unit is the probability of the first input data x i true category y i ; y represents the second predicted classification result obtained by the auxiliary classifier; y i represents the true category of the first input data x i ; x i represents the first input data; L div represents the auxiliary loss.

[0094] The method provided by the present invention further includes: performing network pruning on the pre-trained TCIL-FA network, specifically including:

[0095] (1) Determine all convolutional layers in the pre-trained TCIL-FA network, and respectively determine the number of convolutional kernels corresponding to each convolutional layer;

[0096] (2) Determine the median of the convolutional kernels of all convolutional layers, and calculate the Euclidean distance from all convolutional layers to the median of the convolutional kernels;

[0097] (3) Sort all convolutional layers in ascending order according to their corresponding Euclidean distances, and prune the convolutional layers in the ascending queue according to the preset pruning rate to obtain the pruned pre-trained TCIL-FA network.

[0098] Exemplarily, the present invention applies the Filter Pruning via Geometric Median (FPGM) method to achieve parameter compression of the feature extractor. The feature extractor is a network composed of Resnet18, which contains multiple convolutional kernels. These convolutional kernels will gradually accumulate redundancy during the dynamic expansion of incremental learning - as the number of tasks increases, the feature extractors between different tasks tend to learn similar low-level patterns, resulting in parameter inflation and a decline in computational efficiency. The proposed TCIL-Lite in the present invention first finds the geometric median of the convolutional kernels, then calculates the Euclidean distance from all convolutional kernels to the geometric median, sorts them in ascending order of distance, and directly removes the top 30% of the filters after sorting according to a preset pruning rate (such as 30%), retaining the filters with a relatively large distance. Through this strategy, the model size can be greatly reduced while having a small impact on accuracy.

[0099] The FPGM strategy understands each layer of the network as a multi-dimensional space, and each convolutional kernel as a point. In a network with a total of K convolutional layers, the j-th neural network layer has N convolutional kernels F1,..., F N , and the geometric median of the convolutional kernels is denoted as F c , then:

[0100]

[0101] where F x is the to-be-determined geometric median, i is the convolutional kernel serial number; j is the corresponding value of the convolutional layer.

[0102] See Figure 4 , for a 2×2 convolutional kernel, the Euclidean distance between convolutional kernel 1 and convolutional kernel 2 is:

[0103]

[0104] where (i,j) are the parameter coordinates of the filter.

[0105] The loss function used during the training of the TCIL-FA network provided by the present invention is expressed as:

[0106] L = L clf + αL kd + βL div ;

[0107] where L clf represents the difference loss; α is the first hyperparameter; L kd represents the knowledge distillation loss, where the knowledge distillation loss includes: feature-level knowledge distillation loss and logits-level knowledge distillation loss; β is the second hyperparameter; L div represents the auxiliary loss.

[0108] The calculation formula for the difference loss is as follows:

[0109]

[0110] Among them, D t represents the training data of the t-th batch; p Gt (y = y i |x i ) represents the probability that the first predicted classification result obtained by the classifier in the t-th feature extraction and classification unit for the first input data x i is the true class y i of the first input data x i ; y represents the first predicted classification result obtained by the classifier; y i represents the true class of the first input data x i ; x i represents the first input data; L clf represents the difference loss.

[0111] In a specific embodiment provided by the present invention, the learning process of the present invention is shown in Table 1:

[0112] Table 1 Learning process of the TCIL-FA method

[0113]

[0114]

[0115] The effects of the present invention will be further described below in combination with simulation experiments.

[0116] (1) Experimental data and parameters.

[0117] In this paper, the MSTAR dataset is used as the experimental verification dataset, ensuring the practicality and effectiveness of the research. The MSTAR dataset was organized under a project of the Defense Advanced Research Projects Agency (DARPA), including SAR images in the X-band and HH polarization, with a resolution of 0.3 meters, covering a variety of targets. The MSTAR dataset contains 10 types of targets: 2S1 (cannon), BMP2 (tank), BRDM2 (truck), BTR60 (armored vehicle), BTR70 (armored vehicle), D7 (bulldozer), T62 (tank), T72 (tank), ZIL131 (truck), and ZSU23 / 4 (cannon). These targets were captured at two different pitch angles (15° and 17°). In this paper, the images at the 17° pitch angle are used as the training set, and the images at the 15° pitch angle are used as the test set. The specific optical images and SAR images of vehicle targets are as follows Figure 5 as shown, and the dataset description is shown in Table 1 below.

[0118] Table 1 Class Distribution of MSTAR Dataset

[0119] Target type Training set (17°) Test (15°) 2S1 299 274 BMP2 233 195 BRDM2 298 274 BTR60 256 195 BTR70 233 196 D7 299 274 T62 299 273 T72 232 196 ZIL131 299 274 ZSU23 / 4 299 274

[0120] Incremental class setting: Under the fixed sample memory setting, the memory capacity size is set to 200 images, that is, 20 pictures are randomly reserved for each class. Fixed sample memory means that during the incremental learning process, when the model learns a new task, a small number of training samples of old classes are fixed and reserved. At the t-th step of training the new task, the data of the current new task is mixed with the reserved old class samples and jointly used for model training. Randomly select 4 classes as the initial classes, and the remaining 6 classes as the incremental classes. The incremental step sizes are set to 1, 2, and 3 for incremental training.

[0121] Network parameter setting: The SGD optimizer is adopted, the weight decay parameter is set to 0.0005, and the batch size is set to 128. For the setting of the learning rate, during training, the warmup preheating strategy is adopted in the first 10 epochs, and the learning rate is gradually increased to 0.01 in the initial stage to avoid unstable early training. After that, the staged decay strategy is adopted to finely adjust with a small learning rate, and it decays to 0.001 and 0.0001 at 100 and 120 epochs respectively.

[0122] All models are trained using PyTorch on a workstation equipped with an Nvidia 3090 GPU.

[0123] (2) Measurement metrics.

[0124] The recognition rate is measured by the Mean Incremental Accuracy (MIA). MIA is an evaluation metric that measures the ability of an incremental learning model to maintain and improve the recognition ability of the original knowledge (learned tasks or categories) when dealing with new category data. Specifically, it reflects the average performance of the overall recognition accuracy of the model for all learned tasks or categories (including newly added and previously existing ones) after each incremental learning with new category data. The definition of MIA is as follows:

[0125]

[0126] The FLOPs are used to evaluate the model complexity. FLOPs (Floating-Point Operations) is the core metric for measuring the computational complexity of a model, representing the number of floating-point operations required for the model to complete one forward inference. For example: for the calculation of FLOPs in a convolutional layer, if the number of input channels is C i and the number of output channels is C o , the convolutional kernel size is K×K, and the output feature map size is H×W, the computational amount is:

[0127]

[0128] In the experiments of this section, the Thop tool library is used to automatically calculate the FLOPs and the number of model parameters.

[0129] (3) Validation of the effectiveness of TCIL-FA.

[0130] The contrast analysis method is used to verify the effectiveness of TCIL-FA. The contrast methods include Finetuning, iCaRL, CCIL, TCIL, and TCIL-FA, and their average classification accuracies are shown in Table 2.

[0131] Table 1 Single-step incremental recognition accuracy of the MSTAR dataset

[0132]

[0133] From Table 2 and Figure 6 analysis, it can be seen that:

[0134] 1. Finetuning. As a baseline method, Finetuning directly fine-tunes the model parameters on the new task data. The experimental data shows that its initial accuracy is relatively high (93.06% for category 4), but as the number of categories increases, the performance drops significantly (down to 72.15% for category 10). This indicates that this method does not introduce an anti-forgetting mechanism, resulting in the serious destruction of old category knowledge during the parameter update process, verifying the negative impact of catastrophic forgetting on incremental learning.

[0135] 2. iCaRL. iCaRL retains the mean features of old-class samples through a memory bank. Its performance is close to that of CCIL (both above 90%) when the number of classes is 4 - 6, but the performance degradation accelerates after the number of classes reaches 7 (80.02% at class 10). Its limitations are as follows: (1) Storage limitation: A memory bank with a fixed capacity is difficult to fully represent high-dimensional features. As the number of classes increases, insufficient coverage of old-class samples leads to the accumulation of bias in the estimation of feature means. (2) Static classifier: A linear classifier based on nearest-mean classification has limited ability to model complex class boundaries, resulting in a decrease in classification accuracy when the task complexity increases.

[0136] 3. CCIL. CCIL alleviates forgetting through feature combination and dynamic expansion. Its performance is better than that of iCaRL when the number of classes is 4 - 8 (82.08% vs. 83.29% at class 8), but its accuracy (74.84%) is lower than that of iCaRL at a high number of classes (10 classes). The fluctuations in its performance are attributed to: (1) Combination bias: The accuracy of synthesizing new-class features depends on the completeness of old-class features. As the number of tasks increases, the modeling error of the combination relationship gradually amplifies. (2) Computational overhead: The additional parameters introduced by the dynamic expansion strategy increase the model complexity, which may lead to overfitting (a sharp drop in performance at classes 9 - 10).

[0137] 4. TCIL series methods. TCIL optimizes the classification boundary through multi-level knowledge distillation and attention mechanism. TCIL-FA further integrates global features and achieves the highest accuracy at all numbers of classes (84.72% at class 10). Its advantages are as follows: (1) Knowledge retention: Dual distillation at the feature level and logit level reduces the knowledge conflict between old and new tasks. The attention mechanism dynamically adjusts the feature weights, alleviating the information loss caused by insufficient memory bank capacity. (2) Classification robustness: The cosine similarity classifier combined with the classifier readjustment strategy effectively suppresses the distribution shift of prediction scores between old and new classes. (3) Feature enhancement: TCIL-FA improves the consistency of feature representation through non-local feature fusion and still maintains a small performance degradation when the number of classes increases (a 12.73% decrease from class 4 to 10, better than 14.45% of iCaRL and 19.63% of CCIL).

[0138] (4) Validation of the effectiveness of TCIL-Lite

[0139] Pruning parameter setting: In the network pruning step, this paper prunes all weighted layers at the same pruning rate P i simultaneously. Therefore, only one hyperparameter P i is needed to balance the operation acceleration effect and accuracy. The pruning operation is performed at the end of each training epoch. The experiments test the effectiveness of the TCIL-Lite method using pruning rates of 30% and 40% respectively. The experimental results are shown in Table 3 below:

[0140] Table 2 Pruning Effectiveness Verification

[0141] Incremental learning method Recognition rate (%) <![CDATA[Number of model parameters (×10 6 )]]> <![CDATA[FLOPs(×10 7 ) <!-- 11 -->]]> TCIL-FA 84.72 34.3 4.11 TCIL-Lite (Pruning rate 30%) 83.61 22.3 2.43 TCIL-Lite (Pruning rate 40%) 83.02 18.0 1.87

[0142] Analysis of the data in Table 3 shows that the recognition rate of TCIL-Lite is 83.61% at a pruning rate of 30%, which is only 1.11% lower than that of the original TCIL-FA (84.72%); the recognition rate at a pruning rate of 40% is 83.02%, a decrease of 1.70%. The small loss in accuracy indicates that the pruning strategy based on the geometric median effectively retains the key feature expression ability. Especially in incremental learning, where the model needs to dynamically adapt to the feature distribution of new tasks, the removal of redundant convolutional kernels does not significantly disrupt the continuity of the feature space. In terms of the model compression rate, the number of parameters of TCIL-Lite is 22.3M at a pruning rate of 30%, a 34.98% reduction compared to TCIL-FA (34.3M); the number of parameters at a pruning rate of 40% is 18.0M, a 47.52% reduction. The FLOPs at a pruning rate of 30% is 2.43G, a 40.87% reduction compared to the original model (4.11G); the FLOPs at a pruning rate of 40% is 1.87G, a 54.50% reduction. At a pruning rate of 30%, the model maintains a recognition rate of 83.61% while halving the FLOPs, proving its suitability for deployment on edge devices with limited computing resources; although the model is further compressed at a pruning rate of 40%, the accuracy drops to 83.02%, and the efficiency and accuracy need to be balanced according to the scenario requirements.

[0143] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. All or part of the present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multi-processor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.

[0144] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the present invention; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the present invention.

Claims

1. A task-related incremental learning method based on multi-feature fusion, characterized in that, Including: Obtain the SAR image to be classified; Input the SAR image to be classified into the pre-trained TCIL-FA network to obtain a classification result; wherein, for the TCIL-FA network, first, perform feature extraction on the SAR image to be classified according to the scalable feature extraction module to obtain feature vectors corresponding to each feature extractor, and at the same time, splice the feature vectors corresponding to each feature extractor to obtain a spliced feature vector; then, perform multi-scale feature fusion on the spliced feature vector according to the multi-scale feature fusion module to obtain a refined feature vector; finally, perform feature classification on the refined feature vector according to the classification module based on multi-level knowledge distillation to obtain a classification result.

2. The task-related incremental learning method based on multi-feature fusion according to claim 1, characterized in that, The training process of the TCIL-FA network includes: Obtain the training set data and initialize the TCIL-FA network structure; among them, the training set data includes t batches of training data {D1, …, D t}; the TCIL-FA network structure includes: t parallel feature extraction and classification units; each feature extraction and classification unit includes: a feature extractor, a multi-scale feature fusion device, and a classifier connected in sequence; Determine the number of t feature extraction and classification units according to the data batch; Use the multi-level knowledge distillation method to update the parameters of the feature extractor and the classifier in the t-th feature extraction and classification unit according to the 1st to the (t - 1)-th feature extraction and classification units; Until the training of t batches of data is completed, obtain the pre-trained TCIL-FA network according to the t parallel feature extraction and classification units; wherein, the scalable feature extraction module includes: multiple parallel feature extractors; the multi-scale feature fusion module includes the t-th multi-scale feature fuser; the classification module includes the t-th classifier.

3. The task-related incremental learning method based on multi-feature fusion according to claim 2, wherein Each of the feature extraction and classification units further includes: an auxiliary classifier; The auxiliary classifier is used to reconstruct the category label space corresponding to the training data of the t-th batch; wherein, the parameters of the auxiliary classifier are randomly initialized parameters; The auxiliary classifier is used to classify the feature vector output by the feature extractor in the t-th feature extraction and classification unit to obtain an auxiliary classification result.

4. The task-related incremental learning method based on multi-feature fusion according to claim 2, characterized in that The step of using the multi-level knowledge distillation method to update the parameters of the feature extractor and the classifier in the t-th feature extraction and classification unit according to the 1st to the (t - 1)-th feature extraction and classification units includes: Adopt the feature-level knowledge distillation method to calculate the feature-level knowledge distillation loss of the feature extractor in the t-th feature extraction and classification unit according to the 1st to the (t - 1)-th feature extraction and classification units, and then update the parameters of the feature extractor according to the feature-level knowledge distillation loss; Obtain any image in the training set as the input image, calculate the first Logits value output by the fully connected layer in the t-th feature extraction and classification unit corresponding to the input image, and at the same time calculate the second Logits value output by the fully connected layer in the (t - 1)-th feature extraction and classification unit for the input image; Update the parameters of the feature extractor and the classifier in the t feature extraction and classification units according to the weighted sum of the first Logits value and the second Logits value.

5. The task-related incremental learning method based on multi-feature fusion according to claim 2, characterized in that After using the multi-level knowledge distillation method to update the parameters of the feature extractor and the classifier in the t-th feature extraction and classification unit according to the 1st to the (t - 1)-th feature extraction and classification units, it further includes: Calculate the first weight norm vectors of all categories of the classifier in the (t - 1)-th feature extraction and classification unit, and calculate the second weight norm vectors of all categories of the classifier in the t-th feature extraction and classification unit; Calculate the coefficient for classifier recalibration based on the first weight norm vector and the second weight norm vector; Correct the parameters of the classifier in the t-th feature extraction and classification unit according to the coefficient for classifier recalibration to obtain the corrected parameters of the classifier in the t-th feature extraction and classification unit.

6. The task-related incremental learning method based on multi-feature fusion according to claim 1, characterized in that, It further includes: Perform network pruning on the pre-trained TCIL-FA network, specifically including: Determine all convolutional layers in the pre-trained TCIL-FA network, and respectively determine the number of convolution kernels corresponding to each convolutional layer; Determine the median of the convolution kernels of all convolutional layers, and calculate the Euclidean distance from all convolutional layers to the median of the convolution kernels; Sort all convolutional layers in ascending order according to their corresponding Euclidean distances, and perform pruning on the convolutional layers in the ascending queue according to a preset pruning rate to obtain the pruned pre-trained TCIL-FA network.

7. The task-related incremental learning method based on multi-feature fusion according to claim 3, characterized in that The loss function used during the training of the TCIL-FA network is expressed as: L = L clf + αL kd + βL div ; Among them, L clf represents the difference loss; α is the first hyperparameter; L kd represents the knowledge distillation loss, where the knowledge distillation loss includes: the feature-level knowledge distillation loss and the logits-level knowledge distillation loss; β is the second hyperparameter; L div represents the auxiliary loss.

8. The task-related incremental learning method based on multi-feature fusion according to claim 7, characterized in that The calculation formula for the logits-level knowledge distillation loss is: The calculation formula for the feature-level knowledge distillation loss is expressed as: L f (x) = |||F t (x) - F i (x)|||²; wherein, represents the probability that the prediction result of the input image x obtained by the classifier in the (t - 1)-th feature extraction and classification unit is the c-th class; q c (x) represents the probability that the prediction result of the input image x obtained by the classifier in the t-th feature extraction and classification unit is the c-th class; represents all classes of labels in the training data sets from the 1st batch to the t-th batch; L l (x) represents the logits-level knowledge distillation loss; x represents the input image; F t (x) represents the feature vector obtained by the input image x passing through the feature extractor in the t-th feature extraction and classification unit; F i (x) represents the feature vector obtained by the input image x passing through the feature extractor F in the feature extraction and classification unit corresponding to the data batch to which it belongs i ; L f (x) represents the feature-level knowledge distillation loss; ||·||2 represents the L2 norm.

9. The task-related incremental learning method based on multi-feature fusion according to claim 7, wherein The calculation formula for the difference loss is: Among them, D t represents the training data of the t-th batch; represents the first input data x i The first predicted classification result obtained by the classifier in the t-th feature extraction and classification unit for the first input data x i is the probability of the true class y i ; y represents the first predicted classification result obtained by the classifier; y i represents the true class of the first input data x i ; x i represents the first input data; L clf represents the difference loss.

10. The task-related incremental learning method based on multi-feature fusion according to claim 7, characterized in that The calculation formula for the auxiliary loss is: Among them, D t represents the training data of the t-th batch; represents the first input data x i The second predicted classification result obtained by the auxiliary classifier in the t-th feature extraction and classification unit for the first input data x i is the probability of the true class y i ; y represents the second predicted classification result obtained by the auxiliary classifier; y i represents the first input data x i true class; x i represents the first input data; L div represents the auxiliary loss.