Image classification method and system based on class incremental learning of multiple normalization and dynamic network
By introducing multiple normalized modules and dynamic network expansion and compression strategies in class incremental learning, the problems of model parameter growth and catastrophic forgetting are solved, and efficient learning of the model and task adaptability are achieved.
Patent Information
- Application Number
- CN202510212226.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-25
AI Technical Summary
When the existing class incremental learning method handles new tasks, the number of model parameters increases linearly, resulting in increased training time, computing overhead and memory consumption, and is prone to catastrophic forgetting problems, affecting the performance of old tasks.
A class incremental learning method based on multiple normalization and dynamic networks is adopted. By introducing multiple normalization modules and dynamic network expansion and compression strategies, the model structure and expansion methods are optimized, the catastrophic forgetting problem of the model is alleviated, and the model scale is controlled through model distillation technology.
It significantly improves the learning performance and task adaptability of the model, achieves a good balance between new and old tasks, reduces catastrophic forgetting, and reduces computing overhead and memory requirements.
Smart Images

Figure CN120147711A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer image classification, and particularly relates to an image classification method and system for class incremental learning based on multi-normalization and a dynamic network. Background Art
[0002] Artificial neural networks acquire knowledge through generalized learning at different training stages and have achieved remarkable success in solving specific static classification tasks. However, with the change of training tasks in different real-world scenarios, expanding the static model built in the artificial neural network will cause the problem of catastrophic forgetting, that is, after the model learns a new task, it significantly forgets the knowledge on the old task. In this environment, continual learning emerges, and its goal is to turn the mode of the static network into a network that can continuously accumulate knowledge through different tasks and does not need to start training from scratch every time. As the most typical and realistic scenario in continual learning, class incremental learning aims to train a model with a limited memory size to meet the needs of the real world.
[0003] Currently, the most commonly used methods in the field of class incremental learning are usually based on knowledge distillation. For example, in the invention named "Class Incremental Learning Method Based on Multi-Level Knowledge Distillation" with the application publication number (CN 117494790 A), although the method based on knowledge distillation can align the output of the new model with that of the old model when training the new model, so as to retain the knowledge of the old task and achieve the learning of the new task, there are still some limitations:
[0004] First, since the distillation process depends on a fixed network, the use of a single backbone network often leads to insufficient plasticity of the model and a lack of sufficient learning ability to handle new classes;
[0005] Second, due to the limited access to old data, the features of the old task may degenerate, resulting in catastrophic forgetting and affecting the performance of the old task.
[0006] In recent years, class incremental learning methods based on dynamic networks have achieved good performance. These methods freeze the old modules to retain the classification performance of the old classes and expand new modules on the basis of the original model to improve the learning ability of the new classes. However, with the continuous expansion of the new task modules, the number of parameters of the model increases linearly, resulting in an increase in training time, computational overhead, and memory consumption. If reasonable measures are not taken to optimize the retention of the old model, it may affect the performance of the new model, thereby leading to a decrease in the accuracy of the overall model. Summary of the Invention
[0007] To overcome the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to provide an image classification method and system for class-incremental learning based on multi-normalization and dynamic networks. The present invention aims at the image classification task in the field of class-incremental learning, and through optimizing the model structure and expansion method, it is used to alleviate the problem of catastrophic forgetting of the model and improve the accuracy of the downstream image classification task of the model.
[0008] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0009] An image classification method for class-incremental learning based on multi-normalization and dynamic networks, comprising the following steps;
[0010] Step 1: Introduce a multi-normalization module on the basis of the model backbone network to build a new basic network M base ;
[0011] Step 2: Based on the new basic network, perform training on the initial image classification task D 1 of class-incremental learning to obtain an initialized model, and freeze the initialized model as the old model for the next stage of image classification task D 2 of class-incremental learning;
[0012] Step 3: When dealing with a new class-incremental learning image classification task, based on the new basic network, train an additional new model for the new task, and connect the old model and the new model through a fully connected layer to build a feature enhancement model;
[0013] Step 4: Control the scale of the feature enhancement model through model distillation technology, and distill the feature enhancement model to obtain a target compressed model; the target compressed model is used as the old model when dealing with the next class-incremental task;
[0014] Step 5: On the test samples, use the target compressed model to detect the image classification accuracy.
[0015] Furthermore, in Step 1, the model backbone model adopts a Resnet32 model, and the specific steps for constructing the new multi-normalization module and building the basic network M on this basis are as follows: base Specific steps are:
[0016] Step 1.1: Set the class-incremental learning scenario to perform step-by-step learning in a data stream with new categories. Assume there is a series of K training tasks for image classification {D 1 , D 2 , …, D K}, and there are no overlapping classes between each task, where represents the t-th incremental stage, and this stage contains n tFor each sample, a sample set ε with a limited memory size is constructed to store representative samples of the learned categories. The feature map input to the multi-normalization module is denoted as where B is the mini-batch size, C is the number of channels, W is the width, and H is the height. It is split into two parts along the channel dimension to obtain the first feature map and the second feature map
[0017] Step 1.2: Apply the batch normalization (BN) method to the first feature map p 1 Calculate its mean and variance, and use the affine transformation to obtain the normalized feature map That is:
[0018]
[0019] where, represents the mean, represents the variance, and γ BN is the weighted balance factor;
[0020] Step 1.3: Apply the spatial normalization (SN) method, such as instance normalization (IN) or layer normalization (LN), to the second feature map p 2 During this process, according to the selected type of spatial normalization, calculate the corresponding mean and variance to obtain the feature map
[0021]
[0022] where SN = {IN, LN}. If SN selects IN, calculate the mean along the channel dimension and the variance If SN selects LN, calculate the mean along the layer dimension and the variance
[0023] When using multiple spatial normalization methods, introduce the balance factors w and w ′ , and at the same time, in order to apply softmax later, it is necessary to make ∑ k∈{SN} w k = 1 and ∑ k∈{SN} w′ k = 1, and β SN represents the affine transformation parameter of the spatial normalization method, and γ SN is the weighted balance factor;
[0024] Step 1.4: The feature map obtained in Step 1.2 and the feature map Concatenate along the channel dimension, and use the concatenated features as the input to be provided to the activation layer, thereby building a new basic network M base ;
[0025] Furthermore, the specific steps for building the initialization model in step 2 are as follows:
[0026] Step 2.1: Use the basic network M built in step 1.4 base , and perform training on the initial task D 1 to obtain the initialization model;
[0027] Step 2.2: Freeze the initialization model in step 2.1 and use it as the old model M 1 for training on the next type of incremental learning task D 2 .
[0028] Furthermore, the specific steps for expanding the model scale and building the feature enhancement model in step 3 are as follows:
[0029] Step 3.1: When a new image classification task D t , t ∈ [2, K] arrives, first expand the old model M t-1 from the previous stage, and add a new model structure ΔM t for learning new categories. The adjustment of the network structure is expressed as:
[0030]
[0031] Among them, represents the expansion of the model, M t is the enhanced model after expansion, ΔM t is a model structure of the same scale as M base , and a fully connected layer that can connect the old and new models is built on M t ;
[0032] Step 3.2: After the enhanced model M t is built, in the training, the feature enhancement of the model is realized by adopting the reconstruction loss function. The specific loss function is:
[0033]
[0034] Among them, is the error between the predicted value of the new model M t and the actual value y of the data. D KL (·||·) represents calculating the difference between two probability distributions using the KL divergence, and S(·) represents the softmax operation; is the difference between the predicted value of the old model and the actual value, and γ is a hyperparameter used for weighting; Use knowledge distillation to encourage the new model to have a similar output distribution to the old model on the old classes;
[0035] Step 3.3: Use the constructed loss function on the dataset D t ∪ε to perform class-incremental training on the model M t ;
[0036] Furthermore, the specific steps for compressing the model scale and building the target compressed distillation model in Step 4 are as follows:
[0037] Step 4.1: For the feature-enhanced model M obtained in Step 3 t , its scale is twice that of M base . Here, compression means are taken to reduce its scale to be equal to that of M base :
[0038] M t_new = Com(M t )
[0039] where Com(·) represents model distillation, aiming to remove redundant parameters and unnecessary dimensions in the enhanced model and maintain the scale of the target compressed model to be consistent with M base ;
[0040] Step 4.2: For the Com(·) means in Step 4.1, the balanced distillation algorithm is adopted. Here, the model loss formula is:
[0041]
[0042] where α is a weighted parameter, D KL (·||·) represents the KL divergence, and S(·) represents the softmax operation. Let the loss function be added to the training process to promote the target compressed model to fully learn the knowledge possessed by the enhanced model, thereby reducing catastrophic forgetting;
[0043] Step 4.3: Use the loss mentioned in Step 4.2 to perform the training of M t_new . After the training is completed, use M t_new as the target compressed model M t of this stage;
[0044] Step 4.4: If the current task D t ∈{D 2 ,…,D K-1}, use the target compressed model M t as the old model for the next stage of the task, thus completing the entire class-incremental learning process; if the current task is the final task D K, directly output the target compression model as the result.
[0045] Furthermore, in step 5, image classification accuracy measurement is adopted to conduct performance evaluation on the whole process and in stages of class-incremental learning. The specific steps of classification accuracy prediction are as follows:
[0046] Step 5.1: Since the training environment of the model is class-incremental learning, it is necessary to evaluate the performance of the stage models trained for each task, so as to determine the task performance and classification performance of the model in the overall incremental learning. As can be seen from step 4.3, after the t-th task is trained, the target compression model M of this stage can be obtained. t ;
[0047] Step 5.2: For the t tasks that have been trained currently, the test sample sets of all known classes are where l j represents the number of test samples of the j-th class, is the input sample, is its corresponding true label. Use the trained target compression model M t to predict the test samples of all tasks, and the output of the model is the predicted label For each task, the following formula is used to calculate the classification accuracy Acc j :
[0048]
[0049] where, is the indicator function, which takes the value of 1 when the predicted value is consistent with the true label , and 0 otherwise.
[0050] Step 5.3: Calculate the average classification accuracy of the model on all current tasks to evaluate the overall performance of the model during the incremental learning process:
[0051]
[0052] This value represents the comprehensive classification accuracy of the model on all tasks and is a key indicator for measuring catastrophic forgetting in incremental learning.
[0053] According to the second aspect of the embodiments of the present application, an image classification system for class-incremental learning based on multi-normalization and dynamic network is provided, including a data set unit, a basic network unit, a class-incremental learning unit, and a classification performance evaluation unit;
[0054] For the image classification task under class-incremental learning, the data set unit processes the image classification data set to obtain a training set, a validation set, and a test set;
[0055] The basic network unit constructs a multi-normalization module to optimize the backbone network, realizes the construction of the basic network, and simultaneously uses the training set and validation set transmitted by the dataset unit for model initialization training;
[0056] The class incremental learning unit uses the basic network built by the basic network unit and the training set and validation set from the dataset unit for class incremental learning;
[0057] The final model obtained from the compression and distillation module in the class incremental learning unit is sent to the classification performance evaluation unit, and the performance of the model in the image classification task is evaluated using the test set.
[0058] The specific descriptions of the units are as follows:
[0059] Dataset unit: Used to store and manage the dataset for the class incremental learning image classification task, including the training set, validation set, and test set. This unit supports dynamic loading and splitting of data, can divide the dataset into multiple incremental stages according to task requirements, and provides data augmentation functions to improve the generalization ability of the model.
[0060] Basic network unit: Constructs a basic network containing a multi-normalization module. This unit dynamically selects the normalization method suitable for the task characteristics and completes the training of the basic network in the initial stage, providing strong feature learning ability for subsequent incremental tasks.
[0061] Class incremental learning unit: Includes a dynamic network expansion module and a model distillation and compression module, which are responsible for dynamic network expansion and compression when a new task arrives, enhance the learning ability of the new task by adding new modules, and retain the classification performance of the old task by freezing the old modules. This unit combines knowledge distillation technology to optimize the balance between new and old tasks and effectively alleviates the problem of catastrophic forgetting.
[0062] Classification performance evaluation unit: Evaluates and monitors the classification performance of the model on all incremental tasks by calculating the classification accuracy of the test dataset.
[0063] An image classification device for class incremental learning based on multi-normalization and dynamic network, comprising:
[0064] Memory: Used to store the computer program for implementing the image classification method for class incremental learning based on multi-normalization and dynamic network as described above;
[0065] Processor: Used to implement the image classification method for class incremental learning based on multi-normalization and dynamic network when executing the computer program.
[0066] A computer-readable storage medium, comprising:
[0067] The computer-readable storage medium stores a computer program, which when executed by a processor can implement an image classification method based on multi-normalization and dynamic network for class incremental learning.
[0068] The computer-readable storage medium stores a computer program, which when executed by a processor can implement an image classification method based on multi-normalization and dynamic network for class incremental learning.
[0069] Advantages of the present invention:
[0070] By designing a multi-normalization module, the present invention fully combines the advantages of batch normalization and spatial normalization, providing the model with more flexible and effective feature processing capabilities, enabling the model to dynamically adjust the feature distributions of new and old image classification tasks, and significantly improving the learning performance and task adaptability of the model.
[0071] By retaining the model of the old image classification task to enhance the model's memory ability for old knowledge, and at the same time adding an exclusive module for the new image classification task to enhance the model's learning ability for the new task, the present invention improves the model's performance on the old task while ensuring the performance of the new task, achieving balanced learning of new and old image classification tasks.
[0072] The present invention designs a new loss function, which dynamically adjusts the learning direction of the model from three aspects: new task learning, old knowledge retention, and comparison between new and old models, optimizes the balance between new task learning and old task memory, and ensures that the overall performance of the model does not fluctuate significantly in a multi-task scenario.
[0073] By distilling the feature enhancement model, the present invention reduces redundant parameters and additional overhead, achieves model scale control, and at the same time ensures the model's memory for all tasks, making the model more practical.
[0074] Aiming at the image classification task under class incremental learning, the present invention significantly enhances the stability and accuracy of the model in processing image classification tasks through the collaborative application of multiple technologies, achieving a comprehensive improvement in the model learning efficiency and image classification accuracy.
[0075] The present invention combines dynamic network with knowledge distillation, model compression, and accuracy prediction technologies. The overall system has low computational overhead and memory requirements, is suitable for class incremental learning scenarios, and can effectively alleviate the catastrophic forgetting problem.
[0076] The dynamic network strategy of the present invention is that during the training process, the model scale will be expanded or contracted accordingly. In step 3, it is manifested as: when a new task of class incremental learning arrives, an additional network architecture will be trained for this new task, thus achieving the expansion of the model scale, that is, the old model + the new model, and the model scale is expanded to twice the original model. In step 4, it is manifested as: after the model is expanded, in order to prevent the model scale from continuously increasing (because if a model is trained for each task, it will lead to a linear increase in the number of models and the parameter scale, resulting in excessive consumption of computing resources and memory resources), the method of knowledge distillation is adopted to distill the model scale of twice into one time, achieving the contraction of the model scale. Through "expansion" and "contraction", the dynamic network technology is realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 It is a flowchart provided by the present invention.
[0078] Figure 2 It is a schematic diagram of the multi-normalization module and the basic network provided by the present invention.
[0079] Figure 3 It is a breakdown diagram of the learning steps provided by the present invention.
[0080] Figure 4 It is a line chart comparing the accuracy of incremental learning of B0-5 tasks on the CIFAR100 dataset between the present invention and other methods.
[0081] Figure 5 It is a line chart comparing the accuracy of incremental learning of B0-10 tasks on the CIFAR100 dataset between the present invention and other methods.
[0082] Figure 6 It is a schematic diagram of the system provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0083] The present invention will be further described in detail below with reference to the accompanying drawings.
[0084] Refer to Figure 1 , which is a flowchart of an image classification method for class incremental learning based on multi-normalization and dynamic network provided by the present invention. This method and system can alleviate the catastrophic forgetting problem in the class incremental learning scenario, effectively control the model scale while achieving better model performance. By comprehensively utilizing the advantages of various normalization methods, a multi-normalization module is designed to optimize the network from the module composition level of the model, alleviating the catastrophic forgetting problem. In addition, based on the dynamic network, a two-stage training model of feature enhancement-compression distillation is built. On the one hand, it retains the advantages brought by the model expansion, and on the other hand, it effectively avoids the negative impacts brought by the parameter scale expansion and the retention of the old model. This method consists of the following steps:
[0085] When conducting invention verification, an image dataset CIFAR100 containing 100 categories is used, and experiments are carried out based on the traditional backbone network Resnet32. The number of training iterations for the initial model and the feature enhancement model is 150 times, and the number of iterations for training the compressed model is 130 times. During the training process, the batch size is set to 128, the stochastic gradient descent optimizer is used, the initial learning rate is 0.1, and the cosine annealing is used to adjust the learning rate during training. There are two settings in the experiment, namely B0 and B50, that is, the model does not learn any categories in the initial stage and learns 50 categories in the initial stage. The B0-10 task under the CIFAR100 dataset indicates that no categories are learned in the initial stage, and 10 categories are learned in each of the subsequent 10 task stages.
[0086] Step 1: On the basis of the model backbone network, introduce a multi-normalization module to replace the traditional single batch normalization method, and build a new basic network M base ;
[0087] Step 1.1: The class incremental learning scenario is set to perform step-by-step learning in a data stream with new categories, and a series of K training tasks {D 1 , D 2 , …, D K} are set, and there are no overlapping categories between tasks, where represents the t-th incremental stage, which contains n t samples, and at the same time, a sample set ε with a limited memory size is constructed to store representative samples of the learned categories.
[0088] See Figure 2 , the schematic diagram of the multi-normalization module and the basic network. The left side is the specific composition of the multi-normalization module, and the right side is the specific position of the multi-normalization module in the basic network M base ;
[0089] The multi-normalization module receives the feature map from the convolutional layer. Here, the feature map input to the multi-normalization module is denoted as where B is the mini-batch size, C is the number of channels, W is the width, and H is the height. It is split into two parts along the channel dimension to obtain and
[0090] Step 1.2: Apply the batch normalization (BN) method to the first feature map p 1 , calculate its mean and variance, and use the affine transformation to obtain the normalized feature map That is:
[0091]
[0092] Among them, represents the mean value, represents the variance, and γ BN is the weighted balance factor;
[0093] Step 1.3: Apply the spatial normalization (SN) method, such as instance normalization (IN) or layer normalization (LN), to the second feature map p 2 During this process, according to the selected type of spatial normalization, calculate the corresponding mean value and variance to obtain the feature map
[0094]
[0095] Among them, SN = {IN, LN}. If SN selects IN, calculate and the variance If SN selects LN, calculate the mean value along the layer dimension and the variance When using multiple spatial normalization methods, introduce the balance factors w and w ′ , and at the same time, in order to apply softmax subsequently, it is necessary to make ∑ k∈{SN} w k = 1 and ∑ k∈{SN} w' k = 1, while β SN represents the affine transformation parameter of the spatial normalization method, and γ SN is the weighted balance factor;
[0096] Step 1.4: Concatenate the feature map obtained in Step 1.2 and the feature map obtained in Step 1.3 along the channel dimension and provide it as input to the activation layer, thereby building a new basic network M base ;
[0097] See Figure 3 , Step 1 of the learning step decomposition diagram of the method provided by the present invention. Through Step 1, the basic network M base can be trained on the task D 1 to obtain an initial model.
[0098] Step 2: In the first stage of class-incremental training, based on the network M base perform the training of the initial task D 1 , for example, D in the B0-10 task under the CIFAR100 dataset 1 represents the learning of the first 10 categories to obtain an initial model;
[0099] Step 2.1: Use M built in Step 1.4 base to perform training on the initial task D 1 and obtain an initialized model;
[0100] Step 2.2: Freeze the initialized model in Step 2.1 and use it as the old model M 1 for training the next class of incremental learning tasks D 2 ;
[0101] See Figure 3 , Step 2 of the learning step decomposition diagram of the method provided by the present invention. Through Step 2, a new model structure can be trained for a new image classification task, and at the same time, together with the frozen model in the previous stage, the construction of a feature enhancement model can be realized, so as to improve the accuracy of the model for the image classification task.
[0102] Step 3: When processing a new class incremental task, based on the basic network, train an additional new model for the new task, construct a fully connected layer, and effectively connect the old model and the new model through the fully connected layer to build a feature enhancement model;
[0103] Step 3.1: When the new task D t , t∈[2,K] arrives, first expand the model M t-1 in the previous stage and add a new model structure ΔM t for learning new categories. The adjustment of the network structure is expressed as:
[0104]
[0105] where, represents the expansion of the model, M t is the enhanced model after expansion, ΔM t is a model structure of the same scale as M base , and a fully connected layer that can connect the old and new models is built on M t ;
[0106] Step 3.2: After the enhanced model M t is built, in the training, use the reconstruction loss function to realize the feature enhancement of the model. The specific loss function is:
[0107]
[0108] where, is the error between the predicted value of the new model M t and the actual value y of the data, D KL (·||·) represents calculating the difference between two probability distributions using the KL divergence, and S(·) represents the softmax operation; is the difference between the predicted value of the old model and the actual value, and γ is the hyperparameter used for weighting; Use knowledge distillation to encourage the new model to have a similar output distribution to the old model on the old classes;
[0109] Step 3.3: Use the constructed loss function on the dataset D t ∪ ε to perform class-incremental training on the model M t ;
[0110] See Figure 3 , Step 3 of the learning step decomposition diagram of the method provided by the present invention. Through Step 3, the feature enhancement model can be scaled down using model distillation technology to obtain the target compression model, thereby effectively controlling the model scale without loss of image classification accuracy.
[0111] Step 4: Control the scale of the feature enhancement model through model distillation technology, and distill the feature enhancement model to obtain the target compression model; the target compression model is used as the old model when processing the next class-incremental task;
[0112] Step 4.1: For the feature enhancement model M t obtained in Step 3, its scale is twice that of M base . Here, compression means are taken to reduce its scale to be equal to that of M base :
[0113] M t_new = Com(M t )
[0114] where Com(·) represents model distillation, aiming to remove redundant parameters and unnecessary dimensions in the enhancement model and maintain the scale of the target compression model consistent with M base ;
[0115] Step 4.2: For the Com(·) means in Step 4.1, adopt the balanced distillation algorithm. Here, the model loss formula is:
[0116]
[0117] where α is the weighting parameter, D KL (·||·) represents the KL divergence, and S(·) represents the softmax operation. Add the loss function to the training process to promote the target compression model to fully learn the knowledge possessed by the enhancement model, thereby reducing catastrophic forgetting;
[0118] Step 4.3: Use the loss mentioned in Step 3.2 to perform training on M t_new . After the training is completed, Mt_new As the target compression model M at this stage t ;
[0119] Step 4.4: If the current task D t ∈{D 2 ,…,D K-1}, use the target compression model M t as the old model for the next stage of the task, thus improving the entire class incremental learning process; if the current task is the final task D K , directly output the target compression model as the result;
[0120] Step 5: On the test samples, use the described target compression model to detect the image classification accuracy.
[0121] Step 5.1: Since the training environment of the model is class incremental learning, it is necessary to evaluate the performance of the stage models trained for each task, so as to determine the task performance and classification performance of the model in the overall incremental learning. As can be seen from Step 4.3, after the t-th task is trained, the target compression model M t ;
[0122] Step 5.2: For the t tasks that have been trained currently, the test sample sets of all known classes are where l j represents the number of test samples of the j-th class, is the input sample, is its corresponding true label, and use the trained target compression model M t to predict the test samples of all tasks. The output of the model is the predicted label For each task, use the following formula to calculate the classification accuracy Acc j :
[0123]
[0124] where, is the indicator function, which takes the value of 1 when the predicted value is consistent with the true label , otherwise 0.
[0125] Step 5.3: Calculate the average classification accuracy of the model on all current tasks to evaluate the overall performance of the model in the incremental learning process:
[0126]
[0127] This value represents the comprehensive classification accuracy of the model on all tasks and is a key indicator for measuring catastrophic forgetting in incremental learning.
[0128] To evaluate the specific performance of a proposed class-incremental learning method based on multi-normalization and dynamic networks, the method of this application (Ours) is compared with multiple advanced models, namely iCaRL, BiC, WA, COIL, PODNet, DER, AFC, MEMO, and eTag. Different incremental stage settings and tests are conducted on the CIFAR100 dataset with 100 classes. Specifically, the tests include five settings under B0 and B50:
[0129] Setting 1: B0-5 tasks: B0 indicates that the model has not learned any classes in the initial stage, that is, the initial task contains 0 classes; subsequently, the model gradually learns 5 image classification tasks in the incremental stage, with each task containing 20 classes;
[0130] Setting 2: B0-10 tasks: B0 in the initial stage also indicates that the initial task has 0 classes; subsequently, in the incremental stage, the dataset is decomposed into 10 tasks, with each task containing 10 classes;
[0131] Setting 3: B0-20 tasks: B0 in the initial stage also indicates that the initial task has 0 classes; subsequently, in the incremental stage, the dataset is decomposed into 20 tasks, with each task containing 5 classes;
[0132] Setting 4: B50-5 tasks: B50 in the initial stage indicates that the model has learned 50 classes in the initial task; subsequently, the incremental stage is further decomposed into 5 tasks, and each task learns 10 classes from the remaining 50 classes;
[0133] Setting 5: B50-10 tasks: B50 in the initial stage indicates that the model has learned 50 classes in the initial task; subsequently, the remaining 50 classes are decomposed into 10 tasks in the incremental stage, and each task learns 5 classes.
[0134] The experiment reports the average accuracy of the model at all stages under 4 settings, and the experimental results are shown in Table 1.
[0135] Table 1 Comparison of average accuracy of multiple methods on CIFAR100
[0136]
[0137]
[0138] Referring to Table 1, the method proposed in this application (Ours) performs excellently in Settings 2 - 5 of the CIFAR100 dataset, achieving performance improvement compared to various existing advanced strategies, which fully demonstrates the effectiveness and advantages of the method in this application in the class-incremental learning task. Specifically, by introducing a multi-normalization module, a dynamic network expansion and compression strategy, and an optimized loss function design, the method in this application effectively alleviates the catastrophic forgetting problem while taking into account the new task learning ability, achieving a good balance between old and new tasks.
[0139] Referring to Figure 4 , the line chart of the accuracy comparison of the method of the present invention and other methods during the B0 - 5 task incremental learning on the CIFAR100 dataset. This figure shows the recorded and plotted results of the image classification accuracy for the entire stage under Setting 1. As can be seen from the figure, the method proposed in the present invention shows better overall accuracy than other comparison methods in the incremental learning of each stage, demonstrating the advantages of the present invention.
[0140] Referring to Figure 5 , which shows the line chart of the accuracy comparison of the method of the present invention and other methods during the B0 - 10 task incremental learning on the CIFAR100 dataset. Through the comparison under Experimental Setting 2, it can be further verified that the method of the present invention still maintains relatively stable and high accuracy within a larger task range and continues to lead other existing methods.
[0141] Referring to Figure 6 , which is a schematic diagram of a class-incremental learning system based on multi-normalization and dynamic network provided by an embodiment of the present invention, including the following modules:
[0142] Dataset unit: Used to store and manage the dataset for the class-incremental learning task, including the training set, validation set, and test set. This unit supports dynamic loading and splitting of data, can divide the dataset into multiple incremental stages according to task requirements, and at the same time provides data augmentation functions to improve the generalization ability of the model.
[0143] Basic network unit: Construct a basic network containing a multi-normalization module. This unit dynamically selects the normalization method suitable for the task characteristics and completes the training of the basic network in the initial stage, providing strong feature learning ability for subsequent incremental tasks.
[0144] Class-incremental learning unit: Includes a dynamic network expansion module and a model distillation and compression module, that is, responsible for dynamic network expansion and compression when a new task arrives, enhancing the new task learning ability by adding new modules, and at the same time retaining the classification performance of the old task by freezing the old modules. This unit combines knowledge distillation technology to optimize the balance between old and new tasks and effectively alleviates the catastrophic forgetting problem.
[0145] Classification performance evaluation unit: Evaluates and monitors the classification performance of the model on all incremental tasks by calculating the classification accuracy of the test data set.
Claims
1. An image classification method based on multi-normalization and dynamic network incremental learning, characterized in that: The steps include: Step 1: Introduce the multi-normalization module based on the model backbone network and build a new basic network M base ; Step 2: Perform class incremental learning of the initial image classification task D based on the new base network 1 The training is done to get the initialization model, and the initialization model is frozen as the class incremental learning next stage image classification task D 2 The old model; Step 3: When processing a new class incremental learning image classification task, a new model is additionally trained for the new task based on the new basic network, and the old model and the new model are connected through a fully connected layer to build a feature enhancement model; Step 4: Control the scale of the feature enhancement model through model distillation technology, and distill the feature enhancement model to obtain a target compression model; the target compression model is used as the old model for processing the next class increment task; Step 5: On the test sample, use the target compression model to perform image classification accuracy detection.
2. The image classification method based on multi-normalization and dynamic network incremental learning according to claim 1, characterized in that: In step 1, the model backbone model adopts the Resnet32 model, and the new multi-normalization module is constructed and the basic network M is built on this basis. base The specific steps are: Step 1.1: The incremental learning scenario is set to perform step-by-step learning in a data stream with new categories, with a series of K image classification training tasks {D 1 ,D 2 ,…,D K }, there is no overlapping class between the tasks, where Indicates the tth incremental stage, which contains n t samples, and at the same time construct a sample set ε with limited memory size to store representative samples of the learned categories. The feature map input to the multi-normalization module is recorded as Where B is the batch size, C is the number of channels, W is the width, and H is the height. Split it into two parts along the channel dimension to get the first feature map And the second feature map Step 1.2: Apply the batch normalization (BN) method to the first feature map p1, calculate its mean and variance, and use affine transformation to obtain the normalized feature map Right now: in, represents the mean, represents the variance, γ BN is the weighted balance factor; Step 1.3: Apply the spatial normalization (SN) method to the second feature map p2. According to the selected spatial normalization type, calculate the corresponding mean and variance to obtain the feature map Among them, SN = {IN, IN}, if SN is selected as IN, there is a channel dimension and variance If SN uses LN, there is a mean calculated along the layer dimension and variance When using multiple spatial normalization methods, the balancing factors w and w′ are introduced, and ∑ k∈{SN} w k =1 and∑ k∈{SN} w′ k =1, and β SN represents the affine transformation parameters of the spatial normalization method, γ SN is the weighted balance factor; Step 1.4: The feature map obtained in step 1.2 And the feature map obtained in step 1.3 Splicing along the channel dimension, providing the spliced features as input to the activation layer, and building a new basic network M base .
3. The image classification method based on multi-normalization and dynamic network incremental learning according to claim 2, characterized in that: The specific steps of building the initialization model in step 2 are: Step 2.1: Use the basic network M built in step 1.4 base , perform initial task D 1 Training on and getting the initialization model; Step 2.2: Freeze the initialized model in step 2.1 and use it as the old model M1 for the next incremental learning task D 2 training.
4. The image classification method based on multi-normalization and dynamic network incremental learning according to claim 3, characterized in that: Step 3 expands the model scale, and the specific steps for building a feature enhancement model are: Step 3.1: When a new image classification task D t , when t∈[2,K] arrives, first the old model M in the previous stage t-1 Expand and add a new model structure ΔM t To learn new categories, the adjustment of the network structure is expressed as: in, Represents the extension of the model, M t is the extended enhanced model, ΔM t For M base Model structure of the same scale, M t A fully connected layer that connects the new and old models is built on top; Step 3.2: In the enhanced model M t After the construction is completed, the feature enhancement of the model is achieved by reconstructing the loss function during training. The specific loss function is: in, For the new model M t The error between the predicted value and the actual value y of the data, D KL (·||·) represents the difference between two probability distributions calculated using KL divergence, and S(·) represents the softmax operation; is the difference between the predicted value of the old model and the actual value, γ is the hyperparameter used for weighting; Using knowledge distillation to encourage the new model to have a similar output distribution on old categories as the old model; Step 3.3: Using the constructed loss function In the dataset D t ∪ε for model M t Perform class incremental training.
5. The image classification method based on multi-normalization and dynamic network incremental learning according to claim 4, characterized in that: The step 4 compresses the model scale, and the specific steps of building the target compression distillation model are: Step 4.1: For the feature enhancement model M t Compression measures are taken to reduce its size to the same size as M base Equal scale: M t_new =Com(M t ) Com(·) represents model distillation, which removes redundant parameters and unnecessary dimensions in the enhanced model to maintain the target compression model size and M base Stay consistent; Step 4.2: For the Com(·) method in step 4.1, the equilibrium distillation algorithm is used, where the model loss formula is: Where α is the weighting parameter, D KL (·||·) represents KL divergence, and S(·) represents the softmax operation; Step 4.3: Using the loss mentioned in step 4.2 Carry out M t_new After the training is completed, M t_new As the target compression model M in this stage t ; Step 4.4: If the current task D t ∈{D 2 ,…,D K-1 }, the target compression model M t As the old model for the next stage of tasks, the entire incremental learning process is improved; if the current task is the final task D K , directly output the target compression model as the result.
6. The image classification method based on multi-normalization and dynamic network incremental learning according to claim 1, characterized in that: Step 5 uses image classification accuracy measurement to perform full-process and stage-by-stage performance evaluation of incremental learning. The specific steps of classification accuracy prediction are as follows: Step 5.1: From step 4.3, we can know that after the t-th task training is completed, the target compression model M of this stage is obtained. t ; Step 5.2: For the t tasks that have been trained, the corresponding test sample sets of all known classes are Among them l j represents the number of test samples of the jth category, is the input sample, For its corresponding true label, use the trained target compression model M t Predict the test samples of all tasks, and the output of the model is the predicted label For each task, the classification accuracy Acc is calculated using the following formula: j : in, is the indicator function, when the predicted value With the true label The value is 1 when they are consistent, otherwise it is 0; Step 5.3: Calculate the average classification accuracy of the model on all current tasks to evaluate the overall performance of the model in the incremental learning process: This value represents the comprehensive classification accuracy of the model on all tasks and is a key indicator for measuring catastrophic forgetting in incremental learning.
7. An image classification system based on multi-normalization and dynamic network incremental learning, characterized in that: It includes a data set unit, a basic network unit, a class incremental learning unit, and a classification performance evaluation unit; For the image classification task under incremental learning, the dataset unit performs relevant processing on the image classification dataset to obtain the training set, validation set, and test set; The basic network unit constructs a multi-normalization module to optimize the backbone network and build the basic network. At the same time, the training set and validation set delivered by the data set unit are used to perform model initialization training. The incremental learning unit uses the basic network built by the basic network unit and the training set and verification set from the data set unit to perform incremental learning; The final model obtained by the compression distillation module in the incremental learning unit is sent to the classification performance evaluation unit, and the test set is used to evaluate the performance of the model in the image classification task.
8. The image classification system based on multi-normalization and dynamic network incremental learning according to claim 7, characterized in that: The data set unit is used to store and manage data sets of image classification tasks of class incremental learning, and divide the data sets into multiple incremental stages according to task requirements; The basic network unit dynamically selects a normalization method suitable for the task characteristics and completes the training of the basic network in the initial stage; The quasi-incremental learning unit includes a dynamic network expansion module and a model distillation compression module, which is responsible for dynamic network expansion and compression when a new task arrives; The classification performance evaluation unit evaluates and monitors the classification performance of the model on all incremental tasks by calculating the classification accuracy of the test data set.
9. An image classification device based on multi-normalization and dynamic network class incremental learning, characterized in that: include: Memory: used to store a computer program for implementing the image classification method based on multi-normalization and dynamic network-based incremental learning according to any one of claims 1 to 7; Processor: used to implement the image classification method based on multi-normalization and dynamic network-based incremental learning according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, comprising: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement any one of claims 1-7, an image classification method based on multi-normalization and dynamic network-based incremental learning. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement any one of claims 1-7, an image classification method based on multi-normalization and dynamic network-based incremental learning.
Citation Information
Patent Citations
Distillation type incremental learning method based on multilevel knowledge
CN117494790A
Continuous learning image classification method and device based on deep learning
CN114463605A
Target detection method and device based on lifelong learning
CN115620099A
Picture classification method and system based on knowledge distillation small sample incremental learning
CN116503676A
New and old feature compatible learning method for structure expansion and distillation
CN117934923A
Cited By
Image classification method of class incremental learning based on kernel sparsity and entropy
CN121686094A