Image classification method based on incremental learning of multi-branch network architecture

Through an incremental learning method based on a multi-branch network architecture, combined with the fusion strategy of Wasserstein distance and auxiliary network, the catastrophic forgetting problem arises in multi-task learning by deep neural networks is solved, and efficient image classification and parameter control are achieved.

CN119942234AActive Publication Date: 2025-05-06XIDIAN UNIV

Patent Information

Application Number
CN202510204110.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-06
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

Deep neural networks are prone to catastrophic forgetting when scaling from single-task to multi-task learning, causing the model to perform worse in the old categories. The existing incremental learning methods have problems such as excessive parameter volume, high memory requirements, and serious forgetting of old classes.

Method used

The incremental learning method based on a multi-branch network architecture is adopted to quantify the task similarity by calculating the Wasserstein distance between new and old image classification tasks, and decide whether to expand the new network model. Introduce tasks that are not similar to auxiliary networks, and through knowledge distillation and dynamic structural reorganization strategies, the auxiliary network parameters are integrated into the backbone network to control the parameter scale.

Benefits of technology

It effectively reduces the forgetting of old classes caused by learning new classes, controls the scale of parameter quantity, improves the accuracy of image classification and model adaptability, and is suitable for practical application scenarios where resource limitations are required.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942234A_ABST
    Figure CN119942234A_ABST
Patent Text Reader

Abstract

The invention discloses an incremental learning method based on a multi-branch network architecture. The method comprises the following steps of: 1, for each new task Ti, calculating the similarity between the task and a learned old task, and sorting; 2, comparing the sorted highest similarity value with a set similarity threshold value, if the highest similarity value is smaller than the threshold value, expanding a new network model for learning, otherwise, using the most similar network model for learning and updating; 3, an auxiliary network is introduced into each network model, the auxiliary network is used for learning dissimilar tasks in the network model, the network model learns similar tasks in a knowledge distillation mode, after learning, the auxiliary network is fused into the network model, and a final model after expansion is obtained; and 4, predicting a test sample by using the extended final model, and calculating final classification precision. According to the method, the old class forgetting caused by new class learning is reduced, and the scale of the parameter quantity is also controlled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image classification, and in particular relates to an image classification method based on incremental learning of a multi-branch network architecture. Background Art

[0002] At present, the mainstream machine learning methods are generally learning and reasoning. Learning refers to learning a network model for a specified data set from a specified data set, and then using it for reasoning. Whenever a new task is encountered, the model is generally trained from scratch, which makes repeated training of the model inefficient and time-consuming. Humans have strong memory and continuous learning abilities. Humans can gradually expand from simple task learning to complex task learning. Many scholars also want deep neural networks to imitate this process. Deep neural networks gradually expand from single-task learning to multi-task learning without having to train from scratch every time, which will save a lot of resources and time costs. However, when neural networks expand from single-task learning to multi-task learning, catastrophic forgetting will occur, that is, the network model will perform well on new categories, but will perform worse on old categories. This is because when the deep neural network learns a new task, the network model parameters change, causing the model to forget the previous tasks, limiting the continuous learning ability of the deep neural network.

[0003] Incremental learning methods can be mainly divided into three categories: replay-based methods, regularization-based methods, and architecture-based methods.

[0004] The replay-based method mainly imitates the human memory mechanism, stores the learned knowledge, and learns the stored knowledge and the new task together when learning a new task, so as to reduce catastrophic forgetting. However, storing the original data will have memory limitations and data privacy issues.

[0005] Regularization-based methods mainly use a certain strategy to limit the update amplitude of some weights when updating the weights of the neural network, usually to measure the impact of the sample on the model.

[0006] However, in the scenario of a large number of tasks, the category shift is more serious. Since the weights of important old tasks change little, the model's ability to learn new tasks decreases as the number of tasks increases.

[0007] The architecture-based approach refers to adding corresponding new modules (i.e., network structures) to the neural network for different tasks, so that the modules of the new tasks are independent of the modules of the old tasks and will not interfere with the previously learned knowledge. Although the architecture-based approach solves the problem of catastrophic forgetting to a large extent, it will generate a huge amount of parameters and have high memory requirements.

[0008] The invention is entitled "Image classification method and system based on CLIP category incremental learning" and the application publication number is (CN118506049 A). The invention proposes an image classification method and system based on CLIP category incremental learning. First, according to the similarity between the text features of the newly input training image and the text features of the old category, the new and old category pairs of adjacent categories are screened, and the old category text features in each pair of adjacent categories are sampled using a normal distribution to construct a hinge loss function; after completing the t-th training task, the adapter parameters of the previous training task are fused with the adapter parameters of the current training task to obtain the final adapter parameters of the current training task, and the trained CLIP model is used to obtain the classification result for the processed image, thereby reducing the forgetting of old categories caused by learning new categories and improving the image classification accuracy.

[0009] This method belongs to the regularization algorithm, which mainly focuses on reducing the forgetting of old categories. However, when dealing with new categories, it may have limited learning ability for new categories, especially when the new categories are very different from the old categories, the classification performance may decrease. Summary of the invention

[0010] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide an image classification method based on incremental learning of a multi-branch network architecture, which quantifies the similarity between the new image classification task and the old image classification task by calculating the Wasserstein distance between the new and old image classification tasks, and decides whether to expand the new network model based on the similarity. At the same time, an auxiliary network is introduced into a backbone network to learn dissimilar tasks, which are integrated into the backbone network after learning, so as to reduce the scale of network parameters. This method strategy not only reduces the forgetting of old classes caused by learning new classes, but also controls the scale of parameters.

[0011] In order to achieve the above object, the technical solution adopted by the present invention is:

[0012] An image classification method based on incremental learning of a multi-branch network architecture comprises the following steps;

[0013] Step 1: For each new image classification task T i (For example, classifying a dataset containing cats and dogs), calculating the similarity between the task and the old image classification tasks that have been learned, and ranking them;

[0014] Step 2: Compare the highest similarity value with the set similarity threshold. If it is less than the threshold, expand the new network model for learning. Otherwise, use the most similar network model for learning and updating.

[0015] Step 3: Introduce an auxiliary network into each network model in step 2. The auxiliary network is used to learn dissimilar tasks within the network model. The backbone network of the network model uses knowledge distillation to learn similar tasks. After learning, the auxiliary network is integrated into the backbone network to obtain the expanded final model, which controls the parameter scale and prevents catastrophic forgetting.

[0016] Step 4: Use the expanded final model to predict the test samples and calculate the final classification accuracy.

[0017] The specific process of calculating the task similarity in step 1 is as follows:

[0018] Step 1.1: For T classes incremental image classification task is the t-th class incremental image classification task, including N t sample images, is the i-th sample of the t-th class incremental image classification task, is the label value of the i-th sample of the corresponding t-th class incremental image classification task. Set the image classification task T i The category set is C i , T represents the number of tasks, Ti represents the i-th task, for each category c∈C i , maintain a prototype feature It is defined as the mean of the features of image samples belonging to category c:

[0019]

[0020] Among them, w i Represents sample X i The weights of are used to define according to class balance or sample confidence;

[0021] Step 1.2: Calculate the nonlinear distance between the sample image and the prototype feature based on the kernel function method:

[0022]

[0023] where k(·,·) is a kernel function, such as a Gaussian kernel, defined as:

[0024]

[0025] In addition, the new sample image X l The minimum distance to all category prototypes of the previous image classification task is defined as:

[0026]

[0027] Among them, C j It is task T jA collection of categories, It is task T j The prototype features of category c;

[0028] Step 1.3: Calculate Wasserstein distance, image classification task T i and T j The dissimilarity between the distributions of is measured by the Wasserstein distance, defined as:

[0029]

[0030] Among them, Π(q i ,q j ) represents the joint distribution between task distributions;

[0031] In order to simplify the calculation, the Sinkhorn approximation is used to estimate the Wasserstein distance:

[0032]

[0033] Step 1.4: Design a dynamic mapping function to map the Wasserstein distance to the interval [0,1], the formula is as follows:

[0034] s i,g =1-exp(-α·W ε (q i ,q j ))

[0035] Among them, the new task T i The similarity measure between the old image classification task group g is in the range of [0,1], and α is a scaling factor adjusted through experiments.

[0036] The specific process of the model expansion in step 2 is as follows:

[0037] After the similarity calculation in step 1, we get N t A group of similarity values, mapping the similarity values ​​to the range of [0,1]. t Sort them and select the maximum value T max Compared with the threshold δ, δ is in [0-1]. If the highest similarity value is greater than the threshold, it means that the current task is very similar to the old task. At this time, the task is assigned to the network N corresponding to the learned task with the highest similarity. old Learning; this approach can minimize training costs and alleviate catastrophic forgetting.

[0038] If the highest similarity value is less than the threshold, it means that the current image classification task is not similar to all the learned image classification tasks. At this time, the model needs to be expanded. The mathematical formula is as follows:

[0039]

[0040] Among them, G(T cur ) represents the current image classification task T cur The network assigned to it, the expansion strategy is to copy the network with the highest similarity to the learned task and use the copied network N new Learning new image classification tasks. This approach can reduce the cost of learning from scratch and shorten training time.

[0041] When step 2 decides to expand a new network model, the newly added network will become part of the multi-branch network architecture. Specifically: the newly added network is copied based on the network with the highest similarity to the learned task to reduce the cost of learning from scratch. After learning the new task, the newly added network will be included in the model library of the framework for reuse or expansion of subsequent tasks.

[0042] The step 3 is specifically as follows:

[0043] Step 3.1: Phase 1 (Initial Training Phase)

[0044] Use complete supervised data (training data X1 and label Y1) for standard classification training; train a standard classification model A standard classification model Contains the following components:

[0045] (1) Feature Extractor Responsible for extracting high-dimensional feature representation of samples;

[0046] (2) Classifier g c1 : Map features to classification labels; use feature extractors to extract features for tasks, and then use classifiers to map these extracted features to specific labels, such as cats, dogs, etc. Feature extractors and classifiers are necessary components of a classification model, and they are in a sequential relationship, first feature extraction and then classification using classifiers.

[0047] Use cross entropy loss L ce Train the model and optimize the target; improve classification performance.

[0048] Incremental phase (phase n)

[0049] Current image classification task T n The input data is only the images of the new task Excluding old image classification task samples, using a new feature extractor Extract feature representations for new tasks:

[0050]

[0051] To learn discriminative features for new image classification tasks, we use new task labels (From Y n ) for supervised optimization, using a classifier Map features to label space:

[0052]

[0053] Calculate the cross entropy loss L ce :

[0054]

[0055] At the same time, in order to retain the knowledge of the old tasks, knowledge distillation is used to keep the feature representation of the old category and extract the feature representation of the previous stage n-1:

[0056]

[0057] Calculate the distillation loss L kd :

[0058]

[0059] where F kd represents the Euclidean distance;

[0060] Step 3.2: Since the classification performance of balancing new tasks with old tasks is not considered in step 3.1, step 3.2 further optimizes the feature representation by using the prototype vector and the balance loss to balance the classification performance of new tasks with old tasks. In the NECIL scenario, a prototype vector (Prototype) is stored in the feature space for each category: these prototypes are used to balance the classification performance of new image classification tasks with old image classification tasks; the prototypes are oversampled to match the batch size B to calibrate the classifier:

[0061] p B =U p B(Prototype)

[0062] Among them U p B is the oversampling function;

[0063] Calculate the prototype balance loss L proto :

[0064] L proto =F ce (p B ,yB )

[0065] where y B is the oversampled label set;

[0066] Step 3.3: In order to ensure that the network in steps 3.1 and 3.2 does not destroy the feature representation of the old image classification task when learning the new image classification task, while keeping the model lightweight, in order to learn the new category unbiasedly while keeping the feature representation of the old category, a dynamic structural reorganization strategy is adopted:

[0067] Structural expansion: A residual adapter is inserted into each convolutional block of the fixed feature extractor in the previous stage. The residual adapter only updates the most discriminative part while retaining the old features:

[0068]

[0069] Structural reparameterization: After training is completed, the side branch information is losslessly integrated into the main branch to reduce the number of parameters:

[0070]

[0071] Delete the adapter to keep the network structure unchanged for the next stage and control the parameter scale.

[0072] Step 3.4: Comprehensively optimize steps 3.1 to 3.3, and further reduce feature confusion through prototype selection mechanism and mask strategy to ensure that the feature representation of the new image classification task and the old image classification task can coexist harmoniously; To reduce feature confusion in the distillation process, a prototype selection mechanism based on scalable embedding space is adopted: Calculate the cosine similarity of all new samples with the old prototype:

[0073]

[0074] Set the similarity threshold σ and select the update strategy according to the sample similarity: Dissimilar samples: Participate in the update of the residual adapter to learn new features. Similar samples: Participate in the distillation process to retain the discriminative features of the old category. Optimize different loss functions by adding masks: If the similarity is greater than σ, the distillation loss L is kd Add mask; if the similarity is less than σ, the cross entropy loss L ce Add the mask. The final loss function is defined as:

[0075] L=Mask ce (L ce )+λMask kd (L kd )+γL proto

[0076] where λ and γ are the loss weights.

[0077] Step 4: For a new test sample x, use the pre-trained model Resnet-18 to extract its features, then calculate its similarity value with the old image classification task, decide which model to use for testing, and use the backbone network to test in a single model to map it to the corresponding category. An image classification system based on incremental learning of a multi-branch network architecture, including a task similarity calculation module, a model expansion module, and an auxiliary network module;

[0078] The task similarity calculation module uses Wasserstein distance to measure the image classification task T i and T j The similarity between the distributions of , and the calculated image classification task similarity is used as the basis for the model extension module to expand the model;

[0079] The model extension module adopts the maximum similarity T max Compare with the threshold δ. If it is greater than the threshold, no expansion is required. Otherwise, copy and expand the new model for learning. Use the image classification task similarity calculation module to calculate the similarity of the task to decide whether to expand the model.

[0080] The auxiliary network module realizes structural expansion and reparameterization through the residual adapter, takes into account both new feature learning and old feature retention, and performs structural expansion and reparameterization on the expanded model in the model expansion module, thereby controlling the parameters of the expanded network model while mitigating catastrophic forgetting.

[0081] An image classification device based on incremental learning of a multi-branch network architecture, comprising:

[0082] Memory: used for storing a computer program for implementing the image classification method based on incremental learning of a multi-branch network architecture;

[0083] Processor: used to implement the image classification method based on incremental learning of a multi-branch network architecture when executing the computer program.

[0084] A computer-readable storage medium, comprising:

[0085] The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement an image classification method based on incremental learning of a multi-branch network architecture.

[0086] Beneficial effects of the present invention:

[0087] The present invention quantifies the task similarity by calculating the Wasserstein distance between the new and old tasks, and can intelligently decide whether to expand the new network model or reuse the existing model. This strategy not only reduces the time and computational cost of repeated training, but also enhances the adaptability and flexibility of the model in the face of diverse and dynamically changing task environments. In image classification tasks, when the new task involves image categories similar to existing tasks (such as expanding from the classification of "cats" and "dogs" to the classification of "tigers" and "wolves"), the present invention can effectively reuse existing models and significantly reduce training time.

[0088] The present invention adopts a multi-branch network architecture, and through a dynamic mapping function and a structural reparameterization strategy, the parameters of the auxiliary network are losslessly integrated into the backbone network, thereby achieving learning of new tasks without significantly increasing the number of model parameters. This method effectively controls the parameter scale of the model, reduces the consumption of memory and computing resources, and is suitable for practical application scenarios with limited resources. In image classification tasks, when a new category (such as "birds") needs to be added, the present invention can efficiently complete the learning of new tasks without significantly increasing the complexity of the model.

[0089] The present invention introduces residual adapters and dynamic structural reorganization strategies to ensure that when learning new tasks, the model can focus on the discriminative learning of new features while retaining the important features of old tasks. This balancing mechanism enables the model to fully tap the information of new tasks and consolidate the knowledge of old tasks when processing multiple tasks, thereby achieving comprehensive and balanced performance improvement. In the image classification task, when the "fish" classification task is added, the model can quickly learn the characteristics of "fish" without forgetting the existing "cat" and "dog" classification capabilities, thereby achieving balanced optimization of multiple tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] Figure 1 The figure is a flow chart of the method of the present invention.

[0091] Figure 2 It is a schematic diagram of the task similarity calculation module of the present invention.

[0092] Figure 3 It is a schematic diagram of the model extension module of the present invention.

[0093] Figure 4 It is a schematic diagram of the auxiliary network module of the present invention.

[0094] Figure 5 The classification accuracy curves of the present invention and other SOTA baselines on CIFAR100, TinyImageNet, mini-ImageNet, and subImageNet.

[0095] Figure 6This is a curve chart showing the degradation of classification performance between the present invention and other SOTA baselines on CIFAR100, TinyImageNet, mini-ImageNet, and subImageNet.

[0096] Figure 7 Parameter diagram of different incremental learning methods used in the experiment on SubImageNet 10 for this invention.

[0097] Figure 8 This is a schematic diagram of the results of experiments with different thresholds on SubImageNet 10 conducted in the present invention.

[0098] Fig. 9 It is a schematic diagram of the framework of the present invention. DETAILED DESCRIPTION

[0099] The present invention will be further described in detail below in conjunction with the accompanying drawings.

[0100] See also Figure 1 , an image classification method based on incremental learning of a multi-branch network architecture, comprising the following steps:

[0101] The datasets used are CIFAR100, TinyImageNet, mini-ImageNet, and subImageNet. The categories of the datasets are shuffled by random seeds. In order to compare with the classic method, each image classification task includes the same number of categories. For example, CIFAR100-20 means that it is divided into 20 tasks, each task includes 5 categories. During the training process, batch_size=64, the Adam optimizer is used for optimization, the learning rate is set to 0.01, and the learning rate decay adopts the cosine annealing strategy. The training is 80 epochs.

[0102] See also Figure 2 , Step 1.1: For the T-class incremental image classification task {D 1 ,D 2 ,…,D T}, is the t-th incremental task, including N t samples, is the i-th sample of the t-th class incremental task, is the label value of the i-th sample of the corresponding t-th class incremental task. Assume that task T i The category set is C i , for each category c∈C i , maintain a prototype feature It is defined as the mean of the sample features belonging to category c:

[0103]

[0104] Among them, w i Represents sample X i The weight can be defined according to category balance or sample confidence. In CIFAR10 and CIFAR100 datasets, since the data is balanced, w i The value of is 1, and the features here are extracted using the pre-trained model Resnet-18.

[0105] Step 1.2: Nonlinear distance between sample and prototype. The distance between a new sample and a prototype is calculated by a kernel function-based method:

[0106]

[0107] where k(·,·) is a kernel function, such as a Gaussian kernel, defined as:

[0108]

[0109] σ is the bandwidth parameter of the Gaussian kernel, which controls the width of the kernel function and is set to 1 in the experiment.

[0110] In addition, the new sample X l The minimum distance to all category prototypes of the previous task is defined as:

[0111]

[0112] Among them, C j It is task T j A collection of categories, It is task T j The prototype of category c in .

[0113] Step 1.3: Wasserstein distance, task T i and T j The dissimilarity between the distributions of can be measured by the Wasserstein distance, defined as:

[0114]

[0115] Among them, Π(q i ,q j ) represents the joint distribution between task distributions.

[0116] In order to simplify the calculation, the Sinkhorn approximation is used to estimate the Wasserstein distance:

[0117]

[0118] Step 1.4: Define a dynamic mapping function to map the Wasserstein distance to the interval [0,1], the formula is as follows:

[0119] s i,g =1-exp(-α·W ε (q i ,q j ))

[0120] Among them, this is the new task T i The similarity measure between the existing task group g is in the range of [0,1], and α is set to 0.65 in the experiment.

[0121] In the experiment, after the above steps, a set of similarity values ​​between the new task and the old task will be obtained. According to the similarity values, the model learning can be flexibly selected, which enhances the generalization and adaptability of the model.

[0122] See also Figure 3 , Step 2: For N obtained in step 1 t Group similarity values, for these N t Sort them and select the maximum value T max Compared with the threshold δ, if the highest similarity value is greater than the threshold, it means that the current task is very similar to the task we have learned before. At this time, the task is assigned to the network N corresponding to the learned task with the highest similarity. old This approach can minimize the training cost and alleviate catastrophic forgetting. If the highest similarity value is less than the threshold, it means that the current task is not similar to all learned tasks, and the model needs to be expanded. The mathematical formula is as follows:

[0123]

[0124] Among them, G(T cur ) indicates the current task T cur The expansion strategy we propose is to copy the network with the highest similarity to the learned task and use the copied network N new Learning new tasks. This approach can reduce the cost of learning from scratch and shorten training time.

[0125] In the experiment, the threshold δ is 0.6, which determines the expansion speed and performance of the model.

[0126] See also Figure 4 , Step 3.1: Phase 1 (initial training phase)

[0127] Use the complete supervised data (training data X1 and label Y1) for standard classification training. Train a classification model Contains the following components:

[0128] (1) Feature Extractor Responsible for extracting high-dimensional feature representation of samples.

[0129] (2) Classifier Map features to classification labels.

[0130] Optimization objective: Use cross entropy loss L ce Train the model to improve classification performance.

[0131] Incremental phase (phase n)

[0132] Current Task T n The input data is only the images of the new task Does not include old task samples. Use a new feature extractor Extract feature representations for new tasks:

[0133]

[0134] To learn discriminative features of new tasks, use new task labels (From Y n ) for supervised optimization, using a classifier Map features to label space:

[0135]

[0136] Calculate the cross entropy loss L ce :

[0137]

[0138] At the same time, in order to retain the knowledge of the old tasks, knowledge distillation is used to keep the feature representation of the old category and extract the feature representation of the previous stage n-1:

[0139]

[0140] Calculate the distillation loss L kd :

[0141]

[0142] where F kd represents the Euclidean distance.

[0143] Step 3.2: In the NECIL scenario, a prototype vector is stored in the feature space for each category: these prototypes are used to balance the classification performance of new tasks with that of old tasks.

[0144] Oversample the prototypes to match the batch size B to calibrate the classifier:

[0145] p B =U p B(Prototype)

[0146] Among them U p B is the oversampling function.

[0147] Calculate the prototype balance loss L proto :

[0148] L proto =F ce (p B ,y B )

[0149] where y B is the oversampled label set.

[0150] Step 3.3: In order to learn new categories unbiasedly while maintaining the feature representation of old categories, a dynamic structural reorganization strategy is adopted:

[0151] Structural expansion: A residual adapter is inserted into each convolutional block of the fixed feature extractor in the previous stage. The residual adapter only updates the most discriminative part while retaining the old features:

[0152]

[0153] Structural reparameterization: After training is completed, the side branch information is losslessly integrated into the main branch to reduce the number of parameters:

[0154]

[0155] Remove the adapter, keeping the network structure intact for the next phase.

[0156] Step 3.4: To reduce feature confusion during the distillation process, a prototype selection mechanism based on an extensible embedding space is used: the cosine similarity of all new samples to the old prototypes is calculated:

[0157]

[0158] Set the similarity threshold σ and select the update strategy according to the sample similarity: Dissimilar samples: Participate in the update of the residual adapter to learn new features. Similar samples: Participate in the distillation process to retain the discriminative features of the old category. Optimize different loss functions by adding masks: If the similarity is greater than σ, the distillation loss L is kd Add mask; if the similarity is less than σ, the cross entropy loss L ceAdd the mask. The final loss function is defined as:

[0159] L=Mask ce (L ce )+λMask kd (L kd )+γL proto

[0160] where λ and γ are the loss weights.

[0161] In the experiment, σ=0.45, λ=0.7, γ=0.3, the threshold σ determines the selection of the update strategy, and λ and γ affect the training of the model.

[0162] Step 4: For a new test sample x, use the pre-trained model Resnet-18 to extract its features, calculate its similarity value with the old task, decide which model to use for testing, and use the backbone network for testing within a single model to map it to the corresponding category.

[0163] like Fig. 9 As shown, the above method is based on an incremental learning framework of a multi-branch network architecture, which consists of the following core components:

[0164] Task similarity calculation module: used to calculate the similarity between the new graphic classification task and the learned image classification task, based on Wasserstein distance and dynamic mapping function.

[0165] Model expansion and reuse decision module: Based on the similarity calculation results, decide whether to expand the new network model or reuse the existing model.

[0166] Multi-branch network architecture: After multiple model expansions, there are now multiple network branches, and auxiliary networks are introduced inside each network branch. The backbone network is used to learn similar tasks, and the auxiliary network is used to learn dissimilar tasks. The auxiliary network is integrated into the backbone network through dynamic mapping functions and structural reparameterization strategies.

[0167] Test and Evaluation: Use the expanded final model to predict the test samples and calculate the classification accuracy.

[0168] In order to evaluate the performance of the proposed incremental learning method based on pre-trained models, our method (Ours) is compared with several advanced incremental learning methods, namely EWC, MAS, OWM, Adam-NSCL, iCaRL, PackNet, RPSNet, WSN, Genifer. For the datasets CIFAR100, TinyImageNet, mini-ImageNet, and subImageNet, they are divided into 10 tasks, 20 tasks, and 25 tasks, respectively. The experiments report the average precision and forgetting degree BWT, the experimental results are shown in Table 1.

[0169] Table 1 Comparison of the accuracy of different methods on multiple datasets

[0170]

[0171]

[0172] From the results, it can be seen that the method of the present invention (Ours) achieves the highest classification accuracy on all data sets: CIFAR100-20: 81.53%, TinyImageNet20: 60.89%, Mini-ImageNet-20: 69.21%, SubImageNet-10: 50.79%. Compared with the second best method (Genifer), the accuracy of the method of the present invention on all data sets is significantly improved: CIFAR100-20 is improved by 1.43%, TinyImageNet20 is improved by 1.94%, Mini-ImageNet-20 is improved by 1.76%, and SubImageNet-10 is improved by 3.66%.

[0173] Performance degradation reflects the degree to which the model forgets old tasks when learning new tasks. The closer the value is to zero, the less forgotten. Our method (Ours) has the smallest performance degradation on all datasets: CIFAR100-20: -3.11%, TinyImageNet20: -4.58%, Mini-ImageNet-20: -3.67%, SubImageNet-10: -4.88%, and significantly reduces the performance degradation compared to other methods. For example, on SubImageNet-10, the performance degradation of our method is -4.88%, which is a significant improvement over Genifer's -12.01%.

[0174] See also Figure 5 , Figure 6, while maintaining high classification accuracy, the method of the present invention effectively slows down the model's forgetting of old tasks. Although methods such as EWC and MAS perform well on some data sets, they have a large performance degradation on other data sets, indicating that they are not stable enough in a multi-task environment. Although the iCaRL and Genifer methods perform well in accuracy, they have major problems in performance degradation, especially on Mini-ImageNet-20 and SubImageNet-10. The OWM and PathNet methods have average performance in balancing accuracy and performance degradation, and failed to achieve the best performance on all data sets. The Adam-NSCL and WSN methods perform well on some data sets, but their overall performance is not as good as the method of the present invention.

[0175] See also Figure 7 Since the present invention uses model expansion, it involves the issue of parameter quantity. This application compares the total parameter quantity of the network used by different methods. It can be seen that compared with the methods WSN and RPSNet that use the same architecture, the method of this application (Ours) has a smaller total parameter quantity under the same conditions. Although compared with the regularization-based EWC and MAS, the parameter quantity is larger, the performance of the present invention is much higher, which also proves that our method (Ours) is more suitable for real environments and reduces memory consumption.

[0176] See also Figure 8 Since the present invention uses model expansion, it involves the issue of thresholds. This application compares the impact of different thresholds on the model. As the threshold increases, the ACC and the number of parameters increase. This is because as the threshold increases, the number of tasks smaller than the threshold increases, which means that the number of dissimilar tasks increases, the number of model expansions increases, and therefore the number of parameters increases. The performance of a single model will improve, so the ACC also increases. Considering the ACC and the number of parameters, it is more appropriate to select δ=0.6, which has a smaller number of parameters and higher performance than similar methods.

[0177] An image classification device based on incremental learning of a multi-branch network architecture, comprising:

[0178] Memory: used to store a computer program for implementing the incremental learning method based on a multi-branch network architecture;

[0179] Processor: used to implement the incremental learning method based on a multi-branch network architecture when executing the computer program.

[0180] A computer-readable storage medium, comprising:

[0181] The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement an image classification method based on incremental learning of a multi-branch network architecture.

Claims

1. An image classification method based on incremental learning of a multi-branch network architecture, characterized in that: The steps include: Step 1: For each new image classification task T i , calculate the similarity between the task and the old image classification that has been learned, and sort them; Step 2: Compare the highest similarity value with the set similarity threshold. If it is less than the threshold, expand the new network model for learning. Otherwise, use the most similar network model for learning and updating. Step 3: Introduce an auxiliary network into each network model in step 2. The auxiliary network is used to learn dissimilar tasks within the network model. The backbone network of the network model uses knowledge distillation to learn similar tasks. After learning, the auxiliary network is integrated into the backbone network to obtain the expanded final model and control the parameter scale. Step 4: Use the expanded final model to predict the test samples and calculate the final classification accuracy.

2. The image classification method based on incremental learning of a multi-branch network architecture according to claim 1, characterized in that: The specific process of calculating the task similarity in step 1 is as follows: Step 1.1: For the T-class incremental image classification task {D 1 ,D 2 ,…,D T }, is the t-th class incremental image classification task, including N t sample images, is the i-th sample of the t-th class incremental image classification task, is the label value of the i-th sample of the corresponding t-th class incremental image classification task. Set the image classification task T i The category set is C i , for each category c∈C i , maintain a prototype feature It is defined as the mean of the features of image samples belonging to category c: Among them, w i Represents sample X i The weights of are used to define according to class balance or sample confidence; Step 1.2: Calculate the nonlinear distance between the sample image and the prototype feature based on the kernel function method: Where k(·,·) is the kernel function, defined as: In addition, the new sample image X l The minimum distance to all category prototype features of the previous image classification task is defined as: Among them, C j It is task T j A collection of categories, It is task T j The prototype features of category c; Step 1.3: Calculate Wasserstein distance, image classification task T i and T j The dissimilarity between the distributions of is measured by the Wasserstein distance, defined as: Among them, Π(q i ,q j ) represents the joint distribution between task distributions; The Sinkhorn approximation is used to estimate the Wasserstein distance: Step 1.4: Design a dynamic mapping function to map the Wasserstein distance to the interval [0,1], the formula is as follows: yes i,g =1-exp(-α·W ε (what i ,q j )) Among them, the new image classification task T i The similarity measure between the old image classification task group g is in the range of [0,1], and α is a scaling factor adjusted through experiments.

3. The image classification method based on incremental learning of a multi-branch network architecture according to claim 2, characterized in that: The specific process of step 2 is as follows: After the similarity calculation in step 1, we get N t A group of similarity values, mapping the similarity values ​​to the range of [0,1]. t Sort them and select the maximum value T max Compared with the threshold δ, δ is between [0-1]. If the highest similarity value is greater than the threshold, it means that the current task is very similar to the old task. At this time, the task is assigned to the network N corresponding to the learned task with the highest similarity. old To study; If the highest similarity value is less than the threshold, it means that the current image classification task is not similar to all the learned image classification tasks. At this time, the model needs to be expanded. The formula is as follows: Among them, G(T cur ) represents the current image classification task T cur The network assigned to it, the expansion strategy is to copy the network with the highest similarity to the learned task and use the copied network N new Learning new image classification tasks.

4. The image classification method based on incremental learning of a multi-branch network architecture according to claim 3, characterized in that: The step 3 is specifically as follows: Step 3.1: Initial training phase: Use the complete supervised data (training data X1 and label Y1) for standard classification training; Training a standard classification model A standard classification model Contains the following components: (1) Feature Extractor Responsible for extracting high-dimensional feature representation of samples; (2) Classifier Map features to classification labels; Use a feature extractor to extract features from the task, and then use a classifier to map these extracted features to specific labels; Use cross entropy loss L ce Train the model and optimize the target; Incremental phase: Current image classification task T n The input data is only the images of the new task Excluding old image classification task samples, using a new feature extractor Extract feature representations for new tasks: Leveraging new image classification task labels (From Y n ) for supervised optimization, using a classifier Map features to label space: Calculate the cross entropy loss L ce : At the same time, knowledge distillation is used to maintain the feature representation of the old category and extract the feature representation of the previous stage n-1: Calculate the distillation loss L kd : Among them, F kd represents the Euclidean distance; Step 3.2: Further optimize the feature representation by using the prototype vector and the balance loss to balance the classification performance of the new task with that of the old task. In the NECIL scenario, a prototype vector is stored in the feature space for each category, and these prototypes are used to balance the classification performance of the new image classification task with that of the old image classification task. Step 3.3: Use a dynamic structural reorganization strategy to ensure that the network in steps 3.1 and 3.2 does not destroy the feature representation of the old image classification task when learning the new image classification task, and keeps the feature representation of the old category while learning the new category unbiasedly; Step 3.4: Comprehensively optimize steps 3.1 to 3.3, using a prototype selection mechanism and mask strategy based on a scalable embedding space to reduce feature confusion and ensure that the feature representations of the new image classification task and the old image classification task can coexist.

5. The image classification method based on incremental learning of a multi-branch network architecture according to claim 4, characterized in that: The step 3.2 is specifically as follows: Calibrate the classifier by oversampling the prototypes to match the batch size B: p B =U p B(Prototype) Among them U p B is the oversampling function; Calculate the prototype balance loss L proto : L proto =F ce (p B ,y B ) where y B is the oversampled label set.

6. The image classification method based on incremental learning of a multi-branch network architecture according to claim 4, characterized in that: The step 3.3 is specifically as follows: Structural expansion: A residual adapter is inserted into each convolutional block of the feature extractor. The residual adapter only updates the most discriminative part while retaining the old features: Structural reparameterization: After training is completed, the side branch information is losslessly integrated into the main branch to reduce the number of parameters: Remove the adapter to keep the network structure unchanged for the next stage and control the parameter scale; The step 3.4 is specifically as follows: Calculate the cosine similarity of all new samples to the old prototypes: Set the similarity threshold σ and select the update strategy based on the sample similarity: Dissimilar samples: Participate in the update of the residual adapter to learn new features; resemblance Samples: Participate in the distillation process and retain the discriminative features of the old categories; Optimize different loss functions by adding masks: If the similarity is greater than σ, the distillation loss L kd Add mask; if the similarity is less than σ, add cross entropy loss L ce Adding the mask, the final loss function is defined as: L=Mask ce (L ce )+λMask kd (L kd )+γL proto where λ and γ are the loss weights.

7. The image classification method based on incremental learning of a multi-branch network architecture according to claim 4, characterized in that: In step 4: for a new test sample x, use the pre-trained model Resnet-18 to extract its features, then calculate its similarity value with the old image classification task, decide which model to use for testing, and use the backbone network for testing within a single model to map it to the corresponding category.

8. An image classification system based on incremental learning of a multi-branch network architecture, characterized in that: It includes task similarity calculation module, model extension module and auxiliary network module; The task similarity calculation module uses Wasserstein distance to measure the image classification task T i and T j The similarity between the distributions of , and the calculated image classification task similarity is used as the basis for the model extension module to expand the model; The model extension module adopts the maximum similarity T max Compare with the threshold δ. If it is greater than the threshold, no expansion is required. Otherwise, copy and expand the new model for learning. Use the image classification task similarity calculation module to calculate the similarity of the task to decide whether to expand the model. The auxiliary network module realizes structural expansion and reparameterization through the residual adapter, takes into account both new feature learning and old feature retention, and performs structural expansion and reparameterization on the expanded model in the model expansion module, thereby controlling the parameters of the expanded network model while mitigating catastrophic forgetting.

9. An image classification device based on incremental learning of a multi-branch network architecture, characterized in that: include: Memory: used to store a computer program for implementing the image classification method based on incremental learning of a multi-branch network architecture as described in any one of claims 1 to 7; Processor: used to implement the image classification method based on incremental learning of a multi-branch network architecture when executing the computer program.

10. A computer-readable storage medium, characterized in that: include: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement the image classification method based on incremental learning of a multi-branch network architecture as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Image classification method of class incremental learning based on self-holding representation extension

    CN114677547A

  • Federal incremental learning method based on feature prototype

    CN114861936A

  • Training method for improving new and old category distinction degree of existing category incremental learning

    CN116089883A

  • Incremental hyperspectral image classification method and device based on dynamic learning

    CN117392443A

  • Method for training a classification model using class prototypes of previous classes, and corresponding system

    EP4432170A1

Cited By

  • Incremental learning supported computer vision target identification method and system

    CN120635553A

  • Image recognition method based on few-sample incremental learning and edge updating system

    CN121280807A

  • Marine litter fine granularity identification method and system based on unmanned aerial vehicle image

    CN121963002A