Class incremental learning method, device and equipment based on comparative learning
Through a class incremental learning method based on comparison learning, integrating the feature information of the current and previous tasks, the catastrophic forgetting problem in continuous learning is solved, the storage and computing costs are reduced, the learning efficiency and recognition accuracy of the model are improved, and the adaptability of the model in new tasks is enhanced.
Patent Information
- Application Number
- CN202510410900.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
The existing continuous learning methods have catastrophic forgetting problems when facing dynamically changing data and tasks, and cannot effectively maintain the recognition ability of old tasks. Storing old data increases storage costs and privacy security risks, and the increase in the complexity of the model structure leads to an increase in computing costs.
A class incremental learning method based on contrast learning is adopted, and the pre-trained feature extractor is frozen, and the comparison learning training is used to process data playback and spatial pyramid processing is used to fuse the feature information of the current and previous tasks, and the initial feature classifier is trained to avoid directly storing a large amount of old data, reduce storage and privacy risks, and generate new samples through data fusion and pyramid processing, improving the robustness of the model.
It improves the learning efficiency and recognition accuracy of the model, reduces the computing cost and storage privacy risks, ensures that the model structure does not increase complexity, solves the problem of catastrophic forgetting, and enhances the model's adaptability to new tasks.
Smart Images

Figure CN120339769A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of continuous learning and self-supervised learning, and particularly relates to a class-incremental learning method, device and equipment based on contrast learning. Background Art
[0002] In the field of deep learning, continuous learning is a crucial research direction, aiming to enable a model to learn new knowledge from continuously incoming data while maintaining the memory of old knowledge. However, in the real world, data is often dynamically changing, and its distribution and categories may change over time. Traditional deep learning models are usually trained on a fixed dataset. When faced with new data or tasks, these models often suffer from catastrophic forgetting, that is, during the process of adapting to new tasks, the model will forget the old knowledge learned before, resulting in a significant decline in performance on old tasks. This forgetting problem seriously hinders the continuous learning and adaptation ability of deep learning models in practical applications.
[0003] To address the catastrophic forgetting problem in continuous learning, researchers have explored various methods. Among them, regularization-based methods regularize the weights by estimating and preventing changes in key network weights to limit the update of model parameters, thereby maintaining the memory of old knowledge. In addition, relay-based methods attempt to store old data for relay. These methods directly save the original data and use it together with the current data when learning new tasks to alleviate the forgetting problem. Another approach is architecture-based methods, which tend to continuously expand the network structure for different tasks of incremental learning to accommodate new knowledge and features. As a self-supervised learning method, supervised contrast learning also learns general features by having the model learn the similarity between data points, providing better representations for downstream tasks.
[0004] Although existing continuous learning methods have alleviated the catastrophic forgetting problem to a certain extent, they still have some significant drawbacks. First, many methods lack robustness in feature representation, resulting in the model being unable to effectively maintain the recognition ability of old tasks when migrating to new tasks. Second, relay-based methods need to store a large amount of old data, which not only increases the storage cost but may also cause data privacy and security issues. In addition, as the number of incremental learning tasks continues to increase in architecture-based methods, the model structure becomes more and more complex, which not only increases the computational cost but may also reduce the generalization ability of the model. Therefore, existing continuous learning methods still face many challenges when dealing with dynamically changing data and tasks and require further research and improvement. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides a class-incremental learning method, device and equipment based on contrast learning.
[0006] The technical problem to be solved by the present invention is achieved through the following technical solutions:
[0007] In a first aspect, the present invention provides a class incremental learning method based on contrastive learning, including:
[0008] Obtain the current task sample data; the current task sample data is the sample image data that needs to be class incremented currently;
[0009] Input the current task sample data into the frozen current pre-trained feature extractor to obtain the current sample features;
[0010] Adopt data replay processing to perform data fusion processing on the current sample features and the previous sample features to obtain the current fused features; the previous sample features are the feature information of all tasks before the current task; the current pre-trained feature extractor is obtained through data augmentation processing based on spatial pyramid processing and contrastive learning training;
[0011] Use the current fused features to train the initial feature classifier to obtain the current classification model after class increment; the current classification model is used to classify the image data corresponding to the current task or the image data corresponding to the previous task.
[0012] Optionally, adopting data replay processing to perform data fusion processing on the current sample features and the previous sample features to obtain the current fused features includes:
[0013] Obtain the previous sample features;
[0014] Extract the query features corresponding to the previous task sample data from the previous sample features to form a query feature set;
[0015] Extract the key features corresponding to the previous task sample data from the previous sample features to form a key feature set;
[0016] Based on the query feature set and the key feature set, generate a query feature Gaussian distribution and a key feature Gaussian distribution;
[0017] Randomly sample from the query feature Gaussian distribution and the key feature Gaussian distribution, and perform data fusion processing on the results of the random sampling and the current sample features to obtain the current fused features.
[0018] Optionally, generating a query feature Gaussian distribution and a key feature Gaussian distribution based on the query feature set and the key feature set includes:
[0019] Perform mean processing and standard deviation processing on the query feature set respectively to obtain the mean query features and the standard deviation query features;
[0020] Perform mean processing and standard deviation processing on the key feature set respectively, and correspondingly obtain the mean key feature and the standard deviation key feature;
[0021] Use the mean query feature and the standard deviation query feature to construct a query feature Gaussian distribution;
[0022] Use the mean key feature and the standard deviation key feature to construct a key feature Gaussian distribution.
[0023] Optionally, use the current fusion feature to train the initial feature classifier to obtain the current classification model after class increment, including:
[0024] Use the current fusion feature to train the initial feature classifier on the basis of the frozen current pre-trained feature extractor by using the cross-entropy loss function;
[0025] When the value of the cross-entropy loss function meets the preset cross-entropy loss threshold, the corresponding initial feature classifier is used as the current feature classifier;
[0026] The current pre-trained feature extractor and the current feature classifier together constitute the current classification model.
[0027] Optionally, the training process of the current pre-trained feature extractor includes:
[0028] Obtain the previous task sample data;
[0029] Perform random data fusion processing on the previous task sample data to obtain the previous task enhanced sample data;
[0030] Use spatial pyramid processing to perform data augmentation on the previous task enhanced sample data to obtain the previous task augmented sample data;
[0031] Input the previous task augmented sample data, the previous task sample data, and the previous task enhanced sample data into the initial feature extractor for training to obtain the current pre-trained feature extractor.
[0032] Optionally, input the previous task augmented sample data, the previous task sample data, and the previous task enhanced sample data into the initial feature extractor for training to obtain the current pre-trained feature extractor, including:
[0033] Input the previous task augmented sample data into the query encoder in the initial feature extractor;
[0034] Input the previous task sample data and the previous task enhanced sample data into the key encoder in the initial feature extractor;
[0035] Use contrastive learning and momentum update methods to train the initial feature extractor;
[0036] Use the initial feature extractor that meets the preset stop condition as the current pre-trained feature extractor.
[0037] Optionally, using the initial feature extractor that meets the preset stop condition as the current pre-trained feature extractor includes:
[0038] Use the initial feature extractor corresponding to when the number of iterations is greater than the iteration threshold or the value of the contrast loss is greater than the contrast loss threshold as the current pre-trained feature extractor.
[0039] Optionally, after training the initial feature classifier with the current fused feature to obtain the current classification model after class increment, it further includes:
[0040] Obtain any image to be recognized;
[0041] Use the current classification model to classify any image to be recognized to obtain a recognition result; any image to be recognized belongs to the category of the current task or the category of the task before the current task.
[0042] In a second aspect, the present invention provides a class-incremental learning device based on contrast learning. The class-incremental learning device based on contrast learning includes: an acquisition unit, an input unit, a fusion unit, and a training unit;
[0043] The acquisition unit is used to: acquire current task sample data; the current task sample data is sample image data that currently needs class increment;
[0044] The input unit is used to: input the current task sample data into the frozen current pre-trained feature extractor to obtain the current sample feature;
[0045] The fusion unit is used to: perform data fusion processing on the current sample feature and the previous sample feature by using data replay processing to obtain the current fused feature; the previous sample feature is the feature information of all tasks before the current task; the current pre-trained feature extractor is obtained based on data augmentation processing of spatial pyramid processing and contrast learning training;
[0046] The training unit is used to: train the initial feature classifier with the current fused feature to obtain the current classification model after class increment; the current classification model is used to classify the image data corresponding to the current task or the image data corresponding to the previous task.
[0047] In a third aspect, the present invention provides a class-incremental learning device based on contrast learning, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the class-incremental learning device based on contrast learning runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the class-incremental learning method based on contrast learning as described in the first aspect above.
[0048] The present invention provides a class incremental learning method, device and equipment based on contrastive learning. Among them, a class incremental learning method based on contrastive learning includes: obtaining current task sample data; inputting the current task sample data into a frozen current pre-trained feature extractor to obtain the current sample feature; using data replay processing to perform data fusion processing on the current sample feature and the previous sample feature to obtain the current fusion feature; the previous sample feature is the feature information of all tasks before the current task; the current pre-trained feature extractor is obtained by data augmentation processing based on spatial pyramid processing and contrastive learning training; using the current fusion feature to train the initial feature classifier to obtain the current classification model after class increment. In the present invention, firstly, a pre-trained feature extractor is used to extract features from the current task sample data, avoiding the need to directly store a large amount of old data, thereby significantly reducing the storage cost and the risk of data privacy leakage. At the same time, the feature extractor is obtained by data augmentation processing based on spatial pyramid and contrastive learning training, thereby ensuring that representative features can be efficiently extracted from the input data. Then, the current sample feature is fused with the feature information of the previous task through data replay technology. This method not only solves the problem of catastrophic forgetting, but also enables the model to learn new tasks without increasing complexity. Secondly, using the fused features to fine-tune or train the initial feature classifier instead of constantly adding new sub-networks or modules ensures that the model structure does not become too complex as the incremental learning tasks continue to increase. Ultimately, this approach improves the learning efficiency and recognition accuracy of the model while reducing the increased computational cost caused by the increased model complexity and the privacy and security risks brought by storing the original data.
[0049] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 A flow chart of a quasi-incremental learning method based on contrastive learning provided by an embodiment of the present invention;
[0051] Figure 2 An execution block diagram of a class incremental learning method based on contrastive learning is exemplarily shown;
[0052] Figure 3 Two types of decision boundaries are shown exemplarily;
[0053] Figure 4 A comparison result diagram of a quasi-incremental learning method based on contrastive learning and an existing incremental learning method is shown exemplarily;
[0054] Figure 5Schematic diagram of a class incremental learning device based on contrast learning provided by an embodiment of the present invention;
[0055] Figure 6 Schematic diagram of a class incremental learning device based on contrast learning provided by an embodiment of the present invention. Detailed implementation manners
[0056] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0057] In order to improve the learning efficiency and recognition accuracy of the current classification model, and at the same time reduce the computational cost of the current classification model, an embodiment of the present invention provides a class incremental learning method based on contrast learning. Figure 1 Flowchart of a class incremental learning method based on contrast learning provided by an embodiment of the present invention. As Figure 1 shown, it includes:
[0058] S101. Obtain the current task sample data.
[0059] Among them, the current task sample data is the sample image data that needs class increment currently. Exemplarily, if the current classification model can only recognize bird and cat images currently, and the task goal is to make the current classification model recognize various flower and plant images, then the current task sample data can be various labeled flower and plant image samples.
[0060] S102. Input the current task sample data into the frozen current pre-trained feature extractor to obtain the current sample features.
[0061] It should be noted that the current pre-trained feature extractor in the embodiment of the present invention can specifically adopt the moco network model in contrast learning.
[0062] In this embodiment, the current pre-trained feature extraction is frozen, and only the classifier is trained. When new data pairs arrive, the frozen current pre-trained feature extractor can generate corresponding features. In addition, a single classifier can be trained for the corresponding stage. After obtaining the classifier in the incremental stage, the single classifier can be spliced behind the overall classifier to form the pre-trained current feature classifier.
[0063] S103. Perform data fusion processing on the current sample features and the previous sample features by using data replay processing to obtain the current fusion features.
[0064] Among them, the previous sample features are the feature information of all tasks before the current task; the current pre-trained feature extractor is obtained based on data augmentation processing of spatial pyramid processing and contrast learning training.
[0065] Optionally, S103 may specifically include:
[0066] Obtain previous sample features;
[0067] Extract the query features corresponding to the previous task sample data from the previous sample features to form a query feature set;
[0068] Extract the key features corresponding to the previous task sample data from the previous sample features to form a key feature set;
[0069] Based on the query feature set and the key feature set, generate a query feature Gaussian distribution and a key feature Gaussian distribution;
[0070] Randomly sample from the query feature Gaussian distribution and the key feature Gaussian distribution, and perform data fusion processing on the results of the random sampling and the current sample features to obtain the current fusion feature.
[0071] Optionally, generating a query feature Gaussian distribution and a key feature Gaussian distribution based on the query feature set and the key feature set includes:
[0072] Perform mean processing and standard deviation processing on the query feature set respectively to obtain a mean query feature and a standard deviation query feature;
[0073] Perform mean processing and standard deviation processing on the key feature set respectively to obtain a mean key feature and a standard deviation key feature;
[0074] Use the mean query feature and the standard deviation query feature to construct a query feature Gaussian distribution;
[0075] Use the mean key feature and the standard deviation key feature to construct a key feature Gaussian distribution.
[0076] In the embodiments of the present invention, data replay processing is used to perform data fusion processing on the current sample features and the previous sample features to obtain the current fusion feature. That is, store some representative samples on the old tasks for data reuse in the incremental stage to significantly reduce the catastrophic forgetting problem in incremental learning.
[0077] To minimize the complexity brought by storing data, such as the choice of memory size and the choice of which sample features to store, in the actual storage process, the embodiments of the present invention do not directly store the original data, but store the data distribution M corresponding to different categories t . M t Can be expressed as Where μ c And Are the mean feature and the variance feature respectively. Among them, it should be noted that the mean feature here includes: the mean query feature and the mean key feature, and the variance feature includes: the variance query feature and the variance key feature. ∣Ct ∣ represents the number of categories in task t, and C represents the Cth category. t The categories seen at each stage can be represented and simulated. When a classifier for a new task needs to be trained, random sampling can be done from the stored Gaussian distribution, and the old category features obtained from the sampling can be mixed with the new category features to train the entire classifier. In addition, distillation can be performed to prevent the classification head from forgetting.
[0078] S104: Use the current fusion features to train the initial feature classifier to obtain a current classification model after class increment.
[0079] Among them, the current classification model is used to classify the image data corresponding to the current task or the image data corresponding to the previous task.
[0080] The embodiment of the present invention provides a quasi-incremental learning method based on contrastive learning. First, a pre-trained feature extractor is used to extract features from the sample data of the current task, avoiding the need to directly store a large amount of old data, thereby significantly reducing the storage cost and the risk of data privacy leakage. At the same time, the feature extractor is obtained based on the data augmentation processing and contrastive learning training of the spatial pyramid, thereby ensuring that representative features can be efficiently extracted from the input data. Then, the current sample features are fused with the feature information of the previous task through data replay technology. This method not only solves the problem of catastrophic forgetting, but also enables the model to learn new tasks without increasing complexity. Secondly, the fused features are used to fine-tune or train the initial feature classifier instead of continuously adding new subnetworks or modules, ensuring that the model structure will not become too complex as the incremental learning tasks continue to increase. Ultimately, this method improves the learning efficiency and recognition accuracy of the model, while reducing the increase in computational costs caused by the increase in model complexity, as well as the privacy and security risks brought by storing the original data.
[0081] Optionally, S104 may specifically include:
[0082] Using the current fused features, the initial feature classifier is trained using the cross entropy loss function based on the frozen current pre-trained feature extractor;
[0083] When the value of the cross entropy loss function meets the preset cross entropy loss threshold, the corresponding initial feature classifier is used as the current feature classifier;
[0084] The current pre-trained feature extractor and the current feature classifier together constitute the current classification model.
[0085] Optionally, the training process of the current pre-trained feature extractor includes:
[0086] Get sample data from previous tasks;
[0087] Perform random data fusion processing on the previous task sample data to obtain the enhanced sample data of the previous task;
[0088] Use spatial pyramid processing to perform data augmentation on the enhanced sample data of the previous task to obtain the augmented sample data of the previous task;
[0089] Input the augmented sample data of the previous task, the sample data of the previous task, and the enhanced sample data of the previous task into the initial feature extractor for training to obtain the current pre-trained feature extractor.
[0090] Since the method of the present invention does not require the use of class augmentation in the inference stage (application stage), the classifier is accordingly extended before the basic training stage, that is, random data fusion processing is performed on the previous task sample data to obtain the enhanced sample data of the previous task to expand the class range, and the extra nodes are deleted after training. Thus, the decision boundary expected by the final current classification model has stronger robustness to perturbations.
[0091] Specifically, in the random data fusion processing of the present invention, for any two different previous task sample data, they can be directly added together in different proportions to generate a new image x new , x new =γx1+(1 - γ)x2. The newly generated image x new is considered to belong to a completely new class and is assigned a unique label y new , y new =gen(y1 + y2). y1 represents the label corresponding to the first sample x1 in the previous task sample data, y2 represents the label corresponding to the second sample x2 in the previous task sample data, and γ represents the fusion coefficient.
[0092] In addition, the present invention applies the method of spatial pyramid contrast to the data augmentation processing operation of images. Specifically, the enhanced sample data of the previous task is cut into different sizes on the spatial scale, and then these cut images x q , that is, the augmented sample data of the previous task, are regarded as nodes on the pyramid: where m represents the number of nodes in the pyramid, is the bottom of the pyramid, that is, the original image (any enhanced sample data of the previous task).
[0093] It can be understood that by combining random data fusion and spatial pyramid processing, more diverse new samples with multi-scale features can be generated. These new samples, as additional training data, can further enhance the effect of data augmentation. Additionally, since the new samples combine the features of different samples and contain multi-scale information, the model can learn more complex feature representations and decision boundaries during the training process. This helps improve the robustness of the model when facing complex scenarios such as noise and occlusion. Moreover, through data fusion and spatial pyramid processing, the existing data resources can be effectively utilized to expand the dataset and improve the model performance without increasing the additional data collection cost.
[0094] Therefore, performing random data fusion processing on the sample data of previous tasks and, on this basis, using spatial pyramid processing for data augmentation can significantly increase data diversity, expand data categories, improve the generalization ability and robustness of the model, and optimize resource utilization.
[0095] Optionally, inputting the augmented sample data of previous tasks, the sample data of previous tasks, and the enhanced sample data of previous tasks into the initial feature extractor for training to obtain the current pre-trained feature extractor includes:
[0096] Inputting the augmented sample data of previous tasks into the query encoder in the initial feature extractor;
[0097] Inputting the sample data of previous tasks and the enhanced sample data of previous tasks into the key encoder in the initial feature extractor;
[0098] Training the initial feature extractor using contrastive learning and momentum update methods;
[0099] Taking the initial feature extractor that meets the preset stop condition as the current pre-trained feature extractor.
[0100] Optionally, taking the initial feature extractor that meets the preset stop condition as the current pre-trained feature extractor includes:
[0101] Taking the initial feature extractor corresponding to when the number of iterations is greater than the iteration threshold or the value of the contrastive loss is greater than the contrastive loss threshold as the current pre-trained feature extractor.
[0102] In addition, the contrastive loss in the embodiments of the present invention can be expressed as:
[0103]
[0104]
[0105] where represents the contrastive loss at the bottom of the spatial pyramid, Denote the contrastive loss at the top of the spatial pyramid, λ1 denotes the adjustment weight, P(x) denotes the set of positive samples, and k ′ denotes the feature of a sample in A(x), and k + denotes the feature of a positive sample in P(x), A(x) denotes the concatenation of the key embedding k of sample x and the feature queue Q, T denotes the transpose of the matrix, τ is the temperature parameter, and q i denotes the feature obtained by passing the image of the i-th layer node in the spatial pyramid except the bottom layer through the query network, and q 0 denotes the feature obtained by passing the image of the bottom layer node in the spatial pyramid through the query network, and m denotes the number of layers in the spatial pyramid except the bottom layer.
[0106] Optionally, after S104, it further includes:
[0107] Obtain any image to be recognized;
[0108] Use the current classification model to classify any image to be recognized to obtain a recognition result; any image to be recognized belongs to the category of the current task or the category of the task before the current task.
[0109] To overall illustrate the execution process of the class incremental learning method based on contrastive learning proposed by the present invention, Figure 2 exemplarily shows the execution block diagram of the class incremental learning method based on contrastive learning. As Figure 2 shown, taking the training task of cat and dog category recognition as an example, in the basic task stage (task 0), the initial feature extractor is trained with cat images, cat and dog fusion images, and cat images processed by the spatial pyramid, based on contrastive learning (queue removal and contrast of positive and negative sample pairs during contrastive learning) and momentum update, so as to obtain a pre-trained feature extractor. In the incremental learning stage, that is, the learning stage for tasks 1 - task t, first freeze the pre-trained feature extraction, and train the corresponding classifiers for tasks 1 - task t respectively. Finally, splice the pre-trained classifiers of tasks 1 - task t, and combine with the pre-trained feature extractor to finally obtain the current classification model after class increment. In addition, to optimize the model effect, in some embodiments, data replay processing can also be used to perform data fusion processing on the sample features in the task 0 stage and the current sample features (such as the sample features in the task 1 stage), and use the fused features after fusion for classifier training. Since the data replay processing process has been described in the above embodiments, it will not be elaborated in this embodiment.
[0110] Figure 3Two decision boundaries are shown as examples. The decision boundary on the left is the decision boundary obtained by learning only using cross entropy loss. The decision boundary may be close to the data point, and the classification result is easily affected when facing small disturbances, that is, the model stability is poor. The decision boundary on the right has a wider range of acceptance for disturbances, and can still maintain good classification performance when disturbed, which is more in line with the needs of continuous learning. It is the decision boundary after category enhancement expected by the present invention. Figure 4 The comparison result diagram of the quasi-incremental learning method based on contrastive learning and the existing incremental learning method is shown as an example. Specifically, Figure 4 The effects of different data augmentation methods on model performance in five incremental task stages on the CIFAR100 dataset are shown. The models represented by each curve are as follows: Baseline: This is a basic model that does not use any specific data augmentation technology, which is used as a benchmark for comparison. When studying the effects of different data augmentation methods, Baseline provides a reference standard to highlight the performance changes of other models that use specific augmentation technologies. +Aug: This refers to the model that uses the "Aug" data augmentation method. The "Aug" method is to copy the samples in equal amounts. From the experimental results, this simple data copying method is not effective in improving the classification performance of the model. It cannot help the model learn a more robust decision boundary, nor does it help the model migrate to downstream tasks when the number of tasks gradually increases. +Mixup: This represents a model that uses the Mixup data augmentation technology. Mixup mixes samples by interpolating between two images in proportion, which improves the model performance to a certain extent, but as the number of tasks gradually increases, this improvement effect is relatively limited. +CutMix: This represents a model that uses the CutMix data augmentation strategy. CutMix mixes images by cutting and splicing specific areas of the image, which improves the model performance to a certain extent, but the improvement is also small when the number of tasks increases. +ClassAug: It is a model that uses the incremental learning method of the present invention. This model generates a pyramid structure by performing spatial transformations of images at different scales, and then randomly fuses the data to perform category enhancement to obtain new samples. It can effectively improve the performance of the model in the initial task stage, and this improvement remains effective as the number of tasks gradually increases.
[0111] In summary, the class incremental learning method based on contrastive learning provided by the present invention adopts a spatial pyramid contrast method to capture fine-grained information on images in the contrastive learning branch, and applies class enhancement to images in the supervised classification branch to improve the learning of decision boundaries. The current classification model learns rich knowledge in the basic training stage and has a stronger ability to migrate to downstream tasks. In addition, data replay processing not only overcomes the catastrophic forgetting problem, but also enables the current classification model to learn new tasks without increasing complexity.
[0112] The method provided by the embodiments of the present invention can be applied to an electronic device. Specifically, the electronic device can be: a desktop computer, a portable computer, a smart mobile terminal, a server, etc., which are not limited in the embodiments of the present invention.
[0113] Based on the same inventive concept, the embodiments of the present invention also provide a class incremental learning device based on contrastive learning. Figure 5 It is a schematic structural diagram of a class incremental learning device based on contrastive learning provided by the embodiments of the present invention. As Figure 5 shown, it includes: an acquisition unit 601, an input unit 602, a fusion unit 603, and a training unit 604;
[0114] The acquisition unit 601 is used to: acquire the current task sample data; the current task sample data is the sample image data that needs to be class incremented currently;
[0115] The input unit 602 is used to: input the current task sample data into the frozen current pre-trained feature extractor to obtain the current sample features;
[0116] The fusion unit 603 is used to: perform data fusion processing on the current sample features and the previous sample features by using data replay processing to obtain the current fusion features; the previous sample features are the feature information of all tasks before the current task; the current pre-trained feature extractor is obtained based on data augmentation processing of spatial pyramid processing and contrastive learning training;
[0117] The training unit 604 is used to: train the initial feature classifier by using the current fusion features to obtain the current classification model after class increment; the current classification model is used to classify the image data corresponding to the current task or the image data corresponding to the previous task.
[0118] Figure 6 It is a schematic structural diagram of a class incremental learning device based on contrastive learning provided by the embodiments of the present invention, including: a processor 710, a storage medium 720, and a bus 730. The storage medium 720 stores machine-readable instructions executable by the processor 710. When the class incremental learning device based on contrastive learning runs, the processor 710 communicates with the storage medium 720 through the bus 730, and the processor 710 executes the machine-readable instructions to perform the steps of the above method embodiments. The specific implementation manners and technical effects are similar and will not be elaborated here.
[0119] The storage medium may include a random access memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-Volatile Memory, NVM), such as at least one disk memory. Optionally, the storage medium may also be at least one storage device located far from the aforementioned processor.
[0120] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0121] It should be noted that the terms "first", "second", etc. are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention.
[0122] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0123] Although the present invention has been described in connection with various embodiments herein, however, in the process of implementing the claimed invention, those skilled in the art can understand and implement other variations of the above-described disclosed embodiments by viewing the drawings and the disclosure. In the description of the present invention, the term "including" does not exclude other components or steps, the term "a" or "one" does not exclude a plurality of cases, and the meaning of "a plurality" is two or more, unless otherwise specifically defined. In addition, certain measures are described in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0124] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A class incremental learning method based on contrastive learning, characterized in that include: Acquire current task sample data; the current task sample data is sample image data that currently requires class increment; Inputting the current task sample data into the frozen current pre-trained feature extractor to obtain current sample features; The current sample feature is fused with the previous sample feature by using data replay processing to obtain the current fused feature; The previous sample features are feature information of all tasks before the current task; the current pre-trained feature extractor is obtained by data augmentation processing based on spatial pyramid processing and contrastive learning training; The initial feature classifier is trained using the current fusion features to obtain a current classification model after class increment; the current classification model is used to classify the image data corresponding to the current task or the image data corresponding to the previous task.
2. The method for class incremental learning based on contrastive learning according to claim 1, wherein The data replay process is used to perform data fusion processing on the current sample feature and the previous sample feature to obtain the current fusion feature, including: Obtaining the previous sample features; Extracting query features corresponding to previous task sample data from the previous sample features to form a query feature set; Extract key features corresponding to previous task sample data from the previous sample features to form a key feature set; Based on the query feature set and the key feature set, generate a query feature Gaussian distribution and a key feature Gaussian distribution; Random sampling is performed from the query feature Gaussian distribution and the key feature Gaussian distribution, and data fusion processing is performed on the random sampling result and the current sample feature to obtain the current fusion feature.
3. The class incremental learning method based on contrastive learning according to claim 2, wherein The generating, based on the query feature set and the key feature set, a query feature Gaussian distribution and a key feature Gaussian distribution comprises: Performing mean processing and standard deviation processing on the query feature set respectively, and obtaining mean query features and standard deviation query features accordingly; Performing mean processing and standard deviation processing on the key feature set respectively, and obtaining a mean key feature and a standard deviation key feature respectively; The query feature Gaussian distribution is constructed by using the mean query feature and the standard deviation query feature; The key feature Gaussian distribution is constructed by using the mean key feature and the standard deviation key feature.
4. The class incremental learning method based on contrastive learning according to claim 1, characterized in that The method of training the initial feature classifier using the current fusion feature to obtain the current classification model after the class increment includes: Using the current fused features, based on the frozen current pre-trained feature extractor, the initial feature classifier is trained using a cross entropy loss function; When the value of the cross entropy loss function meets a preset cross entropy loss threshold, the corresponding initial feature classifier is used as the current feature classifier; The current pre-trained feature extractor and the current feature classifier together constitute the current classification model.
5. The class incremental learning method based on contrastive learning according to claim 1, wherein The training process of the current pre-trained feature extractor includes: Get sample data from previous tasks; Performing random data fusion processing on the previous task sample data to obtain previous task enhanced sample data; Performing data augmentation processing on the previous task enhanced sample data by using the spatial pyramid processing to obtain the previous task augmented sample data; Input the augmented sample data of the previous task, the sample data of the previous task, and the enhanced sample data of the previous task into the initial feature extractor for training to obtain the current pre-trained feature extractor.
6. The class incremental learning method based on contrastive learning according to claim 5, wherein The step of inputting the augmented sample data of the previous task, the sample data of the previous task, and the enhanced sample data of the previous task into the initial feature extractor for training to obtain the current pre-trained feature extractor includes: Input the augmented sample data of the previous task into the query encoder in the initial feature extractor; Input the sample data of the previous task and the enhanced sample data of the previous task into the key encoder in the initial feature extractor; Train the initial feature extractor using contrastive learning and momentum update; Use the initial feature extractor that meets the preset stop condition as the current pre-trained feature extractor.
7. The method for class-incremental learning based on contrastive learning according to claim 6, wherein The step of using the initial feature extractor that meets the preset stop condition as the current pre-trained feature extractor includes: Use the initial feature extractor corresponding to when the number of iterations is greater than the iteration threshold or the value of the contrastive loss is greater than the contrastive loss threshold as the current pre-trained feature extractor.
8. The method for class-incremental learning based on contrastive learning according to claim 1, characterized in that After training the initial feature classifier using the current fused feature to obtain the current classification model after class increment, it further includes: Obtain any image to be recognized; Use the current classification model to classify the any image to be recognized to obtain a recognition result; the any image to be recognized belongs to the category of the current task or the category of the task before the current task.
9. An apparatus for class-incremental learning based on contrastive learning, characterized in that, The class-increment learning device based on contrastive learning includes: an acquisition unit, an input unit, a fusion unit, and a training unit; The acquisition unit is configured to: acquire current task sample data; the current task sample data is sample image data that currently needs class increment; The input unit is configured to: input the current task sample data into the frozen current pre-trained feature extractor to obtain the current sample feature; The fusion unit is configured to: perform data fusion processing on the current sample feature and the previous sample feature using data replay processing to obtain the current fused feature; the previous sample feature is the feature information of all tasks before the current task; the current pre-trained feature extractor is obtained based on data augmentation processing using spatial pyramid processing and contrastive learning training; The training unit is configured to: train the initial feature classifier using the current fused feature to obtain the current classification model after class increment; the current classification model is used to classify the image data corresponding to the current task or the image data corresponding to the previous task.
10. An apparatus for class-incremental learning based on contrastive learning, characterized in that, It includes: A processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the class-increment learning device based on contrastive learning runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the class-increment learning method based on contrastive learning according to any one of claims 1-8.