Image classification method and device based on continuous learning

By adopting CutMix and adaptively integrated knowledge distillation module in image classification task, combined with the regularization method of uncertainty evaluation, the balance problem of old knowledge forgetting and new knowledge learning in image classification is solved, and multiple rounds of continuous learning with high accuracy are achieved.

CN114387486BActive Publication Date: 2025-05-06SUN YAT SEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210061145.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-19
Publication Date
2025-05-06
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

The prior art is difficult to effectively improve the classification accuracy of new images while maintaining old knowledge in image classification tasks, and it is unable to adapt to the needs of multiple rounds of continuous learning.

Method used

Using CutMix-based image augmentation technology and adaptive integration knowledge distillation module, the model is regularized through uncertainty evaluation to achieve an adaptive trade-off between model stability and plasticity.

Benefits of technology

While maintaining old knowledge, significantly improve the classification accuracy of new images, adapt to the needs of multiple rounds of continuous learning, and alleviate the problem of catastrophic forgetting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114387486B_ABST
    Figure CN114387486B_ABST
Patent Text Reader

Abstract

The present invention discloses an image classification method and device based on continuous learning, the method comprising: obtaining a target image to be classified, the target image to be classified comprising a portion of stored old category data images and images of a newly added category; inputting the target image to be classified into a pre-constructed image classification model and obtaining a classification feature vector; training image samples according to the classifier and the classification feature vector to obtain a classification result of the target image; deleting and retaining the image samples according to relevant parameters obtained by the training. The present invention can adaptively alleviate the catastrophic forgetting problem of an image classification network based on continuous learning by weighing the stability and plasticity of the model, so as to effectively improve the classification accuracy of new category images while not forgetting old knowledge, and finally enable the model to learn like a human.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to an image classification method and device based on continuous learning. Background Art

[0002] With the continuous breakthroughs in artificial intelligence technology and the increasing popularity of various intelligent terminal devices, the number of images that need to be processed has increased exponentially. In the field of image processing, one of the most common tasks is the image classification task. At present, the effect of image classification tasks has been proven to reach the level of humans. In order to further obtain the continuous learning ability of humans to continuously acquire, adjust and transfer knowledge throughout their lives, current research mainly focuses on task incremental learning or class incremental learning. Task incremental learning (TIL) assumes that tasks are relatively independent of each other, share the same feature extractor but have task-specific model classifier heads. In contrast, class incremental learning (CIL) assumes that the model is for a single task, and a single feature extractor and a single classifier head are continuously updated and learned in multiple rounds of continuous processes. Since TIL assumes that the user knows which task head to use during reasoning, CIL is a relatively more challenging problem.

[0003] At present, many experts and scholars have proposed a variety of methods to solve the CIL problem. Early studies simply adjusted the strategy of TIL to CIL, for example, based on regularization of model parameters during learning of new class data. The basic idea is to estimate the importance of each parameter in the original model and add more penalties to changes in some parameters that are critical to previously learned classes. Typical examples include the elastic weight integration (EWC) method and memory-aware synapses (MAS). Because old knowledge can be represented not only in model parameters but also in responses to input data. Therefore, with new data categories as input, a knowledge transfer (or distillation) strategy based on distillation loss between the logit output of the old model and the corresponding output part of the new model is applied to CIL. By imitating the output response of the old model, the new model is expected to maintain or acquire knowledge of the old class in the process of learning new classes. Extensions of the distillation strategy include using the output of feature extractors (such as UCIR), multi-layer feature maps in deep learning models (such as PODNet), and even transferring old knowledge through the uncertainty of model predictions. In addition to parameter regularization and knowledge distillation, CIL also proposes other strategies, including the expansion of model components and even sub-networks, generating synthetic data for old categories based on generative adversarial networks or model inversion.

[0004] Deep learning models have demonstrated their human-level performance on specific tasks. Most of the time, models are trained offline based on pre-collected training sets. However, in practice, models may need to be continuously updated to learn more and more human knowledge, such as in autonomous stores and smart medical diagnosis. In such continuous learning tasks, when the model is updated to learn new knowledge, it tends to catastrophically forget the old knowledge learned previously. This is due to the well-known stability-plasticity dilemma, where plasticity refers to the ability to integrate new knowledge and stability refers to retaining previous old knowledge. When updating model parameters in the process of continuously learning new knowledge, allowing excessive plasticity often leads to severe forgetting of old knowledge, while enforcing excessive stability hinders the effective learning of new knowledge.

[0005] To mitigate the catastrophic forgetting problem, state-of-the-art methods try to integrate all previously learned old models into the new model or keep a small portion of old data of each previously learned knowledge type. However, integrating old models will quickly expand the model size in multiple rounds of continuous learning, limiting its use, especially in end-user devices. Therefore, most continuous learning research assumes that the model size will not increase significantly. Regularization-based methods can initially retain old knowledge well, but soon find it difficult to balance model stability and plasticity. Knowledge distillation-based methods may not represent old knowledge well for the output response of the old model to new class data, especially considering that the distribution of new class data is often different from that of old classes. For methods that generate synthetic data of old categories based on generative adversarial networks or model inversion, the quality of synthetic images gradually deteriorates in more learning rounds, and the effect is not satisfactory. Experiments have shown that keeping a small portion of old data has been proven to be very effective in preventing old knowledge from being quickly forgotten. The stored old data can not only be used to directly train the new model with the new data category, but also help transfer old knowledge from the old model to the new model more effectively. Existing knowledge transfer strategies include extracting the output of the last layer, feature extractor, or each convolutional block in a convolutional neural network from the old model to the new model, as shown in the famous iCaRL method and End2End method. However, these knowledge transfer strategies do not solve the trade-off between stability and plasticity well and cannot adapt to multiple rounds of continuous learning. Summary of the invention

[0006] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art and to provide an image classification method and device based on continuous learning. The present invention can adaptively alleviate the catastrophic forgetting problem of the image classification network based on continuous learning by balancing the stability and plasticity of the model, so as to effectively improve the classification accuracy of new types of images while not forgetting old knowledge, and ultimately enable the model to learn like a human.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] On one hand, the present invention provides an image classification method based on continuous learning, comprising the following steps:

[0009] Acquire a target image to be classified, wherein the target image to be classified includes a portion of stored old category data images and images of a newly added category;

[0010] The target image to be classified is input into a pre-built image classification model to obtain a classification feature vector; the image classification model is a neural network model obtained by using the image augmentation technology CutMix to solve the problem of image imbalance, relying on the adaptive integrated knowledge distillation module to alleviate the forgetting of old knowledge, and regularizing the model through uncertainty estimation to balance plasticity and stability;

[0011] Training the image samples according to the classifier and the classification feature vector to obtain the classification result of the target image;

[0012] The image samples are deleted and retained according to the relevant parameters obtained through training.

[0013] As a preferred technical solution, the image classification model is constructed as follows:

[0014] Obtaining a target image sample to be classified and an old image classification model from the previous stage; the old image classification model is a neural network model trained based on old class images in the target image sample;

[0015] The sample image is firstly subjected to data augmentation CutMix technology to generate a mixed sample to obtain a target image with a balance between the new and old classes;

[0016] Input the target image of the new and old class balance into the old image classification model to obtain the first sample feature vector and the first classification probability thereof; input the balanced target image into the new image classification model of this stage to obtain the second sample feature vector and the second classification probability thereof; the new image classification model of this stage is a neural network model initialized by the old image classification model of the previous stage;

[0017] According to the first sample feature vector and its first classification probability, the second sample feature vector and its second classification probability, and the preset discriminator, the new image classification model of this stage is trained by using the adaptive integrated knowledge distillation method, and regularized by uncertainty assessment to obtain the new image classification model of this stage.

[0018] As a preferred technical solution, the sample image is firstly subjected to data augmentation CutMix technology to generate a mixed sample to obtain a target image with a balance between the new and old classes, including:

[0019] The modified CutMix strategy is applied to specifically increase the old class image samples that have been trained; specifically, a sample (x 1, y1), randomly select another sample (x2,y2) from the new class data, and then generate a mixed sample:

[0020]

[0021]

[0022] where m is a binary ({0,1}) mask where 0 represents a randomly selected bounding box region with the same aspect ratio as the image representing the randomly selected bounding box region in x1 and represents pixel-wise multiplication; λ initially represents the area ratio of the region with 1 in the mask, which is modified according to the remix strategy, i.e., when the original λ is larger than a pre-defined threshold τ, λ is reset to 1.0.

[0023] As a preferred technical solution, the target image of the new and old class balance is input into the old image classification model to obtain a first sample feature vector and its first classification probability; the equalized target image is input into the new image classification model of this stage to obtain a second sample feature vector and its second classification probability; the new image classification model of this stage is a neural network model initialized by the old image classification model of the previous stage, including:

[0024] Calculating the mean square error loss between the first sample feature graphs at different levels and the second sample feature graphs at each level through adaptive integration as a first objective function; the first objective function is used to improve the similarity between the output results of the old image classification model and the new image classification model;

[0025] Calculating the cosine similarity between the first sample feature vector and the second sample feature vector as a second objective function; the second objective function is used to improve the similarity between the output results of the old image classification model and the new image classification model;

[0026] Calculating the uncertainty regularized cross entropy loss between the second classification probability and the true classification result corresponding to the equalized target image as the third objective function; the third objective function is used to balance the stability and plasticity of the new image classification model while improving the similarity between its output result and the true classification result corresponding to the sample image, so as to ensure its effect on the old class target image and improve its training effect on the new class target image;

[0027] The new image classification model is trained according to the first objective function, the second objective function and the third objective function to obtain a final image classification model based on continuous learning.

[0028] As a preferred technical solution, the mean square error loss between the first sample feature maps at different levels and the second sample feature maps at each level is calculated through adaptive integration as a first objective function; the first objective function is used to improve the similarity between the output results of the old image classification model and the new image classification model, including:

[0029] Convert the first sample feature atlas outputted by each convolution block in the old image classification model into a new first sample feature atlas through a specific convolution layer;

[0030] The spatial size of all the converted first sample feature graph sets is made the same as the second sample feature graph output by the block in the new image classification model after downsampling or upsampling respectively;

[0031] The first sample feature maps after calculating the downsampling and upsampling are aggregated and input into the convolution layer to generate a set of block-level first attention maps;

[0032] The aggregated first sample feature map is weighted and added block by block based on the first attention map;

[0033] The weighted summed first sample feature map is input into the last convolutional layer, and the shape of its output is the same as the shape of the specific block in the new image classification model.

[0034] As a preferred technical solution, it is characterized in that the uncertainty regularized cross entropy loss between the second classification probability and the true classification result corresponding to the equalized target image is calculated as the third objective function; the third objective function is used to balance the stability and plasticity of the new image classification model while improving the similarity between its output result and the true classification result corresponding to the sample image, so as to ensure its effect on the old class target image and improve its training effect on the new class target image, including:

[0035] Calculating the prediction uncertainty u(x) of the new image classification model for the old class image samples in the continuous learning stage, specifically: calculating the kl divergence between the first classification probability and the second classification probability, treating it as the difference between the two, and obtaining the prediction uncertainty u(x) of the new image classification model for the old class image samples;

[0036] The new image classification model is trained by minimizing the uncertainty regularized cross entropy loss by considering both prediction error and uncertainty.

[0037]

[0038] Among them, L c (x) is the regular cross entropy loss of the input image x, and u(x) is the prediction uncertainty of the old knowledge estimated using the input data, regardless of the class of the input data; this regularization loss is used to allow the new classifier to dynamically weigh the prediction error L c (x) and prediction uncertainty u(x).

[0039] As a preferred technical solution, the new image classification model is trained according to the first objective function, the second objective function and the third objective function to obtain the final image classification model based on continuous learning, including:

[0040] The first objective function, the second objective function and the third objective function are weightedly summed, and the new image classification model is trained by minimizing the sum value to obtain a final image classification model based on continuous learning.

[0041] Another aspect of the present invention provides an image classification device based on continuous learning, which is applied to the image classification method based on continuous learning, and is characterized in that it includes a data acquisition unit, an initialization unit, a data training unit and a data processing unit;

[0042] The data acquisition unit is used to acquire a target image to be classified, wherein the target image to be classified includes a portion of stored old category data images and images of a newly added category;

[0043] The initialization unit is used to input the target image to be classified into a pre-built image classification model and obtain a classification feature vector; the image classification model is a neural network model obtained by solving the problem of image imbalance through the image augmentation technology CutMix, alleviating the forgetting of old knowledge by relying on the adaptive integrated knowledge distillation module, and regularizing the model through uncertainty estimation to balance plasticity and stability;

[0044] The data training unit is used to train the image samples according to the classifier and the classification feature vector to obtain the classification result of the target image;

[0045] The data processing unit is used to delete and retain the image samples according to the relevant parameters obtained through training.

[0046] Another aspect of the present invention provides an electronic device, the electronic device comprising:

[0047] at least one processor; and,

[0048] a memory communicatively connected to the at least one processor; wherein,

[0049] The memory stores computer program instructions executable by the at least one processor, the computer program

[0050] The instructions are executed by the at least one processor so that the at least one processor can perform the image classification method based on continuous learning.

[0051] In another aspect, the present invention provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the image classification method based on continuous learning.

[0052] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0053] The embodiment of the present invention first obtains the target image, constructs the image classification model and obtains the classification feature vector, then trains the balanced data samples according to the image classification model and the classification feature vector, and then deletes and retains the data samples according to the relevant parameters obtained by the training. The existing methods usually fix the stability and plasticity of the model, so that the model cannot make a good trade-off between learning new knowledge and retaining old knowledge. The present invention proposes a feature map adaptive module to achieve weighted fusion of the knowledge of different convolution blocks of the old image classification model according to a learnable weight map, and flexibly transfers it to each convolution block of the new image classification model. While achieving model stability, it also allows the plasticity of the model at a specific stage. Secondly, in the field of continuous learning, the present invention innovatively introduces uncertainty assessment into the normal image classification cross entropy loss for the first time, and evaluates the prediction uncertainty of the new model for the old knowledge by measuring the prediction difference between the new and old image classification models for the old knowledge, so that the new image classification model automatically adjusts the learning of new knowledge and the retention of old knowledge during the training stage. Through the above two modules, the present invention realizes adaptively weighing the plasticity and stability of the model, and alleviates catastrophic forgetting while learning new types of images well. At the same time, in order to alleviate the imbalance problem of new and old class images, the present invention improves the previous two-stage re-balanced sampling training method, adopts an improved CutMix strategy, specifically increases the amount of old class data, and achieves the problem of alleviating the imbalance of new and old class data in end-to-end training. The present invention evaluates the effectiveness of the method on CIFAR100, ImageNet100 and ImageNet1000, and achieves the best effect overall. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0055] Figure 1 A flowchart of an image classification method based on continuous learning according to an embodiment of the present invention;

[0056] Figure 2 It is a schematic diagram of a process of building an image classification model based on continuous learning provided by an embodiment of the present invention;

[0057] Figure 3 is a schematic diagram of a network structure for constructing an image classification model based on continuous learning provided by an embodiment of the present invention;

[0058] Figure 4is a schematic diagram of the structure of an adaptive integrated knowledge distillation module provided by an embodiment of the present invention;

[0059] Figure 5 is a structural schematic diagram of an image classification device based on continuous learning provided by an embodiment of the present invention;

[0060] Figure 6 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0061] In order to enable those skilled in the art to better understand the present application, the technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments and drawings in the present application. It should be understood that the drawings are only used for exemplary illustration and cannot be understood as limitations on this patent. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of this application.

[0062] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments.

[0063] Deep learning models have demonstrated their human-level performance on specific tasks. Most of the time, models are trained offline based on pre-collected training sets. However, in practice, models may need to be continuously updated to learn more and more human knowledge, such as in autonomous stores and smart medical diagnosis. In such continuous learning tasks, when the model is updated to learn new knowledge, it tends to catastrophically forget the old knowledge learned previously. This is due to the well-known stability-plasticity dilemma, where plasticity refers to the ability to integrate new knowledge and stability refers to retaining previous old knowledge. When updating model parameters in the process of continuously learning new knowledge, allowing excessive plasticity often leads to severe forgetting of old knowledge, while enforcing excessive stability hinders the effective learning of new knowledge.

[0064] In order to alleviate the catastrophic forgetting problem, the current image classification methods based on continuous learning are mainly:

[0065] One is to integrate old models, which is to try to integrate all previously learned old models into the new model, but this approach will quickly expand the model size in multiple rounds of continuous learning, limiting its use, especially in end-user devices. Therefore, most continuous learning research assumes that the model size will not increase significantly.

[0066] The second is regularization-based methods, which initially preserve old knowledge well but quickly become unable to strike a balance between model stability and plasticity.

[0067] The third is the method based on knowledge distillation, where the output response of the old model to the new class data may not represent the old knowledge well, especially considering that the distribution of the new class data is often different from that of the old class.

[0068] Fourth, the method of generating synthetic data of old categories based on generative adversarial networks or model inversion gradually degrades the quality of synthetic images in more learning rounds, and the effect is not satisfactory.

[0069] Fifth, a small portion of old data of previously learned knowledge types is retained for each.

[0070] Experiments have shown that retaining a small portion of old data has been shown to be very effective in preventing old knowledge from being quickly forgotten. The stored old data can not only be used to train the new model directly together with the new data category, but also help transfer the old knowledge from the old model to the new model more effectively. Existing knowledge transfer strategies include extracting the output of the last layer, feature extractor, or each convolutional block in the convolutional neural network from the old model to the new model. However, these knowledge transfer strategies do not solve the trade-off between stability and plasticity well, and cannot adapt to multiple rounds of continuous learning.

[0071] See also Figure 1 , is a flow chart of an image classification method based on continuous learning provided by an embodiment of the present invention, the method comprising the following steps:

[0072] S101: Acquire a target image.

[0073] In this implementation example, any image to be classified is defined as a target image. It should be noted that this embodiment does not limit the type of target image to be identified, and any image can be obtained from image classification datasets such as CIFAR-100 and ImageNet as the target image. After further obtaining the target image, the solution provided in this embodiment can be used to perform image classification based on continuous learning on the target image to classify the target image into the real category to which it belongs.

[0074] S102: Input the target image to be classified into a pre-built image classification model and obtain a classification feature vector; the image classification model is a neural network model that solves the problem of image imbalance through image augmentation technology CutMix, relies on an adaptive integrated knowledge distillation module to alleviate the forgetting of old knowledge, and regularizes the model through uncertainty estimation to balance plasticity and stability.

[0075] In this implementation example, after the target image to be classified is obtained through step S101, in order to quickly and accurately classify the target image, the target image can be further input into a pre-built image classification model based on continuous learning, so as to obtain a feature map and feature vector of the target image to execute the subsequent step S103.

[0076] It should be noted that the specific formats of the feature map and feature vector of the target image can be set according to actual conditions (such as the architecture of the selected image classification model, etc.), and this embodiment does not limit this. For example, the feature map of the target image can be a 32*32*16-dimensional feature map, and the target vector can be a 256-dimensional vector, etc.

[0077] S103: Training the image samples according to the classifier and the classification feature vector to obtain a classification result of the target image.

[0078] In this embodiment, after the image classification model and the corresponding feature map and feature vector are obtained through step S102, the model training phase is entered. This embodiment proposes a new method for adaptively balancing model stability and plasticity consisting of two modules. The first is to adaptively integrate multiple levels of old knowledge and transfer it to each block level in the new model. The second is to use the prediction uncertainty of the old knowledge to naturally adjust the importance of learning new knowledge during the model training process. In addition, this embodiment specifically applies the modified CutMix to increase the data of the old knowledge, balance the number of new and old class images, and further alleviate the catastrophic forgetting problem.

[0079] S104: Deleting and retaining the image samples according to the relevant parameters obtained through training.

[0080] In this embodiment, after obtaining the trained image classification model based on continuous learning and the results through step S103, all the images of all categories that have been learned so far will be selected and stored in the memory for use in the next stage of model training. The method of selecting image samples is to input the image samples into the image classification model trained in this stage, obtain the corresponding feature vectors and the class centers of each category, and select a certain number of image samples whose feature vectors are closest to the class center. It should be noted that this embodiment does not limit the number, for example, it can be fixed to 20 image samples for each category, or it can be a total of 2000 image samples for all categories.

[0081] Next, this embodiment will introduce the construction process of the image classification model based on continuous learning. Figure 2 As shown, it shows a schematic diagram of the process of building an image classification model based on continuous learning, which includes the following steps AD:

[0082] Step A: Obtain target image samples and the old image classification model of the previous stage; the old image classification model is a neural network model trained based on old class images in the target image samples.

[0083] In the present embodiment, in order to construct an image classification model based on continuous learning, a large amount of preparatory work needs to be performed in advance. First, a large number of target image samples need to be collected, including obtaining old class image samples from the memory, new class image samples that have not been learned by the previous stage model, and samples obtained after the modified CutMix strategy used in the present embodiment. The target image samples here can be natural image data sets such as CIFAR-100 and ImageNet. Secondly, the old image classification model of the previous stage is obtained, which is actually trained according to the images of the categories currently stored in the memory, but the number of images of each category is different from that stored in the memory now. The number of images stored in the memory now is obtained by sampling the training samples of the old image classification model.

[0084] Among them, the implementation process of training the old image classification model includes: first, selecting the image classification model that has been trained in the previous stage as the original image classification model and initializing the model parameters. It should be noted that this embodiment does not limit the specific network structure of the original model, and it can be any form of neural network, for example, it can be a CNN structure, etc.

[0085] Then, each sample image is used to perform the current round of training on the original image recognition model, and the loss function proposed in this embodiment is used as the objective function to construct an old image classification model, and the network parameters of the model are updated to improve the classification accuracy of the model for the sample images and alleviate the catastrophic forgetting of the model. After multiple rounds of parameter updates (that is, after the training end conditions are met, such as reaching the preset number of training rounds or the change in cross entropy loss is less than the preset threshold, etc.), the old image classification model can be trained.

[0086] Specifically, an optional implementation method is that during the training process, a number of sample images (such as 128 images) are randomly sampled from the training set in each iteration, and a fixed size (such as 224x224) original input image is generated in combination with the data enhancement strategy, and a small batch stochastic gradient descent optimization strategy is used to complete the training of the old image classification model. After the old image classification model is trained, the model will no longer be updated.

[0087] Step B: The sample image is first subjected to data augmentation CutMix technology to generate a mixed sample to obtain a target image with a balance between the new and old classes.

[0088] In this example, after obtaining the target image through step A, in order to alleviate the class imbalance problem during continuous learning, the modified CutMix strategy is applied to increase the training data related to the old class to improve the classification effect of the image classification model. The specific operation is: each time a sample (x1, y1) is randomly selected from the stored old class data, and another sample (x2, y2) is randomly selected from the new class data, and then a mixed sample is generated:

[0089]

[0090]

[0091] Where m is a binary ({0,1}) mask, where 0 represents a randomly selected bounding box region that has the same aspect ratio as the image representing the randomly selected bounding box region in x1 and represents pixel-level multiplication. λ originally represents the area ratio of the region that is 1 in the mask, and is modified here according to the remix strategy, that is, when the original λ is greater than the pre-defined threshold τ, λ is reset to 1.0 (τ=0.6 in this study). In this way, more augmented images have labels of the old category, further alleviating the category imbalance problem.

[0092] Step C: Input the equalized image into the old image classification model to obtain the first sample feature map and vector and the first classification probability thereof; input the equalized target image into the new image classification model of this stage to obtain the second sample feature map and vector and the second classification probability thereof; the new image classification model of this stage is a neural network model initialized by the old image classification model of the previous stage.

[0093] In this embodiment, after the target image is obtained through step B, it can be further input into the old image classification model and the new image classification model to obtain the feature map, feature vector and feature probability corresponding to the image sample, such as Figure 3 As shown in the figure, the new image classification model in this stage is initialized from the old image classification model, and the dimension of the fully connected classifier is adjusted and increased. It should be noted that there is no limit on the size of the increase in the dimension of the last layer classifier. It is designed according to the number of new classes added in the continuous learning increment stage, which can be 2, 5, or 10.

[0094] Step D: Based on the first sample image and feature vector and their first classification probability, the second sample feature image and vector and their second classification probability, and the preset discriminator, the new image classification model of this stage is trained using the adaptive integrated knowledge distillation method, and regularized through uncertainty assessment to obtain the new image classification model of this stage.

[0095] In this embodiment, the first feature map, the first feature vector and the first classification probability belonging to the old image classification model, the second feature map, the second feature vector and the second classification probability belonging to the new image classification model, and the preset classifier are obtained through step C, and the new image classification model of this stage is trained by using the adaptive integrated knowledge distillation method, and regularized by uncertainty assessment to obtain the new image classification model of this stage.

[0096] Specifically, an optional implementation method is that the implementation process of this step D may include the following steps D1-D4:

[0097] D1: calculating the mean square error loss between the first sample feature graphs at different levels and the second sample feature graphs at each level through adaptive integration as a first objective function; the first objective function is used to improve the similarity between the output results of the old image classification model and the new image classification model;

[0098] It should be noted that this embodiment proposes an adaptive integrated knowledge distillation strategy, such as Figure 4As shown, by extracting old knowledge from multiple blocks of the old feature extractor to each block of the new feature extractor. In the proposed adaptive integration module ('AIM'), in order to refine the old knowledge in the old feature extractor into a specific (l-th) convolutional block in the new feature extractor, the output feature atlas of each convolutional block in the old feature extractor is transformed into a new feature atlas through a specific convolutional layer, and then all the transformed feature atlases are downsampled or upsampled respectively to the same spatial size as the output feature maps of the blocks in the new feature extractor. Then, the downsampled and upsampled feature maps are aggregated and input to the convolutional layer to generate a set of block-level attention maps. Based on the attention map, the aggregated feature maps are weighted and added block by block. Finally, the weighted summed feature map is input to the last convolutional layer, and its output shape is the same as the shape of the l-th block in the new feature extractor. The attention map and the last convolutional layer are used to integrate the knowledge from multiple blocks of the old feature extractor. Formally, with F t-1 (x) represents the set of original feature maps of all blocks in the old feature, denoted by A l (F t-1 (x)) represents the output of the AIM module used to extract knowledge from the old feature extractor to the lth block of the new feature extractor. The knowledge distillation from the old feature extractor to the new feature extractor can then be obtained by minimizing the loss

[0099]

[0100] Where L is the number of convolutional blocks in the two feature extractors. Note that the AIM module and the new classifier are trainable in each round of continuous learning, and their parameters are trainable, which are omitted in Eq. (1) for simplicity.

[0101] D2: Calculate the cosine similarity between the first sample feature vector and the second sample feature vector as a second objective function; the second objective function is used to improve the similarity between the output results of the old image classification model and the new image classification model;

[0102] It should be noted that the second objective function is defined as

[0103]

[0104] where f t-1 (x) and f t (x) represent the outputs of the old feature extractor and the new feature extractor respectively. The distance metric <.,.> in this embodiment represents the definition of 1-cos(f t-1 (x), f t (x)).

[0105] D3: Calculate the uncertainty regularized cross entropy loss between the second classification probability and the true classification result corresponding to the equalized target image as the third objective function; the third objective function is used to balance the stability and plasticity of the new image classification model while improving the similarity between its output result and the true classification result corresponding to the sample image, so as to ensure its effect on the old class target image and improve its training effect on the new class target image;

[0106] It is important to note that when updating a new classifier during the continuous learning process, flexible updates in the feature extractor can easily lead to the updated classifier having accurate predictions for every available training data. However, as observed in extensive studies, classifiers are often overconfident in their predictions on both training and test data, suggesting that the classifiers are not actually sure about their predictions. Intuitively, learning or retaining a particular type of knowledge well with a classifier usually corresponds to lower uncertainty in predictions about such knowledge during inference. This means that, in the process of continuously learning new knowledge, forcing the new classifier to reduce the uncertainty of old knowledge can help the new classifier retain the old knowledge. If the new classifier somehow knows its own prediction uncertainty, especially its predictions for old classes during continuous learning of new classes, then the new classifier may be trained in a way that considers both prediction error and uncertainty, such as Figure 3 As shown, by minimizing the uncertainty regularized cross entropy loss:

[0107]

[0108] Among them, L c (x) is the regular cross entropy loss of the input image x, and u(x) is the prediction uncertainty estimated using the old knowledge of the input data, regardless of the class of the input data. With this regularized loss, the new classifier can dynamically weigh the prediction error L c (x) and prediction uncertainty u(x). In particular, when the uncertainty is large, the contribution of the first loss term in the equation becomes smaller, so the learning adaptively aims to reduce the second term u(x), and vice versa. In this way, in the process of learning new knowledge, the prediction uncertainty of the old knowledge is well minimized, that is, the prediction of the new classifier on the old knowledge is generally certain, indicating that the new classifier has a lot of knowledge about the old knowledge and therefore retains the old knowledge well.

[0109] In the continuous learning task, u(x) can be calculated based on the logit g of the old classifier. t-1 The difference between (x) (i.e., the input of the softmax operator) and the corresponding logit, reasonably defines the new classifier For the prediction uncertainty u(x) part of the old knowledge. If the difference is large, it means that the new classifier is more updated than the old classifier, so the new classifier may forget more old knowledge and the prediction of the old classification data may be more uncertain. Based on the research, this example recommends but does not limit u(x) to be defined as

[0110]

[0111] The difference measure D can be designed flexibly. For example, two logits g t-1 (x) and The cross entropy or KL divergence between .

[0112] D4: Training the new image classification model according to the first objective function, the second objective function and the third objective function to obtain a final image classification model based on continuous learning.

[0113] In the implementation method, in order to comprehensively optimize the overall network parameters of the new image classification model, the three objective functions in (1), (2), and (3) above can be integrated to obtain the final model optimization objective function as shown in the following diagram:

[0114]

[0115] in is the available training set for the current t-th round of continuous learning, i.e., the initial target image obtained. is an enhanced training set based on the CutMix strategy, i.e., a balanced target image, especially to alleviate the class imbalance problem. a λ u λ f is a constant that balances the three loss terms.

[0116] In another embodiment of the present invention, an image classification device based on continuous learning will be introduced. For related content, please refer to the above method embodiment.

[0117] See also Figure 5 , is a schematic diagram of the structure of an image classification device based on continuous learning provided in this embodiment, the device 500 includes:

[0118] The data acquisition unit 501 is used to acquire a target image, where the target image includes a portion of the stored old category data and an image of a newly added category;

[0119] Initialization unit 502, used to build an image classification model and obtain a classification feature vector;

[0120] A data training unit 503, configured to train the target image according to the image classification model and the classification feature vector;

[0121] The data processing unit 504 is used to delete and retain the data samples according to the relevant parameters obtained through training.

[0122] In a first possible implementation of this example, the initialization unit 502 includes:

[0123] An initial data acquisition unit acquires a target image sample and an old image classification model from a previous stage; the old image classification model is a neural network model trained based on old class images in the target image sample;

[0124] A balanced data acquisition unit, which first generates a mixed sample by using the data augmentation CutMix technology to obtain a target image with a balance between the new and old classes;

[0125] The obtaining unit inputs the balanced data into the old image classification model to obtain a first sample feature vector and the first classification probability thereof; inputs the balanced target image into the new image classification model of this stage to obtain a second sample feature vector and the second classification probability thereof; the new image classification model of this stage is a neural network model initialized by the old image classification model of the previous stage;

[0126] The training unit trains the new image classification model of this stage by using an adaptive integrated knowledge distillation method according to the first sample feature vector and its first classification probability, the second sample feature vector and its second classification probability, and a preset discriminator, and regularizes it through uncertainty assessment to obtain the new image classification model of this stage.

[0127] In a second possible implementation manner of this example, the balance data acquisition unit includes:

[0128] The modified CutMix strategy is applied to specifically increase the old class image samples that have been trained. Specifically, each time a sample (x1, y1) is randomly selected from the stored old class data, another sample (x2, y2) is randomly selected from the new class data, and then a mixed sample is generated:

[0129]

[0130]

[0131] Where m is a binary ({0,1}) mask, where 0 represents a randomly selected bounding box region that has the same aspect ratio as the image representing the randomly selected bounding box region in x1 and represents pixel-wise multiplication. λ originally represents the area ratio of the region that is 1 in the mask, and here it is modified according to the remix strategy, that is, when the original λ is greater than a pre-defined threshold τ, λ is reset to 1.0 (τ=0.6 in this study). In this way, more enhanced images have labels of the old category.

[0132] In a third possible implementation of this example, the training unit includes:

[0133] A first operator unit, calculating the mean square error loss between the first sample feature graphs at different levels and the second sample feature graphs at each level through adaptive integration as a first objective function; the first objective function is used to improve the similarity between the output results of the old image classification model and the new image classification model;

[0134] A second operator unit is configured to calculate the cosine similarity between the first sample feature vector and the second sample feature vector as a second objective function; the second objective function is used to improve the similarity between the output results of the old image classification model and the new image classification model;

[0135] The third operator unit calculates the uncertainty regularized cross entropy loss between the second classification probability and the true classification result corresponding to the equalized target image as the third objective function; the third objective function is used to balance the stability and plasticity of the new image classification model while improving the similarity between its output result and the true classification result corresponding to the sample image, so as to ensure its effect on the old class target image and improve its training effect on the new class target image;

[0136] The first training subunit trains the new image classification model according to the first objective function, the second objective function and the third objective function to obtain a final image classification model based on continuous learning.

[0137] In a fourth possible implementation of this example, the first operator unit includes:

[0138] A conversion unit, converting the first sample feature atlas outputted by each convolution block in the old image classification model into a new first sample feature atlas through a specific convolution layer;

[0139] An alignment unit, which downsamples or upsamples all the converted first sample feature graphs to have the same spatial size as the second sample feature graph output by the block in the new image classification model;

[0140] An attention map generating unit, which calculates the first sample feature maps after downsampling and upsampling, and then aggregates them and inputs them into the convolution layer to generate a set of block-level first attention maps;

[0141] A weighted summing unit, weighting and adding the aggregated first sample feature map block by block based on the first attention map;

[0142] The output unit inputs the weighted summed first sample feature map into the last convolutional layer, and the shape of its output is the same as the shape of the specific block in the new image classification model.

[0143] In a fifth possible implementation of this example, the second operator unit includes:

[0144] A prediction uncertainty evaluation unit, which calculates the prediction uncertainty u(x) of the new image classification model for the old class image samples in the continuous learning phase;

[0145] Regularization unit, which considers both prediction error and uncertainty at the same time, trains the new image classification model by minimizing the cross entropy loss of uncertainty regularization.

[0146]

[0147] Among them, L c (x) is the regular cross entropy loss of the input image x, and u(x) is the prediction uncertainty of the old knowledge estimated using the input data, regardless of the class of the input data; this regularization loss is used to allow the new classifier to dynamically weigh the prediction error L c (x) and prediction uncertainty u(x).

[0148] In a sixth possible implementation of this example, the prediction uncertainty assessment unit includes:

[0149] The difference calculation unit calculates the kl divergence between the first classification probability and the second classification probability, regards it as the difference between the two, and obtains the prediction uncertainty u(x) of the new image classification model for the old class image sample.

[0150] In a seventh possible implementation manner of this example, the first training subunit includes:

[0151] The first objective function, the second objective function and the third objective function are weightedly summed, and the new image classification model is trained by minimizing the sum value to obtain a final image classification model based on continuous learning.

[0152] It should be noted that the image classification device based on continuous learning of the present invention corresponds one-to-one to the image classification method based on continuous learning of the present invention. The technical features and beneficial effects described in the embodiment of the above-mentioned image classification method based on continuous learning are all applicable to the embodiment of the image classification device based on continuous learning. For specific contents, please refer to the description in the embodiment of the method of the present invention. It will not be repeated here. This is hereby declared.

[0153] In addition, in the implementation of the continuous learning-based image classification device in the above-mentioned embodiment, the logical division of each program module is only an example. In actual applications, the above-mentioned functions can be assigned to different program modules as needed, for example, for the configuration requirements of the corresponding hardware or the convenience of software implementation. That is, the internal structure of the continuous learning-based image classification device is divided into different program modules to complete all or part of the functions described above.

[0154] like Figure 5 As shown, in one embodiment, an electronic device of an image classification method based on continuous learning is provided, and the electronic device 600 may include a first processor 601, a first memory 602 and a bus, and may also include a computer program stored in the first memory 602 and executable on the first processor 601, such as an image classification program 603 based on continuous learning.

[0155] Among them, the first memory 602 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (for example: SD or DX memory, etc.), magnetic memory, disk, optical disk, etc. In some embodiments, the first memory 602 can be an internal storage unit of the electronic device 600, such as a mobile hard disk of the electronic device 600. In other embodiments, the first memory 602 can also be an external storage device of the electronic device 600, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device 600. Further, the first memory 602 can also include both an internal storage unit of the electronic device 600 and an external storage device. The first memory 602 can not only be used to store application software and various types of data installed in the electronic device 600, such as the code of the image classification program 603 based on continuous learning, but also can be used to temporarily store data that has been output or is to be output.

[0156] In some embodiments, the first processor 601 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips, etc. The first processor 601 is the control core (Control Unit) of the electronic device, and uses various interfaces and lines to connect various components of the entire electronic device, and executes or executes programs or modules (such as federated learning defense programs, etc.) stored in the first memory 602, and calls data stored in the first memory 602 to execute various functions of the electronic device 600 and process data.

[0157] Figure 6 Only an electronic device with components is shown, and those skilled in the art will understand that Figure 6 The structure shown does not constitute a limitation on the electronic device 600 , and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.

[0158] The image classification program 603 based on continuous learning stored in the first memory 602 of the electronic device 600 is a combination of multiple instructions. When running in the first processor 601, it can achieve:

[0159] Acquire a target image to be classified, wherein the target image to be classified includes a portion of stored old category data images and images of a newly added category;

[0160] The target image to be classified is input into a pre-built image classification model to obtain a classification feature vector; the image classification model is a neural network model obtained by using the image augmentation technology CutMix to solve the problem of image imbalance, relying on the adaptive integrated knowledge distillation module to alleviate the forgetting of old knowledge, and regularizing the model through uncertainty estimation to balance plasticity and stability;

[0161] Training the image samples according to the classifier and the classification feature vector to obtain the classification result of the target image;

[0162] The image samples are deleted and retained according to the relevant parameters obtained through training.

[0163] Furthermore, if the module / unit integrated in the electronic device 600 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).

[0164] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0165] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0166] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. An image classification method based on continuous learning, characterized in that: The steps include: Acquire a target image to be classified, wherein the target image to be classified includes a portion of stored old category data images and images of a newly added category; The target image to be classified is input into a pre-built image classification model to obtain a classification feature vector; the image classification model is a neural network model obtained by using the image augmentation technology CutMix to solve the problem of image imbalance, relying on the adaptive integrated knowledge distillation module to alleviate the forgetting of old knowledge, and regularizing the model through uncertainty estimation to balance plasticity and stability; Training the image samples according to the image classification model and the classification feature vector to obtain the classification result of the target image; Deleting and retaining the image samples according to the relevant parameters obtained through training; The image classification model is constructed as follows: Obtaining a target image sample to be classified and an old image classification model from the previous stage; the old image classification model is a neural network model trained based on old class images in the target image sample; The target image to be classified is firstly subjected to data augmentation CutMix technology to generate mixed samples to obtain a target image with a balance between new and old classes; Input the target image of the new and old class balance into the old image classification model to obtain a first sample feature vector and a first classification probability thereof; input the balanced target image into the new image classification model of this stage to obtain a second sample feature vector and a second classification probability thereof; the new image classification model of this stage is a neural network model initialized by the old image classification model of the previous stage; According to the first sample feature vector and the first classification probability thereof, the second sample feature vector and the second classification probability thereof, and the preset discriminator, the new image classification model of this stage is trained by using the adaptive integrated knowledge distillation method, and regularized by uncertainty evaluation to obtain the new image classification model of this stage; Input the target image of the new and old class balance into the old image classification model to obtain a first sample feature vector and a first classification probability thereof; input the balanced target image into the new image classification model of this stage to obtain a second sample feature vector and a second classification probability thereof; The new image classification model in this stage is a neural network model initialized by the old image classification model in the previous stage, including: Calculating the mean square error loss between the first sample feature graphs at different levels and the second sample feature graphs at each level through adaptive integration as a first objective function; the first objective function is used to improve the similarity between the output results of the old image classification model and the new image classification model; Calculating the cosine similarity between the first sample feature vector and the second sample feature vector as a second objective function; the second objective function is used to improve the similarity between the output results of the old image classification model and the new image classification model; The uncertainty regularized cross entropy loss between the second classification probability and the true classification result corresponding to the equalized target image is calculated as the third objective function; the third objective function is used to balance the stability and plasticity of the new image classification model while improving the similarity between its output result and the true classification result corresponding to the sample image, so as to ensure its effect on the old class target image and improve its training effect on the new class target image; The new image classification model is trained according to the first objective function, the second objective function and the third objective function to obtain a final image classification model based on continuous learning; Calculating the mean square error loss between the first sample feature graphs of different levels and the second sample feature graphs of each level through adaptive integration as the first objective function; the first objective function is used to improve the similarity between the output results of the old image classification model and the new image classification model, including: Convert the output first sample feature atlas of each convolution block in the old image classification model into a new first sample feature atlas through a specific convolution layer; All converted first sample feature maps are respectively downsampled or upsampled to have the same spatial size as the second sample feature map output by the block in the new image classification model; The first sample feature maps after calculating the downsampling and upsampling are aggregated and input into the convolution layer to generate a set of block-level first attention maps; The aggregated first sample feature map is weighted and added block by block based on the first attention map; The weighted summed first sample feature map is input into the last convolutional layer, and the shape of its output is the same as the shape of the specific block in the new image classification model.

2. The image classification method based on continuous learning according to claim 1, characterized in that: The target image to be classified is firstly subjected to data augmentation CutMix technology to generate mixed samples to obtain a target image with a balance between new and old classes, including: Apply the modified CutMix strategy to specifically increase the old class image samples that have been trained; specifically, each time randomly select a sample (x1, y1) from the stored old class data, and randomly select another sample (x2, y2) from the new class data, and then generate a mixed sample: where m is a binary ({0,1}) mask where 0 represents a randomly selected bounding box region with the same aspect ratio as the image representing the randomly selected bounding box region in x1 and represents pixel-wise multiplication; λ initially represents the area ratio of the region with 1 in the mask, which is modified according to the remix strategy, i.e., when the original λ is larger than a pre-defined threshold τ, λ is reset to 1.

0.

3. The image classification method based on continuous learning according to claim 2, characterized in that: The uncertainty regularized cross entropy loss between the second classification probability and the true classification result corresponding to the equalized target image is calculated as the third objective function; the third objective function is used to balance the stability and plasticity of the new image classification model while improving the similarity between its output result and the true classification result corresponding to the sample image, so as to ensure its effect on the old class target image and improve its training effect on the new class target image, including: Calculating the prediction uncertainty u(x) of the new image classification model for the old class image samples in the continuous learning stage, specifically: calculating the kl divergence between the first classification probability and the second classification probability, treating it as the difference between the two, and obtaining the prediction uncertainty u(x) of the new image classification model for the old class image samples; The new image classification model is trained by minimizing the uncertainty regularized cross entropy loss by considering both prediction error and uncertainty. Among them, L c (x) is the regular cross entropy loss of the input image x, and u(x) is the prediction uncertainty of the old knowledge estimated using the input data, regardless of the class of the input data; this regularization loss is used to allow the new classifier to dynamically weigh the prediction error L c (x) and prediction uncertainty u(x).

4. The image classification method based on continuous learning according to claim 1, characterized in that: The new image classification model is trained according to the first objective function, the second objective function and the third objective function to obtain a final image classification model based on continuous learning, including: The first objective function, the second objective function and the third objective function are weightedly summed, and the new image classification model is trained by minimizing the sum value to obtain a final image classification model based on continuous learning.

5. An image classification device based on continuous learning, characterized in that: The image classification method based on continuous learning applied to any one of claims 1 to 4, characterized in that it comprises a data acquisition unit, an initialization unit, a data training unit and a data processing unit; The data acquisition unit is used to acquire a target image to be classified, wherein the target image to be classified includes a portion of stored old category data images and images of a newly added category; The initialization unit is used to input the target image to be classified into a pre-built image classification model and obtain a classification feature vector; the image classification model is a neural network model obtained by solving the problem of image imbalance through the image augmentation technology CutMix, alleviating the forgetting of old knowledge by relying on the adaptive integrated knowledge distillation module, and regularizing the model through uncertainty estimation to balance plasticity and stability; The data training unit is used to train the image samples according to the image classification model and the classification feature vector to obtain the classification result of the target image; The data processing unit is used to delete and retain the image samples according to the relevant parameters obtained through training.

6. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the image classification method based on continuous learning as described in any one of claims 1-4.

7. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the image classification method based on continuous learning described in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Method for intensively improving graph classification precision under continuous learning based on comparison categories

    CN113554078A