Multi-label class incremental learning method and system based on explicit known and unknown knowledge
Through dynamic feature purification module and distribution prior pseudo-labeling technology, combined with the use of detecting unknown modules, the problems of missing tags, future category interference and feature overlap in multi-label incremental learning are solved, and more efficient knowledge recall and feature mining are achieved, and the performance and stability of the model are improved.
Patent Information
- Application Number
- CN202510190260.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-17
AI Technical Summary
When existing multi-label incremental learning methods deal with the problems of missing tags, interference of future categories on the model, and overlapping of historical and current category features, it is difficult to effectively balance the model's goals in historical, current and future learning, resulting in catastrophic forgetting and feature confusion.
A multi-label class incremental learning method based on explicitly known and unknown knowledge is proposed. The class-aware features of fine-grained class perception are captured through dynamic feature purification modules, the old knowledge is recalled using the distribution prior pseudo-labeling technology, and the features related to future categories are mined by detecting unknown modules.
This method can clearly distinguish and specify known and unknown knowledge, alleviate the contradiction between historical, current and future learning goals in the incremental process, prevent feature confusion, and improve the model's performance in multi-label incremental learning.
Smart Images

Figure CN120164013A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, and particularly relates to a multi-label class incremental learning method and system based on explicit known and unknown knowledge. Background Art
[0002] Multi-label class incremental learning (MLCIL) has become an important research direction in the field of incremental learning in recent years. It aims to enable the model to efficiently accept and learn new class information on the basis of existing class knowledge while maintaining the recognition accuracy of old classes. In traditional incremental learning, the model may forget the knowledge of existing tasks when learning new tasks. Especially in multi-label tasks, the mutual relationship and dependence between labels make this problem more complex. With the continuous expansion of the data scale and the increasing diversity of application scenarios, significant progress has been made in this field. The current mainstream methods include knowledge storage methods based on replay, which avoid catastrophic forgetting by saving old data or generating data; using regularization methods to protect the knowledge of old tasks by constraining the model weights; meanwhile, modeling the correlation between labels is also widely used.
[0003] The existing multi-label class incremental learning methods have the following problems:
[0004] 1. In multi-label incremental learning, images often contain multiple labels, and historical and future labels are usually not available for annotation. Traditional anti-forgetting methods (such as knowledge distillation and sample replay) fail to effectively handle the target conflict caused by label missing. These methods fail to solve the contradiction between maintaining old knowledge, learning the current task, and preparing for future tasks, resulting in the model being unable to balance these three key goals during the learning process.
[0005] 2. The existing methods fail to effectively cope with the interference of future classes to the model. Since the current model does not know future classes, the model is prone to misactivate the features of future classes as the features of known classes, resulting in confusion and overlap in the model representation.
[0006] 3. In multi-label tasks, the features of historical and current classes may cross and overlap. This "feature confusion" problem exacerbates catastrophic forgetting and affects the recall of the learned knowledge. Without sufficient fine-grained class awareness ability, the model is difficult to distinguish the boundaries between various classes, resulting in blurred knowledge boundaries between tasks. Summary of the Invention
[0007] The objective of the present invention is to clarify the known and unknown content, adapt to historical, current, and future knowledge in multi-label class increment, and propose a new framework called HCP. To clarify the known knowledge, a dynamic feature purification module is first proposed to capture fine-grained class-aware features and prevent feature aliasing across tasks. Then, the old known knowledge is effectively recalled through pseudo-labeling technology with distribution priors, thus alleviating the problem of large inter-class forgetting differences. To detect unknown knowledge, knowledge is mined from images covering historical, current, and future categories to develop features related to future categories in preparation for future learning.
[0008] The technical solution adopted by the present invention is as follows:
[0009] A multi-label class incremental learning method based on clarifying known and unknown knowledge, comprising the following steps:
[0010] Extract visual features from the input image;
[0011] Utilize the class embeddings of historical and current classes and the extracted visual features to obtain fine-grained class-aware features;
[0012] Synthesize unknown class features using the class-aware features;
[0013] Predict the probabilities of old classes using a stability classifier, predict the probabilities of new and unknown classes using a plasticity classifier, and merge the probabilities of old classes and new and unknown classes for image classification;
[0014] Supplement the pseudo-labels of old classes for the current image according to the classes predicted by the historical model.
[0015] Furthermore, the step of utilizing the class embeddings of historical and current classes and the extracted visual features to obtain fine-grained class-aware features includes: Assigning a learnable class embedding S to all historical and current classes learned for the current task i , and the embeddings of all known classes constitute the overall class embedding S = S 1:t-1 |S t , where S t is the class embedding being learned in the current task t, and S 1:t-1 represents the class embedding learned in the previous task; Concatenate S with the visual features P of the image for attention calculation, aggregate object information, and extract fine-grained class-aware features O S .
[0016] Furthermore, the step of synthesizing unknown class features using the class-aware features is to generate unknown class features by interpolation using the class-aware features.
[0017] Further, the stability classifier remains frozen to retain old knowledge; for each next incremental task, the stability classifier and the plasticity classifier are merged into a new stability classifier, and a new plasticity classifier is attached for learning new classes.
[0018] Further, supplementing pseudo-labels of old classes for the current image according to the classes predicted by the historical model includes: formulating a pseudo-label threshold for each class according to the prediction distribution of the historical model for each class; for each old class, if the prediction probability of the current model for this class is greater than or equal to the pseudo-label threshold, it is considered that this class exists in the image, then the binary label of this class is changed from 0 to 1 as supplementary supervision; if the prediction probability of the current model for this class is less than the pseudo-label threshold, it is considered that this class does not exist in the image, and the binary label of this class remains 0.
[0019] Further, under the supervision of the ground-truth label and the pseudo-label, optimize the training by calculating the weighted asymmetric loss function L WASL to optimize the training.
[0020] A multi-label class incremental learning system based on explicitly known and unknown knowledge, which includes:
[0021] A feature extraction module for extracting visual features from the input image;
[0022] A feature purification module for obtaining fine-grained class-aware features by using the class embeddings of historical classes and current classes and the extracted visual features;
[0023] An unknown detection module for synthesizing unknown class features by using class-aware features;
[0024] A stability-plasticity classifier module for predicting the probabilities of old classes by using the stability classifier, predicting the probabilities of new classes and unknown classes by using the plasticity classifier, and merging the probabilities of old classes and the probabilities of new classes and unknown classes for image classification;
[0025] A recall enhancement module for supplementing pseudo-labels of old classes for the current image according to the classes predicted by the historical model.
[0026] The beneficial effects of the present invention are as follows:
[0027] Compared with existing methods, the advantages of the present invention lie in being able to clearly distinguish and specify known and unknown knowledge, alleviating the contradiction between historical, current, and future learning objectives during the incremental process of the model. By introducing a dynamic feature purification module, the present invention precisely captures fine-grained category-aware features, thus preventing feature confusion among historical, current, and future categories. This innovation enables the model to clearly define the boundaries of each category, avoiding interference and aliasing between categories. For historical knowledge, the present invention effectively recalls old knowledge by means of pseudo-labeling with distribution priors. For future categories, the present invention enables the model to learn a richer feature set and enhance the discriminability of current features by effectively mining potential category knowledge. Experiments show that the present invention can achieve better performance on existing data sets, effectively alleviating the catastrophic forgetting phenomenon in multi-label class incremental learning. Brief Description of the Drawings
[0028] Figure 1 is the HCP framework structure of the present invention. To clarify known knowledge, a dynamic feature purification module is designed to capture fine-grained category-aware features to avoid cross-session feature aliasing, and distribution priors are used to effectively retain historical known knowledge. To explore unknown knowledge, known features are interpolated into future categories to enrich the feature set and promote future learning.
[0029] Figure 2 is the visualization of the attention map of each true label in the feature purification module. Calculating the average value of each head attention shows that the class embedding focuses on specific local regions of the image. The text below the image is the true label of the input image. Detailed Description of the Invention
[0030] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below through specific embodiments and drawings.
[0031] The present invention proposes a new framework HCP (Historical-Current-Prospective), as Figure 1As shown in the figure. The core idea of HCP is to clarify the known and unknown content in the current incremental session, enabling the model to accommodate historical, current, and future knowledge simultaneously, thereby alleviating the conflicts between learning objectives. Specifically, to clarify the known knowledge, the HCP framework first introduces a dynamic feature purification module, where each category embedding focuses on fine-grained class-aware features rather than covering multiple categories, thus avoiding cross-session feature aliasing. It can flexibly introduce new classes in incremental learning by continuously adding new category embeddings. In addition, the HCP framework enhances the recall of historical knowledge by effectively utilizing the prior knowledge of the previous model while alleviating the problem of large differences in inter-class forgetting. To detect unknown knowledge, the HCP framework interpolates category features to generate features of unknown classes, thereby pushing away the features of all other non-target classes from the generated features, making the features of known classes more compact and conducive to future learning. The entire framework consists of several parts: a feature extraction module, a feature purification module, an unknown detection module, a stability-plasticity classifier module, and a recall enhancement module.
[0032] 1. Feature Extraction Module
[0033] The feature extraction module consists of a TResNet-M network and is used to extract rich visual features P from the input images. Among them, visual features refer to the information extracted from images that can represent the content of the images, including low-level features (basic features such as edges and textures), middle-level features (local shapes, partial object structures, etc.), and high-level features (complete object and scene semantics).
[0034] 2. Feature Purification Module
[0035] The feature purification module is implemented based on the attention mechanism. Purification means extracting fine-grained class-aware features from entangled multi-class global features. Due to the complex background and the existence of multiple object categories in multi-label images, the global features contain a large amount of redundant information, and the features of different categories interfere with each other, resulting in a decrease in the inter-class distinguishability of the model. Purification extracts fine-grained features highly relevant to the target category by removing the interference of noise and non-target features, thereby improving the performance of the model.
[0036] Specifically, a learnable embedding S is assigned to all known categories (all categories learned in the current task, including historical and current categories) i , representing the semantic features of this category. The embeddings of all known categories constitute the overall category embedding S = S 1:t-1 |S t , where S t is the embedding of the category being learned in the current task t, and S 1:t-1 represents the embedding of the categories learned in the previous task. S is concatenated with the visual features P of the image for attention calculation, aggregating object information and extracting fine-grained class-aware features OS , for detecting unknown processes.
[0037] By calculating the attention weights between each visual feature and the class embeddings through self-attention, the model can focus on the correlations between different class embeddings and visual features. Specifically, each region (or local region) of the image interacts with all known class embeddings of the current task, enabling the model to identify which class features are most relevant to the local region in the image, and then aggregating the key information of different classes to obtain fine-grained class-aware features O S .
[0038] 3. Unknown Detection Module
[0039] Since multi-label images simultaneously cover objects of historical, current, and future classes, the present invention utilizes the unknown detection module to mine features related to future classes, helping the model learn richer features and enhancing the discriminability of current features. Specifically, the unknown detection module utilizes the extracted fine-grained class-aware features O S to synthesize unknown class features. For the classes present in the image, the attention of the class embeddings aggregates on the corresponding objects, while the attention of the non-existent classes is freely distributed in the background regions of the image, which contain the classes to be learned in the future. Using these class-aware features O S new features are generated by interpolation, namely unknown class features O V , which are learned as unknown classes. This enriches the feature set, can promote the compact representation of real class features, make the boundaries of real class features clearer, and is conducive to reserving space for future learning.
[0040] 4. Stability-Plasticity Classifier Module
[0041] The stability-plasticity classifier consists of a stability classifier and a plasticity classifier. Stability refers to the ability of the model to maintain the memory and accuracy of old knowledge when learning new knowledge, preventing catastrophic forgetting. Plasticity refers to the adaptability of the model when learning new knowledge, that is, the learning ability of the model for new tasks or new data. The old class embeddings are input into the stability classifier to predict the classification probability p of the old classes 1:t-1 , and the new class embeddings and unknown class embeddings are input into the plasticity classifier to obtain the prediction probabilities p of the new classes and unknown classes t +1, and the probabilities of the old classes and new classes are concatenated as p 1:t +1 for image classification. Among them, the stability classifier remains frozen to maintain old knowledge. For each next incremental task, the stability classifier and the plasticity classifier are combined into a new stability classifier, and at the same time, a new plasticity classifier is added for the learning of new classes.
[0042] Among them, the stability classifier and the plasticity classifier can be implemented using a linear classifier. Keeping frozen means freezing the parameters during training so that they are not updated or adjusted during training.
[0043] 5. Recall Enhancement Module
[0044] The recall enhancement module refers to supplementing the current image with pseudo-labels of old classes according to the prediction output of the old model (historical model). Due to the large differences in the learning difficulty and forgetting degree between classes, this module selects class-specific thresholds for pseudo-labeling (pseudo-labels) according to the prediction distribution N i (·) of each class by the old model, so as to increase the recall of old class labels and reduce the false labeling rate. Where N represents the Gaussian distribution, and the prediction distributions of all classes form a distribution queue {N1(·), N2(·), …, N 1:t-1 (·)}.
[0045] The entire process of the method of the present invention is divided into the following steps:
[0046] 1. The input picture extracts visual features through the feature extraction module.
[0047] 2. The extracted visual features obtain fine-grained class-aware features through the feature purification module, and then are input into the unknown detection module.
[0048] 3. In order to detect the unknown, the embedded features of the missing classes in the image are freely distributed in the background area of the image, and these areas contain future classes. The features of the unknown classes are synthesized through interpolation operations using the missing class features, and are input into the classifier together with the real class features.
[0049] 4. Use the stability classifier to predict the probability of old classes, use the plasticity classifier to predict the probability of new classes and unknown classes, and classify after combining the probability of old classes and the probability of new classes and unknown classes into a complete probability. Among them, the old class embeddings and the stability classifier are frozen to retain old knowledge, and the new class embeddings and the plasticity classifier are updated to adapt to new data.
[0050] 5. Recall old knowledge using the class probability predicted by the historical model to generate pseudo-labels of old classes. Considering the large differences in forgetting between classes, class-specific pseudo-label thresholds are formulated here using the prediction distribution of classes to reduce the false alarm rate and increase the recall rate.
[0051] Each old class has a confidence distribution (assumed to be a normal distribution here where μ i 、 (representing the mean and variance of the confidence distribution respectively), which is obtained statistically based on the confidence level of the model's prediction for this category in previous training tasks. Since the learning difficulty and forgetting degree of each category vary greatly, for each old category, the mean of the confidence distribution of each category is selected as the threshold of the pseudo-label for this category. If the prediction probability of the current model for this category is greater than or equal to the pseudo-label threshold, it is considered that this category exists in the image sample, and then the binary label of this category is changed from 0 to 1 as supplementary supervision. If the prediction probability of the current model for this category is less than the pseudo-label threshold, it is considered that this category does not exist in the image, and the binary label of this category remains 0. By specifying a class-specific pseudo-label threshold according to the historical confidence distribution of each category, the quality of the pseudo-label and the stability of training can be effectively improved, and the performance of the model in incremental learning can be enhanced.
[0052] 6. Under the supervision of the true label and the pseudo-label, by calculating the weighted asymmetric loss function L WASL to optimize the training.
[0053] Among them, the formula of the weighted asymmetric loss function is as follows:
[0054]
[0055] Among them, K is the total number of known categories, p k is the classification probability of the k-th category. y k is the binary label of the k-th category, 1 indicates that this category exists in the image, and 0 indicates that this category does not exist in the image. γ + and γ - are used to amplify the influence of positive and negative samples. w k represents the weight of the category, and the weight w k of the new category is set to The weight w k of the old category is set to 1, C t is the current number of categories, and C 1:t is the total number of all categories known up to the current task.
[0056] 7. Use the trained feature extraction module, feature purification module, unknown detection module, stability-plasticity classifier module, and recall enhancement module to classify the image to be classified.
[0057] Key points of the present invention:
[0058] 1. Reveal the challenges of learning objective conflicts in multi-label class incremental learning, and accommodate historical, current, and future knowledge by clarifying known and unknown knowledge
[0059] 2. To clarify known knowledge, a dynamic feature purification module is proposed, which focuses on the fine-grained category-aware features of known categories. A recall enhancement module with distribution priors is designed to effectively retain the known knowledge of old categories.
[0060] 3. To explore unknown knowledge, features related to future categories are mined and generated, and learned as unknown categories, improving the discriminative ability of the model.
[0061] Specific application scenarios of the present invention:
[0062] The multi-label class incremental learning method of the present invention has important application values in scenarios such as sentiment analysis and sentiment classification, semantic sentiment analysis, and medical diagnosis. In the text sentiment analysis task, multiple sentiment labels may appear in the same text, and over time, the scope of sentiment labels may gradually expand. Incremental learning can help the system continuously learn new sentiment categories without forgetting the previous sentiment labels. Specific scenarios include social media or online review analysis. With the emergence of emerging topics and sentiment labels, the model needs to continuously learn new sentiment types, such as "boring" sentiment, "satisfied" sentiment, etc., while maintaining accurate judgments of previous sentiment categories.
[0063] Effects of the present invention:
[0064] Experiments were conducted on two commonly used datasets, MS-COCO and PASCAL VOC, and various incremental settings to evaluate the effectiveness of HCP. The MS-COCO dataset contains 82,081 training images and 40,137 test images, with a total of 80 categories, and each image has an average of 2.9 category labels. The PASCAL VOC dataset contains 5,011 training images and 4,952 test images, with a total of 20 categories, and each image has an average of 1.6 category labels.
[0065] Table 1 shows the effect comparison between each module of the model, and the results prove that the new framework proposed by the present invention can bring obvious improvements. Table 2 shows the effect comparison between the present invention and other mainstream methods on the MS-COCO test dataset, and Table 3 shows the performance comparison between the present invention and other methods on the VOC test dataset. It can be seen that the present invention achieves the best performance in multiple incremental settings, proving the effectiveness of the present invention. Figure 2 The visualization of the attention weights of the present invention is shown, and it can be found that the present invention can focus on category-specific regions and alleviate the phenomenon of feature aliasing.
[0066] Comparison experiments of each module in Table 1
[0067]
[0068] Table 2 Comparison between HCP and Other Methods on MS-COCO Dataset
[0069]
[0070] Table 3 Comparison between HCP and Other Methods on VOC Dataset
[0071]
[0072] Another embodiment of the present invention provides a computer device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing each step in the method of the present invention.
[0073] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a magnetic disk, an optical disk). When the computer program stored in the computer-readable storage medium is executed by a computer, each step of the method of the present invention is implemented.
[0074] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention is subject to the scope defined by the claims.
Claims
1. A multi-label incremental learning method based on explicit known and unknown knowledge, characterized in that: The following steps are involved: Extract visual features from the input image; Utilize the category embeddings of historical and current categories and the extracted visual features to obtain fine-grained category-aware features; Synthesize unknown class features using class-aware features; The stability classifier is used to predict the probability of the old class, the plasticity classifier is used to predict the probability of the new class and the unknown class, and the old class probability and the new class and the unknown class probability are combined to perform image classification; Supplement the pseudo labels of the old categories for the current image based on the categories predicted by the historical model.
2. The method according to claim 1, characterized in that The method uses the category embeddings of the historical categories and the current category and the extracted visual features to obtain fine-grained category-aware features, including: assigning a learnable category embedding S to all the historical categories and the current category learned in the current task i , the embeddings of all known categories constitute the overall category embedding S = S 1:t-1 |S t , where S t is the category embedding being learned in the current task t, S 1:t-1 Represents the category embedding learned in the previous task; concatenates S with the visual features P of the image, performs attention calculation, aggregates object information and extracts fine-grained class-aware features O S .
3. The method according to claim 1, characterized in that The method of synthesizing unknown class features by using class-aware features is to generate unknown class features by interpolation using class-aware features.
4. The method according to claim 1, characterized in that The stability classifier remains frozen to retain old knowledge; each time the next incremental task is performed, the stability classifier and the plasticity classifier are merged into a new stability classifier, and a new plasticity classifier is attached for learning new classes.
5. The method according to claim 1, characterized in that The method of supplementing the pseudo-label of the old class for the current image according to the category predicted by the historical model includes: formulating the pseudo-label threshold of the class according to the predicted distribution of each category by the historical model; for each old class, if the predicted probability of the class by the current model is greater than or equal to the pseudo-label threshold, it is considered that the class exists in the image, and the binary label of the class is changed from 0 to 1 as supplementary supervision; if the predicted probability of the class by the current model is less than the pseudo-label threshold, it is considered that the class does not exist in the image, and the binary label of the class is maintained at 0.
6. The method according to claim 1, characterized in that Under the supervision of the true value label and the pseudo label, the weighted asymmetric loss function L is calculated. WASL to optimize training.
7. The method according to claim 6, characterized in that The weighted asymmetric loss function L WASL The formula is as follows: Where K is the total number of known categories, p k is the classification probability of the kth class; y k is the binary label of the kth class, 1 means that the class exists in the image, and 0 means that the class does not exist in the image; γ + and γ - Used to amplify the impact of positive and negative samples; w k Represents the weight of the category, the weight of the new category w k Set to The weight of the old class w k Set to 1, C t is the current category number, C 1:t is the number of all categories known up to the current task.
8. A multi-label incremental learning system based on explicit known and unknown knowledge, characterized in that: include: A feature extraction module is used to extract visual features from an input image; Feature purification module, which is used to obtain fine-grained class-aware features using the category embeddings of historical and current classes and the extracted visual features; The unknown detection module is used to synthesize unknown class features using class-aware features; A stability-plasticity classifier module is used to predict the probability of the old class using the stability classifier, predict the probability of the new class and the unknown class using the plasticity classifier, and combine the probability of the old class and the probability of the new class and the unknown class for image classification; The recall enhancement module is used to supplement the pseudo labels of old categories for the current image based on the categories predicted by the historical model.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Factory data anomaly detection method and system based on big data analysis
CN120744783A
Target incremental learning model training method and device, equipment and storage medium
CN121415192A
Class incremental learning method for high-adaptation fine tuning and pseudo sample playback under subspace integration
CN121480616A