Long-tail multi-label image classification method based on neural collapse
By adopting a neural collapse-based method in the multi-label image classification model, and using the label embedding and contrast learning technology of the isogonal tight framework, the problem of insufficient tail class recognition capability on the long-tail distribution data set is solved, significantly improving the prediction accuracy and model robustness.
Patent Information
- Application Number
- CN202411751632.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-12-02
AI Technical Summary
The multi-label image classification model performs poorly on the long-tail distribution dataset, especially the weak recognition ability of tail classes, resulting in a reduced prediction accuracy.
The long-tail multi-label image classification method based on neural collapse is adopted, and feature learning is guided through label embedding of the isogonal tight frame, and the features are collapsed into the class prototype using contrast learning, and combined in the binary classifier of the isogonal tight frame to enhance the generalization ability of the model.
The prediction accuracy of multi-label image classification is significantly improved, especially on the long-tail distribution dataset, which enhances the recognition ability of tail classes and improves the robustness and generalization ability of the model.
Smart Images

Figure CN119580003B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of multi-label image classification and relates to a long-tail multi-label image classification method based on neural collapse. Background Art
[0002] Multi-label image classification is an important task in computer vision, which aims to simultaneously identify multiple objects, attributes or actions in an image. Unlike traditional single-label classification, multi-label classification requires the model to be able to handle complex scenes in the image, which often involves learning multiple labels in the same image. To solve this task, researchers have introduced a variety of deep learning techniques, especially Convolutional Neural Networks (CNN). CNN networks have significantly improved the accuracy of multi-label classification through efficient feature extraction and classification mechanisms.
[0003] However, in practical applications, multi-label image classification faces the problem of long-tail distribution. Long-tail distribution means that in a dataset, the majority class (head class) has a large number of instances, while the minority class (tail class) has only a small number of instances. In the field of multi-label image classification, it is common for the head class and the tail class to appear simultaneously in an image. This unbalanced class distribution causes the model to tend to learn the features of the head class during training, and thus performs poorly in identifying the tail class.
[0004] This phenomenon is common in many real-world scenarios, such as object detection and image annotation, which results in a decrease in the generalization ability of the model, and in turn leads to a decrease in the prediction accuracy of the model when performing multi-label image classification. Summary of the invention
[0005] In response to the problems of long-tail distribution of labels, co-occurrence of head and tail classes, label dependence, etc. in long-tail multi-label learning, the present invention proposes a long-tail multi-label image classification method based on neural collapse to improve the model's attention to the tail classes, avoid the deviation of the model caused by unbalanced data sets, and maximize the generalization ability of the model on long-tail multi-label data sets, thereby improving the prediction accuracy of the model when performing multi-label image classification.
[0006] In order to achieve the above object, the present invention adopts the following technical scheme:
[0007] The long-tail multi-label image classification method based on neural collapse includes the following steps:
[0008] Step 1. Obtain a multi-label dataset with a long-tail distribution of labels ;
[0009] in, Represented by the image, i.e., the sample and tags The dataset consists of is the total number of samples, ;
[0010] Step 2. Build the deep learning model MOL, which includes the feature extractor , Projection layer and classifier ; Initialize label embedding of equiangular tight frame , time step And hyperparameters ;
[0011] Among them, the label is embedded , Indicates the number of tags, Indicates Tags are embedded, ;Will Samples in Input feature extractor In the above example, we can obtain high-dimensional features. ;
[0012] Step 3. Calculate label embedding using spatial attention mechanism and high-dimensional features The position weight between
[0013] Step 4. Use tag embedding and high-dimensional features The position weight between the samples is calculated Zhongyu Corresponding features ;
[0014] Step 5. Design alignment features With tag embedding The loss function If the sample Contains tags , then the alignment feature With tag embedding , so that the cosine similarity approaches 1, otherwise, the feature With tag embedding There are differences between With tag embedding The cosine similarity to , so that it presents an equiangular tight frame;
[0015] Step 6. Add features Input into the projection layer and normalized to obtain the normalized features after projection ;
[0016] Step 7. According to , calculate each class in the current Characteristic archetype of the moment , the loss is calculated using the contrastive learning method , so that the projection feature of each sample Collapse to Class prototype calculated at all times ;
[0017] Step 8. Initialize the binary classifier of the equiangular tight frame , the sample Features Combining features , and put it into the classifier In the logit , and then go through The function obtains the final prediction result;
[0018] Step 9. Design bidirectional loss , using the two-way loss Alleviate the imbalance problem of positive and negative samples in long-tail multi-label classification, improve the confidence of positive label prediction, and reduce the confidence of negative label prediction;
[0019] Step 10. Convert the loss function , and The weighted sum is used to obtain the final loss , back-propagation updates the parameters of the model MOL, and continues to iterate the training task until the training endpoint is reached;
[0020] The final loss , hyperparameters and is the weight coefficient corresponding to the loss term;
[0021] Step 11. Use the trained model MOL to perform multi-label prediction on the input image.
[0022] The present invention has the following advantages:
[0023] As shown above, the present invention relates to a long-tail multi-label image classification method based on neural collapse. The method guides the learning of different category features by using label embedding of an equiangular tight frame, so that the features of different categories meet the neural collapse property NC2. NC2 requires that the features between different categories be separated as much as possible in the embedding space, thereby enhancing the differences between categories, reducing the interference of the head class on the tail class, and ensuring the full learning of the tail class features. Inspired by the neural collapse property NC1, the present invention also adopts a contrastive learning method to cluster the features of each category and collapse them to the prototype of the class, thereby enhancing the robustness of the features. NC1 emphasizes the aggregation of sample features within the same category, ensuring the consistency within the category, thereby improving the discrimination ability of the model. Finally, influenced by the neural collapse property NC3, the present invention combines the sample features extracted by the model between different categories and inputs them into a binary classifier that conforms to the equiangular tight frame, making full use of the multidimensional features of the sample and avoiding the negative impact of the imbalance of the data set on the training process. NC3 emphasizes the effective combination of multi-category features and the synergy of the binary classifier, thereby improving the generalization ability of the model. The method of the present invention significantly improves the prediction accuracy of the multi-label image classification method, especially on the data set with long tail distribution, and can effectively improve the recognition ability of the tail class. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flow chart of a long-tail multi-label image classification method based on neural collapse in an embodiment of the present invention;
[0025] Figure 2 is a schematic diagram of long-tail multi-label data distribution in an embodiment of the present invention;
[0026] Figure 3 It is a structural diagram of the deep learning model constructed in an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0028] This embodiment describes a long-tail multi-label image classification method based on neural collapse, which aims to utilize the phenomenon of neural collapse in the multi-classification field, that is, the phenomenon of an equiangular tight frame that appears in the last layer of features and classifier vectors at the end of training of a balanced data set (when the training loss is zero). The method generally has the following steps: aligning the label semantic embedding of the equiangular tight frame with the category feature, and using the label semantic embedding of the equiangular tight frame to guide the feature learning of different categories; projecting the category feature to a low dimension, and calculating the feature prototypes of different categories; using the contrastive learning method, collapsing the projected category feature to the feature prototype and expanding the distance with the feature prototypes of other categories; combining the collapsed features and putting them into a classifier with an equiangular tight frame to obtain logit; adding the loss of the label semantic embedding and feature vector alignment, the contrastive learning loss with the feature prototype, and the classifier loss in proportion, back-propagating to update the model parameters, and iterating the training. The method of the present invention utilizes the idea of neural collapse to further solve the problems of category interdependence, co-occurrence of head and tail classes, and long-tail distribution of labels in long-tail multi-label datasets. It can effectively improve the accuracy and robustness of deep learning models, thereby improving the accuracy of multi-label image classification prediction.
[0029] like Figure 1 As shown, the long-tail multi-label image classification method based on neural collapse in this embodiment includes the following steps:
[0030] Step 1. Obtain a multi-label dataset with a long-tail distribution of labels ; Among them, multi-label dataset Represented by image and tags A dataset consisting of is the total number of samples.
[0031] Label Representative samples The label vector.
[0032] in, Representative samples Include tags ,on the contrary, , then it represents the sample Does not contain tags .
[0033] like Figure 2 As shown in Figure 1, a distribution diagram of long-tail multi-label data is shown. Figure 2 As can be seen in (a) and (c), there is often a co-occurrence phenomenon of head classes and tail classes in long-tail multi-label datasets, e.g. Figure 2 (a) "man" and "horse" and Figure 2 (c) "car" and "sheep". In addition, Figure 2(a) and (b) also show the dependency between labels, for example, "people" usually ride on "horses" and "cats" usually lie on "sofas". Therefore, there is usually a certain dependency between the labels of "horse" and "people" and between the labels of "cat" and "sofa". By observing the distribution of the number of labels, it can be found that the labels present a long-tail distribution, with the head class samples occupying the vast majority, while the number of tail class labels is large, but the samples of each class are extremely scarce. Long-tail distribution is widely present in many practical scenarios. The long-tail data problem often leads to the model's weak classification ability for the tail class, which requires special strategies in model design to ensure effective learning and classification of the tail class.
[0034] Step 2. Build the deep learning model MOL, which includes the feature extractor , Projection layer and classifier ,like Figure 3 Initialize label embedding of equiangular tight frame , time step , and hyperparameters .
[0035] Among them, the label is embedded , Indicates the number of tags, Indicates Tags are embedded, .
[0036] And the multi-label dataset Samples in Input feature extractor In the above example, we can obtain high-dimensional features. .
[0037] The specific steps of step 2 are:
[0038] Step 2.1. According to the formula Initialize label embeddings that satisfy an equiangular tight frame .
[0039] in is an orthogonal matrix, represents the dimension of the vector, are the identity matrix and the matrix of all 1s respectively.
[0040] For any unequal integers ,satisfy is a unit vector, and .
[0041] Step 2.2. Using feature extractor , for example, ResNet50 extracts samples Features .
[0042] in , are the width, height and number of channels of the feature map respectively.
[0043] In order to obtain the spatial information of the feature map, the feature Decompose into , .
[0044] Step 3. Calculate label embedding using spatial attention mechanism and high-dimensional features The position weight between .
[0045] The specific steps of step 3 are:
[0046] Step 3.1. Extract the sample Spatial characteristics With tag embedding Calculate the cosine similarity and get:
[0047] ;
[0048] Step 3.2. Based on cosine similarity And SoftMax function, calculate the weights of different label embeddings for different spatial positions of features , the formula is as follows:
[0049] .
[0050] Step 4. Use tag embedding and high-dimensional features The position weights between samples are calculated Zhongyu Corresponding features .
[0051] Specifically, using label embedding and the eigenvector The weights of the samples are obtained by weighting In the example, tags are embedded in Corresponding features .
[0052] Step 5. Design alignment features With tag embedding The loss function If the sample Contains tags , then the alignment feature With tag embedding , so that the cosine similarity approaches 1, otherwise, the feature With tag embedding There are differences between With tag embedding The cosine similarity to , so that it presents an isometric tight frame.
[0053] Specifically, the loss function The calculation formula is as follows:
[0054] ;
[0055] in, Representation characteristics With tag embedding The cosine similarity of .
[0056] If the sample Contains tags ,Right now , then the loss function Alignment Features and tag embedding , so that its cosine similarity tends to 1; otherwise, it makes its cosine similarity tend to .
[0057] pass , , ,therefore, , , because the tag embeds ,So .
[0058] This makes the features of the classes extracted by the feature extractor present an equiangular tight framework, fully ensuring the independent learning of features of all categories and the differences in features between classes, thereby better ensuring the feature extraction of the tail class.
[0059] Step 6. Add features Input to projection layer After normalization, the normalized features after projection are obtained .
[0060] The step 6 is specifically as follows:
[0061] Step 6.1. Through the projection function The characteristics Project to low latitude and get features :
[0062] ;
[0063] in, are the parameters of the model projection layer.
[0064] Step 6.2. Projected features Normalize and get the normalized features :
[0065] ;
[0066] in , is the dimension after projection.
[0067] Step 7. According to , calculate each class in the current Characteristic archetype of the moment , the loss is calculated using the contrastive learning method , so that the projection feature of each sample Collapse to Class prototype calculated at all times .
[0068] The step 7 is specifically as follows:
[0069] Step 7.1. Based on the normalized features after projection , calculate the feature prototype at the current time t :
[0070] ;
[0071] in Contains tags The number of samples, is the indicator function, Indicated in Moment Sample Features ;
[0072] Step 7.2. Design a loss function based on the contrastive learning loss NT-Xent :
[0073] ;
[0074] in It is a sample exist The feature corresponding to the category c after the normalization of the moment projection with Class prototype of moment The cosine similarity of .
[0075] Representation characteristics With class prototype The cosine similarity of , All are temperature parameters.
[0076] Through loss , if the sample Contains categories , then it makes Collapse to Class Prototype , that is, to improve the cosine similarity of the two, and at the same time make Stay away from other prototypes , that is, reducing the cosine similarity.
[0077] Step 8. Initialize the binary classifier of the equiangular tight frame , the feature Combining features And put it into the classifier In the example, the logit of each class is , and then go through The function obtains the final prediction result.
[0078] The specific steps of step 8 are:
[0079] Step 8.1. According to the formula Initialize the binary classifier that satisfies the equiangular tight frame ;
[0080] ;
[0081] in, , , is the dimension after projection.
[0082] Step 8.2. Transfer the sample Features Combine to get the final features of the sample :
[0083] , , Indicates a connection.
[0084] Step 8.3. Add features Do the inner product with the classifier vector to get the logit of each class ;
[0085] ;
[0086] in, For the model to sample No. logit of class prediction, is a trainable scaling parameter, is the classifier vector.
[0087] Afterwards, Input to In the function, the final binary classification prediction result of the model is obtained.
[0088] Step 9. Design bidirectional loss , using the two-way loss Alleviate the imbalance problem of positive and negative samples in long-tail multi-label classification, improve the confidence of positive label prediction, and reduce the confidence of negative label prediction.
[0089] To address the imbalance of positive and negative samples in multi-label samples, loss functions are designed from the perspective of samples and labels. : .
[0090] in, is the number of samples in a batch in each model training;
[0091] ;
[0092] ;
[0093] in, is the temperature parameter, ,and , Represents the index of two different samples in a batch, represents the activation function, , Represents logarithm.
[0094] Step 10. Convert the loss function , and The weighted sum is used to obtain the final loss , back-propagation updates the model parameters and continues to iterate the training task until the training endpoint is reached.
[0095] Final loss ,in, and is the weight coefficient corresponding to the loss term.
[0096] Step 11. Use the trained model MOL to perform multi-label prediction on the input image.
[0097] The method of the present invention utilizes the idea of neural collapse to effectively solve the problems of category interdependence, co-occurrence of head and tail classes, and long-tail distribution of labels in long-tail multi-label data sets, and can effectively improve the accuracy and robustness of deep learning models.
[0098] Of course, the above description is only a preferred embodiment of the present invention, and the present invention is not limited to the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any technician familiar with the field under the guidance of this specification fall within the essential scope of this specification and should be protected by the present invention.
Claims
1. A long-tail multi-label image classification method based on neural collapse, characterized in that: The steps include: Step 1. Obtain a multi-label dataset with a long-tail distribution of labels ; in, Represented by the image, i.e., the sample and tags The dataset consists of is the total number of samples, ; Step 2. Build the deep learning model MOL, which includes the feature extractor , Projection layer and classifier ; Initialize label embedding of equiangular tight frame , time step And hyperparameters ; in, , Indicates the number of tags, Indicates Tags are embedded, ; Will Samples in Input feature extractor In the above example, we can obtain high-dimensional features. ; Step 3. Calculate label embedding using spatial attention mechanism and high-dimensional features The position weight between Step 4. Use tag embedding and high-dimensional features The position weight between the samples is calculated Zhongyu Corresponding features ; Step 5. Design alignment features With tag embedding The loss function If the sample Contains tags , then the alignment feature With tag embedding , so that the cosine similarity approaches 1, otherwise, the feature With tag embedding There are differences between With tag embedding The cosine similarity to , so that it presents an equiangular tight frame; Step 6. Add features Input into the projection layer and normalized to obtain the normalized features after projection ; Step 7. According to , calculate each class in the current Characteristic archetype of the moment , the loss is calculated using the contrastive learning method , so that the projection feature of each sample Collapse to Class prototype calculated at all times ; Step 8. Initialize the binary classifier of the equiangular tight frame , the sample Features Combining features , and put it into the classifier In the logit , and then go through The function obtains the final prediction result; Step 9. Design bidirectional loss , using the two-way loss Alleviate the imbalance problem of positive and negative samples in long-tail multi-label classification, improve the confidence of positive label prediction, and reduce the confidence of negative label prediction; Step 10. Convert the loss function , and The weighted sum is used to obtain the final loss , back-propagation updates the parameters of the model MOL, and continues to iterate the training task until the training endpoint is reached; The final loss , hyperparameters and is the weight coefficient corresponding to the loss term; Step 11. Use the trained model MOL to perform multi-label prediction on the input image.
2. The long-tail multi-label image classification method based on neural collapse according to claim 1 is characterized in that: In the step 1, Representative samples The label vector of ; in, Representative samples Include tags ,on the contrary, , then it represents the sample Does not contain tags .
3. The long-tail multi-label image classification method based on neural collapse according to claim 1 is characterized in that: The step 2 is specifically as follows: Step 2.
1. According to the formula Initialize label embeddings that satisfy the equiangular tight frame ; in is an orthogonal matrix, represents the dimension of the vector, are the identity matrix and the all-1 matrix respectively; For any unequal integers ,satisfy is a unit vector, and ; Step 2.
2. Using feature extractor , extract samples Features ; in , are the width, height and number of channels of the feature map respectively; In order to obtain the spatial information of the feature map, the feature Decompose into , .
4. The long-tail multi-label image classification method based on neural collapse according to claim 3 is characterized in that: The step 3 is specifically as follows: Step 3.
1. Extract the sample Spatial characteristics With tag embedding Calculate the cosine similarity and get: ; Step 3.
2. Based on cosine similarity And SoftMax function, calculate the weights of different label embeddings for different spatial positions of features , the formula is as follows: 。 5. The long-tail multi-label image classification method based on neural collapse according to claim 4 is characterized in that: In step 4, using label embedding and the eigenvector The weights of Zhongyu Corresponding features: 。 6. The long-tail multi-label image classification method based on neural collapse according to claim 1, characterized in that: In step 5, the loss function The calculation formula is as follows: ; in, Representation characteristics With tag embedding The cosine similarity of If the sample Contains tags ,Right now , then the loss function Alignment Features and tag embedding , so that its cosine similarity tends to 1; otherwise, it makes its cosine similarity tend to .
7. The long-tail multi-label image classification method based on neural collapse according to claim 1, characterized in that: The step 6 is specifically as follows: Step 6.
1. Through the projection function The characteristics Project to low latitude and get features : ; in, are the parameters of the model projection layer; Step 6.
2. Projected features Normalize and get the normalized features : ; in , is the dimension after projection.
8. The long-tail multi-label image classification method based on neural collapse according to claim 1, characterized in that: The step 7 is specifically as follows: Step 7.
1. Based on the normalized features after projection , calculate the feature prototype at the current time t : ; in Contains tags The number of samples, is the indicator function, Indicated in Moment Sample Features ; Step 7.
2. Design a loss function based on the contrastive learning loss NT-Xent : ; in It is a sample exist The feature corresponding to the category c after the normalization of the moment projection with Class prototype of moment The cosine similarity of Representation characteristics With class prototype The cosine similarity of , All are temperature parameters; Through loss , if the sample Contains categories , then it makes Collapse to Class Prototype , that is, to improve the cosine similarity of the two, and at the same time make Stay away from other prototypes , that is, reducing the cosine similarity.
9. The long-tail multi-label image classification method based on neural collapse according to claim 1, characterized in that: The step 8 is specifically as follows: Step 8.
1. According to the formula Initialize the binary classifier that satisfies the equiangular tight frame ; ; in, , , It is the dimension after projection; Step 8.
2. Transfer the sample Features Combine to get the final features of the sample : , , Indicates connection; Step 8.
3. Add features Do the inner product with the classifier vector to get the logit of each class ; ; in For the model to sample No. logit of class prediction, is a trainable scaling parameter, is the classifier vector; Afterwards, Input to In the function, the final binary classification prediction result of the model is obtained.
10. The long-tail multi-label image classification method based on neural collapse according to claim 9, characterized in that: In step 9, the loss function is designed from the perspective of samples and labels respectively. : ; in, is the number of samples in a batch in each model training; ; ; in, is the temperature parameter, ,and , Represents the index of two different samples in a batch, represents the activation function, , Represents logarithm.
Citation Information
Patent Citations
Semi-supervised generative adversarial network image classification method based on local manifold regularization
CN111046900A
Multi-modal data-based heavy balance long-tail image data classification method
CN115205592A