Class Imbalance Multi-Label Image Classification Method and System Based on Sample Representativeness

By using a dynamic loss function representative sample in multi-label image classification, the problem of category imbalance is solved, the classification accuracy of tail categories is improved, and the better multi-label image classification effect is achieved.

CN115984607BActive Publication Date: 2025-07-18SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211554147.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-06
Publication Date
2025-07-18
Estimated Expiration
2042-12-06

AI Technical Summary

Technical Problem

In the multi-label image classification with unbalanced categories, the model is prone to biasing the head class, resulting in underfitting the tail class, and the existing methods fail to effectively utilize the representativeness of the tail class samples, resulting in poor classification results.

Method used

A dynamic loss function representative of the sample is adopted. By combining the classification weights and dynamic focal loss, taking into account the correlation between categories, positive and negative weighting of categories are classified, and a representative coordination function is designed to enhance the weight of the tail class samples and dynamically adjust the γ parameter of the loss function.

Benefits of technology

It effectively solves the problem of category imbalance and improves the accuracy of multi-label image classification, especially in the tail category, and the experimental effect is significantly better than the existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984607B_ABST
    Figure CN115984607B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for class-imbalanced multi-label image classification based on sample representativeness. In the method, the dynamic loss of sample representativeness is composed of a classification weight and a dynamic focal loss. The classification weight is obtained by inputting the co-occurrence rate of the current classification category and the labels of other categories of the sample and the number of categories into a representativeness coordination function. The dynamic focal loss is obtained by combining the logits output by the classifier and the parameters calculated for each sample for each category by the classification weight. This method takes into account the correlation between categories, conducts positive and negative weighted classification discussions on categories, realizes a more reasonable weighted design for negative categories, emphasizes the representativeness of samples for categories, and is used to deal with some difficult samples with a large number of categories, effectively solving the problem of class imbalance existing in the dataset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, relates to the field of computer vision processing, and mainly relates to a class-imbalanced multi-label image classification method and system based on sample representativeness. Background Art

[0002] Multi-label image recognition, as a traditional problem where each image needs to be recognized with multiple labels, has always been a popular research topic in the industry. To keep up with the trend of the times, there have been many new changes in multi-label classification in recent years: by combining emerging technologies with more complex models, such as graph neural networks (GNNs), attention mechanisms, transformer models, and curriculum learning, this field has reached a new depth; by exploring some more specific sub-tasks, such as partial multi-label classification (PML), multi-label classification with missing labels (MLML), and semi-supervised multi-label classification, this field has reached a new breadth. And class-imbalanced multi-label image classification, as a newly emerging multi-label sub-task in the past two years, is very challenging and practically significant.

[0003] The class imbalance problem, simply put, is that the number of positive samples in each class in the training set varies greatly, with only a small number of classes having a large number of samples, while other classes only account for a small number of samples. Among them, the classes with a large number of samples are called head classes, and conversely, the classes with a small number of samples are called tail classes. When affected by class imbalance, network training is often dominated by negative labels. This imbalance between classes and the number of samples will cause the model to be biased towards head classes during training and is prone to underfitting in tail classes, ultimately leading to a decline in training results.

[0004] Existing methods usually adopt resampling and reweighting to balance the distribution of different classes, but due to label co-occurrence, most of them cannot be applied to class-imbalanced multi-label classification. Label co-occurrence means that an image contains multiple labels, and these labels may have strong correlations between classes. In the class imbalance problem, resampling these multi-label images may not necessarily make the distribution of the dataset more balanced, and may even lead to intra-class imbalance. Most reweighting methods are only applicable to single-label classification. Although some improved versions are used for class-imbalanced multi-label classification, these methods still underestimate the impact of label co-occurrence, resulting in limited improvement in multi-label image classification.

[0005] In the class imbalance problem, some tail classes always appear together with their parent classes or similar classes. For example, images containing the tail class "apple" are often labeled as the head class "fruit", and the label "fruit" may help the model predict the tail class "apple". Therefore, the few samples of the tail class "apple" become more representative of their tail class and guide the classifier to make correct judgments. To make full use of the few samples of the tail class, it is necessary to assign more weights to the representative samples when calculating the loss. Nevertheless, previous methods often ignored these samples, resulting in poor performance on the tail class, and there are loopholes in the design of some existing loss functions. Their reweighting methods have no explanation for negative classes and cannot handle some difficult samples with a large number of classes. Therefore, it is very necessary to design a new loss function to solve this problem. Summary of the Invention

[0006] The present invention precisely aims at the problems existing in the prior art and provides a method and system for class-imbalanced multi-label image classification based on sample representativeness. In the method, the dynamic loss of sample representativeness is composed of a classification weight and a dynamic focal loss. The classification weight is calculated by inputting the co-occurrence rate of the current classification category and the labels of other categories of the sample and the number of classes into a representativeness coordination function. The dynamic focal loss is obtained by combining the logits output by the classifier and the parameters calculated for each sample for each category by the classification weight. This method considers the correlation between categories, conducts positive and negative weighted classification discussions on categories, realizes a more reasonable weighted design for negative categories, emphasizes the representativeness of samples for categories, and copes with some difficult samples with a large number of classes, effectively solving the class imbalance problem existing in the dataset.

[0007] To achieve the above object, the technical solution adopted by the present invention is: a method for class-imbalanced multi-label image classification based on sample representativeness. In this method, the dynamic loss of sample representativeness is composed of a classification weight and a dynamic focal loss. The classification weight is calculated by inputting the co-occurrence rate of the current classification category and the labels of other categories of the sample and the number of classes into a representativeness coordination function. The dynamic focal loss is obtained by combining the logits output by the classifier and the parameters calculated for each sample for each category by the classification weight.

[0008] As an improvement of the present invention, when the current classification category is a positive category for the sample, the representativeness coordination function is a decreasing function, specifically including a linear function, a sigmoid function, a cosine function or a reciprocal form function;

[0009] When the current classification category is a negative category for the sample, the representativeness coordination function is an increasing function, specifically including a linear function, a sigmoid function, a logarithmic function or an exponential function.

[0010] As an improvement of the present invention, the class-imbalanced multi-label image classification method based on sample representativeness includes the following steps:

[0011] S1: For the input class-imbalanced image training set, given the number of classes and the number of samples, count the number of samples in each class and the label co-occurrence rate r between classes ij ;

[0012] S2: Perform data augmentation on the input image training set;

[0013] S3: After dividing the training set enhanced in step S2 into multiple batches according to the quantity B, successively input it into the feature extractor to obtain feature maps, and input the feature maps into the classifier to obtain logits. The classifier outputs a logit for each class of each sample;

[0014] S4: Calculate the representative coordination function. The input of the function is the co-occurrence rate r between the current classification category i and the labels of the sample ij , specifically as follows:

[0015]

[0016] where a, b, c, d, e, f are all parameters;

[0017] S5: Calculate the reweighted form ER that emphasizes sample representativeness. The calculation form of the weight of the i-th class of the k-th sample is specifically as follows:

[0018]

[0019]

[0020] where n i represents the number of samples in class i; Counts all positive classes in the k-th sample, and uses class j to represent a certain positive class; n j is the number of samples corresponding to these positive classes; r ij is the label co-occurrence rate between class i and the positive class j of the sample; f(r ij ) is the representative coordination function obtained in step S4; is the result after smooth mapping of the weight ; α, β, μ are all smoothing parameters;

[0021] S6, Calculate the dynamic focal loss DF. The dynamic focal loss is divided into two parts l + and l - according to the positive and negative of the current classification category i, and the specific form is as follows:

[0022]

[0023]

[0024] Among them, represents the result after the classifier output passes through the sigmoid function σ; is the γ parameter in dynamic form, which receives the weight obtained from step S5 According to whether the category i is a positive category or a negative category, the value ranges of are controlled between (0.5, 1) and (3, 5) respectively;

[0025] S7. Loss calculation. Taking the k-th sample x k as an example, its loss function is as follows:

[0026]

[0027] Among them, there are a total of C categories; y k represents the label information of the sample x k ; is the weight of the k-th sample for the i-th category; l + and l - represent the dynamic focal loss when i is a positive category and a negative category respectively; the final loss form is the sum of the losses of all samples:

[0028]

[0029] S8. Perform backpropagation on the network according to the final loss obtained in step S7, update the model parameters, and cycle E epochs to form an optimal model;

[0030] S9. Input the data into the optimal model obtained after the update cycle in step S8 to achieve multi-label image classification.

[0031] As another improvement of the present invention, in step S1, the number of categories is C, the number of samples is K, and the number of samples in each category is denoted as Then the label co-occurrence rate between categories is denoted as Its calculation formula is r ij = n i∩j / n i , where the label information of the k-th sample is denoted as

[0032] As another improvement of the present invention, the image data augmentation in step S2 includes at least random cropping and resizing the image to 224×224, random Gaussian blur, and color distortion processing, and the batch size B is set to 32.

[0033] As another improvement of the present invention, in the step S6 The γ parameter in dynamic form has different design forms for the positive or negative categories of class i. The specific design forms are as follows:

[0034]

[0035] As another improvement of the present invention, in the step S8, during the backpropagation process, the epoch was trained for 10 rounds in total. The SGD with a momentum of 0.9 and a weight decay exponent of 1×10 -4 was used as the optimizer, the learning rate was initialized to 0.02, and decayed at 5 to 7 epochs.

[0036] To achieve the above object, the technical solution adopted by the present invention is also: a class-imbalanced multi-label image classification system based on sample representativeness, including a computer program, and when the computer program is executed by a processor, it implements the steps of any one of the above methods.

[0037] Compared with the prior art, the present invention has the following beneficial effects: It provides a class-imbalanced multi-label image classification method and system based on sample representativeness. When calculating the loss, the correlation between classes is further considered, and the positive and negative weighted classification of classes is discussed, emphasizing the representativeness of samples for classes. This method can better handle those difficult samples with a large number of labels, and balance the difference in the size of positive and negative losses during training, effectively solving the class-imbalance problem existing in multi-label data sets. Compared with existing methods, the experimental results (mAP index) on multiple class-imbalanced multi-label image data sets have been significantly improved. In addition, this method does not additionally affect the structure of the model, nor does it increase the training time and computational overhead of the model, and has more practical value. And it is simple and easy to implement, and can be more conveniently combined with other methods to achieve a synergistic effect. Description of the Drawings

[0038] Figure 1 is the step flow chart of the method of the present invention;

[0039] Figure 2 is the distribution of three data sets in Embodiment 1 of the present invention;

[0040] Figure 3 is the heat map of the label co-occurrence rate of three data sets in Embodiment 1 of the present invention;

[0041] Figure 4 is the comparison result diagram of the mAP index of the ER-DF of this method and other methods on the COCO-MLT-1 and VOC-MLT data sets;

[0042] Figure 5 This figure shows the comparison results of the mAP metric between the ER-DF of this method and other methods on the COCO-MLT dataset with two different imbalance rates.

[0043] Figure 6 This figure shows the specific classification effect comparison between the ER-DF of this method and other methods on difficult samples. Detailed implementation manners

[0044] The present invention will be further illustrated below in conjunction with the accompanying drawings and specific implementation manners. It should be understood that the following specific implementation manners are only used to illustrate the present invention and not to limit the scope of the present invention.

[0045] Example 1

[0046] A multi-label image classification method for class imbalance based on sample representativeness. Through this method, the model can better handle the multi-label image classification problem with class imbalance, more effectively utilize the label dependence in the problem, pay attention to the representativeness of samples, and better handle some difficult samples.

[0047] The method in this case emphasizes the dynamic loss of sample representativeness (ERDF). ERDF is divided into two parts: (1) The reweighting form that focuses on sample representativeness (ER). This part receives the number of class samples and the label co-occurrence rate statistically obtained in the training set, and separately designs a set of functions to reflect the representativeness of samples for the current class: the representativeness coordination function. Its input is the label co-occurrence rate between the current classification class and other classes of the sample, and then the output of the function is combined with the number of classes, so as to design different classification weights for each sample for each class; (2) Dynamic focal loss (DF). This part receives the logits output by the classifier and the previously obtained weights, calculates different γ parameters for each sample for each class using the weights, then uses this set of dynamic γ parameters to replace the fixed γ parameter in the traditional focal loss, and combines the logits to obtain the dynamic loss. Finally, the weights and the dynamic loss are combined to obtain ERDF. This method has achieved good results on multiple long-tail versions of multi-label image datasets, effectively solving the class imbalance problem existing in the dataset. The step process of this method is as Figure 1 shown:

[0048] Step S1: For the input class-imbalanced image training set, given C classes, K samples, and the label information of each sample. Specifically, the label information of the kth sample is denoted as Count the number of samples in each class, denoted as Its calculation formula is Count the label co-occurrence rate between classes, denoted as Its calculation formula is r ij =ni∩j / n i (n i∩j (representing the number of samples that simultaneously include categories i and j).

[0049] In this embodiment, the class-imbalanced image training sets are imbalanced subsets collected from MS-COCO and PASCAL VOC respectively, with a total of three imbalanced data sets, namely COCO-MLT-1, COCO-MLT-2, and VOC-MLT. Figure 2 The sample distribution of 3 data sets is shown. The imbalance rates of the three imbalanced data sets COCO-MLT-1, COCO-MLT-2, and VOC-MLT are 188, 153, and 193.75 respectively. The calculation formula for the imbalance rate is ρ = max i (n i ) / min i (n i ), that is, dividing the number of samples of the most numerous category in the data set by the number of samples of the least numerous category. The COCO data set has 80 categories, and VOC has only 20 categories. COCO-MLT-2 has more samples than COCO-MLT-1 and is relatively balanced.

[0050] Figure 3 The label co-occurrence rate of 3 data sets is shown. The dependence between labels in the two COCO-MLT data sets is more obvious, while the dependence between labels within the VOC-MLT data set varies significantly. In addition, each data set also divides the categories into head categories, middle categories, and tail categories according to the distribution for subsequent result display.

[0051] Step S2: For the input image training set, perform data augmentation processing; the data augmentation of the image includes randomly cropping and resizing to a size of 224×224, random Gaussian blur, color distortion, etc. The batch size B is set to 32.

[0052] Step S3: After dividing the training set into multiple batches of size B, continuously input it into the feature extractor to obtain a feature map, and then input the feature map into the classifier to obtain logits. The classifier outputs a logit for each category of each sample. Specifically, the logit for the i-th category of the k-th sample is denoted as The feature extractor includes neural networks such as Resnet-50, Resnet-101, CNN, and VGG, and the classifier is C binary classifiers.

[0053] Step S4: Calculate using the representative coordination function f(r ij ). When calculating the loss of the i-th category for the k-th sample, its input is the label co-occurrence rate r between category i and the positive category j to which the sample belongs.ij The representative coordination function designs different forms according to whether the sample k belongs to the positive or negative class for category i, as follows:

[0054]

[0055] Among them, a, b, c, d, e, and f are all parameters.

[0056] The design form of the representative coordination function is not limited to the above forms, but it must satisfy that the positive class part is a decreasing function and the negative class part is an increasing function. From the perspective of the convexity and concavity of the function, the positive class part functions include linear functions, sigmoid functions, cosine functions, and reciprocal form functions; the negative class part functions include linear functions, sigmoid functions, logarithmic functions, and exponential functions. Among them, the form with the positive class being in reciprocal form and the negative class being in logarithmic form has the best effect.

[0057] Step S5: Calculate the reweighted form ER that emphasizes sample representativeness. The weight of the k-th sample for the i-th category is calculated as follows:

[0058]

[0059]

[0060] where n i represents the number of samples in category i; represents all the positive classes of the k-th sample; n j is the number of samples corresponding to these positive classes; f(r ij ) is the representative coordination function obtained in step S4, and its input is the label co-occurrence rate r ij ; is the result after the weight is smoothly mapped by the above formula, and α, β, and μ are all smoothing parameters.

[0061] In this step weight design scheme, the purpose of weighting is to increase the weight of representative samples and decrease the weight of non-representative samples, which is mainly reflected in the denominator part of the weight. The denominator part is obtained by summing the products of multiple representative coordination functions and the reciprocals of the number of positive labels of the samples. For each positive label of sample k, there is one term in the denominator. If the number of labels of sample k is too large, it means that the images are mixed and not representative, the denominator becomes larger, and the weight is reduced. If class i is a positive class and has a high co-occurrence rate with the positive label of the sample, it means that the two are highly similar, and the positive label of the sample will not affect the classification of class i. Therefore, the sample is a representative sample of class i, and the weight is increased. So, at this time, the representative coordination function needs to be a decreasing function. If class i is a negative class and has a high co-occurrence rate with the positive label of the sample, it means that the two are highly similar, and the sample has a high correlation with this negative class. Even class i may be a misclassified class. Therefore, the sample is not a representative sample of class i, and the weight is reduced. So, at this time, the representative coordination function needs to be an increasing function.

[0062] Step S6: Calculate the dynamic focal loss DF. The loss is divided into two parts l + and l - according to whether class i is positive or negative, and the specific form is as follows:

[0063]

[0064]

[0065] Among them, represents the result after the output of the classifier passes through the sigmoid function σ. is the γ parameter in dynamic form, which receives the weight obtained from step S5 and also has different design forms for class i being positive or negative class. The specific design forms are as follows:

[0066]

[0067] In the original focal loss, the γ parameter can be regarded as a restriction on the output logits. Therefore, the method in this case hopes to increase the restriction on the logits of non-representative samples and decrease the restriction on the logits of representative samples. Therefore, it directly borrows the weight obtained in step S5 and controls the value ranges of positive and negative classes to be between (0.5, 1) and (3, 5) respectively.

[0068] Step S7: Calculate the loss. Taking the k-th sample x k as an example, its loss function is as follows:

[0069]

[0070] Among them, there are a total of C classes; yk The label information of the representative sample x k ; is the weight of the k-th sample for the i-th category. For details, see Step 5; l + and l - respectively represent the dynamic focal loss when i is the positive class and the negative class. For details, see Step S6. The final loss form is the sum of the losses of all samples:

[0071]

[0072] Step S8: Perform backpropagation on the network based on the final loss obtained in Step S7, update the model parameters, and loop for E epochs to form the optimal model. There are a total of 10 epochs in the backpropagation process. SGD with a momentum of 0.9 and a weight decay exponent of 1×10 -4 is used as the optimizer. The learning rate is initialized to 0.02 and decays at 5 to 7 epochs.

[0073] Step S9: Input the data into the optimal model obtained after the update loop in Step S8 to achieve multi-label image classification.

[0074] Test and verification

[0075] Compare and test the optimal model obtained according to the method of this case with existing multi-label classification algorithms and class imbalance algorithms on the dataset. The comparison methods include:

[0076] (1) RW (Re-weighting): The weight is inversely proportional to the square root of the number of class instances;

[0077] (2) RS (Re-sampling);

[0078] (3) ML-GCN: A multi-label classification method based on the graph convolutional network (GCN);

[0079] (4) OLTR: It sets a new benchmark for class-imbalanced single-label classification and is modified in this experiment to be applicable to multi-label classification;

[0080] (5) Focal: γ is fixed at 2;

[0081] (6) ASL: Set two different γ for the positive and negative class losses based on the Focal loss + = 0, γ - = 4 (but it is a fixed value, not the dynamic form in this paper);

[0082] (7) CB: A re-weighting method based on the number of valid samples per class;

[0083] (8) LDAM: A reweighting method that pushes the classification boundary to the head classes;

[0084] (9) DB: It introduces rebalancing weighting and suppresses the negative sample gradient, but there is no explanation for the negative class weighting and it does not recognize the importance of representative samples, performing poorly in the tail classes. In addition, we also conducted combined tests, such as combining DB and Focal to form DB-Focal.

[0085] The comparative verification of the above-mentioned numerous methods is attached as Figures 4 - 6 shown. Figure 4 It shows the comparison results of the mAP metric of the proposed method ER-DF and other methods on the VOC-MLT and COCO-ML-1 datasets. It can be seen that the proposed method ER-DF and the two sub-methods DF and ER-Focal in this paper have higher mAP metrics compared to other methods. Figure 5 It is the comparison result of the mAP metric of the proposed method ER-DF and other methods on the COCO-MLT dataset with two different imbalance rates. Since COCO-MLT-2 is more balanced, the overall performance of each method is better, and the proposed method ER-DF and the two sub-methods DF and ER-Focal in this paper also have higher mAP metrics compared to other methods. Figure 6 It shows the specific comparison of the classification effects of the proposed method ER-DF and other methods on difficult samples. It can be seen that the proposed method ER-DF can predict more correct classes compared to other methods and has better performance.

[0086] In summary, this case provides a class imbalance multi-label image classification method based on sample representativeness, further considering the correlation between classes when calculating the loss, discussing the positive and negative weighting classification of classes, and emphasizing the representativeness of samples for classes. This method can better handle those difficult samples with a large number of labels, effectively solve the class imbalance problem existing in multi-label datasets, and has a significant improvement in the experimental results (mAP metric) on multiple class imbalance multi-label image datasets.

[0087] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.

Claims

1. A class-imbalanced multi-label image classification method based on sample representativeness, characterized in that: In this method, the dynamic loss of sample representativeness is composed of the combination of classification weights and dynamic focal loss. The classification weights are obtained by inputting the co-occurrence rate of the current classification category and the labels of other categories of the sample and the number of categories into the representativeness coordination function. The dynamic focal loss is obtained by combining the logits output by the classifier and the parameters calculated for each sample for each category by the classification weights. The specific steps are as follows: S1: For the input imbalanced image training set, given the number of categories and the number of samples, count the number of samples in each category and the label co-occurrence rate r between categories ij ; S2: Perform data augmentation processing on the input image training set; S3: After dividing the training set enhanced in step S2 into multiple batches of size B, successively input it into the feature extractor to obtain feature maps, and input the feature maps into the classifier to obtain logits. The classifier outputs a logit for each category of each sample; S4: Representative coordination function calculation. The input of the function is the co-occurrence rate r of the current classification category i and the label to which the sample belongs ij , which is specifically as follows: Among them, a, b, c, d, e, and f are all parameters; S5: Calculate the reweighted form ER that emphasizes sample representativeness. The calculation formula for the weight of the \(i\)-th category of the \(k\)-th sample is as follows: Specifically, it is as follows: Among them, n i represents the number of samples of category i; Counts all positive categories in the k-th sample, and uses category j to represent a certain positive category; n j is the number of samples corresponding to the positive category; r ij is the co-occurrence rate of the labels of category i and the positive category j of the sample; f(r ij ) is the representative coordination function obtained in step S4; is the weight The result after smooth mapping; ɑ, β, μ are all smooth parameters; S6, Dynamic Focal Loss DF calculation. The dynamic focal loss is divided into two parts \(l\) and \(l\) according to the positive and negative of the current classification category \(i\). The specific form is as follows: + and \(l\) - ​ Among them, represents the result after the output of the classifier passes through the sigmoid function σ; is the γ parameter in dynamic form, which receives the weight obtained from step S5 According to whether the class i is a positive class or a negative class, the value ranges of are respectively controlled between (0.5, 1) and (3, 5); S7, Loss calculation, taking the k-th sample x k as an example, its loss function is as follows: Among them, there are a total of C categories; y k represents the label information of the sample x k ; is the weight of the k-th sample for the i-th category; l + and l - represent the dynamic focal losses for the positive and negative categories of i respectively; the final loss form is the sum of the losses of all samples: S8: Perform backpropagation on the network according to the final loss obtained in step S7, update the model parameters, and cycle for E epochs to form an optimal model; S9: Input the data into the optimal model obtained after the update cycle in step S8 to implement multi-label image classification.

2. The method for class-imbalanced multi-label image classification based on sample representativeness according to claim 1, wherein: When the current classification category is a positive category for the sample, the representativeness coordination function is a decreasing function, specifically including a linear function, a sigmoid function, a cosine function, or a reciprocal form function; When the current classification category is a negative category for the sample, the representativeness coordination function is an increasing function, specifically including a linear function, a sigmoid function, a logarithmic function, or an exponential function.

3. The method for class-imbalanced multi-label image classification based on sample representativeness according to claim 2, wherein: In the said step S1, the number of categories is C, the number of samples is K, and the number of samples in each category is denoted as Then, the co-occurrence rate of labels between categories is denoted as Its calculation formula is r ij = n i∩j / n i , where the label information of the k-th sample is denoted as 4. The method for class-imbalanced multi-label image classification based on sample representativeness according to claim 3, wherein: The image data augmentation in step S2 at least includes randomly cropping and resizing the image to 224×224, random Gaussian blur, and color distortion processing; the batch size B in step S3 is set to 32.

5. The method for class-imbalanced multi-label image classification based on sample representativeness according to claim 4, wherein: In the step S6 The γ parameter in dynamic form has different design forms for the positive or negative category of class i. The specific design forms are as follows:

6. The method for class-imbalanced multi-label image classification based on sample representativeness according to claim 5, wherein: In step S8, during the backpropagation process, a total of 10 epochs were trained, and SGD with a momentum of 0.9 and a weight decay exponent of 1×10 -4 was used as the optimizer. The learning rate was initialized to 0.02 and decayed at 5 to 7 epochs.

7. A class-imbalanced multi-label image classification system based on sample representativeness, including a computer program, characterized in that: When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Training corpus screening method based on various loss fusion text classification model results

    CN114116969A

  • Network training method for class unbalanced data set

    CN114332539A