Feature Cluster Compression Method for Long-Tail Distribution Problem in Image Recognition
By linearly compressing the features output by the deep neural network, a tighter feature cluster is formed, which solves the bias problem of deep neural networks when training on the long-tail distribution data set, and significantly improves the accuracy of identification of tail classes.
Patent Information
- Application Number
- CN202310252180.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-16
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-03-16
AI Technical Summary
When deep neural networks are trained on long-tail distribution datasets, they are prone to bias against categories that occupy most of the data, and have low accuracy in identifying categories that occupy a small part of the data, resulting in poor classification performance.
A feature cluster compression method is proposed. In the training stage, the original features output from the backbone network are linearly compressed. By multiplying by the preset scaling factor, a tighter feature cluster is formed, and the original features are directly inputted in the test stage for classification.
By compressing the backbone feature clusters, the feature density is increased, especially the sparse cluster compression effect of tail-classes, making the boundary features difficult to cross the decision boundary, thereby improving the model's performance on the long-tail distribution dataset.
Smart Images

Figure CN116310608B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to the problem of long-tail distribution in the field of image recognition. Background Art
[0002] Deep neural networks (DNNs) have achieved great success in various visual tasks, such as object detection and visual recognition. This is not only because neural networks have powerful learning capabilities, but also inseparable from large-scale balanced datasets. However, the datasets collected from actual scenarios are usually unbalanced and follow a long-tail distribution, where a few categories occupy most of the data, while many categories are underrepresented. DNNs trained on such datasets usually show biases towards the categories that occupy most of the data, and have low recognition accuracy for the categories that occupy a small part of the data.
[0003] In recent years, many measures have been proposed to solve the above problems. For example, resampling methods aim to balance the data distribution by designing different sampling strategies, and reweighting methods aim to reduce the advantages of multi-data categories by assigning different weights to each category, etc. However, they all ignore the impact of backbone feature density on this problem. In practical applications, due to sparsity, the backbone features of real samples cannot be mapped close enough, scattered in the feature space, making the points especially located at the boundaries of feature clusters easily cross the decision boundary, resulting in poor classification performance. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a feature cluster compression method for the long-tail distribution problem in image recognition, so as to increase the backbone feature density by compressing the backbone feature clusters, and this method can be combined with existing long-tail methods. Extensive experiments show that the present invention has achieved significant performance improvement for the long-tail distribution.
[0005] To achieve the above object, the present invention provides a feature cluster compression method for the long-tail distribution problem in image recognition, including:
[0006] In the training stage, multiply the original features output by the backbone network by a preset specific factor to obtain the multiplied features, so that the original features are linearly compressed relative to the multiplied features, and input the multiplied features into the classifier for training; in the testing stage, directly input the original features into the trained classifier; because the original feature clusters are linearly compressed, the features are closer to each other, especially the compression effect on the sparse clusters of the tail classes is more significant, making the boundary features less likely to cross the decision boundary, thereby solving the long-tail distribution problem on the dataset in image recognition.
[0007] Optionally, the training stage is: the training stage of the deep neural network using the dataset.
[0008] Optionally, the feature cluster is compressed into:
[0009]
[0010] where and are the multiplied feature and the original feature of the i-th class respectively, and τ i is the scaling factor of the i-th class.
[0011] Optionally, the way to compress the original features output by the backbone network is: adopting arithmetic compression;
[0012] The arithmetic compression is:
[0013]
[0014] where τ i is the scaling factor of the i-th class, γ is the scaling hyperparameter, C is the number of classes, and i is the class index.
[0015] Optionally, the test phase is: the test phase of using the dataset to test the trained deep neural network.
[0016] Optionally, based on the feature cluster compression, the original features form denser clusters than the multiplied features, and a linear compression relationship is established between the original features and the multiplied features, and the multiple of the linear compression depends on the value of the scaling factor.
[0017] Compared with the prior art, the present invention has the following advantages and technical effects:
[0018] Multiply the original features output by the backbone network by a specific factor (τ>1) and input them into the classifier. The original feature clusters are linearly compressed relative to the multiplied feature clusters. This relationship will force the separator to train the decision boundary under larger-sized feature clusters. In the test phase, the more closely mapped original features are directly input into the trained classifier, making the feature points closer to each other, and the boundary points will be brought back inside the decision boundary, thereby improving the model performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0020] Figure 1 is the schematic diagram of feature cluster compression according to the embodiment of the present invention; wherein, (a) is the schematic diagram of feature cluster compression in a two-dimensional space, and (b) is the visualization diagram of the original feature cluster and the multiplied feature cluster;
[0021] Figure 2Schematic diagram of the comparison of the accuracy of using different compression strategies with the Baseline and the recall rate of each category with the Baseline under the equal difference compression strategy in the embodiments of the present invention; wherein, (a) is the schematic diagram of the comparison of the equal difference compression strategy and other compression strategies, and (b) is the schematic diagram of the comparison of the FCC method and the Baseline;
[0022] Figure 3 Schematic diagram of the structure of the FC network in the embodiments of the present invention;
[0023] Figure 4 Schematic diagram of the relationship between plane η and plane η′ and feature points in the two-dimensional geometric space in the embodiments of the present invention; wherein, (a) is the schematic diagram of the relationship when the feature point is higher than plane η and plane η′ is lower than plane η, (b) is the schematic diagram of the relationship when the feature point is higher than plane η and plane η′ is higher than plane η, (c) is the schematic diagram of the relationship when the feature point is lower than plane η and plane η′ is higher than plane η, and (d) is the schematic diagram of the relationship when the feature point is lower than plane η and plane η′ is lower than plane η;
[0024] Figure 5 Schematic diagram of the result visualization of using FCC for three different binary classification datasets in the embodiments of the present invention;
[0025] Figure 6 Schematic diagram of the relative accuracy of using ResNet-10 and ResNet-32 networks respectively starting from the 0-70th Epochs on 100 and 200 Epochs in the embodiments of the present invention; wherein, (a) is the schematic diagram of the accuracy of ResNet-32 network on 200 Epochs, (b) is the schematic diagram of the accuracy of ResNet-10 network on 200 Epochs, (c) is the schematic diagram of the accuracy of ResNet-32 network on 100 Epochs, and (d) is the schematic diagram of the accuracy of ResNet-10 network on 100 Epochs. Detailed implementation manners
[0026] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0027] It should be noted that the steps shown in the flowchart of the drawings may be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0028] The present invention provides a feature cluster compression method for the long-tail distribution problem in image recognition, including:
[0029] In the training stage, the original features output by the backbone network are multiplied by a preset specific factor to obtain the multiplied features, so that the original features are linearly compressed relative to the multiplied features, and the multiplied features are input into a classifier for training; in the testing stage, the original features are directly input into the trained classifier; because the original feature clusters are linearly compressed, the features are closer to each other, especially for the sparse clusters of the tail classes, the compression effect is more significant, making the boundary features less likely to cross the decision boundary, thus solving the long-tail distribution problem on the dataset in image recognition.
[0030] Further, the training stage is: the training stage of the deep neural network using the dataset.
[0031] Further, the feature cluster compression is as follows:
[0032]
[0033] where and are the first feature and the original feature of the i-th class respectively, and τ i is the scaling factor of the i-th class.
[0034] Further, the way to perform feature cluster compression on the original features output by the backbone network is: adopt arithmetic compression;
[0035] The arithmetic compression is as follows:
[0036]
[0037] where τ i is the scaling factor of the i-th class, γ is the scaling hyperparameter, C is the number of classes, and i is the class index.
[0038] Further, the testing stage is: the testing stage of the trained deep neural network using the dataset.
[0039] Further, based on the feature cluster compression, the original features form a denser cluster than the first features, and a linear compression relationship is established between the original features and the first features, and the multiple of the linear compression depends on the value of the scaling factor.
[0040] Embodiment
[0041] In the feature space, deep neural networks (DNNs) can map backbone features into the form of feature clusters. For datasets with long-tailed distributions, the head classes present dense clusters, while the tail classes present sparse clusters. Inspired by this, in this embodiment, a feature cluster compression method for the long-tailed distribution problem in the field of image recognition is proposed to increase the density of backbone network features, making the feature points map closer to their clusters, and further pulling the boundary features back into the decision boundary to solve the problem of limited accuracy of the model on long-tailed distribution datasets in the field of image recognition. This method is called FCC (Feature Clusters Compression), and it is completed in two parts. First, in the training stage, the features output by the backbone network are multiplied by a scaling factor τ (τ > 1), and the multiplied features are used as the input to the classifier. Then, in the testing stage, the original features are directly input to the classifier. This will establish a linear compression relationship between the original feature cluster and the multiplied feature cluster.
[0042] By compressing the backbone feature clusters, the density of the backbone neural network features is increased. Especially for the sparse clusters of the tail classes, the compression effect is more significant, and the features are closer to each other, making it difficult for the boundary features to cross the decision boundary, thus solving the long-tailed distribution problem in image recognition datasets.
[0043] First, to illustrate the compression principle of the first part, this embodiment uses a square in a two-dimensional space as an example to explain. As Figure 1 (a) shows, assume a square ABCD in a two-dimensional space. When its vertex coordinates are multiplied by the scaling factor τ (τ = 2), the square ABCD will change into a square A'B'C'D'. It can be easily observed that the square ABCD is linearly compressed relative to the square A'B'C'D', and the distances between points (A, B, C, D) are shortened by τ times, and the density increases to τ 2 times. Similarly, extended to an n-dimensional space, the distance will also be shortened by τ times, and the density increases by τ n times (n is the dimension of the space).
[0044] In this embodiment, the backbone features are regarded as points in an n-dimensional space, where n represents the number of parameters in the backbone features, and each parameter represents a dimension. After each batch of data is input to the backbone network in the training stage, the original backbone features of each class (similar to the square ABCD) are multiplied by a specific scaling factor τ (τ > 1), and further the multiplied features (similar to the square A'B'C'F') are input to the classifier. During the process of feature space mapping, this operation makes the original features form denser clusters than the multiplied features, and a linear compression relationship is established between the original feature cluster and the multiplied feature cluster. The multiple of linear compression depends on the value of τ. To more clearly show the compression results, the original feature cluster and the multiplied feature cluster are visualized using principal component analysis, as Figure 1As shown in (b), it can be seen that the original feature clusters are compressed and their feature points are closer to each other.
[0045] The formula for feature cluster compression is defined as follows:
[0046]
[0047] Where, and represent the multiplied feature and the original feature of the i-th class respectively, and τ i represents the scaling factor of the i-th class. To control the compression degree of each class, three compression strategies are defined in this embodiment as follows:
[0048] Equal compression, setting the same τ for each class i , and the formula is defined as follows:
[0049] τ i = 1 + γ
[0050] Arithmetic progression compression, τ i decreases sequentially from the majority class to the minority class, and the formula is defined as follows:
[0051]
[0052] Half compression, only compressing the first half or the second half of the classes. The compressed half is achieved by the arithmetic progression compression strategy, and the τ of the other half i are all set to 1, and the formula is defined as follows:
[0053]
[0054] Where γ > 0 is the scaling hyperparameter, C is the number of classes, i ∈ [0, C) is the class index. When the parameter in is negative, is equal to 1, otherwise is equal to 0. When compressing the first half of the classes, β = 0, and when compressing the second half of the classes, β = 1.
[0055] In the test stage, the original features are directly used as the input of the classifier, so that through the already trained backbone network, these feature points can be mapped closer to each other and are not likely to cross the decision boundary, resulting in misclassification.
[0056] In this embodiment, experiments are conducted on the CIFAR-100-LT-100 dataset using each compression strategy, and are respectively compared with Baseline (a model without using FCC, and the network settings are all ResNet32) to evaluate the performance of the models using different compression strategies, as Figure 2As shown in (a), the use of the arithmetic compression strategy is superior to other compression strategies. When γ = 1, the model performance is the highest, and the accuracy reaches 41.1%. Based on this result, the compression strategy of FCC is set to arithmetic compression, and the recall rate results of each category of Baseline and FCC are further recorded respectively. As Figure 2 shown in (b), the use of the FCC method proposed in this embodiment (marked in black) has a significant improvement in the minority categories (tail classes) compared with Baseline (marked in light gray).
[0057] Because the multiplied features and the original features are used as the inputs of the classifier in the training stage and the testing stage respectively, this embodiment takes binary classification as an example here to prove that the classifier can work properly with different inputs, and the result of this proof is also applicable to the multi-classification task. Here, it is necessary to prove that for the same sample as the input, in the training stage, if the multiplied feature is recognized as class 1 by the classifier, then in the testing stage, the original feature will also be recognized as class 1 by the classifier.
[0058] The classifier network structure is set as follows:
[0059] As Figure 3 shown, the classifier (fully connected network, called the FC network in the proof process) includes an input layer with 3 neurons, a hidden layer with 3 neurons {a1, a2, a3}, and an output layer with 2 neurons {o1, o2}. The multiplied features are {τx1, τx2, τx3}, and the original features are {x1, x2, x3}, and they all belong to class 1, and the scaling factor of class 1 is τ (τ > 1). Through the processing of the FC network, the outputs of the multiplied features are {y1, y2}, and the outputs of the original features are {y′1, y′2}. The weights and biases of neuron ai (i ∈ {1, 2, 3}) are {w i1 , w i2 , w i3} and b i , and the weights and biases of neuron o j (j ∈ {1, 2}) are {n j1 , n j2 , n j3} and z j .
[0060] The proof objectives are as follows:
[0061] If the FC network can work properly, then the classification results of the original features and the multiplied features should be the same, that is, when y1 > y2, y'1 > y'2.
[0062] The output expression formula of the multiplied features is as follows:
[0063] y1 = n 11 (τw 11 x1 + τw 12 x2 + τw 13 x3) + n 11 b1 + n 12 (τw 21 x1 + τw 22 x2 + τw 23 x3) + n 12 b2 + n 13 (τw 31 x1 + τw 32 x2 + τw 33 x3) + n 13 b3 + z1
[0064] y2 = n 21 (τw 11 x1 + τw 12 x2 + τw 13 x3) + n 21 b1 + n 22 (τw 21 x1 + τw 22 x2 + vw 23 x3) + n 22 b2 + n 23 (τw 31 x1 + τw 32 x2 + τw 33 x3) + n 23 b3 + z2
[0065] In this embodiment, (y1 - y2) is denoted as (w 11 x1 + w 12 x2 + w 13 x3) is denoted as X1, (w 21 x1 + w 22 x2 + w 23 x3) is denoted as X2, (w 31 x1 + w 32 x2 + w 33 x3) is denoted as X3, (n 11 b1 + n 12 b2 + n 13 b3 + z1) - (n 21 b1 + n 22 b2 + n 23 b3 + z2) is denoted as B.
[0066] Furthermore, it is converted into the following formula:
[0067]
[0068] where k i is (n 1i -n 2i ), i ∈ {1, 2, 3}. It can be seen that when , is a plane in the geometric space.
[0069] Since y1 > y2, so Equation (1) can be derived as follows:
[0070]
[0071] where d i is -k i / B, i ∈ 1, 2, 3. In the set space, when B < 0, the point (X1, X2, X3) is above this plane, when B > 0, the point (X1, X2, X3) is below this plane, and when B = 0, it will be discussed later. By the same principle, (y'1 - y′2) is denoted as The equation derivation is as follows:
[0072]
[0073] When , Equation (3) can be derived into the following equation:
[0074] d1X1 + d2X2 + d3X3 = 1 (4)
[0075] When , is also a plane. It can be observed that and are respectively the intercepts of plane and plane , and the intercept of plane is τ times that of plane . So plane and plane are parallel in the geometric space. At the same time, according to the size of the intercepts, it is inferred that plane is above or below plane .
[0076] Next, this embodiment explores the relationship between plane and plane and the characteristic points in the geometric space to explore whether y'1 > y'2 holds when y1 > y2.
[0077] In the first case, when B < 0, according to Equation (2), it can be known that the point (X1, X2, X3) is in plane Above. If the plane is in the plane below, the same points will also be above the plane , as Figure 4 (a) shows. So according to formula (4), d1X1 + d2X2 + d3X3 > 1. According to formulas (3) and (4), it can be obtained that that is, y'1 > y'2, which indicates that the FC network can work properly at these points. If the plane is in the plane above, these points may be above or below the plane , as Figure 4 (b) shows. When these points are above the plane , because d1X1 + d2X2 + d3X3 > 1, these points can also be correctly classified. But when these points are below the plane , d1X1 + d2X2 + d3X3 < 1, and y'1 < y'2, which means that the FC network will misclassify these points.
[0078] In the second case, when B > 0, by the same principle, these points are below the plane . As Figure 4 (c) shows, if the plane is in the plane above, y'1 > y'2. As Figure 4 (d) shows, if the plane is in the plane below, and these points are above the plane , y'1 < y'2. When B = 0, the plane coincides with the plane , and when y1 > y2, y′1 > y′2.
[0079] In summary, the classifier can work properly when the original features are input in the test stage, except for the points that fall between the plane and the plane . This area is called the "misclassification area", as Figure 4 shown by "misclassified area" in
[0080] According to formulas (2) and (4), it is observed that the "misclassification area" is positively correlated with τ. Therefore, the value of τ cannot be set too large to avoid the classifier from failing. In fact, when we set τ to a reasonable value during the actual application process, the "misclassification area" will not affect the overall performance because only very few points will fall into this area.
[0081] In this embodiment, experiments were conducted on the CIFAR10 / 100-LT-100 dataset using the FCC method (setting different τ values), as shown in Table 1; the network was set to ResNet-32, and different τ values were obtained by the parameter γ, where γ ∈ {0.1, 0.5, 1, 2, 3}. Experiments were carried out to illustrate that the "misclassification region" does not affect the overall performance.
[0082] Table 1 Analysis of the original features falling into the "misclassification region"
[0083]
[0084] Among them, NMF (Number of Multiplied Features) represents the number of points correctly classified by the classifier, NOF (Number of Original Features) represents the number of points falling into the "misclassification region", and Ratio (NOF / NMF) represents the ratio of misclassified points to correctly classified points.
[0085] When γ is set to 0.1, the original features hardly fall into the "misclassification region". When γ is set to 0.5 or 1, Ratio < 3%, and the results are acceptable. However, γ cannot be set too large (such as 2 or 3), which will cause a large number of original features to fall into the "misclassification region".
[0086] This embodiment can also explain this problem from another perspective. As Figure 1 (a) shows, during the compression process, the original feature cluster (square ABCD) will move towards the origin (point P in the figure) relative to the multiplied feature cluster (square A'B'C'D'), and the moving distance d can be expressed as follows:
[0087] d = D i *(1 - 1 / τ i )
[0088] where D i represents the distance from the center of the multiplied features of class i to the origin (the feature point where all parameters are 0 in the feature space). There is also a positive correlation between d and τ. The larger τ is, the farther the original feature cluster moves, which may cause them to cross the decision boundary and lead to misclassification. Based on this result, different τ values were set for each class during the previous compression strategy selection, and smaller τ values were set for the minority classes because the feature clusters of the minority classes are closer to the decision boundary and are more likely to cross the decision boundary and cause misclassification.
[0089] To illustrate that the FCC can bring boundary points back inside the decision boundary, this embodiment visually demonstrates the results of the FCC on three other imbalanced datasets (with an imbalance factor of 10). These three datasets are created based on common datasets in scikit-learn, including two circles, two blobs, and two moons. As Figure 5 shown, the top row of three figures shows the results without using the FCC, and the bottom row of three figures shows the results using the FCC. The majority class is represented by light gray, and the minority class is represented by black. The results show that the FCC can compress the feature clusters and bring the boundary points of the minority class back inside the decision boundary. In some cases, this embodiment observes that some feature points of the majority class move towards the origin during the compression process and even cross the decision boundary (resulting in misclassification), but more feature points of the minority class are indeed brought back inside the decision boundary, and extensive experiments will be conducted next to verify that the overall performance has been improved.
[0090] The experiments and results are shown as follows:
[0091] Dataset
[0092] In the embodiments of the present invention, three long-tail baseline datasets are used, including CIFAR-10-LT, CIFAR-100-LT, and ImageNet-LT. Among them, CIFAR-10 / 100-LT is established by downsampling the training samples of each class in the original CIFAR-10 / 100 dataset. These datasets contain the same classes as the original datasets and balanced test sets. The imbalance factor (IF) of the long-tail datasets is defined as the number of training samples in the largest class divided by the number of training samples in the smallest class. By changing IF ∈ {50, 100}, this embodiment creates four long-tail datasets, namely CIFAR-10-LT-50, CIFAR-10-LT-100, CIFAR-100-LT-50, and CIFAR-100-LT-100. The long-tail ImageNet (ImageNet-LT) dataset is created by artificially truncating the original ImageNet2012. ImageNet-LT contains 1000 classes and the number of training images in each class ranges from 5 to 1280.
[0093] Implementation
[0094] All networks are trained from scratch. For the CIFAR-10 / 100-LT datasets, ResNet-32 is used as the backbone network and momentum of 0.9 and weight decay of 2*10 -4It is trained using Stochastic Gradient Descent (SGD). The number of training epochs is 200 and the batch size is 128. The learning rate is initialized to 0.1 and divided by 100 at the 160th and 180th epochs respectively. Warm-up is used for the first 5 epochs. For ImageNet-LT, ResNet-10 is used as the backbone network. The number of training epochs is 200 and the batch size is 64. Momentum of 0.9 and weight decay of 1*10 -4 It is trained using Stochastic Gradient Descent (SGD). The learning rate is initialized to 0.2 and divided by 10 at the 160th and 180th epochs respectively. Warm-up is not set. The random seed is set to 42. All networks are trained on 2 NVIDIA RTX3090 GPUs using the PyTorch toolkit. The TOP-1 error rate is used to compare the experimental results.
[0095] To explore the impact of the hyperparameter γ on FCC on different long-tailed datasets, in this embodiment, γ ∈ {0.1, 0.5, 1, 2, 3} is set, and experiments are conducted on the long-tailed CIFAR and ImagNet-LT datasets using the ResNet-32 / ResNet-10 network respectively. As shown in Table 2, on the CIFAR-10-LT dataset, when γ is set to 0.5, the improvement compared with the original method exceeds 3%; on the CIFAR-100-LT dataset, when γ is set to 1, the best result is obtained, with an improvement of about 2%; on the ImageNet-LT dataset, when γ is set to 0.1, the best result is obtained. We observe that the optimal hyperparameter γ values for different datasets are inconsistent, but generally 0.1, 0.5, and 1 are the best choices. As the value of γ decreases, the performance improvement of FCC gradually decreases. This is because the smaller the compression degree, the fewer the boundary points brought back to the decision boundary. However, too large a γ (e.g., γ = 3) will damage the performance because the original feature clusters may cross the decision boundary.
[0096] Table 2 Impact of the hyperparameter γ on FCC on the long-tailed datasets CIFAR and ImageNet-LT
[0097]
[0098] Based on the above analysis, in all the following experiments, the FCC equal difference compression strategy is used. If there is no specific description, the γ values of FCC are set to 0.5, 1, and 0.1 on the CIFAR-10-LT, CIFAR-100-LT, and ImageNet-LT datasets respectively.
[0099] To explore when to use FCC in the training stage, this embodiment conducted experiments on the long-tailed CIFAR dataset using ResNet-32 and ResNet-10 networks for 100 and 200 epochs respectively. FCC was used starting from the 0th, 10th, 20th, 30th, 40th, 50th, 60th, and 70th epochs respectively. As Figure 6 shown, the relative accuracy (accuracy from the same dataset minus a specific constant) was used as the ordinate to clearly show the impact of using FCC at which epoch on the accuracy. For 200 training epochs, starting to use FCC from the 50th epoch usually produces the best results. It is speculated that after ordinary training, the feature clusters will be separated from each other. On this basis, further using FCC can make the distance between them farther and easier to identify. However, for 100 training epochs, starting to use FCC from the 0th epoch usually produces the best results, which means that in the case of fewer training epochs, FCC requires sufficient epochs to train the compressed feature clusters. At the same time, it was observed that in the case of the same dataset, the ResNet-10 network using FCC starting from the 0th epoch has better performance than starting from the 50th epoch, suggesting that the weaker network FCC requires more epochs.
[0100] Based on the above experiments, this embodiment uses FCC starting from the 0th training epoch on the ImageNet-LT dataset and starting from the 50th training epoch on other datasets. Without specific instructions, the backbone networks used on the long-tailed CIFAR dataset and the ImageNet-LT dataset are ResNet-32 and ResNet-10 respectively.
[0101] Comparison methods
[0102] To evaluate the effectiveness and generality of FCC, in this embodiment, FCC is widely applied to the current best 31 methods, which come from four categories of methods for solving long-tailed distribution datasets, namely reweighting, resampling, two-stage training, and multi-expert methods. In this embodiment, the latest released and most representative methods in these categories are selected. For example, focal loss in reweighting and NCL in multi-expert. For two-stage training, resampling and reweighting methods are used for DRS and DRW respectively. At the same time, hybrid training methods (such as Input Mixup, Manifold Mixup, and Remix) are also introduced because they have good performance in long-tailed visual recognition. For comparison, the results of the original methods and the results of applying FCC are shown respectively. Tables 3 and 4 list the experimental results of the CIFAR-10 / 100-LT and ImageNet-LT datasets respectively.
[0103] Table 3 Comparison of experimental results of applying FCC to existing methods (CIFAR dataset), where * indicates that the γ value is set to 0.1
[0104]
[0105] Table 4 Comparison of experimental results of applying FCC to existing methods (ImageNet dataset), where + and * indicate that the backbone networks are ResNet-50 and ResNeXt-50 respectively
[0106]
[0107] Result Analysis
[0108] For the CIFAR-10 / 100-LT datasets, the results of the baseline (ResNet-32), reweighting, resampling, mixup training, two-stage training, and multi-expert methods are shown in Table 2 respectively. Among these 98 experimental groups, the proposed FCC improved 94 of the results, with an average improvement of 1.52% (highest 4.89%, lowest 0.04%). FCC successfully refreshed the top-1 error rates on these four datasets, reducing CIFAR-10-LT-50, CIFAR-10-LT-100, CIFAR-100-LT-50, and CIFAR-100-LT-100 to 12.72%, 14.20%, 41.56%, and 45.49% respectively. Extensive experiments fully demonstrate the effectiveness of the method in this embodiment for long-tailed recognition. At the same time, these results also show the generality of FCC, which can be combined with existing methods very friendly to further improve them, especially the mixup training method. For the multi-expert method, setting γ to 0.1 has better results than 0.5 and 1. It is speculated in this embodiment that for such complex networks, smaller γ is more suitable because they are more sensitive to the movement of feature clusters. It is observed in this embodiment that FCC only fails to achieve enhanced effects in four experimental groups. It is speculated in this embodiment that the decision boundaries drawn by these methods are too close to their feature clusters, so that these clusters can easily cross the decision boundaries.
[0109] For ImageNet-LT, the results are shown in Table 4. FCC improved the introduced methods by an average of 0.9% (highest 1.98%, lowest 0.06%). In particular, applying FCC to SADE refreshed the current best performance (top-1 error rate 39.47%). Applying FCC to simple tricks (such as Focal loss and CBCE) has little improvement (less than 0.5%), but applying FCC to complex networks (such as SADE and NCL) has significant improvement (more than 1.5%). It is speculated in this embodiment that this is because complex networks can create a better basis for compression. At the same time, the backbone networks (BBN, NCL, SADE) of some methods use ResNet-50 and ResNeXt-50, which does not affect the conclusion of the effectiveness of our method. Their training settings are consistent with the training settings of FCC.
[0110] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A feature cluster compression method for the long-tail distribution problem in image recognition, characterized in that Including: In the training stage, multiply the original features output by the backbone network by a preset specific factor to obtain the multiplied features, so that the original features are linearly compressed relative to the multiplied features, and input the multiplied features into the classifier for training; In the testing stage, directly input the original features into the trained classifier; because the original feature clusters are linearly compressed, the features are closer to each other, and the compression effect on the sparse clusters of the tail classes is more significant, making it difficult for the boundary features to cross the decision boundary, thus solving the long-tailed distribution problem on the dataset in image recognition; The compression of the feature clusters is as follows: Among them, and are the multiplicand feature and the original feature of the i-th category respectively, and τ i is the scaling factor of the i-th category; The method for performing feature cluster compression on the original features output by the backbone network is: adopting arithmetic progression compression; The arithmetic progression compression is: where τ i is the scaling factor for the i-th class, γ is the scaling hyperparameter, C is the number of classes, and i is the class index.
2. The feature cluster compression method for the long-tailed distribution problem in image recognition according to claim 1, wherein The training stage is: the training stage of the deep neural network using the dataset.
3. The feature cluster compression method for the long-tail distribution problem in image recognition according to claim 1, wherein The testing stage is: the testing stage of the trained deep neural network using the dataset.
4. The feature cluster compression method for the long-tail distribution problem in image recognition according to claim 1, wherein Based on the feature cluster compression, the original features form denser clusters than the multiplied features, a linear compression relationship is established between the original features and the multiplied features, and the multiple of the linear compression depends on the value of the scaling factor.
Citation Information
Patent Citations
Long-tail image recognition method based on self-supervision and self-distillation
CN113837238A
Improving Deep Neural Networks Using Prototype Factoring
DE102021212086A1