Optimization Method and System for a Long-Tail Data Recognition Model
Through the optimization model of a few class supersampling and class-related weighted significant graphs, the problem of low recognition accuracy of long-tail data sets is solved, the characteristic significance of tail class samples is enhanced, and the recognition accuracy and robustness are improved.
Patent Information
- Application Number
- CN202310386005.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-11
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-04-11
AI Technical Summary
The existing deep learning models have low recognition accuracy of tail class samples on long-tail data sets, mainly due to insufficient feature extraction, and the model tends to be head classes, which reduces the recognition accuracy.
Mixed images are generated through a few class supersampling strategy, class-related weighted significant graphs are calculated and label weights are assigned to the images, and the initial recognition model is optimized to enhance the significance of class-related regional features.
The model's recognition accuracy of long-tail data sets is improved, the feature extraction ability of tail-like samples is enhanced, and the robustness and prediction accuracy of the recognition model are improved.
Smart Images

Figure CN116310590B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and particularly relates to an optimization method and system for a long-tail data recognition model. Background Art
[0002] In recent years, deep neural networks have dominated the learning of visual representations and demonstrated powerful performance on various downstream tasks. However, this success highly depends on artificially created relatively balanced dataset classes, such as ImageNet, MSCoco, CIFAR-100 / 10, etc. Since large-scale balanced datasets often require substantial resource costs to construct, training models with real-world data has become an important research direction in deep learning. Real-world data often exhibits a long-tail distribution, that is, there are more samples in the head classes and fewer samples in the tail classes. Models trained on such datasets tend to focus on the head classes, resulting in low recognition accuracy for tail-class samples.
[0003] A key reason for the low recognition accuracy of long-tail datasets is that the insufficient number of samples leads to the model being unable to extract sufficient significant features, that is, the features extracted by the long-tail model cannot well represent the test images. And this imbalance misguides the classifier towards the head classes, further reducing the recognition accuracy of the model. Summary of the Invention
[0004] The present invention provides an optimization method and system for a long-tail data recognition model to solve the problem of low recognition accuracy of the recognition model for long-tail datasets.
[0005] In a first aspect, the present invention provides an optimization method for a long-tail data recognition model, the method comprising the following steps:
[0006] Obtain an initial long-tail dataset and construct an initial recognition model based on the initial long-tail dataset;
[0007] Generate mixed images based on the initial long-tail dataset through a minority class oversampling strategy;
[0008] Calculate the class-related weighted saliency map of the mixed images;
[0009] Assign label weights to the mixed images according to the class-related weighted saliency map;
[0010] Optimize the initial recognition model through the class-related weighted saliency map so that the initial recognition model enhances the saliency of class-related region features.
[0011] Optionally, the calculation formula of the class-related weighted saliency map is as follows:
[0012]
[0013] Where: y c represents the final score relative to class C; A d represents the feature map of the d-th channel; i and j respectively represent the coordinates in the feature map; represents A d 's weight; represents the data at the position of coordinate ij in channel d of A; Z represents the width * height of the feature map.
[0014] Optionally, assigning label weights to the mixed image according to the class-related weighted saliency map includes the following steps:
[0015] Set the salient partition in the class-related weighted saliency map according to a preset saliency threshold;
[0016] Assign label weights to the mixed image based on the partition ratio of the salient partition.
[0017] Optionally, training the initial recognition model through the class-related weighted saliency map to enhance the saliency of class-related region features includes the following steps:
[0018] Substitute the class-related weighted saliency map into the initial recognition model;
[0019] Activate the class-related weighted saliency map through the activation function in the initial recognition model to obtain an activation result;
[0020] Calculate an enhancement mask by combining the activation result and a preset feature threshold;
[0021] Optimize the initial recognition model through the enhancement mask to enhance the saliency of class-related region features.
[0022] Optionally, after activating the class-related weighted saliency map through the activation function in the initial recognition model to obtain an activation result, the following steps are further included:
[0023] Obtain the number of training samples of class C in the initial recognition model;
[0024] Calculate a loss threshold based on the number of training samples;
[0025] Calculate a discard mask by combining the activation result and the loss threshold;
[0026] Optimize the initial recognition model through the discard mask to make the initial recognition model discard the most class-specific features.
[0027] In a second aspect, the present invention also provides an optimization system for a long-tail data recognition model, and the system includes:
[0028] A model creation module, configured to obtain an initial long-tail data set and construct an initial recognition model based on the initial long-tail data set;
[0029] An image generation module, configured to obtain the initial long-tail data set from the model creation module, and generate a mixed image based on the initial long-tail data set through a minority class oversampling strategy;
[0030] A saliency calculation module, preset with a calculation formula for a class-related weighted saliency map, configured to obtain the mixed image generated by the image generation module and calculate the class-related weighted saliency map of the mixed image;
[0031] A weight assignment module, configured to obtain the class-related weighted saliency map calculated by the saliency calculation module and assign label weights to the mixed image according to the class-related weighted saliency map;
[0032] A model optimization module, configured to obtain the class-related weighted saliency map calculated by the saliency calculation module and optimize the initial recognition model through the class-related weighted saliency map, so that the initial recognition model enhances the saliency of class-related region features.
[0033] Optionally, the calculation formula for the class-related weighted saliency map preset in the saliency calculation module is as follows:
[0034]
[0035] In the formula: y c represents the final score relative to class C; A d represents the feature map of the d-th channel; i and j respectively represent the coordinates in the feature map; represents the weight of A d ; represents the data at the position of coordinate ij in channel d of A; Z represents the width * height of the feature map.
[0036] Optionally, the weight assignment module includes:
[0037] A partition setting unit, preset with a saliency threshold, configured to obtain the class-related weighted saliency map calculated by the saliency calculation module and set a salient partition in the class-related weighted saliency map according to the saliency threshold;
[0038] A weight assignment unit, configured to assign label weights to the mixed image according to the partition ratio of the salient partition set by the partition setting unit.
[0039] Optionally, the model optimization module includes:
[0040] An activation unit, configured to substitute the class-related weighted saliency map into the initial recognition model, and activate the class-related weighted saliency map through an activation function in the initial recognition model to obtain an activation result;
[0041] A mask calculation unit, preset with a feature threshold, configured to calculate an enhanced mask by combining the activation result and the feature threshold;
[0042] An enhancement and optimization unit, configured to optimize the initial recognition model through the enhanced mask, so that the initial recognition model enhances the saliency of class-related region features.
[0043] Optionally, the system further includes a feature extension module, and the feature extension module includes:
[0044] A data acquisition unit, configured to acquire the number of training samples of class C in the initial model;
[0045] A first calculation unit, configured to calculate a loss threshold according to the number of training samples;
[0046] A second calculation unit, connected to the activation unit and the first calculation unit, configured to calculate a discard mask by combining the activation result and the loss threshold;
[0047] A discard optimization unit, configured to optimize the initial recognition model through the discard mask, so that the initial recognition model discards the most class-discriminative features.
[0048] The beneficial effects of the present invention are:
[0049] The method of the present invention includes the following steps: acquiring an initial long-tail data set, and constructing an initial recognition model based on the initial long-tail data set; generating a mixed image based on the initial long-tail data set and through a minority class oversampling strategy; calculating a class-related weighted saliency map of the mixed image; assigning label weights to the mixed image according to the class-related weighted saliency map; optimizing the initial recognition model through the class-related weighted saliency map, so that the initial recognition model enhances the saliency of class-related region features. By using the class-related saliency map to enhance the saliency of the class-related region and restricting the saliency of other features, the accuracy of the recognition model for recognizing the long-tail data set is improved. Description of the Drawings
[0050] Figure 1 It is a schematic flowchart of an optimization method for a long-tail data recognition model in one embodiment of the present application.
[0051] Figure 2 It is a schematic flowchart of an optimization method for a long-tail data recognition model in one embodiment of the present application.
[0052] Figure 3 Schematic diagram of the optimization method for the long-tail data recognition model in one embodiment of the present application.
[0053] Figure 4 Schematic diagram of the result of experimenting with the feature threshold in one embodiment of the present application.
[0054] Figure 5 Schematic diagram of the optimization method for the long-tail data recognition model in one embodiment of the present application.
[0055] Figure 6 System structure diagram of the optimization system for the long-tail data recognition model in one embodiment of the present application. Detailed implementation mode
[0056] The present invention discloses an optimization method for a long-tail data recognition model.
[0057] In one embodiment, referring to Figure 1 , the optimization method for the long-tail data recognition model mainly includes the following steps:
[0058] S101. Obtain an initial long-tail data set and construct an initial recognition model based on the initial long-tail data set.
[0059] S102. Generate mixed images based on the initial long-tail data set and through the minority class oversampling strategy.
[0060] Among them, through the minority class oversampling strategy (Context-rich Minority Oversampling, CMO), one picture of the minority class in the initial long-tail data set is pasted onto multiple majority class pictures with rich context, and the majority class pictures are used as the background, that is, mixed images are generated.
[0061] S103. Calculate the class-related weighted saliency map of the mixed images.
[0062] Among them, the calculation result is obtained by calculating the class-related weighted saliency of the mixed images, and then the calculation result is superimposed on the original mixed images through interpolation and other methods to obtain the final visualized result, that is, the class-related weighted saliency map.
[0063] S104. Assign label weights to the mixed images according to the class-related weighted saliency map.
[0064] S105. Optimize the initial recognition model through the class-related weighted saliency map so that the initial recognition model enhances the saliency of the class-related region features.
[0065] The implementation principle of this embodiment is:
[0066] Generate a mixed image through a minority class oversampling strategy, calculate the class-related weighted saliency map of the mixed image, and then assign label weights to the mixed image according to the class-related weighted saliency map, so as to adopt a method of discriminant of similar features to assign more reasonable label weights to the mixed image, and then use the class-related weighted saliency map to optimize the initial recognition model to enhance the saliency of the class-related features of the initial recognition model and improve the feature representation ability of the initial recognition model, thereby improving the recognition accuracy of the recognition model for the long-tailed dataset.
[0067] In one implementation manner of the embodiment of the present invention, the calculation formula of the class-related weighted saliency map is as follows:
[0068]
[0069] In the formula: y c represents the final score relative to class C; A d represents the feature map of the d-th channel; i and j respectively represent the coordinates in the feature map; represents the weight of A d ; represents the data at the position of coordinate ij in channel d of A; Z represents the width * height of the feature map.
[0070] In one implementation manner of the embodiment of the present invention, referring to Figure 2 , step S104, that is, assigning label weights to the mixed image according to the class-related weighted saliency map, specifically includes the following steps:
[0071] S201. Set the significant partition in the class-related weighted saliency map according to a preset significant threshold.
[0072] Among them, in this implementation manner, the mixed image is generated by the CutMix data augmentation method in the CMO method. The goal of the CutMix method is to generate a new mixed image by combining the foreground image (x b , y b ) and the background image (x f , y f ). CMO assigns the mixed label of the generated image as a linear combination of y b and y f . The definition of generating a mixed image by CutMix is as follows:
[0073]
[0074] In the formula: M is a binary mask, and M = g(λ), λ ~ Beta(1, 1), and M is a randomly sampled rectangle under the constraint of the formula .
[0075] When assigning labels to the generated mixed images, first calculate the class-related saliency maps S of the mixed images according to the foreground image and background image labels respectively through the calculation formula of the class-related weighted saliency map. c (A, y f ) and S c (A, y b ). In this embodiment, the preset saliency threshold ζ is 0.6. Count the number of pixels w c (A, y b ) in S b with saliency greater than ζ, and count the number of pixels w c (A, y f ) in S f with saliency greater than ζ. Then set the pixels in the class-related weighted saliency map that are more salient than the saliency threshold as the salient partitions. Finally, assign the target score to the mixed image according to the partition ratio of the salient partitions corresponding to the labels.
[0076] S202. Assign label weights to the mixed image based on the partition ratio of the salient partitions.
[0077] Among them, the formula for assigning label weights in this embodiment is as follows:
[0078]
[0079] In the formula: w f and w b are both the number of pixels counted in step S201.
[0080] The implementation principle of this embodiment is:
[0081] The label weight assignment method adopted in this embodiment can assign label weights to the mixed image more accurately compared with the original mixing ratio assignment method of the CMO method. In the mixing ratio assignment method, the mixing ratio is only determined by the size of the binary mask, resulting in that even if there are significant changes in the generated mixed image, the label weight assignment still remains the original ratio.
[0082] In one implementation manner of the embodiment of the present invention, referring to Figure 3 , step S105, that is, optimizing the initial recognition model through the class-related weighted saliency map to enhance the saliency of the class-related region features specifically includes the following steps:
[0083] S301. Substitute the class-related weighted saliency map into the initial recognition model.
[0084] S302. Activate the class-related weighted saliency map through the activation function in the initial recognition model to obtain the activation result.
[0085] In this embodiment, the class-related weighted saliency map is activated by the sigmoid layer in the initial recognition model. The specific formula is as follows:
[0086] F = sigmoid(S c (A,y b )+S c (A,y f ))
[0087] S303. Calculate an enhancement mask by combining the activation result and a preset feature threshold.
[0088] In this embodiment, the enhancement mask M aug The calculation is done using the following formula:
[0089]
[0090] Where: i and j are the coordinates in the feature map; θ is the preset feature threshold.
[0091] Different feature thresholds produce different results. In this embodiment, the feature threshold θ is tested on the CIFAR100-LT dataset. The test results are shown in Table 1. Figure 4 , Figure 4 The horizontal axis is the value of the feature threshold θ, and the vertical axis is the classification accuracy of the recognition model for the CIFAR100-LT data set. Therefore, in this embodiment, when the feature threshold θ is 0.7, the effect is best, so 0.7 is the best value of the feature threshold θ in this embodiment.
[0092] S304. Optimize the initial recognition model by enhancing the mask so that the initial recognition model enhances the saliency of the class-related regional features.
[0093] In this embodiment, the characteristics of the initial recognition model enhancement-related regional features are as follows:
[0094] A′=A⊙sigmoid(A)⊙M aug
[0095] The implementation principle of this embodiment is:
[0096] First, the class-weighted saliency map is obtained by calculation, and then the enhancement mask M is calculated based on the class-weighted saliency map aug Enhancement mask M aug The initial recognition model class-related regional features are given a lower weight, guiding the model to enhance the significance of class-related regional features so that the model has better feature representation capabilities. This enables the model to extract more meaningful features corresponding to the correct category of the sample, thereby significantly improving the classification accuracy of long-tail data sets.
[0097] In one implementation manner of the embodiment of the present invention, referring to Figure 5 , after step S105, that is, optimizing the initial recognition model through the class-related weighted saliency map to enhance the saliency of the class-related region features, the following steps are further included:
[0098] S401. Obtain the number of training samples of class C in the initial recognition model.
[0099] S402. Calculate the loss threshold based on the number of training samples.
[0100] Wherein, in this implementation manner, the calculation formula of the loss threshold γ is as follows:
[0101] γ = sigmoid(exp(-num c ))
[0102] In the formula: num c is the number of training samples of class C.
[0103] S403. Calculate the discard mask by combining the activation result and the loss threshold.
[0104] Wherein, in this implementation manner, the discard mask M drop can be expressed as:
[0105]
[0106] In the formula: F is the activation result; i and j are the coordinates in the feature map; γ is the loss threshold.
[0107] S404. Optimize the initial recognition model through the discard mask so that the initial recognition model discards the most class-discriminative features.
[0108] The implementation principle of this implementation manner is:
[0109] Since the ability of the initial recognition model to extract features of the tail-class images in the long-tail dataset is weaker than that of the head-class images, the initial recognition model can only detect the most class-discriminative features in the image, while ignoring a large number of other key features, reducing the robustness and prediction accuracy of the model. By obtaining the number of training samples of class C in the initial recognition model, calculating the loss threshold and the discard mask, and then optimizing the initial recognition model through the discard mask, the model discards the most class-discriminative features from the features, forcing the model to learn more features, thereby effectively improving the accuracy of the model during prediction.
[0110] The embodiment of the present invention also discloses an optimization system for a long-tail data recognition model.
[0111] Referring to Figure 6 , in one implementation manner, the system includes:
[0112] A model creation module, configured to obtain an initial long-tail data set and construct an initial recognition model based on the initial long-tail data set;
[0113] An image generation module, connected to the model creation module, configured to obtain the initial long-tail data set from the model creation module, and generate mixed images based on the initial long-tail data set by means of a minority class oversampling strategy;
[0114] A saliency calculation module, connected to the image generation module and pre-set with a calculation formula for class-related weighted saliency maps, configured to obtain the mixed images generated by the image generation module and calculate the class-related weighted saliency maps of the mixed images; the saliency calculation module obtains a calculation result by calculating the class-related weighted saliency of the mixed images, and then superimposes the calculation result on the original mixed images by means of interpolation and other methods to obtain the finally visualized result, i.e., the class-related weighted saliency map.
[0115] A weight assignment module, connected to the image generation module and the saliency calculation module, configured to obtain the class-related weighted saliency maps calculated by the saliency calculation module and the mixed images generated by the image generation module, and the weight assignment module assigns label weights to the mixed images according to the class-related weighted saliency maps;
[0116] A model optimization module, connected to the weight assignment module and the model creation module, configured to obtain the class-related weighted saliency maps calculated by the saliency calculation module, and optimize the initial recognition model by means of the class-related weighted saliency maps, so as to enhance the saliency of the class-related region features of the initial recognition model.
[0117] The implementation principle of this embodiment is as follows:
[0118] The mixed images are generated by the image generation module by means of a minority class oversampling strategy, the saliency calculation module calculates the class-related weighted saliency maps of the mixed images, and then the weight assignment module assigns label weights to the mixed images according to the class-related weighted saliency maps, so as to assign more reasonable label weights to the mixed images by means of a similar feature discrimination method. Finally, the model optimization module optimizes the initial recognition model by means of the class-related weighted saliency maps, so as to enhance the saliency of the class-related features of the initial recognition model and improve the feature representation ability of the initial recognition model, thereby improving the recognition accuracy of the recognition model for the long-tail data set.
[0119] In one of the embodiments, the calculation formula for the class-related weighted saliency maps pre-set in the saliency calculation module is as follows:
[0120]
[0121] Where: y c represents the final score with respect to class C; A dDenote the feature map of the d-th channel; i and j respectively represent the coordinates in the feature map; Denote A d 's weight; Denote the data of A at the position of coordinate ij in channel d; Z represents the width * height of the feature map. In the calculation formula of the class-related weighted saliency map, is to backpropagate through the predicted score y of the predicted class C c and then calculate the importance degree of each channel d of the feature layer A by using the gradient information A` inverted to the feature layer A. Then, the data of each channel of the feature layer A is weighted and summed by α, and finally the class-related weighted saliency map is obtained through the ReLU activation function.
[0122] In one implementation, the weight allocation module includes:
[0123] A partition setting unit, connected to the saliency calculation module and preset with a saliency threshold, is used to obtain the class-related weighted saliency map calculated by the saliency calculation module and set the salient partition in the class-related weighted saliency map according to the saliency threshold;
[0124] In this implementation, in this implementation, the image generation module generates a mixed image through the CutMix data augmentation method in the CMO method. The goal of the CutMix method is to combine the foreground image (x b , y b ) and the background image (x f , y f ) to obtain a new mixed image CMO assigns the mixed label of the generated image as y b and y f 's linear combination. The definition of generating a mixed image by CutMix is as follows:
[0125]
[0126] In the formula: M is a binary mask, and M = g(λ), γ ~ Beta(1, 1), and M is a randomly sampled rectangle under the constraint of the formula .
[0127] When assigning labels to the generated mixed image, the partition setting unit first calculates the class-related saliency maps S c (A, y f ) and s c (A, y b ) of the mixed image respectively according to the foreground image and background image labels and through the calculation formula of the class-related weighted saliency map. The preset saliency threshold ζ in this implementation is 0.6, and count S c (A, y bThe number of pixels w with significance greater than ζ in b , count S c (A, y f ) The number of pixels w with significance greater than ζ in f , so that the pixels in the class-related weighted saliency map that are more salient than the saliency threshold are set as the salient partition, and finally the target score is assigned to the mixed image according to the partition ratio of the salient partition corresponding to the label.
[0128] The weight assignment unit, connected to the partition setting unit and the image generation unit, is used to assign label weights to the mixed image according to the partition ratio of the salient partition set by the partition setting unit. The formula for the weight assignment unit to assign label weights in this embodiment is as follows:
[0129]
[0130] In the formula: w f and w b are both the number of pixels counted in step S201.
[0131] The implementation principle of this embodiment is:
[0132] In this embodiment, the method of assigning label weights through the partition setting unit and the weight assignment unit can assign label weights to the mixed image more accurately compared with the original mixing ratio assignment method of the CMO method. In the mixing ratio assignment method, the mixing ratio is only determined by the size of the binary mask, resulting in the generated mixed image having the same original ratio of label weight assignment even if there are significant changes in the image.
[0133] In one of the embodiments, the model optimization module includes:
[0134] The activation unit is used to substitute the class-related weighted saliency map into the initial recognition model and activate the class-related weighted saliency map through the activation function in the initial recognition model to obtain the activation result; in this embodiment, the activation function is the sigmoid function, and the specific formula for activation by the sigmoid function is as follows:
[0135] F = sigmoid(S c (A, y b ) + S c (A, y f ))
[0136] The mask calculation unit is preset with a feature threshold and is used to calculate the enhanced mask by combining the activation result and the feature threshold;
[0137] In this embodiment, the mask calculation unit calculates the enhanced mask M using the following formula aug :
[0138]
[0139] where: i and j are coordinates in the feature map; θ is a preset feature threshold.
[0140] An enhancement and optimization unit for optimizing the initial recognition model through an enhancement mask to enhance the saliency of class-related region features in the initial recognition model.
[0141] In this embodiment, the characteristics of enhancing the class-related region features of the initial recognition model by the enhancement and optimization unit are as follows:
[0142] A′ = A ⊙ sigmoid(A) ⊙ M aug
[0143] The implementation principle of this embodiment is:
[0144] First, a class-weighted saliency map is calculated by the saliency calculation module, and the enhancement mask M is calculated by the activation unit and the mask calculation unit according to the class-weighted saliency map aug . The enhancement and optimization unit uses the enhancement mask M aug to assign lower weights to the class-related region features of the initial recognition model, guiding the model to enhance the saliency of the class-related region features so that the model has better feature representation ability. Thereby enabling the model to extract more meaningful features corresponding to the correct class of the samples, and then significantly improving the classification accuracy of the long-tail dataset.
[0145] In one of the embodiments, the optimization system of the long-tail data recognition model further includes a feature expansion module, and the feature expansion module specifically includes:
[0146] A data acquisition unit for acquiring the number of training samples of class C in the initial model;
[0147] A first calculation unit, connected to the data acquisition unit, for acquiring the number of training samples acquired by the data acquisition unit and calculating a loss threshold according to the number of training samples; in this embodiment, the calculation formula of the loss threshold γ is as follows:
[0148] γ = sigmoid(exp(-num c ))
[0149] where: num c is the number of training samples of class C.
[0150] A second calculation unit, connected to the activation unit and the first calculation unit, for calculating a discard mask by combining the activation result and the loss threshold; in this embodiment, the discard mask M drop can be expressed as:
[0151]
[0152] In the formula: F is the activation result; i and j are coordinates in the feature map; γ is the loss threshold.
[0153] The discard optimization unit is used to optimize the initial recognition model through a discard mask, so that the initial recognition model discards the most categorical features.
[0154] The implementation principle of this embodiment is as follows:
[0155] Since the initial recognition model has a weaker ability to extract features of tail-class images in the long-tail dataset than that of head-class images, the initial recognition model can only detect the most categorical features in the image, while ignoring a large number of other key features, reducing the robustness and prediction accuracy of the model. Therefore, the number of training samples of class C in the initial recognition model can be obtained through the data acquisition unit, and the loss threshold and the discard mask can be calculated through the first calculation unit and the second calculation unit, and then the initial recognition model can be optimized through the discard optimization unit according to the discard mask, so that the model discards the most categorical features from the features, forcing the model to learn more features, thereby effectively improving the accuracy of the model during prediction.
[0156] Those of ordinary skill in the art should understand that: the discussion of any above embodiment is only exemplary, and is not intended to imply that the protection scope of the present application is limited to these examples; under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments in the present application as above, which are not provided in detail for the sake of brevity.
[0157] One or more embodiments in the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the present application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments in the present application shall be included within the protection scope of the present application.
Claims
1. An optimization method for a long-tail data recognition model, characterized in that, The method includes the following steps: Obtain an initial long-tail dataset and construct an initial recognition model based on the initial long-tail dataset; Generate mixed images based on the initial long-tail dataset and through a minority class oversampling strategy; Calculate the class-related weighted saliency map of the mixed images. The calculation formula of the class-related weighted saliency map is as follows: , Wherein: represents the final score relative to class C; represents the feature map of the d-th channel; i and j respectively represent the coordinates in the feature map; represents the weight of; represents the data of A at the position of coordinate ij in channel d; Z represents the width * height of the feature map; Assign label weights to the mixed images according to the class-related weighted saliency map, including the following steps: Set the salient partitions in the class-related weighted saliency map according to a preset saliency threshold; Assign label weights to the mixed images based on the partition ratio of the salient partitions; Optimize the initial recognition model through the class-related weighted saliency map so that the initial recognition model enhances the saliency of class-related region features, including the following steps: Substitute the class-related weighted saliency map into the initial recognition model; Activate the class-related weighted saliency map through the activation function in the initial recognition model to obtain an activation result; Calculate an enhancement mask by combining the activation result and a preset feature threshold; Optimize the initial recognition model through the enhancement mask so that the initial recognition model enhances the saliency of class-related region features.
2. The optimization method of the long-tail data recognition model according to claim 1, characterized in that, After activating the class-related weighted saliency map through the activation function in the initial recognition model to obtain an activation result, the following steps are further included: Obtain the number of training samples of class C in the initial recognition model; Calculate a loss threshold based on the number of training samples; Calculate a discard mask by combining the activation result and the loss threshold; Optimize the initial recognition model through the discard mask so that the initial recognition model discards the most class-discriminative features.
3. An optimization system for a long-tail data recognition model, characterized in that, The system includes: A model creation module, configured to obtain an initial long-tail dataset and construct an initial recognition model based on the initial long-tail dataset; An image generation module, configured to obtain the initial long-tail dataset from the model creation module, and generate mixed images based on the initial long-tail dataset and through a minority class oversampling strategy; A saliency calculation module, pre-set with a calculation formula for the class-related weighted saliency map, configured to obtain the mixed images generated by the image generation module and calculate the class-related weighted saliency map of the mixed images; The calculation formula of the class-related weighted saliency map pre-set in the saliency calculation module is as follows: , In the formula: represents the final score relative to category C; represents the feature map of the d-th channel; i and j respectively represent the coordinates in the feature map; represents the weight of; represents the data at the position of coordinates ij in channel d of A; Z represents the width * height of the feature map; A weight assignment module, configured to obtain the class-related weighted saliency map calculated by the saliency calculation module and assign label weights to the mixed images according to the class-related weighted saliency map; The weight assignment module includes: A partition setting unit, pre-set with a saliency threshold, configured to obtain the class-related weighted saliency map calculated by the saliency calculation module and set the salient partitions in the class-related weighted saliency map according to the saliency threshold; A weight assignment unit, configured to assign label weights to the mixed images according to the partition ratio of the salient partitions set by the partition setting unit. A model optimization module, configured to obtain the class-related weighted saliency map calculated by the saliency calculation module, and optimize the initial recognition model through the class-related weighted saliency map, so that the initial recognition model enhances the saliency of class-related region features; The model optimization module includes: An activation unit, configured to substitute the class-related weighted saliency map into the initial recognition model, and activate the class-related weighted saliency map through the activation function in the initial recognition model to obtain an activation result; A mask calculation unit, preset with a feature threshold, configured to calculate an enhancement mask by combining the activation result and the feature threshold; An enhancement optimization unit, configured to optimize the initial recognition model through the enhancement mask, so that the initial recognition model enhances the saliency of class-related region features.
4. The optimization system of the long-tail data recognition model according to claim 3, wherein The system further includes a feature expansion module, and the feature expansion module includes: A data acquisition unit, configured to acquire the number of training samples of class C in the initial recognition model; A first calculation unit, configured to calculate a loss threshold according to the number of training samples; A second calculation unit, connected to the activation unit and the first calculation unit, configured to calculate a discard mask by combining the activation result and the loss threshold; A discard optimization unit, configured to optimize the initial recognition model through the discard mask, so that the initial recognition model discards the most class-discriminative features.
Citation Information
Patent Citations
Long-tail learning image classification and training method and device based on mixed batch normalization
CN114863193A
Deep learning-based weakly supervised salient object detection method and system
WO2019136946A1