Weakly Supervised Object Localization Method Based on Adversarial Erasure and Background Suppression
Through the method of anti-erase and background suppression, weak-supervised target positioning technology is improved, and the semantic information and intermediate feature information of high-level feature maps are used to solve the problem of inaccurate positioning in the prior art, achieving more accurate target area coverage.
Patent Information
- Application Number
- CN202211192618.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-09-28
AI Technical Summary
Existing weak-supervised target positioning technology cannot accurately locate image targets, especially ignore the position, texture, and edge information of the intermediate features of the network, resulting in poor positioning effect.
Using the method of anti-erase and background suppression, the pooling layer and fully connected layer of the convolutional neural network are replaced into the global average pooling layer and the fully connected layer, and feature erasing and background suppression are performed in combination with the category activation map and erasing map. The semantic information of the high-level feature map is used to guide the network to activate more target areas and suppress background areas to achieve accurate coverage.
Improve the accuracy of target positioning, make full use of intermediate feature information, reduce background activation, and generate more accurate target area positioning maps.
Smart Images

Figure CN115546471B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of target localization, and particularly relates to a weakly supervised target localization method based on adversarial erasure and background suppression. Background Art
[0002] Object detection and object localization are key technologies in the field of computer vision and are widely used in various life scenarios, such as the detection and localization of people, vehicles, objects, and industrial product defects. In recent years, with the increase in data volume and computing performance, object detection and object localization technologies based on deep learning have developed rapidly and shown excellent results in various fields. However, the deep models for object detection and localization require a large amount of data and accurate instance-level annotations during the training process, and the acquisition cost is extremely high. In the actual application process, existing algorithms often can only be trained based on a small amount of fully annotated data and cannot achieve ideal results. Therefore, researchers have proposed weakly supervised target localization technologies, hoping to train object detection and localization models only using low-cost image-level annotations. Existing weakly supervised target localization technologies need to first train a deep classification convolutional network through image category labels, and then obtain a class activation map (CAM) and perform threshold segmentation. However, the localization map required for class activation mapping is obtained from a deep classification model supervised by a single-class loss function. The model tends to distinguish various objects through the significant spatial features of each object. The results obtained by this method usually only locate the significant parts of the object (such as the wings of a bird, the head of a dog) rather than the entire object. Therefore, the localization effect is poor. This is the main problem existing in the weakly supervised target localization model based on CAM.
[0003] Yang Jinfu et al. provided "A Method for Adversarial Elimination Weakly Supervised Object Detection Based on Channel Selection" in the Chinese invention open patent CN110569901A. It uses an adversarial elimination method to remove the channels that contribute the most to the classification task in the feature map and inputs them into another branch network two with the same structure but without shared parameters, so as to force the network to capture less discriminative target spatial features during the process of learning classification, and thus be able to activate a more complete target area. However, the two branches in the network structure of this solution do not share parameters, which will lead to an increase in network parameters and increase the learning difficulty of the network; and the data streams cannot be shared between the branches, and the information of the two data streams is not fully utilized; and the purpose of this solution is to capture less discriminative spatial features, and the method is to adversarially eliminate the most discriminative channel features. The spatial features and channel features are not in an equivalent relationship, so the role of erasure cannot be fully exerted.
[0004] In addition, a convolutional neural network is composed of multiple stacked convolutional layers. The convolutional network in the lower layer recognizes primary image features, while the higher-layer network contains abstract semantic information. Existing methods first train a deep convolutional network through image class labels, and then obtain a class activation map from the high-level feature map of the network, ignoring the fact that the intermediate features in the network contain information about the location, texture, and edges of the target, and this information is very helpful for correctly locating the foreground target. Summary of the Invention
[0005] The present invention aims to improve the problem that existing weakly supervised object localization algorithms cannot accurately locate image objects, and provides a weakly supervised object localization method based on adversarial erasing and background suppression, which can guide the neural network to discover general object regions that are easily suppressed in the classification task, and reduce the activation of the background as a non-target region, and finally generate a localization map that accurately covers the complete target region, effectively realizing the weakly supervised object localization task.
[0006] To achieve the object of the present invention, the weakly supervised object localization method based on adversarial erasing and background suppression provided by the present invention includes the following steps:
[0007] Use a convolutional neural network as the backbone network, replace the pooling layer and fully connected layer at its end with a global average pooling layer and a fully connected layer, and the network finally directly outputs a vector composed of the prediction probabilities of each category;
[0008] Input the training set images into the convolutional neural network with the replaced structure, perform the first forward propagation, and after propagating to the last convolutional layer, output to obtain the feature map F, which contains n channels (f1~f n ), f n is the nth channel in the feature map F. After passing through the global average pooling layer, a one-dimensional eigenvalue V (v1~v n ) with a length of n is obtained, v n is the channel value of the nth channel. Each eigenvalue corresponds to a channel in the feature map F. After the one-dimensional eigenvalue V is input into the fully connected layer, a one-dimensional vector category prediction result with a length of C is obtained;
[0009] Let the parameters in the fully connected layer corresponding to the true label of the image be w, with n values (w1~w n ), corresponding to the n values of the one-dimensional eigenvalue V and the n channels of the feature map F respectively. Use the parameter w to evaluate the contribution degree of each channel in the feature map F to the correct category, and calculate the class activation map CAM, and calculate the two-dimensional map CAM norm and the erasing map M e ;
[0010] Take the erasing map M eThe intermediate feature F of the l-th layer of the convolutional neural network with the feature map F obtained from the first forward propagation l Perform an erasing operation, input the erased feature into the layer after the l-th layer of the convolutional neural network to continue the second forward propagation, and obtain the second category prediction result;
[0011] Respectively take the intermediate features of the m-th layer (m > l) in the convolutional neural network during the first forward propagation and the second forward propagation, perform channel average pooling on each of them and then pass through the activation function to obtain their respective importance maps, denoted as
[0012] Iteratively train the convolutional neural network based on the total loss function to obtain the trained convolutional neural network;
[0013] Input the image to be detected into the trained convolutional neural network to obtain the target localization result.
[0014] Furthermore, before inputting the training set images into the convolutional neural network, first perform normalization processing on the training set images.
[0015] Furthermore, the calculation formula of the class activation map CAM is
[0016]
[0017] In the formula, CAM is the class activation map.
[0018] Furthermore, the erasing map M e The calculation formula of is
[0019]
[0020] In the formula, γ is a preset erasing threshold.
[0021] Furthermore, the activation function is the Sigmoid function.
[0022] Furthermore, the total loss function includes the cross-entropy loss between the two category predictions and the true category labels and the background suppression loss.
[0023] Furthermore, the calculation formula of the cross-entropy loss is
[0024]
[0025] In the formula, C represents there are C categories, y i represents the true label, represents the prediction result.
[0026] Furthermore, the calculation formula of the background suppression loss is
[0027]
[0028] In the formula, L bs is the background suppression loss, and S is the spatial scale size of the output features of the m-th convolutional layer of the convolutional neural network.
[0029] Furthermore, the error backpropagation algorithm is used to train the convolutional neural network.
[0030] Furthermore, in the test stage, the test set images are input into the trained convolutional neural network, and only one forward propagation is performed. Finally, the network output obtains the class prediction result of the input image. Among them, for the two-dimensional CAM norm threshold segmentation is performed to obtain the target localization result.
[0031] Compared with the prior art, the beneficial effects that the present invention can achieve are at least as follows:
[0032] (1) The present invention provides a weakly supervised target localization method based on adversarial erasing and background suppression, which can use the semantic information of the high-level feature map for erasing, introduce multiple data streams in the convolutional neural network, prompt the network to activate more target regions, and can make full use of the intermediate feature information to realize the activation suppression of the non-target background region, so that the activation region of the network more accurately covers the target, effectively improving the localization accuracy.
[0033] (2) After the first forward propagation of the present invention, the generated erasing features are input back into the same network again, that is, there is only one branchless network from beginning to end, but there are two data streams in this network. In this way, not only no new parameters are added to the network, but all network parameters can be updated by the two data streams at the same time.
[0034] (3) The present invention directly performs spatial scale erasing based on the most discriminative target spatial features extracted from the first forward propagation, which can further prompt the network to learn the less discriminative target spatial features.
[0035] (4) Regarding "ignoring that the intermediate features in the network contain information about the position, texture, and edges of the target": the present invention proposes a background activation suppression to standardize the expression of the intermediate features, suppress the expression of their background regions, and make their activation regions concentrated in the target region range to improve the subsequent classification and localization effects. Description of the Drawings
[0036] Figure 1 It is a schematic diagram of the overall network architecture provided by the embodiment of the present invention.
[0037] Figure 2 It is a schematic diagram of the calculation method of the background suppression loss provided by the embodiment of the present invention.
[0038] Figure 3A flowchart of the network training stage provided by the embodiments of the present invention.
[0039] Figure 4 A flowchart of the network test application stage provided by the embodiments of the present invention. Detailed implementation manners
[0040] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts are within the scope of protection of the present invention.
[0041] Before introducing the specific solutions, some terms will be introduced first.
[0042] 1. Convolutional neural network: A deep feedforward neural network containing convolutional calculations, one of the basic models of deep learning, which is very suitable for application in the field of computer vision. Its basic structure includes a convolutional layer, a pooling layer, and a fully connected layer, which are used to extract image features, perform data downsampling, and output results respectively.
[0043] 2. Class Activation Map: Also known as class activation mapping, class heatmap, saliency map, etc. The class activation map is obtained by weighted superposition of the feature map corresponding to the class on the original image, and can visualize the areas that the deep neural network focuses on when predicting the class. It is the same size as the input original image, and the larger the pixel value in the map, the greater the contribution of the corresponding position in the original image to the prediction of the deep neural network.
[0044] 3. Thresholding segmentation: The thresholding segmentation method is a region-based image segmentation technology. It uses the difference in gray values between the target region and the background to divide the image into several types of images by setting different thresholds. The advantage of thresholding segmentation is that it is simple to implement, has a small amount of calculation, and has relatively stable performance. Now it has become the most basic and widely used segmentation technology in image segmentation.
[0045] 4. Image normalization: A process of performing a series of standard processing transformations on an image to transform it into a fixed standard form.
[0046] 5. Cross Entropy: A method for measuring the difference information between two probability distributions.
[0047] 6. Learning rate: An important hyperparameter in supervised learning and deep learning, representing the update speed of the model, which determines whether the objective function can converge to the local minimum and when it converges to the minimum.
[0048] 7. Error Back Propagation: An algorithm for training multi-layer neural networks. By calculating the gradients of the loss function with respect to each weight parameter, it obtains the optimization direction, with advantages such as solid theoretical basis, rigorous derivation process, clear physical concepts, and strong generality.
[0049] 8. Forward propagation: Generally refers to the process in which input data sequentially undergoes pre-set operations in the model and finally obtains the output value.
[0050] 9. Gradient Descent: An iterative method commonly used in machine learning algorithms. By continuously optimizing and approaching the optimal solution of the problem in the direction of the gradient descent, it can be used to solve the least squares problem.
[0051] 10. Intersection over Union: An evaluation metric widely used in tasks such as segmentation and object detection. The specific calculation method is the ratio of the overlapping part of two regions to the union part of the two regions.
[0052] The weakly supervised object localization method based on adversarial erasing and background suppression provided by the present invention includes the following steps:
[0053] Step 1: Normalize the training set images so that the values of each pixel in the RGB three channels are in the range of 0 to 1.
[0054] Step 2: Use a pre-trained convolutional neural network as the backbone network, and replace the pooling layer and fully connected layer at its end with a global average pooling layer and a fully connected layer to retain the target space information in the features for facilitating subsequent localization tasks. The network finally directly outputs a vector composed of the predicted probabilities of each category (assuming there are C categories).
[0055] Step 3: Input the training set images into the pre-trained convolutional neural network in Step 2, perform the first forward propagation. After propagating to the last convolutional layer, output the feature map F, which contains n channels (f1~f n ), the shape of the feature map F is three-dimensional H*W*n, H*W is the spatial scale size of the feature, n is the number of channels, f nRefers to a certain channel among the n channels in the feature map F (three-dimensional), that is, a two-dimensional feature. After passing through the global average pooling layer, a one-dimensional eigenvalue V (v1 to v n ) of length n is obtained. The n channels in the feature map F correspond to n two-dimensional features H*W. After performing spatial pooling on these n two-dimensional feature maps, a vector of length n is obtained, and v n is a certain value in this vector. Each eigenvalue corresponds to a channel in the feature map F. After the one-dimensional eigenvalue V is input into the fully connected layer, a one-dimensional vector class prediction result of length C is obtained.
[0056] Step 4: Let the parameters in the fully connected layer corresponding to the true image label be w, which has n values (w1 to w n ), corresponding to the n values of the one-dimensional eigenvalue V and the n channels of the feature map F mentioned in Step 3 respectively. Among them, the one-dimensional feature V (v1 to v n ) has a length of n. After passing through the fully connected layer, a one-dimensional vector of length C is obtained (this vector of length C is the prediction probability for each category). This fully connected layer has a total of n*C parameters. If only the parameters of the fully connected layer corresponding to the correct category are concerned, there are only n*1 parameters. w is the vector composed of these n parameters, and w n is a certain parameter in w. As shown in formula (1), the parameter w is used to evaluate the contribution degree of each channel in the feature map F to the correct category, and the class activation map CAM is calculated. After maximum-minimum normalization, CAM norm is obtained. Set the erasing threshold to γ, and the erasing map M is calculated by formula (2) e .
[0057]
[0058]
[0059]
[0060] f i is the i-th channel among the n channels (f1 to f n ) of the feature map F, w i is a certain value among the n values (w1 to w n ) of w, The i and j in refer to the position of a certain pixel. CAM min and CAM max are respectively the minimum and maximum values among the values corresponding to all pixel points in the two-dimensional graph CAM; Formula (3) means that if the value of the position i, j of a certain pixel point in CAM norm ≤γ, then M eThe value at its pixel position i, j is set to 1. In other words, it is the two-dimensional map CAM norm The values at the positions where the values in norm are less than or equal to γ are all set to 1, and the values at all other pixel positions are set to 0, thus obtaining a two-dimensional erasure map M with the same shape but only values of 0 and 1 e .
[0061] In some embodiments of the present invention, the erasure threshold γ is taken as 0.6. Of course, in other embodiments, the value of γ can be adjusted according to different data sets to obtain the best positioning performance
[0062] Step 5: Take the erasure map M obtained in Step 4 e and the intermediate feature F of the l-th layer of the convolutional neural network obtained in Step 3 l Perform a dot product operation, that is, perform an erasure operation. Input the erased feature into the layer after the l-th layer of the convolutional neural network to continue the second forward propagation, and finally, like in Step 3, obtain the second class prediction result
[0063] Step 6: Respectively take the intermediate features of the m-th layer (m > l) in the convolutional neural network in the first forward propagation in Step 3 and the second forward propagation in Step 5, perform channel average pooling on each of them and then pass through the activation function to obtain their respective importance maps, denoted as
[0064] In some embodiments of the present invention, the activation function uses the Sigmoid function. The Sigmoid function constrains the values of each pixel in the importance map between 0 and 1; the gradient propagation effect is better
[0065] Step 7: Construct the total loss function
[0066] In some embodiments of the present invention, the class prediction results obtained in Step 3 and Step 5 are respectively calculated with the true class labels. As shown in formula (3), their respective cross-entropy losses are obtained where y i represents the true label (if the i-th class is the true class of the input image, then y i is 1, otherwise it is 0), represents the prediction result (the predicted probability that the input image is the i-th class). The importance map obtained in Step 6 is calculated according to formula (4) to obtain the background suppression loss L bs , where S is the spatial scale size of the output feature of the convolutional layer of the m-th layer of the convolutional neural network (that is, the number of pixels in a single channel of the feature map). and L bs together constitute the total loss function L
[0067]
[0068]
[0069]
[0070] Step 8: Use the error backpropagation algorithm to train the convolutional neural network.
[0071] In some embodiments of the present invention, based on the total loss function L in Step 7, the partial derivatives of the parameters in the network are obtained using the gradient descent method, and the parameters are updated using the product of the partial derivative values and the learning rate, and the iteration is repeated until the total loss function L no longer decreases significantly.
[0072] Step 9: In the test phase, input the test set images into the convolutional neural network trained through the above steps, and only perform one forward propagation. Finally, the network output obtains the class prediction result of the input image. Then, calculate the CAM corresponding to the input image according to the description in Step 4 norm . Perform threshold segmentation on the two-dimensional map CAM norm . The threshold value is in the range of 0 to 1. Pixel positions greater than the threshold are assigned a value of 1, otherwise 0. Finally, draw a rectangular box that exactly contains all the pixel points with a value of 1 as the localization result. The selected threshold should make the overlap degree between the predicted localization box and the true localization box as large as possible.
[0073] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A weakly supervised object localization method based on adversarial erasure and background suppression, characterized in that Including the following steps: Using a convolutional neural network as the backbone network, replacing the pooling layer and fully connected layer at its end with a global average pooling layer and a fully connected layer, and the network finally directly outputs a vector composed of the prediction probabilities of each category; Input the training set images into the convolutional neural network with the changed input structure, perform the first forward propagation, and output the feature map after propagating to the last convolutional layer. , which contains channels . is the -th channel in the feature map. After passing through the global average pooling layer, a one-dimensional eigenvalue of length is obtained. . is the channel value of the -th channel. Each eigenvalue corresponds to a channel in the feature map . After the one-dimensional eigenvalue is input into the fully connected layer, a one-dimensional vector class prediction result of length C is obtained. Let the parameters in the fully connected layer corresponding to the true labels of the images be , with n values , corresponding to n values of the one-dimensional eigenvalue and n channels of the feature map respectively. Use the parameter to evaluate the contribution degree of each channel in the feature map to the correct category, and calculate the class activation map CAM. Based on the class activation map CAM, calculate the two-dimensional map Erasure map ; Erase the said erasure graph and the feature map obtained from the first forward propagation the intermediate features of the layer of the convolutional neural network are subjected to an erasure operation, and the erased features are input into the layer after the layer of the convolutional neural network to continue the second forward propagation to obtain a second class prediction result; Respectively take the intermediate features of the m-th layer in the convolutional neural network during the first forward propagation and the second forward propagation, perform channel average pooling on each of them, and obtain their respective importance maps after passing through the activation function, denoted as , , where ; Iteratively training the convolutional neural network based on the total loss function, where the total loss function includes background suppression loss, to obtain the trained convolutional neural network; Inputting the image to be detected into the trained convolutional neural network to obtain the target localization result.
2. The weakly supervised object localization method based on adversarial erasure and background suppression according to claim 1, wherein Before inputting the training set images into the convolutional neural network, first perform normalization processing on the training set images.
3. The weakly supervised object localization method based on adversarial erasure and background suppression according to claim 1, wherein The calculation formula of the class activation map CAM is In the formula, is the class activation map.
4. The weakly supervised object localization method based on adversarial erasure and background suppression according to claim 1, wherein The erasure graph has the following calculation formula Wherein, Preset erasure threshold value.
5. The weakly supervised object localization method based on adversarial erasure and background suppression according to claim 1, wherein The activation function is the Sigmoid function.
6. The weakly supervised object localization method based on adversarial erasure and background suppression according to claim 1, characterized in that, The total loss function includes the cross-entropy loss of two times of class prediction and the true class label, as well as background suppression loss.
7. The weakly supervised object localization method based on adversarial erasure and background suppression according to claim 6, wherein The calculation formula of the cross-entropy loss is In the formula, indicates that there are C categories, represents the true label, represents the predicted result.
8. The weakly supervised object localization method based on adversarial erasure and background suppression according to claim 6, characterized in that The calculation formula of the background suppression loss is Wherein, is the background suppression loss, and S is the spatial scale size of the output features of the m-th convolutional layer of the convolutional neural network.
9. The weakly supervised object localization method based on adversarial erasure and background suppression according to claim 1, wherein Using the error backpropagation algorithm to train the convolutional neural network.
10. The weakly supervised object localization method based on adversarial erasure and background suppression according to any one of claims 1-9, characterized in that, In the test phase, the test set images are input into the trained convolutional neural network, and only one forward propagation is performed. Finally, the network output obtains the class prediction result of the input image. Among them, for the two-dimensional graph Threshold segmentation is performed to obtain the target localization result.
Citation Information
Patent Citations
Adversarial elimination weak supervision target detection method based on channel selection
CN110569901A
Feature weak supervision image semantic segmentation method based on hierarchical joint convolutional network
CN110929744A
Weak supervision time sequence behavior positioning method based on hierarchical category model
CN113221633A