A target localization method and system based on multi-scale attention constraints

By augmenting and scaling the training images in multiple scale data, combined with the prediction classification of neural networks and the fusion of class activation maps, the problem of incomplete and inaccurate target positioning areas in the prior art is solved, especially when dealing with pictures with inconsistent target scales, higher positioning accuracy is achieved.

CN114202649BActive Publication Date: 2025-05-13NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111536372.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-15
Publication Date
2025-05-13
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

The existing weakly supervised target positioning methods have shortcomings in the problem of incomplete and inaccurate positioning areas, especially when dealing with input images to be tested with inconsistent target scales.

Method used

By augmenting and scaling the training images, multiple input images of different scales are generated, and neural networks are used to fusion of predictive classification and class activation graphs. Combining the cross entropy loss and divergence loss function, the neural network is trained by the stochastic gradient descent method to obtain the trained neural network for target positioning.

Benefits of technology

It improves the accuracy of target positioning, can better deal with the problem of inconsistent target scales in the real world, enhances the attention of the target area and suppresses the wrong attention of the background area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114202649B_ABST
    Figure CN114202649B_ABST
Patent Text Reader

Abstract

The invention relates to a target positioning method and system based on multi-scale attention constraint. The method comprises: performing data augmentation and scaling on training images to obtain a plurality of input images of different scales; performing prediction classification by using a neural network according to the input images to obtain classification categories; determining a cross entropy loss function according to the classification categories; fusing a plurality of input images of different scales to determine a plurality of class activation maps of different scales; determining a divergence loss function of attention and a positioning result according to the plurality of class activation maps of different scales; taking the plurality of input images of different scales as input, taking the classification categories as output, taking the cross entropy loss function and the divergence loss function as loss functions, and training the parameters of the neural network by using a stochastic gradient descent method to obtain a trained neural network; inputting a test image into the trained neural network to obtain positioning information. The invention improves the accuracy of target positioning by processing the image scale.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target positioning, and in particular to a target positioning method and system based on multi-scale attention constraints. Background Art

[0002] In the field of computer vision, target localization is a very important and basic research topic. Usually, in order to train the target localization model, a large number of images with target location information are required. However, labeling the target location in the image requires a lot of human resources. The weakly supervised target localization method does not require expensive bounding box annotation information, but only uses the category annotation information of the image, which can greatly reduce the labor cost. Therefore, the target localization method based on weak supervision has great research significance and practical value.

[0003] Since weakly supervised target localization methods only require coarse-grained image category information and do not require precise target location information, weakly supervised target localization has attracted increasing attention from scholars at home and abroad in recent years. Some people have proposed the concept of Class Activation Maps (CAM), which adds a global average pooling layer to the network, calculates the network's attention map on the predicted target, and locates the target area through threshold segmentation. Others have proposed an adversarial complementary learning method, which locates different and complementary target areas by training a network model with two adversarial complementary branches, and finally merges them to generate a complete target area. Others have proposed a lightweight and efficient module ADL (Attention-Based Dropout Layer) based on the attention mechanism. This module generates an attention map for convolutional features, and then randomly selects the area corresponding to the higher value in the attention map to erase the corresponding area of ​​the original image, or uses the attention map to generate an importance coefficient and multiplies it with the original feature, so that the network can achieve the target localization function while maintaining a certain classification accuracy. Others have proposed a geometric constraint network model, which directly outputs a rectangular or elliptical area containing the target by training an object detector, and then uses this area to erase the input image to obtain the target area and background area respectively. Finally, a multi-task loss function is combined to constrain the geometric shape of the output of the object detector to achieve the target positioning function. Others have proposed a method that reduces the network's sensitivity to image prediction category errors and image changes, so that the network's attention is fully activated within the target range, thereby achieving the target positioning function.

[0004] There is a dual-gradient method in the prior art to achieve target positioning under weak supervision; some prior arts use multiple component perception modules to capture the attention maps of multiple components in the target, and obtain the importance weights of multiple components, and achieve weak supervision target positioning through the attention maps of multiple components and the corresponding importance weights; there are also prior arts that reversely transfer the gradient from the network output layer to the input layer layer by layer, and calibrate the gradients of the fully connected layer, convolutional layer and other layers in the network, so as to achieve weak supervision target positioning. There are also prior arts that use difference divergence learning and hierarchical divergence learning to mine the positioning information of the target from different angles, and finally achieve the weak supervision target positioning function. There are also prior arts that generate difficult samples by occluding the discriminative area of ​​the target for model training to achieve weak supervision target positioning.

[0005] Although the above methods can obtain good target positioning results, there are still problems such as incomplete and inaccurate positioning areas. In addition, the input images to be tested in real life often have the problem of different target scales. Summary of the invention

[0006] The purpose of the present invention is to provide a target positioning method and system based on multi-scale attention constraints, which improves the accuracy of target positioning by processing the image scale.

[0007] To achieve the above object, the present invention provides the following solutions:

[0008] A target localization method based on multi-scale attention constraints, comprising:

[0009] Perform data augmentation and scaling on the training images to obtain multiple input images of different scales;

[0010] Perform prediction classification using a neural network according to the input image to obtain a classification category;

[0011] Determining a cross entropy loss function according to the classification category;

[0012] Using a neural network to fuse the multiple input images of different scales to determine multiple class activation maps of different scales;

[0013] Determining a divergence loss function and a positioning result of attention according to a plurality of class activation maps of different scales;

[0014] Taking the input images of the plurality of different scales as input, taking the classification category as output, taking the cross entropy loss function and the divergence loss function as loss functions, and using the stochastic gradient descent method to train the parameters of the neural network to obtain a trained neural network;

[0015] The test image is input into the trained neural network to obtain positioning information; the positioning information includes classification category and positioning result.

[0016] Optionally, the expression of the cross entropy loss function is:

[0017]

[0018] Among them, L ce represents the cross entropy loss, p c Represents the predicted classification, y c Represents the labeled category, C represents the number of categories, and c represents the predicted category.

[0019] Optionally, the fusing the multiple input images of different scales by using a neural network to determine the multiple class activation maps of different scales specifically includes:

[0020] Generating first-class activation maps of different scales using a neural network according to the plurality of input images of different scales;

[0021] Fusing the first class activation maps of different scales to obtain a fused class activation map;

[0022] The fused class activation map is optimized using a fully connected conditional random field to obtain multiple class activation maps of different scales.

[0023] Optionally, the expression of the fusion class activation map is:

[0024] CAM fusion =(CAM' 0.5 +CAM1+CAM' 1.5 +CAM'2) / 4

[0025] Among them, CAM fusion Represents the fusion class activation map, CAM' 0.5 represents the first-class activation map with a scaling ratio of 0.5, CAM1 represents the first-class activation map with a scaling ratio of 1, and CAM' 1.5 represents the first-class activation map with a scaling factor of 1.5, and CAM'2 represents the first-class activation map with a scaling factor of 2.

[0026] Optionally, the expression of the class activation graph is:

[0027] CAM crf =Q(X;CAM fusion )

[0028] Among them, CAM crf Represents class activation maps of different scales, CAM fusion represents the fusion class activation map, Q represents the fully connected conditional random field, and X represents the training image.

[0029] Optionally, the expression of the attention divergence loss function is:

[0030]

[0031] Among them, L KL represents the divergence loss, n represents the number of input images, u represents each spatial position in the class activation map, c represents the predicted classification, X represents the input image with a scaling factor of 1, and f u,c (X) represents the class activation map CAM1 extracted from the network, CAM crf Represents a class activation map.

[0032] A target localization system based on multi-scale attention constraints, comprising:

[0033] The data augmentation and scaling module is used to perform data augmentation and scaling on training images to obtain multiple input images of different scales;

[0034] A prediction and classification module, used to perform prediction and classification based on the input image using a neural network to obtain a classification category;

[0035] A cross entropy loss determination module, used to determine a cross entropy loss function according to the classification category;

[0036] A fusion module, used to fuse the multiple input images of different scales using a neural network to determine multiple class activation maps of different scales;

[0037] A divergence loss determination module, used to determine a divergence loss function and a positioning result of attention according to a plurality of class activation maps of different scales;

[0038] A training module, used to take the input images of the plurality of different scales as input, take the classification category as output, take the cross entropy loss function and the divergence loss function as loss functions, and use the stochastic gradient descent method to train the parameters of the neural network to obtain a trained neural network;

[0039] The positioning result determination module is used to input the test image into the trained neural network to obtain positioning information; the positioning information includes classification categories and positioning results.

[0040] Optionally, the expression of the cross entropy loss function is:

[0041]

[0042] Among them, L ce represents the cross entropy loss, p c Represents the predicted classification, y cRepresents the labeled category, C represents the number of categories, and c represents the predicted category.

[0043] Optionally, the fusion module specifically includes:

[0044] A generating unit, configured to generate first-type activation maps of different scales using a neural network according to the plurality of input images of different scales;

[0045] A fusion unit, used for fusing the first class activation maps of different scales to obtain a fused class activation map;

[0046] The optimization unit is used to optimize the fused class activation map using a fully connected conditional random field to obtain multiple class activation maps of different scales.

[0047] Optionally, the expression of the fusion class activation map is:

[0048] CAM fusion =(CAM' 0.5 +CAM1+CAM' 1.5 +CAM'2) / 4

[0049] Among them, CAM fusion Represents the fusion class activation map, CAM' 0.5 represents the first-class activation map with a scaling ratio of 0.5, CAM1 represents the first-class activation map with a scaling ratio of 1, and CAM' 1.5 represents the first-class activation map with a scaling factor of 1.5, and CAM'2 represents the first-class activation map with a scaling factor of 2.

[0050] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0051] The present invention performs data augmentation and scaling on training images to obtain multiple input images of different scales; performs prediction classification using a neural network according to the input images to obtain classification categories; determines cross entropy loss according to the classification categories; fuses multiple input images of different scales using a neural network to determine multiple class activation maps of different scales; determines attention divergence loss and positioning results according to the multiple class activation maps of different scales; and uses multi-scale fusion technology in the training stage to make the model easier to handle the problem of inconsistent target scales in images to be tested in the real world, thereby improving the accuracy of target positioning. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0053] Figure 1 A flowchart of a target localization method based on multi-scale attention constraints provided by the present invention;

[0054] Figure 2 A flowchart of the target localization method based on multi-scale attention constraints provided by the present invention in practical application;

[0055] Figure 3 Class activation maps and fusion result maps obtained for images of different scales;

[0056] Figure 4 Visualization of the prediction results for the ImageNet dataset;

[0057] Figure 5 A visual comparison of the positioning results of the target localization method based on multi-scale attention constraints and ADL and CAM.

[0058] Figure 6 A schematic diagram of a target localization system based on multi-scale attention constraints provided by the present invention. DETAILED DESCRIPTION

[0059] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0060] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0061] like Figure 1 As shown, the present invention provides a target positioning method based on multi-scale attention constraints, characterized by comprising:

[0062] Step 101: Perform data augmentation and scaling on the training images to obtain multiple input images of different scales.

[0063] Step 102: Use a neural network to perform prediction classification based on the input image to obtain a classification category.

[0064] Step 103: Determine a cross entropy loss function according to the classification category. The expression of the cross entropy loss is:

[0065]

[0066] Among them, L ce represents the cross entropy loss, p c Represents the predicted classification, y c Represents the labeled category, C represents the number of categories, and c represents the predicted category.

[0067] Step 104: fuse the multiple input images of different scales using a neural network to determine multiple class activation maps of different scales. Step 104 specifically includes:

[0068] A neural network is used to generate first-class activation maps of different scales according to the multiple input images of different scales.

[0069] The first class activation maps of different scales are fused to obtain a fused class activation map. The expression of the fused class activation map is:

[0070] CAM fusion =(CAM' 0.5 +CAM1+CAM' 1.5 +CAM'2) / 4

[0071] Among them, CAM fusion Represents the fusion class activation map, CAM' 0.5 represents the first-class activation map with a scaling ratio of 0.5, CAM1 represents the first-class activation map with a scaling ratio of 1, and CAM' 1.5 CAM0.5, CAM1, CAM1.5, CAM2 are four activation maps of different scales and sizes, and cannot be directly fused. Therefore, they must be scaled to the same size before fusion. The present invention scales the three activation maps CAM0.5, CAM1.5, and CAM2 to the size of CAM1. Therefore, CAM' 0.5 ,CAM' 1.5 ,CAM'2 The three prime characters represent the scaled activation map.

[0072] The fused class activation map is optimized using a fully connected conditional random field to obtain multiple class activation maps of different scales.

[0073] Wherein, the expression of the class activation graph is:

[0074] CAM crf =Q(X;CAMfusion )

[0075] Among them, CAM crf Represents class activation maps of different scales, CAM fusion represents the fusion class activation map, Q represents the fully connected conditional random field, and X represents the training image.

[0076] Step 105: Determine the attention divergence loss function and positioning results according to the plurality of class activation maps of different scales. The expression of the attention divergence loss is:

[0077]

[0078] Among them, L KL represents the divergence loss, n represents the number of input images, u represents each spatial position in the class activation map, c represents the predicted classification, X represents the input image with a scaling factor of 1, and f u,c (X) represents the class activation map CAM1 extracted from the network, CAM crf Represents a class activation map.

[0079] The process of determining the positioning result is as follows:

[0080] The activation map output by the convolutional layer of the neural network is used for weighted summation to obtain the object class activation map;

[0081] The object class activation map is segmented by a threshold value to obtain a positioning result. The region in the object class activation map whose value is greater than a set threshold value is determined as the positioning result of the object.

[0082] Step 106: Taking the multiple input images of different scales as input, taking the classification category as output, taking the cross entropy loss function and the divergence loss function as loss functions, and using the stochastic gradient descent method to train the parameters of the neural network to obtain a trained neural network.

[0083] Step 107: Input the test image into the trained neural network to obtain positioning information; the positioning information includes classification category and positioning result.

[0084] like Figure 2 As shown, the present invention also provides a workflow of a target localization method based on multi-scale attention constraints in practical applications. The network in the present invention is any classification network, and the convolution output features of the convolution layer in the classification network are classified through a global average pooling layer, a fully connected layer, and a softmax layer. The steps are as follows:

[0085] (1) Perform data augmentation on training images.

[0086] (2) Scale the training images at different ratios to obtain multiple input images of different scales.

[0087] (3) Input the input image with a scaling factor of 1 into the network to obtain the predicted classification and calculate the cross entropy loss L ce .

[0088] (4) Input the multiple input images of different scales into the network to generate corresponding class activation maps of different scales.

[0089] (5) Fusion the class activation maps of different scales.

[0090] (6) Input the fused class activation map into the fully connected conditional random field for optimization to obtain a better class activation map CAM crf .

[0091] (7) CAM crf As auxiliary constraint information of network attention, calculate the attention constraint loss L KL .

[0092] (8) Calculate the cross entropy loss L ce and KL divergence loss L kl The weighted sum of is taken as the total loss L.

[0093] (9) Use the total loss L and the stochastic gradient descent algorithm to update the network parameters and jump to (2) to process the next training image.

[0094] (10) During the test, the test image is processed through the network to obtain the predicted category, and the positioning result is obtained using the CAM method. That is, the positioning result is determined by the class activation map, and the specific determination method is the existing technology. After the test image is input into the network, the output of the last layer of the network is the predicted category. The positioning result is obtained by extracting the information of the middle layer of the network (the output of the convolutional layer) and then using the CAM method.

[0095] Optionally, the step (2) includes:

[0096] Each training image is scaled at four scales {0.5, 1, 1.5, 2} to obtain input images A of four scales corresponding to the training image. 0.5 ,A1,A 1.5 ,A2.

[0097] Optionally, the step (3) includes:

[0098] Input sample A1 with a scaling factor of 1 into the network to obtain the predicted classification;

[0099] Calculate the cross entropy loss based on the labeled classification of the training image and the predicted classification.

[0100] Optionally, the step (4) includes:

[0101] Input image A with 4 scaling ratios {0.5, 1, 1.5, 2} 0.5 ,A1,A 1.5 ,A2, respectively input the network for forward operation and obtain the corresponding class activation map CAM 0.5 ,CAM1,CAM 1.5 ,CAM2.

[0102] Optionally, the step (5) includes:

[0103] (5.1) First, CAM 0.5 ,CAM 1.5 ,CAM2 is scaled according to the size of CAM1 to obtain CAM' 0.5 ,CAM' 1.5 ,CAM'2, its size is the same as CAM1;

[0104] (5.2) According to formula (1), the four scaled CAMs are fused. The specific calculation method is to find the average value by element, which is expressed as follows: CAM fusion =(CAM' 0.5 +CAM1+CAM' 1.5 +CAM'2) / 4 ⑴

[0105] Optionally, step (6) includes: inputting the fused class activation map and the original training image into a fully connected conditional random field module for optimization, and the optimized output is CAM crf .

[0106] Optionally, the step (7) includes: crf As auxiliary supervision information, CAM1 and CAM are calculated crf The KL divergence loss between them is used as the network attention constraint loss.

[0107] Compared with the existing inventions, the present invention has the following advantages: 1) The multi-scale fusion technology is used in the training stage to make the model easier to handle the problem of different target scales in the test pictures in the real world. 2) The fusion of multi-scale class activation maps can effectively enhance the attention of the target area while suppressing the wrong attention of the background area. Using the fused and optimized class activation map as the constraint condition of the network target attention can make the network's attention to the target area more complete and accurate. 3) The method does not require complex changes to the model structure and can be easily applied to various existing models.

[0108] like Figure 2As shown in the figure, the training image input into the network is used for forward operation to obtain the predicted category, and the cross entropy loss is calculated with the real category of the image annotation. At the same time, the training image is scaled to different proportions, and the corresponding class activation maps of different scales are obtained after forward operation of the network. The class activation maps of different scales are fused and optimized using fully connected conditional random fields to obtain a more accurate class activation map. The fused and optimized class activation map is used as auxiliary supervision information to calculate the KL divergence loss of the class activation map and the network attention, so as to achieve the purpose of constraining the network attention and make the network attention pay more attention to the target area in the picture.

[0109] The specific steps are as follows:

[0110] The first step is to standardize and augment the data. The data of the training image samples is standardized and the input images are augmented to increase the diversity of the training data.

[0111] Specifically, for the images X1, X2…X in the training data set n , standardize its pixel values ​​so that the numerical distribution of each pixel value is a standard normal distribution. Then perform data augmentation operations such as color transformation and random flipping on the image to enrich the training set samples. Finally, the size of the image after the augmentation operation is uniformly adjusted to 256×256 and used as the final training image.

[0112] The second step is to calculate the classification cross entropy loss. The training image is input into the network for forward operation to obtain the predicted category, and the cross entropy loss is calculated with the labeled category corresponding to the training image.

[0113] The cross entropy loss is calculated as:

[0114]

[0115] Among them, p c Represents the predicted classification, y c represents the label category, and C represents the number of categories.

[0116] In the third step, the training images are scaled according to four scaling ratios, and then input into the network to obtain the corresponding class activation maps of different scales and perform fusion optimization.

[0117] Specifically, the 256×256 training image obtained in the first step is first reduced and enlarged, with the scaling ratios being {0.5, 1, 1.5, 2}, to obtain four images of different scales corresponding to the training image, namely the input images. Among them, the image with a scaling ratio of 1 is essentially the training image itself. The input images of the four scales are respectively forward-operated through the network to obtain the corresponding class activation maps of different scales and fuse them, such as Figure 3As shown in Figure 2, it can be found that the target activation area after fusion is more accurate and complete.

[0118] The specific method of fusion is as follows:

[0119] For the four scales of class activation maps obtained, CAM 0.5 ,CAM1,CAM 1.5 , CAM2. Among them, CAM1 can represent network attention. By scaling the CAM 0.5 ,CAM 1.5 The size of CAM2 is adjusted to be consistent with CAM1, and they are respectively denoted as CAM' 0.5 ,CAM' 1.5 , CAM'2, and then fuse them through formula (1) to obtain CAM fusion .

[0120] Furthermore, we use the fully connected conditional random field to CAM fusion Optimize and obtain the final optimized class activation map CAM crf . It is expressed as follows:

[0121] CAM crf =Q(X;CAM fusion )⑶

[0122] Among them, Q represents the fully connected conditional random field and X represents the training image.

[0123] The fourth step is to use the class activation map CAM after multi-scale fusion optimization crf As auxiliary supervision information, calculate the KL divergence loss between it and the network attention. crf As auxiliary supervision information for network attention, that is, calculating the KL divergence loss between it and CAM1.

[0124] The specific calculation method is:

[0125]

[0126] Where n is the number of input images, u is each spatial position in the class activation map, c is the predicted classification, X is the input image with a scaling factor of 1, and f is u,c (X) represents the class activation map CAM1 extracted from the network.

[0127] In the fifth step, the total loss Loss is calculated based on the cross entropy loss calculated in the third step and the KL divergence loss calculated in the fourth step, as shown below:

[0128] Loss = αL ce +βL kl ⑸

[0129] Among them, α and β are weight factors, which keep the cross entropy loss and KL divergence loss at the same order of magnitude. The total loss and stochastic gradient descent algorithm are used to update the network parameters, so as to learn a positioning model that can more accurately locate the target area.

[0130] Figure 4 The following is a visualization of the positioning results of the method of the present invention on the authoritative dataset ImageNet in this field. 4(a) is the positioning result of the network for a test image containing a "bowl". 4(b) is the positioning result of the network for a test image containing a "cantaloupe". 4(c) and 4(d) are both the positioning results of the network for a test image containing a "dog". 4(e) is the positioning result of the network for a test image containing a "bird". In the above test results, the dark gray box represents the real bounding box of the target, and the light gray box represents the bounding box predicted by the network for the target. It can be found that the method of the present invention can locate targets relatively accurately on both large and small targets.

[0131] Figure 5 By comparing the positioning effects of the method of the present invention with those of the classical methods CAM and ADL in the art, it can be found that the positioning effect of the method of the present invention is better.

[0132] like Figure 6 As shown, the present invention provides a target positioning system based on multi-scale attention constraints, comprising:

[0133] The data augmentation and scaling module 201 is used to perform data augmentation and scaling on the training images to obtain multiple input images of different scales.

[0134] The prediction and classification module 202 is used to perform prediction and classification based on the input image using a neural network to obtain a classification category.

[0135] The cross entropy loss determination module 203 is used to determine a cross entropy loss function according to the classification category.

[0136] The fusion module 204 is used to fuse the multiple input images of different scales using a neural network to determine multiple class activation maps of different scales.

[0137] The divergence loss determination module 205 is used to determine the divergence loss function and positioning results of attention according to the multiple class activation maps of different scales.

[0138] The training module 206 is used to take the input images of the multiple different scales as input, take the classification category as output, take the cross entropy loss function and the divergence loss function as loss functions, and use the stochastic gradient descent method to train the parameters of the neural network to obtain a trained neural network.

[0139] The positioning result determination module 207 is used to input the test image into the trained neural network to obtain positioning information; the positioning information includes classification category and positioning result.

[0140] In practical applications, the expression of the cross entropy loss is:

[0141]

[0142] Among them, L ce represents the cross entropy loss, p c Represents the predicted classification, y c Represents the labeled category, C represents the number of categories, and c represents the predicted category.

[0143] In practical applications, the fusion module 204 specifically includes:

[0144] A generating unit is used to generate first-class activation maps of different scales using a neural network according to the multiple input images of different scales.

[0145] The fusion unit is used to fuse the first-class activation maps of different scales to obtain a fused class activation map.

[0146] The optimization unit is used to optimize the fused class activation map using a fully connected conditional random field to obtain multiple class activation maps of different scales.

[0147] In practical applications, the expression of the fusion class activation map is:

[0148] CAM fusion =(CAM' 0.5 +CAM1+CAM' 1.5 +CAM'2) / 4

[0149] Among them, CAM fusion Represents the fusion class activation map, CAM' 0.5 represents the first-class activation map with a scaling ratio of 0.5, CAM1 represents the first-class activation map with a scaling ratio of 1, and CAM' 1.5 represents the first-class activation map with a scaling factor of 1.5, and CAM'2 represents the first-class activation map with a scaling factor of 2.

[0150] The present invention discloses that in the training stage, each training image is scaled at different ratios to obtain input images of different scales. Among these input images, the input image with a scaling ratio of 1 is subjected to network operation to obtain a predicted category for calculating the cross entropy loss. At the same time, the category label of the training image and the class activation map (CAM) generation method are used to obtain class activation maps of different scales corresponding to these input images. These class activation maps of different scales are fused and then optimized using a fully connected conditional random field. The fused and optimized class activation map can locate a more complete and accurate target area than the class activation map of the original scale. The fused and optimized class activation map is used as auxiliary supervision information, and the KL divergence loss between it and the class activation map of the original scale is calculated to constrain the network attention, and finally the positioning ability of the network is improved. By fusing the target attention map at multiple scales during the training process and using it as the auxiliary constraint information of the network, the model can predict a more complete and accurate target positioning result.

[0151] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0152] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A target localization method based on multi-scale attention constraints, characterized in that: include: Perform data augmentation and scaling on the training images to obtain multiple input images of different scales; Perform prediction classification using a neural network according to the input image to obtain a classification category; Determining a cross entropy loss function according to the classification category; Using a neural network to fuse the multiple input images of different scales to determine multiple class activation maps of different scales; The attention divergence loss function and the positioning result are determined according to the multiple class activation maps of different scales; the expression of the attention divergence loss function is: Among them, L KL represents the divergence loss, n represents the number of input images, u represents each spatial position in the class activation map, c represents the predicted classification, X represents the input image with a scaling factor of 1, and f u,c (X) represents the class activation map CAM1 extracted from the network, CAM crf represents the class activation map; Taking the input images of the plurality of different scales as input, taking the classification category as output, taking the cross entropy loss function and the divergence loss function as loss functions, and using the stochastic gradient descent method to train the parameters of the neural network to obtain a trained neural network; Inputting the test image into the trained neural network to obtain positioning information; the positioning information includes classification category and positioning result; The step of fusing the plurality of input images of different scales by using a neural network to determine a plurality of class activation maps of different scales specifically includes: generating first class activation maps of different scales by using a neural network according to the plurality of input images of different scales; fusing the first class activation maps of different scales to obtain a fused class activation map; optimizing the fused class activation map by using a fully connected conditional random field to obtain a plurality of class activation maps of different scales; The expression of the fusion class activation map is: ORANGE fusion =(CAM' 0.5 +CAM1+CAM' 1.5 +CAM'2) / 4; Among them, CAM fusion Represents the fusion class activation map, CAM' 0.5 represents the first-class activation map with a scaling ratio of 0.5, CAM1 represents the first-class activation map with a scaling ratio of 1, and CAM' 1.5 represents the first-class activation map with a scaling ratio of 1.5, and CAM'2 represents the first-class activation map with a scaling ratio of 2; The expression of the class activation graph is: CAM crf =Q(X;CAM fusion ); Among them, CAM crf Represents the class activation map, CAM fusion represents the fusion class activation map, Q represents the fully connected conditional random field, and X represents the training image.

2. The target localization method based on multi-scale attention constraints according to claim 1, characterized in that: The expression of the cross entropy loss function is: Among them, L ce represents the cross entropy loss, p c Represents the predicted classification, y c Represents the labeled category, C represents the number of categories, and c represents the predicted category.

3. A target localization system based on multi-scale attention constraints, characterized in that: include: The data augmentation and scaling module is used to perform data augmentation and scaling on training images to obtain multiple input images of different scales; A prediction and classification module, used to perform prediction and classification based on the input image using a neural network to obtain a classification category; A cross entropy loss determination module, used to determine a cross entropy loss function according to the classification category; A fusion module, used to fuse the multiple input images of different scales using a neural network to determine multiple class activation maps of different scales; The divergence loss determination module is used to determine the divergence loss function and positioning result of the attention according to the multiple class activation maps of different scales; the expression of the divergence loss function of the attention is: Among them, L KL represents the divergence loss, n represents the number of input images, u represents each spatial position in the class activation map, c represents the predicted classification, X represents the input image with a scaling factor of 1, and f u,c (X) represents the class activation map CAM1 extracted from the network, CAM crf represents the class activation map; A training module, used to take the input images of the plurality of different scales as input, take the classification category as output, take the cross entropy loss function and the divergence loss function as loss functions, and use the stochastic gradient descent method to train the parameters of the neural network to obtain a trained neural network; A positioning result determination module, used to input the test image into the trained neural network to obtain positioning information; the positioning information includes classification category and positioning result; The fusion module specifically includes: a generation unit, which is used to generate first-class activation maps of different scales using a neural network according to the multiple input images of different scales; a fusion unit, which is used to fuse the first-class activation maps of different scales to obtain a fused class activation map; an optimization unit, which is used to optimize the fused class activation map using a fully connected conditional random field to obtain multiple class activation maps of different scales; The expression of the fusion class activation map is: ORANGE fusion =(CAM' 0.5 +CAM1+CAM' 1.5 +CAM'2) / 4; Among them, CAM fusion Represents the fusion class activation map, CAM' 0.5 represents the first-class activation map with a scaling ratio of 0.5, CAM1 represents the first-class activation map with a scaling ratio of 1, and CAM' 1.5 represents the first-class activation map with a scaling factor of 1.5, and CAM'2 represents the first-class activation map with a scaling factor of 2.

4. The target positioning system based on multi-scale attention constraints according to claim 3 is characterized in that: The expression of the cross entropy loss is: Among them, L ce represents the cross entropy loss, p c Represents the predicted classification, y c Represents the labeled category, C represents the number of categories, and c represents the predicted category.

Citation Information

Patent Citations

  • Small target detection method based on multi-scale images and weighted fusion loss

    CN111461110A

  • Weak supervision semantic segmentation method based on adaptive affinity and category allocation

    CN112668579A