A method, device and electronic device for generating pseudo-label boxes
By generating high-precision pseudo-label boxes and sparse candidate boxes, the target detection model is optimized, and the problem of high cost and low accuracy of object detection model training under weak supervised learning is solved, and more efficient and accurate object detection is achieved.
Patent Information
- Application Number
- CN202210464331.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-04-29
AI Technical Summary
The existing object detection model is costly and has low accuracy under weak supervised learning, making it difficult to effectively apply in actual scenarios, especially due to insufficient detection efficiency and accuracy due to enumeration-intensive and redundant candidate boxes.
By generating a limited number of pseudo-label boxes with high precision, the attention map is used to significantly present the category targets in the target image, and combined with sparse candidate boxes to refine the subnet, the training process of the object detection model is optimized, and the enumeration and noise impact of redundant candidate boxes are reduced.
The training speed and accuracy of the object detection model are improved, the calculation amount is reduced, and the accuracy and efficiency of object detection are improved.
Smart Images

Figure CN114973064B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, and electronic device for generating pseudo-label boxes. Background Art
[0002] The object detection task is to find the objects of interest in an image or video, and at the same time detect their positions and sizes. Different from the image classification task, object detection not only has to solve the classification problem, but also has to solve the positioning problem.
[0003] Although object detection algorithms have made great progress in the past, these progress seriously rely on the big data-driven supervised learning mode. The explosive growth of the data scale and the high cost of obtaining supervision information severely restrict the application of deep learning models in actual scenarios. At present, weakly supervised learning algorithms are often used to train object detection models for object detection. However, the current cost of obtaining an object detection model based on weakly supervised learning algorithms is relatively high. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, vehicle, computer storage medium, and computer program product for generating pseudo-label boxes, which can generate a limited number of pseudo-label boxes with high precision, and improve the speed and accuracy of subsequent object detection model training.
[0005] In a first aspect, this application provides a method for generating pseudo-label boxes. The method includes: determining the categories of each object in the target image to obtain N categories; based on the N categories, processing the target image to obtain N attention maps, where each attention map is associated with one of the N categories, and each attention map is used to significantly present the objects of one category in the target image; based on the N attention maps, obtaining the pseudo-label boxes of each object in the target image.
[0006] In this way, based on the number of categories of the objects in the target image, N attention maps equal to the number are obtained, and at least one object belonging to the same category in the target image can be significantly presented on each attention map. From the obtained attention maps, the pseudo-label boxes of each object in the target image can be obtained. Since the number of obtained attention maps is limited, the number of candidate boxes obtained from the attention maps is also limited, so that it is not necessary to enumerate dense, redundant, and low-precision pseudo-label boxes, which is convenient for subsequent object detection model training and improves object detection accuracy.
[0007] In a possible implementation, the target image is processed to obtain N attention maps, which specifically includes: based on C first markers and through an attention mechanism, the target image is processed to obtain C attention maps and classification scores for C categories. Each first marker is used to learn the semantics of one category, and C ≥ N; based on the classification scores for C categories, N attention maps are selected from the C attention maps, where the classification score for each category associated with the N attention maps is higher than a preset score threshold.
[0008] In a possible implementation, after obtaining the pseudo-label boxes for each target in the target image based on the N attention maps, the method further includes: for any one of the pseudo-label boxes, adjusting the size of any one of the pseudo-label boxes in at least one direction to obtain a target pseudo-label box, and the target pseudo-label box contains a complete target. Thereby, the noise in the pseudo-label box is filtered out, and the accuracy of subsequent model training is improved.
[0009] In a possible implementation, obtaining the pseudo-label boxes for each target in the target image based on the N attention maps specifically includes: performing binarization processing on each of the N attention maps, and processing the binarized image using the connected component method to obtain the pseudo-label boxes for each target in the target image.
[0010] In a possible implementation, the method further includes: detecting the targets included in the target image to obtain a set of predicted label boxes, where the set of predicted label boxes includes the predicted label boxes for each target in the target image; training the target detection model based on the set of pseudo-label boxes and the set of predicted label boxes, where the set of pseudo-label boxes includes the pseudo-label boxes for each target in the target image. Thereby, a target detection model is obtained, and then target detection can be performed based on this target detection model.
[0011] In a possible implementation, training the target detection model based on the set of pseudo-label boxes and the set of predicted label boxes specifically includes: based on each pseudo-label box in the set of pseudo-label boxes, x predicted label boxes are selected from the set of predicted label boxes, where the value of x is equal to the number of pseudo-label boxes, and each label box in the x predicted label boxes is associated with one pseudo-label box in the set of pseudo-label boxes; based on the pseudo-label boxes in the set of pseudo-label boxes and the x predicted label boxes, the network parameters of the first network and the second network in the target detection model are updated. The first network is used to obtain the set of pseudo-label boxes, and the second network is used to obtain the set of predicted label boxes. Thereby, through this one-to-one matching method, the subsequent calculation amount is reduced, and the speed of model training is improved.
[0012] In a second aspect, the present application provides a pseudo-label box generation device, the device includes: a determination module, configured to determine the categories of each target in the target image to obtain N categories; a processing module, configured to process the target image based on the N categories to obtain N attention maps, where each attention map is associated with one of the N categories, and each attention map is used to prominently present the targets of one category in the target image; the processing module is further configured to obtain pseudo-label boxes of each target in the target image based on the N attention maps.
[0013] In a possible implementation manner, when the processing module processes the target image to obtain N attention maps, it is specifically configured to: based on C first markers, and through an attention mechanism, process the target image to obtain C attention maps and classification scores of C categories, each first marker is used to learn the semantics of one category, C≥N; based on the classification scores of the C categories, screen out N attention maps from the C attention maps, where the classification score of each category associated with the N attention maps is higher than a preset score threshold.
[0014] In a possible implementation manner, after the processing module obtains the pseudo-label boxes of each target in the target image based on the N attention maps, it is further configured to: for any one pseudo-label box, adjust the size of any one pseudo-label box in at least one direction to obtain a target pseudo-label box, and the target pseudo-label box contains a complete target.
[0015] In a possible implementation manner, when the processing module obtains the pseudo-label boxes of each target in the target image based on the N attention maps, it is specifically configured to: perform binarization processing on each of the N attention maps, and process the binarized image by using the connected component method to obtain the pseudo-label boxes of each target in the target image.
[0016] In a possible implementation manner, the processing module is further configured to: detect the targets included in the target image to obtain a set of predicted label boxes, where the set of predicted label boxes includes the predicted label boxes of each target in the target image; train the target detection model based on the set of pseudo-label boxes and the set of predicted label boxes, and the set of pseudo-label boxes includes the pseudo-label boxes of each target in the target image.
[0017] In a possible implementation, when the processing module trains the object detection model based on the pseudo-label box set and the predicted label box set, it is specifically used to: based on each pseudo-label box in the pseudo-label box set, select x predicted label boxes from the predicted label box set, where the value of x is equal to the number of pseudo-label boxes, and each label box among the x predicted label boxes is associated with one pseudo-label box in the pseudo-label box set; based on the pseudo-label boxes in the pseudo-label box set and the x predicted label boxes, update the network parameters of the first network and the second network in the object detection model, where the first network is used to obtain the pseudo-label box set, and the second network is used to obtain the predicted label box set.
[0018] In a third aspect, the present application provides an electronic device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory; wherein, when the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0019] In a fourth aspect, the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program runs on a processor, it causes the processor to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0020] In a fifth aspect, the present application provides a computer program product, characterized in that when the computer program product runs on a processor, it causes the processor to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0021] In a sixth aspect, the present application provides a chip, characterized in that it includes at least one processor and an interface; at least one processor obtains program instructions or data through the interface; at least one processor is used to execute the program line instructions to implement the method described in the first aspect or any possible implementation manner of the first aspect.
[0022] It can be understood that the beneficial effects of the above second aspect to the sixth aspect can refer to the relevant descriptions in the first aspect above, and will not be elaborated here. Description of the Drawings
[0023] Figure 1 is a schematic diagram of the network structure of an object detection model provided by an embodiment of the present application;
[0024] Figure 2 is a schematic diagram of the training steps of an object detection model provided by an embodiment of the present application;
[0025] Figure 3 is a schematic diagram of an object image provided by an embodiment of the present application;
[0026] Figure 4 It is a schematic diagram of the training process of an object detection model provided by an embodiment of the present application;
[0027] Figure 5 It is a schematic diagram of the process of generating pseudo-label boxes provided by an embodiment of the present application;
[0028] Figure 6 It is a schematic diagram of candidate boxes on an object image provided by an embodiment of the present application;
[0029] Figure 7 It is a schematic flowchart of a method for generating pseudo-label boxes provided by an embodiment of the present application;
[0030] Figure 8 It is a schematic structural diagram of a device for generating pseudo-label boxes provided by an embodiment of the present application;
[0031] Figure 9 It is a schematic structural diagram of a chip provided by an embodiment of the present application. Detailed implementation manners
[0032] In this article, the term "and / or" is an association relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In this article, the symbol " / " represents an "or" relationship between associated objects. For example, A / B represents A or B.
[0033] In the description of the embodiments of the present application, terms such as "first" and "second" in the specification and claims are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first response message and the second response message are used to distinguish different response messages, rather than to describe the specific order of the response messages.
[0034] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0035] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" refers to two or more. For example, a plurality of processing units refers to two or more processing units; a plurality of elements refers to two or more elements.
[0036] Exemplarily, for the task of object detection using weakly supervised algorithms (abbreviated as "weakly supervised object detection (WSOD)"), due to the lack of object location annotations, the weakly supervised detection algorithm needs to simultaneously estimate the object's location and learn the object detector. To achieve this goal, the WSOD algorithm often adopts the "enumerate-and-select" paradigm. For example, two-stage "enumerate-and-select" methods and end-to-end "enumerate-and-select" methods. Under this paradigm, the WSOD algorithm first enumerates the possible locations where the object may appear, and then selects the box that can best distinguish the image category as the prediction of the object location.
[0037] Among them, for the two-stage "enumerate-and-select" method, generally, prior information such as color, texture, and contour in the image is used to enumerate the object's location, and finally more than 2000 object candidate boxes are generated for each image. Based on the above candidate boxes, taking the multiple instance learning (MIL) algorithm as the basis, each image is regarded as a "bag", and the candidate boxes in the image are regarded as "instances". Combining with a deep neural network, the candidate box that locates the object is selected under the drive of the multiple instance learning loss. However, due to the large number of determined candidate boxes and the huge positional redundancy between the candidate boxes, it is difficult to select accurate candidate boxes, resulting in a low accuracy of the subsequent trained object detection model. In addition, since this method requires an additional stage to generate candidate boxes, its detection efficiency will be greatly reduced.
[0038] For the end-to-end "enumerate-and-select" method, compared with the two-stage method, the end-to-end method uses a candidate box generation network to improve the detection efficiency. In order to enable the fusion of the candidate box generation network and the deep framework, this type of method cannot use the underlying visual prior information in the traditional candidate box generation method. Therefore, the end-to-end method uses a window sweeping method to enumerate the object's location in a traversal manner, and deletes low-confidence candidate boxes using weakly supervised information. However, the candidate boxes generated by this method are very dense, which greatly increases the difficulty of candidate box screening, making it difficult for this type of method to select accurate candidate boxes, resulting in a low accuracy of the subsequent trained object detection model.
[0039] In order to obtain a target detection model with relatively high accuracy and reduce the annotation cost, an example embodiment of this application provides a method for generating pseudo-label boxes. This method mainly obtains attention maps of the number of targets (equal to the number of categories of the targets annotated by the user) based on the target image, and at least one target belonging to the same category in the target image can be significantly presented on each attention map. Then, candidate boxes (i.e., pseudo-label boxes) can be obtained from the obtained attention maps. Since the number of attention maps is limited, the number of candidate boxes obtained from the attention maps is also limited, so that it is not necessary to enumerate dense, redundant, and low-precision pseudo-label boxes, which is convenient for subsequent training of the target detection model and improves the target detection accuracy.
[0040] Exemplarily, Figure 1 shows the structure of a target detection model. It can be understood that this model can be configured in any device, equipment, platform, or device cluster with computing and processing capabilities. Figure 1 The shown target detection model is mainly obtained based on the transformer model. As Figure 1 shown, the target detection model 100 may include: a backbone network 110, a seed candidate box generation (SPG) subnet 120, and a sparse candidate box refinement (SPR) subnet 130.
[0041] Among them, the backbone network 110 is mainly used for feature extraction of the image, and it can be established based on the CaiT model of the transformer. Continuing to refer to Figure 1 , the backbone network 110 may include a convolutional layer 111 and an attention layer 112. The convolutional layer 111 is mainly used for feature extraction of the image, which is equivalent to dividing the picture information into independent (w*h) small slices. The attention layer 112 is mainly used to determine the correlation between the slices after segmentation by the convolutional layer 111 based on the self-attention mechanism, so as to avoid weakening the relationship between the slices and solve the long-distance dependence problem.
[0042] The convolutional layer 111 may include multiple convolutional operators. A convolutional operator is also called a kernel, and its role in image processing is equivalent to a filter that extracts specific information from the input image matrix. Essentially, a convolutional operator can be a weight matrix, which is usually predefined. During the convolution operation on the image, the weight matrix usually processes pixel by pixel (or two pixels by two pixels... depending on the value of the stride) along the horizontal direction of the input image, so as to complete the work of extracting specific features from the image.
[0043] The attention layer 112 may include a backbone module 1121 connected to the convolutional layer 111, a branch module 1122, and a branch module 1123. Among them, both the branch module 1122 and the branch module 1123 are connected to the backbone module 1121. The backbone module 1121, the branch module 1122, and the branch module 1123 may all be composed of one or more self-attention blocks. Exemplarily, the multiple self-attention blocks in any one of the backbone module 1121, the branch module 1122, and the branch module 1123 may be connected in series. In some embodiments, the image features processed by the backbone module 1121 can be decoupled through the branch module 1122 and the branch module 1123, so that the features output by the two branches are different, thereby improving the accuracy of subsequent processing and the generalization ability of the model.
[0044] The seed candidate box generation SPG subnet 120 is mainly used to generate seed candidate boxes, that is, obtain pseudo-label boxes, based on the image features processed by the branch module 1122 in the backbone network 110; in addition, it can also be used to classify images. Continuing to refer to Figure 1 , the SPG subnet 120 may include an image classification module 121 and a candidate box generation module 122. The image classification module 121 is mainly used to classify images based on the image features output by the backbone network 110, and / or generate an attention matrix.
[0045] The image classification module 121 may mainly be composed of two class-attention blocks and two fully connected (FC) layers. Among them, the class-attention blocks are mainly used to obtain the attention matrix, and the FC layers are mainly used to obtain the classification of the image. When processing the image features through the class-attention blocks, a learnable class token and C (for example, C may be the maximum value of the required categories defined in advance) learnable semantic category perception tokens can be introduced. In this way, finally, the class-attention blocks can output a token for labeling the original image, a token for the category included in the original image, and the classification scores corresponding to each token. Furthermore, these tokens can be used in combination with the loss function to optimize the model parameters in the SPG subnet 120, and the attention map of the required category can be selected from the classification scores. Exemplarily, each semantic category perception token can be used to learn the semantics of a category.
[0046] In some embodiments, the categories corresponding to the labels of the classes included in the output original image may be included in the C classes predefined by the user. For example, when the C classes predefined by the user are "person", "vehicle", and "plant", if the original image contains the two classes of "person" and "vehicle", then the labels of the two classes of "person" and "vehicle" may be output, but not limited to this.
[0047] The candidate box generation module 122 mainly generates seed candidate boxes, that is, pseudo-label boxes, based on the attention matrix output by the image classification module 121.
[0048] The sparse candidate box refinement SPR subnet 130 mainly performs detection on the image to be detected based on the image features processed by the branch module 1123 in the backbone network 110, and outputs the sparse candidate boxes detected in the image, that is, the object boxes of the objects included in the image. In addition, the SPR subnet 130 can be trained based on the seed candidate boxes output by the SPG subnet 120, so that the SPR subnet 130 can reach the state of performing detection as required. Continuing to refer to Figure 1 , the SPR subnet 130 may include an encoder 131, a decoder 132, and a one-to-one candidate box matching module 133.
[0049] The encoder 131 and the decoder 132 together are mainly used to output the sparse candidate boxes predicted from the image to be detected. Among them, the encoder 131 may be constituted by, but not limited to, at least one self-attention block. The decoder 132 may be constituted by, but not limited to, at least one cross-attention block and a feedforward neural network (FNN).
[0050] The one-to-one candidate box matching module 133 is mainly used to optimize the model parameters in the SPR subnet 130 based on the sparse candidate boxes detected by the decoder 132 and the seed candidate boxes output by the SPG subnet 120, so that the SPR subnet 130 meets the requirements of object detection.
[0051] The above is the introduction to the object detection model 100 provided by the embodiments of the present application. Object detection can be performed through the object detection model 100.
[0052] For ease of understanding, the training process of the object detection model 100 will be described below by taking an image as an example.
[0053] Exemplarily, Figure 2 shows the training process of the object detection model 100. As Figure 2As shown, the training process may include the following steps:
[0054] In S201, determine the categories of each target in the target image to obtain a category dataset (I) related to the categories, and design Figure 1 the network parameters G(θ) of the SPG subnet 120 and the network parameters G(γ) of the SPR subnet 130 as shown. Exemplarily, when the target image is Figure 3 the image shown, people and motorcycles can be labeled. For example, the category of a person can be labeled as "person", and the category of a motorcycle can be labeled as "mbike". In some embodiments, in S201, it is also possible to set Figure 1 1 class label required in the SPG subnet 120 and C semantic category perception labels as shown, and set K sparse candidate box labels defined in the SPR subnet 130. Exemplarily, during labeling, it can be manually labeled or automatically labeled by a machine, which can be determined according to the actual situation and is not limited here.
[0055] In S202, input the target image into the target network model 100, and use the SPG subnet 120 to generate a set (A) of pseudo-label boxes for the pseudo-label boxes of the category dataset (I).
[0056] In some embodiments, as Figure 1 and Figure 4 shown, after inputting the target image into the target network model 100, the convolutional layer 111 in the target network model 100 can extract the image features of the target image. For example, the convolutional layer 111 can divide the target image into (w*h) image patches and label each image patch. Then, the backbone module 1121 and the branch module 1122 can sequentially process the image features processed by the convolutional layer 111 based on the self-attention mechanism and input the processed image features into the SPG subnet 120. In addition, the backbone module 1121 and the branch module 1123 can sequentially process the image features processed by the convolutional layer 111 based on the self-attention mechanism and input the processed image features into the SPR subnet 130.
[0057] In the SPG subnet 120, a learnable class label t c ∈R 1×D and C semantic category perception labels t s ∈R C×D can be introduced, where C represents the maximum value of the required categories defined in advance, and D represents the feature dimension. Then, t c and t s can be input into the class attention block in the image classification module 121 of the SPG subnet 120, and the obtained image features can be calculated through the class attention mechanism to obtain a newly generated class encoding And semantic perception encoding Meanwhile, a corresponding attention matrix A ∈ R (C+1)×(C+N+1) .
[0058] Subsequently, the class encoding can be input into an FC layer in the image classification module 121. After being processed by this FC layer, the classification score of the target image can be obtained; and the semantic perception encoding can be input into another FC layer in the image classification module 121. After being processed by this FC layer, the score of each category annotated in S201 can be obtained, thereby numericalizing each category; among them, when one of the predefined C categories appears in the target image, the score of this category is high, otherwise the score is low. After obtaining the classification score of the target image and the classification scores of each category included in the target image, based on these classification scores and a preset loss function, the gap between the prediction result and the true result of the SPG subnet 120 can be determined. In addition, the required attention map can also be filtered out based on these classification scores.
[0059] Exemplarily, the loss function sampled in the SPG subnet 120 can be as follows:
[0060]
[0061] where l BCE (*) represents the binary sigmoid cross-entropy loss, w c and w s are the parameters of the above two FC layers for image classification, and y is the label of the category.
[0062] When the image classification module 121 in the SPG subnet 120 generates the attention matrix A ∈ R (C+1)×(C+N+1) , the candidate box generation module 122 can generate seed candidate boxes based on this attention matrix A ∈ R (C+1)×(C+N+1) , that is, obtain pseudo-label boxes. Among them, the candidate box generation module 122 can obtain the semantic perception attention matrix A (C+1)×(C+N+1) ∈ R * by indexing the first C rows and the middle N columns (i.e., the columns except the first C columns and the last 1 column) of the attention matrix A ∈ R C×N . Then, the candidate box generation module 122 can reshape the C-th row in A * ∈ R C×N , convert the C-th row into a (w * h)-dimensional matrix, and adjust its size to be the same as the resolution of the target image, so as to obtain the attention map A C of the C-th category. In this way, by processing the attention matrix A * ∈ R C×NBy operating on each row, C attention maps can be obtained, where each attention map corresponds to one of the C predefined categories.
[0063] Next, the candidate box generation module 122 can screen out the required attention maps from the C attention maps based on the classification scores of each category contained in the target image determined by the FC layer. The number of these attention maps can be equal to the number of categories contained in the target image, that is to say, one category contained in the target image corresponds to one attention map. Among them, at least one target belonging to the same category in the target image can be significantly presented on each of the screened attention maps. Exemplarily, as Figure 5 shown, Figure 5 (A) in is the target image. After processing this target image, the attention map shown in Figure 5 (B) can be obtained. In Figure 5 (B), the target (i.e., "airplane") in the target image can be significantly presented.
[0064] For example, when the C (C = 3) categories predefined by the user are "person", "vehicle", and "plant", if the target image contains the two categories of "person" and "vehicle", the classification scores of the two categories of "person" and "vehicle" can be output as 1, while the classification score of the category of "plant" is 0; from the classification scores, it can be determined that the categories contained in the target image are "person" and "vehicle". Therefore, the attention maps corresponding to the two categories of "person" and "vehicle" can be selected from the attention maps of the C categories as the required attention maps. One of these two attention maps can significantly present the "person" in the target image, and the other attention map can significantly present the "vehicle" in the target image.
[0065] Finally, the candidate box generation module 122 can perform binarization processing on the screened attention maps, and process the binarized image by using the method of connected components (such as the Two-Pass algorithm, etc.) to generate the seed candidate boxes contained in each attention map, so as to generate the pseudo-label box set (A) of the pseudo-label boxes of the category dataset (I). In addition, the candidate box generation module 122 can also mark the category of the target corresponding to each seed candidate box. In some embodiments, when using the method of connected components, a constraint condition can be set to filter out the noise in the attention map and improve the accuracy of generating seed candidate boxes. Exemplarily, the constraint condition can be that the area of each connected region in an attention map should be greater than N times the area of the largest connected region in the attention map, where 0 < N < 1. Exemplarily, continue to refer to Figure 5 for, Figure 5After binarizing and performing connected component processing on the attention map shown in (B), one can obtain Figure 5 the pseudo-label box shown in (C), and this pseudo-label box can at least enclose Figure 5 a part of the target (i.e., "airplane") in the target image shown in (A).
[0066] It can be understood that since the class label t c cannot distinguish semantic information, it cannot generate an attention map for each semantic category. By adding C semantic category perception labels t s , it is possible to generate an attention map for each semantic category, and thus one can obtain attention maps equal in number to the number of categories of the targets annotated by the user. From the obtained attention maps, a finite number of candidate boxes (i.e., pseudo-label boxes) can be obtained, so that it is not necessary to enumerate dense, redundant, and low-precision pseudo-label boxes, which is beneficial to the subsequent training of the target detection network and improves the target detection accuracy.
[0067] In some embodiments, the seed candidate boxes generated from the attention map may contain localization noise, which will affect the subsequent model training. For example, as Figure 6 shown in (A), the generated seed candidate box 61 does not fully cover the target at this time. To alleviate this problem, the size of the seed candidate box can be increased through the "candidate box jitter" strategy, which mainly generates bounding boxes with random jitter in four directions, thereby achieving the refinement of the seed candidate box and the improvement of the detection performance.
[0068] Exemplarily, the candidate box jitter process of the seed candidate box b i =(t x ,t y ,t w ,t h ) is defined as:
[0069] Γb i =(t x ,t y ,t w ,t h )±(ε x t x ,ε y t y ,ε w t w ,ε h t h )
[0070] where the jitter coefficients ε x ,ε y ,ε w ,ε h are from the uniform distribution U(-δ aug, +δ aug ) randomly sampled, δ aug can be a very small value to ensure that the enhanced seed candidate box Γb i is near the seed candidate box b i .
[0071] By applying the "candidate box jitter" strategy to the seed candidate box, the seed candidate box can be extended to an enhanced seed candidate box, where the class label of the enhanced seed candidate box Γb i is the same as the class label of the seed candidate box b i . Thus, noise in the seed candidate box can be corrected through seed candidate box enhancement. Exemplarily, continue to refer to Figure 6 , after enhancing the seed candidate box 61 in (A) of Figure 6 , the enhanced seed candidate box 62 shown in (B) of Figure 6 can be obtained.
[0072] In some embodiments, part or all of the processes performed by the candidate box generation module 122 in the SPR subnet 120 can also be performed in the SPG subnet 130, which can be determined according to the actual situation and is not limited here.
[0073] After the SPG subnet 120 generates the seed candidate box (i.e., the set of pseudo-label boxes (A)), the SPG subnet 120 can send the generated seed candidate box to the SPR subnet 130.
[0074] In S203, the SPR subnet 130 in the target detection module 100 is used to generate a set of predicted label boxes (B) of the predicted label boxes of the category data set (I).
[0075] In some embodiments, continue to refer to Figure 1 and 4 , after the SPR subnet 130 obtains the image features processed by the backbone module 1121 and the branch module 1123 in sequence based on the self-attention mechanism for the image features processed by the convolutional layer 111, the encoder 131 can first be used to encode the obtained image features, and then the decoder 132 can process the encoded features to obtain the set of predicted label boxes (B). Among them, in the encoder 132, a set of sparse candidate box markers t p ∈R K ×D can be defined, where K is the maximum value of the number of predefined object detections. The sparse candidate box marker t p can perform conditional cross-attention with the features encoded by the encoder 131 to obtain the encoded Subsequently, Input into the FFN to predict K sparse candidate boxes and the corresponding categories of each sparse candidate box, so as to obtain the predicted label box set (B).
[0076] In S204, the SPR subnet 130 updates the network parameters G(θ) of the SPG subnet 120 and the network parameters G(γ) of the SPR subnet 130 in reverse based on the predicted label box set (B) and the pseudo-label box set (A).
[0077] In some embodiments, the one-to-one candidate box matching module 133 in the SPR subnet 130 can use the seed candidate boxes in the pseudo-label box set (A) as pseudo-targets, and use the bipartite graph matching algorithm, such as the Hungarian algorithm, to perform the best bipartite matching on the seed candidate boxes in the pseudo-label box set (A) and the sparse candidate boxes in the predicted label box set (B), so as to select the same number of prediction results with the highest similarity between the predicted label box set (B) and the pseudo-label box set (A). For example: if the pseudo-label box set (A) contains pseudo-label boxes a0 and a1, and the predicted label box set (B) contains sparse candidate boxes b0, b1, and b2, then b0 with the highest similarity to a0 can be selected from the predicted label box set (B) with a0 as the target, and the matching result is (a0, b0); then, with a1 as the target, b1 with the highest similarity to a2 can be selected from the predicted label box set (B), and the matching result is (a1, b1); among them, in the selection process with a1 as the target, b0 can be excluded from the predicted label box set (B) first and then screened, so as to reduce the subsequent calculation amount.
[0078] Then, based on the pre-set loss function, the matching result can be processed to determine the loss of the SPR subnet 130. Finally, the network parameters G(θ) of the SPG subnet 120 and the network parameters G(γ) of the SPR subnet 130 can be updated in reverse based on the determined loss. Exemplarily, the loss function in the SPR subnet 130 can be shown as follows:
[0079]
[0080] where l FL (*), l L1 (*) and l GIoU (*) are Focal Loss, L1 loss, and Generalized IoU loss respectively, λ FL , λ L1 and λ GIoU are regularization factors, o i is the i-th semantic classification, b i is the seed candidate box corresponding to the i-th semantic classification, is the Semantic classification of a seed candidate box Indicates that the i-th sparse candidate box matches the m-th seed candidate box Is the m-th seed candidate box
[0081] In some embodiments, S202 to S204 can be repeated until a preset number of training times is reached, or the loss determined by the SPR subnet 130 based on the predicted label box set (B) and the pseudo-label box set (A) is lower than the preset loss.
[0082] After training the target detection model 100 into the required model, the target detection model 100 can be used for target detection. When using the target detection model 100 for target detection, the image to be detected can be input into the target network model 100. The convolutional layer 111, backbone module 1121, branch module 1123 in the target network model 100, the encoder 131 and decoder 132 in the SPR subnet 130 can process the image in sequence. Finally, the encoder 132 in the SPR subnet 130 outputs the detection result. Exemplarily, the detection result can include the object box and category of the detected target.
[0083] Exemplarily, when performing target detection, only the SPR subnet 130 in the target detection model 100 and the modules related to the SPR subnet 130 in the backbone network 110 can be used, without using the SPG subnet 120 and the modules related to the SPG subnet 120 in the backbone network 110, thereby reducing the computational amount of the target detection model 100.
[0084] In some embodiments, during the processing of the SPR subnet 130, the scores of the categories of each object included in the image can also be determined, and the object boxes corresponding to the detected objects (i.e., the detected candidate boxes) can be screened based on the scores of the categories of each object, thereby screening out the required object boxes. Exemplarily, when the score of the category of an object is greater than a preset threshold (such as 0.3), it can be determined that the object box corresponding to the object is the required object box.
[0085] Next, based on the content described above, a method for generating pseudo-label boxes provided by an embodiment of the present application will be introduced. It can be understood that this method is proposed based on the content described above, and part or all of the content in this method can refer to the description above.
[0086] Please refer to Figure 7 , Figure 7It is a schematic flowchart of a method for generating pseudo-label boxes provided by an embodiment of the present application. It can be understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. For example, Figure 7 As shown, the method for generating pseudo-label boxes may include:
[0087] In S701, determine the categories of each target in the target image to obtain N categories. Exemplarily, the categories of each target in the target image can be determined by manual or automatic machine annotation to obtain N categories. Exemplarily, the target image can be, but is not limited to, the Figure 3 image shown in.
[0088] In S702, based on the N categories, process the target image to obtain N attention maps, where each attention map is associated with one of the N categories, and each attention map is used to prominently present the targets of one category in the target image. Exemplarily, the target detection model 100 shown above Figure 1 can be used to process the target image to obtain N attention maps; among them, the processing process can refer to the process described in S202 above Figure 2 .
[0089] As a possible implementation, first, based on C first markers and through an attention mechanism, process the target image to obtain C attention maps and classification scores for C categories. Each first marker is used to learn the semantics of one category, and C≥N. Then, based on the classification scores of the C categories, screen out N attention maps from the C attention maps, where the classification score of each category associated with the N attention maps is higher than a preset score threshold. Thus, N attention maps are obtained. Exemplarily, the first marker can be the aforementioned semantic category perception marker, and the attention map can be the Figure 5 image shown in (B) of.
[0090] In S703, based on the N attention maps, obtain the pseudo-label boxes of each target in the target image. Exemplarily, each of the N attention maps can be binarized, and the binarized image can be processed using the connected component method to obtain the pseudo-label boxes of each target in the target image. Among them, this process can refer to the process described in S202 above Figure 2 .
[0091] Thus, based on the number of categories of the objects in the target image, an equal number of attention maps as that number are obtained, and at least one object belonging to the same category in the target image can be significantly presented on each attention map. From the obtained attention maps, the pseudo-labeling boxes of each object in the target image can be obtained. Since the number of the obtained attention maps is limited, the number of candidate boxes obtained from the attention maps is also limited, so that it is not necessary to enumerate dense, redundant, and low-precision pseudo-labeling boxes, which is convenient for subsequent training of the object detection model and improves the object detection accuracy.
[0092] In some embodiments, for any one of the obtained pseudo-labeling boxes, the size of any one of the pseudo-labeling boxes in at least one direction can be adjusted to obtain a target pseudo-labeling box, and a complete object is included in the target pseudo-labeling box. Thus, the noise in the pseudo-labeling box can be filtered out, and the accuracy of subsequent model training is improved. Among them, this strategy can be equivalent to the aforementioned "candidate box jitter" strategy. Exemplarily, continue to refer to Figure 6 , Figure 6 the box 61 shown in (A) of Figure 6 can be a pseudo-labeling box, and the box 62 shown in (B) of
[0093] In some embodiments, the Figure 7 method shown in can also detect the objects included in the target image to obtain a set of predicted labeling boxes, and the set of predicted labeling boxes includes the predicted labeling boxes of each object in the target image. And, based on the set of pseudo-labeling boxes and the set of predicted labeling boxes, the object detection model is trained, and the set of pseudo-labeling boxes includes the pseudo-labeling boxes of each object in the target image. Thus, an object detection model can be obtained, and then object detection can be performed based on the object detection model. Exemplarily, this process can be the process described in S203 and S204 in Figure 2 above.
[0094] As a possible implementation manner, when training the object detection model based on the set of pseudo-labeling boxes and the set of predicted labeling boxes, x predicted labeling boxes can be selected from the set of predicted labeling boxes based on each pseudo-labeling box in the set of pseudo-labeling boxes, where the value of x is equal to the number of pseudo-labeling boxes, and each labeling box in the x predicted labeling boxes is associated with one pseudo-labeling box in the set of pseudo-labeling boxes. Then, based on the pseudo-labeling boxes in the set of pseudo-labeling boxes and the x predicted labeling boxes, the network parameters of the first network and the second network in the object detection model are updated. The first network is used to obtain the set of pseudo-labeling boxes, and the second network is used to obtain the set of predicted labeling boxes. Thus, by this one-to-one matching method, the subsequent calculation amount is reduced, and the speed of model training is improved. Exemplarily, the first network can be Figure 1The SPG subnet 120 shown in Figure 1 the SPR subnet 130 shown in
[0095] Based on the method in the above embodiments, an embodiment of the present application provides a pseudo-label box generation device. Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a pseudo-label box generation device provided by an embodiment of the present application.
[0096] As Figure 8 shown, the pseudo-label box generation device 800 may include: a determination module 801 and a processing module 802. Among them, the determination module 801 may be used to determine the categories of each target in the target image to obtain N categories. The processing module 802 may be used to process the target image based on the N categories to obtain N attention maps, where each attention map is associated with one of the N categories, and each attention map is used to significantly present the targets of one category in the target image. In addition, the processing module 802 may also be used to obtain the pseudo-label boxes of each target in the target image based on the N attention maps.
[0097] In some embodiments, when the processing module 802 processes the target image to obtain N attention maps, it specifically is used for: based on C first markers, and through an attention mechanism, process the target image to obtain C attention maps and classification scores of C categories, each first marker is used to learn the semantics of one category, C≥N; based on the classification scores of the C categories, screen out N attention maps from the C attention maps, where the classification scores of each category associated with the N attention maps are higher than a preset score threshold.
[0098] In some embodiments, after the processing module 802 obtains the pseudo-label boxes of each target in the target image based on the N attention maps, it is further used for: for any one of the pseudo-label boxes, adjust the size of any one of the pseudo-label boxes in at least one direction to obtain a target pseudo-label box, and the target pseudo-label box contains a complete target.
[0099] In some embodiments, when the processing module 802 obtains the pseudo-label boxes of each target in the target image based on the N attention maps, it specifically is used for: perform binarization processing on each of the N attention maps, and use the method of connected components to process the binarized image to obtain the pseudo-label boxes of each target in the target image.
[0100] In some embodiments, the processing module 802 is further configured to: detect the targets included in the target image to obtain a set of predicted label boxes, where the set of predicted label boxes includes the predicted label boxes of each target in the target image; based on the set of pseudo-label boxes and the set of predicted label boxes, train the target detection model, and the set of pseudo-label boxes includes the pseudo-label boxes of each target in the target image.
[0101] In some embodiments, when the processing module 802 trains the target detection model based on the set of pseudo-label boxes and the set of predicted label boxes, it is specifically configured to: based on each pseudo-label box in the set of pseudo-label boxes, select x predicted label boxes from the set of predicted label boxes, where the value of x is equal to the number of pseudo-label boxes, and each label box in the x predicted label boxes is associated with a pseudo-label box in the set of pseudo-label boxes; based on the pseudo-label boxes in the set of pseudo-label boxes and the x predicted label boxes, update the network parameters of the first network and the second network in the target detection model, where the first network is used to obtain the set of pseudo-label boxes, and the second network is used to obtain the set of predicted label boxes.
[0102] It should be understood that the above device is used to execute the method in the above embodiments. For the corresponding program modules in the device, the implementation principles and technical effects are similar to those described in the above method. The working process of the device can refer to the corresponding process in the above method, and will not be elaborated here.
[0103] Based on the method in the above embodiments, an electronic device is provided in an embodiment of the present application. The electronic device may include: at least one memory for storing a program; at least one processor for executing the program stored in the memory; wherein, when the program stored in the memory is executed, the processor is configured to execute the method in the above embodiments.
[0104] Based on the method in the above embodiments, a computer-readable storage medium is provided in an embodiment of the present application. The computer-readable storage medium stores a computer program, and when the computer program runs on a processor, the processor is caused to execute the method in the above embodiments.
[0105] Based on the method in the above embodiments, a computer program product is provided in an embodiment of the present application. It is characterized in that when the computer program product runs on a processor, the processor is caused to execute the method in the above embodiments.
[0106] Based on the method in the above embodiments, a chip is further provided in an embodiment of the present application. Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a chip provided in an embodiment of the present application. As Figure 9 shown, the chip 900 includes one or more processors 901 and an interface circuit 902. Optionally, the chip 900 may further include a bus 903. Wherein:
[0107] The processor 901 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 901 or instructions in the form of software. The above-mentioned processor 901 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods and steps disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0108] The interface circuit 902 can be used for sending or receiving data, instructions, or information. The processor 901 can utilize the data, instructions, or other information received by the interface circuit 902 for processing, and can send the processed information through the interface circuit 902.
[0109] Optionally, the chip 900 further includes a memory. The memory may include a read-only memory and a random access memory, and provides operation instructions and data to the processor. A part of the memory may also include a non-volatile random access memory (NVRAM).
[0110] Optionally, the memory stores executable software modules or data structures. The processor can execute corresponding operations by calling the operation instructions stored in the memory (the operation instructions can be stored in the operating system).
[0111] Optionally, the interface circuit 902 can be used to output the execution result of the processor 901.
[0112] It should be noted that the respective functions corresponding to the processor 901 and the interface circuit 902 can be implemented through hardware design, can also be implemented through software design, or can be implemented through a combination of software and hardware. There is no limitation here.
[0113] It should be understood that each step of the above method embodiments can be completed by the logic circuit in the form of hardware in the processor or instructions in the form of software.
[0114] It can be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. In addition, in some possible implementation manners, the steps in the above embodiments can be selectively executed according to the actual situation, can be partially executed, or can be fully executed. There is no limitation here.
[0115] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0116] The method steps in the embodiments of the present application may be implemented in a hardware manner or by a processor executing software instructions. The software instructions may be composed of corresponding software modules, and the software modules may be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may be located in an ASIC.
[0117] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
[0118] It can be understood that the various digital numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.
Claims
1. A method for generating pseudo-label boxes, characterized in that, The method includes: Determining the categories of each target in the target image to obtain N categories; the N categories are obtained by annotating the target image; Processing the target image based on the N categories to obtain N attention maps, where each attention map is associated with one of the N categories, and each attention map is used to present the targets of one category in the target image; Based on the N attention maps, obtaining pseudo-label boxes for each target in the target image.
2. The method according to claim 1, characterized in that, The processing of the target image to obtain N attention maps specifically includes: Based on C first markers and through an attention mechanism, processing the target image to obtain C attention maps and classification scores for C categories, where each first marker is used to learn the semantics of one category, and C≥N; Based on the classification scores of the C categories, screening out the N attention maps from the C attention maps, where the classification score of each category associated with the N attention maps is higher than a preset score threshold.
3. The method according to claim 1 or 2, characterized in that, After obtaining the pseudo-label boxes for each target in the target image based on the N attention maps, the method further includes: For any one of the pseudo-label boxes, adjusting the size of the any one of the pseudo-label boxes in at least one direction to obtain a target pseudo-label box, and the target pseudo-label box contains a complete target.
4. The method according to claim 1 or 2, characterized in that, The obtaining of the pseudo-label boxes for each target in the target image based on the N attention maps specifically includes: Performing binarization processing on each of the N attention maps, and processing the binarized image by using the connected component method to obtain the pseudo-label boxes for each target in the target image.
5. The method according to claim 1 or 2, characterized in that, The method further includes: Detecting the targets included in the target image to obtain a set of predicted label boxes, where the set of predicted label boxes includes the predicted label boxes for each target in the target image; Training a target detection model based on the set of pseudo-label boxes and the set of predicted label boxes, where the set of pseudo-label boxes includes the pseudo-label boxes for each target in the target image.
6. The method according to claim 5, wherein The training of the target detection model based on the set of pseudo-label boxes and the set of predicted label boxes specifically includes: Based on each pseudo-label box in the set of pseudo-label boxes, selecting x predicted label boxes from the set of predicted label boxes, where the value of x is equal to the number of pseudo-label boxes, and each label box in the x predicted label boxes is associated with one pseudo-label box in the set of pseudo-label boxes; Based on the pseudo-label boxes in the set of pseudo-label boxes and the x predicted label boxes, updating the network parameters of the first network and the second network in the target detection model, where the first network is used to obtain the set of pseudo-label boxes, and the second network is used to obtain the set of predicted label boxes.
7. A pseudo-label box generation device, characterized in that, The device includes: A determination module, configured to determine the categories of each target in the target image to obtain N categories; the N categories are obtained by annotating the target image; A processing module, configured to process the target image based on the N categories to obtain N attention maps, where each attention map is associated with one of the N categories, and each attention map is used to present an object of one category in the target image; The processing module is further configured to obtain pseudo-label bounding boxes of each object in the target image based on the N attention maps.
8. The device according to claim 7, characterized in that, When the processing module processes the target image to obtain N attention maps, it is specifically configured to: Based on C first markers and through an attention mechanism, process the target image to obtain C attention maps and classification scores of C categories. Each first marker is used to learn the semantics of one category, and C≥N; Based on the classification scores of the C categories, screen out the N attention maps from the C attention maps, where the classification score of each category associated with the N attention maps is higher than a preset score threshold.
9. The device according to claim 7 or 8, characterized in that, After the processing module obtains the pseudo-label bounding boxes of each object in the target image based on the N attention maps, it is further configured to: For any one of the pseudo-label bounding boxes, adjust the size of the any one of the pseudo-label bounding boxes in at least one direction to obtain a target pseudo-label bounding box, and the target pseudo-label bounding box contains a complete object.
10. The device according to claim 7 or 8, characterized in that When the processing module obtains the pseudo-label bounding boxes of each object in the target image based on the N attention maps, it is specifically configured to: Perform binarization processing on each of the N attention maps, and process the binarized image by means of connected components to obtain the pseudo-label bounding boxes of each object in the target image.
11. The device according to claim 7 or 8, characterized in that, The processing module is further configured to: Detect the objects included in the target image to obtain a set of predicted label bounding boxes, where the set of predicted label bounding boxes includes the predicted label bounding boxes of each object in the target image; Train a target detection model based on the set of pseudo-label bounding boxes and the set of predicted label bounding boxes, where the set of pseudo-label bounding boxes includes the pseudo-label bounding boxes of each object in the target image.
12. The device according to claim 11, characterized in that, When the processing module trains the target detection model based on the set of pseudo-label bounding boxes and the set of predicted label bounding boxes, it is specifically configured to: Based on each pseudo-label bounding box in the set of pseudo-label bounding boxes, select x predicted label bounding boxes from the set of predicted label bounding boxes, where the value of x is equal to the number of pseudo-label bounding boxes, and each label bounding box in the x predicted label bounding boxes is associated with one pseudo-label bounding box in the set of pseudo-label bounding boxes; Based on the pseudo-label bounding boxes in the set of pseudo-label bounding boxes and the x predicted label bounding boxes, update the network parameters of the first network and the second network in the target detection model. The first network is used to obtain the set of pseudo-label bounding boxes, and the second network is used to obtain the set of predicted label bounding boxes.
13. An electronic device, characterized in that, Including: At least one memory for storing programs; At least one processor for executing the programs stored in the memory; Wherein, when the programs stored in the memory are executed, the processor is configured to execute the method according to any one of claims 1-6.
14. A computer-readable storage medium storing a computer program, which, when running on a processor, causes the processor to execute the method according to any one of claims 1-6.
15. A computer program product, characterized in that, When the computer program product runs on a processor, it causes the processor to execute the method according to any one of claims 1-6.
16. A chip, characterized in that, Comprising at least one processor and an interface; The at least one processor obtains program instructions or data through the interface; The at least one processor is configured to execute the program line instructions to implement the method according to any one of claims 1-6.
Citation Information
Patent Citations
Image detection model training method and device
CN111563541A