An Active Learning Image Object Detection Method and Device with Adaptive Annotation Types
Through adaptive selection of annotation targets and joint training models, the problem of high-cost annotation in the existing technology is solved, and efficient object detection and model generalization are achieved.
Patent Information
- Application Number
- CN202111435129.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-11-29
AI Technical Summary
The existing technology relies on large-scale annotation data sets, and manual annotation costs are high, and the model generalization ability is limited when random sampling annotation is used.
An active learning method for label type adaptive is proposed. Through the detection model decoupling classification and positioning tasks, the detection model is decoupled, and the fully supervised and weakly supervised data are used to train together, and a semi-supervised detection model is designed for iterative training.
Significantly save labeling costs, improve the target detection algorithm's ability to judge target categories and locations, and enhance the generalization ability of model.
Smart Images

Figure CN114155398B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of adaptive active learning, and particularly to an active learning image object detection method and device with adaptive annotation types. Background Art
[0002] In the related art, object detection methods based on convolutional neural networks mainly rely on large-scale data sets and full supervision training, including mainly two-stage detectors based on candidate boxes, single-stage detectors based on anchor boxes, and box-free detectors based on feature points.
[0003] Generally, two-stage detection first extracts candidate boxes by means of selective search or region extraction network, then extracts the image features of the candidate boxes, and makes class and position predictions. Girshick et al. first used a convolutional neural network to extract the features of candidate boxes, and classification and localization were respectively implemented by a support vector machine and a regression model; the spatial pyramid pooling model maps the candidate boxes to the feature map, and only one forward calculation is required for the whole image, and the pooling layer is inserted before the last fully connected layer of the network, and a fixed-length image representation can be obtained without scaling the candidate boxes. Also, in the related art, the accuracy of single-stage detection methods has reached the level of two-stage methods, but a large number of background anchor boxes limit the performance of the network.
[0004] However, these algorithms still rely on data sets with large scale, diverse patterns, and detailed annotations, and the cost of manual annotation becomes more time-consuming and complex. It is feasible to only select a part of representative data for annotation. If image data is randomly sampled for annotation, sufficient rich information can only be guaranteed when the sampling scale is large enough, otherwise the generalization ability of the model will be severely affected. Summary of the Invention
[0005] The present invention aims to solve at least one of the technical problems in the related art to some extent.
[0006] To this end, the first object of the present invention is to propose an active learning image object detection method with adaptive annotation types, to achieve the decoupling of multiple objects in a single image and the decoupling of classification and localization tasks, and to use the joint training of full supervision and weak supervision data to maximize the savings in annotation costs.
[0007] The second object of the present invention is to propose an active learning image object detection device with adaptive annotation types.
[0008] The third object of the present invention is to propose a non-transitory computer-readable storage medium.
[0009] The fourth object of the present invention is to propose a computer program product.
[0010] To achieve the above object, an embodiment of the first aspect of the present invention provides a method, including:
[0011] Detect the target detection object with the detection model to obtain the positioning information and classification information corresponding to the target object
[0012] For the target object whose classification information meets the first preset condition, select the most valuable object for annotation according to the quantization index quota to obtain the category label corresponding to the target detection object. For the target detection object whose positioning information meets the second preset condition, select the most valuable detection object for annotation according to the quantization index quota to obtain the supplementary bounding box label corresponding to the detection target object;
[0013] Generate the first annotation data of the target object according to the category label and the supplementary bounding box label, and add the first annotation data to the annotation data set, wherein the second annotation data is pre-stored in the annotation data set;
[0014] Retrain the semi-supervised detection model according to the annotation data set to obtain the iteratively updated target semi-supervised detection model until the model reaches the expected performance or the annotation quantity reaches the budget.
[0015] Optionally, in an embodiment of the present application, the initial semi-supervised detection model is designed in the following manner:
[0016] Extract the multi-scale feature map of the image, use the center point estimation as the first branch, and the weakly supervised global average pooling as the second branch. The first branch and the second branch share the parameters of part of the multi-scale feature map;
[0017] In the first branch, perform convolution on the multi-scale feature map to obtain the predicted position information;
[0018] In the second branch, perform convolution on the multi-scale feature map to obtain a response map that can be supervised by the image-level label. Optionally, in an embodiment of the present application, detecting the target image to obtain the positioning information and classification information corresponding to the target object includes:
[0019] Predict the target image through the initial semi-supervised detection model to obtain the positioning information and classification information corresponding to the target object.
[0020] Optionally, in an embodiment of the present application, the first preset condition includes a target object with a classification information amount higher than a first specific threshold, and the second preset condition includes a target object with a positioning information amount higher than a second specific threshold.
[0021] Optionally, in an embodiment of the present application, target objects whose classification information meets the first preset condition are labeled to obtain class labels corresponding to the target objects, and target objects whose positioning information meets the second preset condition are labeled to obtain supplementary bounding box labels corresponding to the target objects, including:
[0022] Measure the class information quantity of the target with entropy:
[0023]
[0024] Wherein, is to measure the class information quantity of the target with entropy, represents the class prediction probability at the center point coordinate , and c is the total number of candidate classes.
[0025] is to calculate the positioning information quantity at the center point coordinate . First, calculate the local probability distribution expectation of the scale compensation matrix :
[0026]
[0027] Wherein, r defines the local neighborhood radius.
[0028] Secondly, measure the mutual information between the data distribution and the model prediction distribution with the difference between the entropy of the local average prediction value and the mean of the prediction value entropy, and use as an estimate of the positioning information quantity:
[0029]
[0030] Wherein, calculate the information entropy, which is defined here as:
[0031]
[0032] Similarly, obtain the size information quantity at the center point coordinate Use to represent the total positioning information quantity:
[0033]
[0034] Set thresholds ∈ c , ∈ l for classification and positioning respectively. When the class information quantity of the target exceeds the corresponding threshold, adaptively provide the corresponding type of label.
[0035] To achieve the above object, an embodiment of the second aspect of the present invention proposes an active learning image target detection device with adaptive annotation types, including:
[0036] A detection module, configured to detect a target detection object using a detection model to obtain positioning information and classification information corresponding to the target object;
[0037] An evaluation module, for target objects whose classification information meets the first preset condition, selects the most valuable object for annotation according to the quantization index quota to obtain a category label corresponding to the target detection object, and for target detection objects whose positioning information meets the second preset condition, selects the most valuable detection object for annotation according to the quantization index quota to obtain a supplementary bounding box label corresponding to the detection target object;
[0038] A labeling module, configured to generate first labeling data of the target object according to the category label and the supplementary bounding box label, and add the first labeling data to a labeling data set, wherein second labeling data is pre-stored in the labeling data set;
[0039] A training module, configured to retrain a semi-supervised detection model according to the labeling data set to obtain an iteratively updated target semi-supervised detection model until the model reaches the expected performance or the labeling quantity reaches the budget.
[0040] Optionally, in an embodiment of the present application, the initial semi-supervised detection model is designed in the following manner:
[0041] Extract multi-scale feature maps of an image, use center point estimation as the first branch, and weakly supervised global average pooling as the second branch, and the first branch and the second branch share part of the parameters of the multi-scale feature maps;
[0042] In the first branch, perform convolution on the multi-scale feature maps to obtain predicted position information;
[0043] In the second branch, perform convolution on the multi-scale feature maps to obtain a response map that can be supervised by an image-level label.
[0044] Optionally, in an embodiment of the present application, the detection module is further configured to:
[0045] Perform prediction on the target image through the initial semi-supervised detection model to obtain positioning information and classification information corresponding to the target object.
[0046] Optionally, in an embodiment of the present application, the first preset condition includes target objects with classification information amount higher than a first specific threshold, and the second preset condition includes target objects with positioning information amount higher than a second specific threshold.
[0047] To achieve the above object, an embodiment of the third aspect of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the active learning image target detection method with adaptive annotation types described in the embodiment of the first aspect of the present application is implemented.
[0048] To achieve the above object, an embodiment of the fourth aspect of the present application provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the active learning image target detection method with adaptive annotation types described in the embodiment of the first aspect of the present application is implemented.
[0049] In summary, for the method, device, computer device, and non-transitory computer-readable storage medium of active learning image target detection with adaptive annotation types according to the embodiments of the present invention, an active detection iteration process is proposed, which includes five steps: model inference, target retrieval, information quantity evaluation, adaptive annotation, and semi-supervised training. This method can respectively estimate the classification information quantity and localization information quantity of each target in the image, select valuable targets to adaptively add class labels or bounding box annotations, and at the same time design a detection model that can jointly train full-supervised and weakly-supervised data. Additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0051] Figure 1 is a flowchart of an active learning image target detection method with adaptive annotation types provided by an embodiment of the present invention.
[0052] Figure 2 is a structural diagram of a device for an active learning image target detection method with adaptive annotation types provided by an embodiment of the present invention.
[0053] Figure 3 is a structural diagram of a target detection model for joint full-supervised and weakly-supervised training provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.
[0055] The method and device for actively learning image object detection with adaptive annotation types according to the embodiments of the present invention will be described below with reference to the accompanying drawings.
[0056] Figure 1 It is a schematic flowchart of a method for actively learning image object detection with adaptive annotation types provided by an embodiment of the present invention.
[0057] As Figure 1 shown, the method for actively learning image object detection with adaptive annotation types includes the following steps:
[0058] Step S1, detecting the object to be detected by a detection model to obtain the location information and classification information corresponding to the target object.
[0059] In an embodiment of the present application, before step S1, it is first necessary to design a semi-supervised detection model, including:
[0060] Extracting the multi-scale feature map of the image, taking the center point estimation as the first branch, and the weakly supervised global average pooling as the second branch, and the first branch and the second branch share the parameters of part of the multi-scale feature map;
[0061] In the first branch, convolving the multi-scale feature map to obtain the predicted position information;
[0062] In the second branch, convolving the multi-scale feature map to obtain a response map that can be supervised by the image-level label.
[0063] Specifically, in an embodiment of the present application, extracting the multi-scale feature map of the image, taking the center point estimation as the first branch, and the weakly supervised global mean average pooling as the second branch, and the first branch and the second branch share the parameters of part of the multi-scale feature map, including:
[0064] First, use a feature pyramid network to extract the multi-scale feature map F i , with the size of W i ×H i ×D i , where i∈{1,2,3} represents three feature map resolutions that increase sequentially, and W, H, and D respectively represent the width, height, and depth of the feature map. The center point estimation and the weakly supervised branch share part of the parameters. In this branch, the feature map is compressed by a 3×3 convolution to where C is the detection category. The response map is supervised at the pixel level and obtains a class prediction of length C after global average pooling that can be supervised by the image-level label. Another branch is responsible for predicting the position information, and the feature map is compressed by a 3×3 convolution to indicating the position compensation of the center point in two dimensions. Similarly, Indicates the size estimation of the bounding box length and width.
[0065] Step S2: For the target objects whose classification information meets the first preset condition, select the most valuable objects for annotation according to the quantization index quota to obtain the class labels corresponding to the target detection objects. For the target detection objects whose positioning information meets the second preset condition, select the most valuable detection objects for annotation according to the quantization index quota to obtain the supplementary bounding box labels corresponding to the detected target objects.
[0066] In an embodiment of the present application, detecting the target image to obtain the positioning information and classification information corresponding to the target object includes:
[0067] Predicting the target object through the initial semi-supervised detection model to obtain the positioning information and classification information corresponding to the target object.
[0068] Moreover, in an embodiment of the present application, the first preset condition includes target objects with classification information amount higher than the first specific threshold, and the second preset condition includes target objects with positioning information amount higher than the second specific threshold.
[0069] Specifically, the above model is trained with the existing labeled dataset. If it is the first training, randomly select a part of the images (such as 2,000 commonly used VOC07 images as the initial training set) for complete target annotation to form the initial training set to train the above detection model.
[0070] Model prediction Indicates the probability that the target represented by the key point at (x, y) on the feature map belongs to class c. Denote the center point of each manually annotated target as p ∈ R 2 , predict and calculate the loss on the low-resolution feature map with the downsampling ratio of R, then Map the ground truth to the heat map Y ∈ [0, 1] using the following Gaussian kernel W×H×C as follows:
[0071]
[0072] where, σ p is the standard deviation of the adaptive target size.
[0073] The loss function of the target center point estimation is pixel-level logistic regression, denoted as L k :
[0074]
[0075] where, α and β are hyperparameters, and N is the number of targets in the image.
[0076] To recover the discretization error, Local coordinate compensation is predicted for each center point, and the training of the model is supervised with the L1 loss. The loss function is denoted as L off :
[0077]
[0078] Denote the size of each manually annotated object as where k represents the number of object targets in an image. For the size of each target object perform regression, and use the L1 loss as the loss function to supervise the regression of the size. The loss function is denoted as L size :
[0079]
[0080] Use the one-hot vector g to represent the image-level label of the sample. When there is at least one target belonging to the c-th class in the sample, the corresponding position is marked as 1:
[0081]
[0082] where is the indicator function, c (k) is the class of the k-th target, and g c is the true label of the target belonging to the c-th class. This branch is supervised with the multi-label cross-entropy loss function, denoted as L cls :
[0083]
[0084] Based on the above, the overall training objective of the model is:
[0085] L = L k + λ off L off + λ size L size + λ cls L cls .
[0086] where λ off , λ size , λ cls are hyperparameters that control the weights of each branch during training.
[0087] Step S3, generate the first annotation data of the target object according to the class label and the supplementary bounding box label, and add the first annotation data to the annotation dataset, where the second annotation data is pre-stored in the annotation dataset.
[0088] In one embodiment of the present application, labeling a target object whose classification information satisfies a first preset condition to obtain a category label corresponding to the target object, and labeling a target object whose positioning information satisfies a second preset condition to obtain a supplementary bounding box label corresponding to the target object includes:
[0089] Use entropy to measure the amount of category information of the target:
[0090]
[0091] in, To use entropy to measure the amount of information about the target category, Indicates the center point coordinates The category prediction probability at , c is the total number of candidate categories.
[0092] To calculate the center point coordinates The amount of positioning information at the location, first scale the compensation matrix The local probability distribution expectation of is:
[0093]
[0094] Here, r defines the local neighborhood radius.
[0095] Secondly, the difference between the entropy of the local average predicted value and the mean of the predicted value entropy is used to measure the mutual information between the data distribution and the model prediction distribution. As an estimate of the amount of positioning information:
[0096]
[0097] in, Calculate the information entropy, which is defined here as:
[0098]
[0099] Similarly, get the center point coordinates The amount of size information at use Indicates the total amount of positioning information:
[0100]
[0101] Set thresholds ∈ for classification and localization respectively c ,∈ l ,When the amount of information of a category of the target exceeds the corresponding threshold, the corresponding type of annotation is adaptively provided.
[0102] Step S4, retraining the semi-supervised detection model according to the labeled data set to obtain an iteratively updated target semi-supervised detection model until the model reaches the expected performance or the number of annotations reaches the budget.
[0103] Technical effects of the present application: The characteristics that classification and localization in the detection task can be decoupled and multiple targets can be separated are fully utilized. The classification information amount and localization information amount of each target in the image are estimated respectively, and valuable targets are selected to adaptively add category labels or bounding box annotations. At the same time, a detection model that can jointly train full-supervised and weakly-supervised data is designed, which can not only significantly save the annotation cost, but also specifically improve the judgment of the detection algorithm on the target category and position.
[0104] To achieve the above embodiments, the present invention also proposes an active learning image target detection device with adaptive annotation types.
[0105] Figure 2 It is a schematic structural diagram of an active learning image target detection device with adaptive annotation types provided by an embodiment of the present invention.
[0106] As Figure 2 shown, the active learning image target detection device with adaptive annotation types includes:
[0107] A detection module, configured to detect a target detection object with a detection model to obtain the corresponding localization information and classification information of the target object;
[0108] An evaluation module, configured to, for a target object whose classification information meets the first preset condition, select the most valuable object for annotation according to the quantization index quota to obtain the category label corresponding to the target detection object, and for a target detection object whose localization information meets the second preset condition, select the most valuable detection object for annotation according to the quantization index quota to obtain the supplementary bounding box label corresponding to the detection target object;
[0109] A labeling module, configured to generate first annotation data of the target object according to the category label and the supplementary bounding box label, and add the first annotation data to the annotation data set, wherein the second annotation data is pre-stored in the annotation data set;
[0110] A training module, configured to retrain the semi-supervised detection model according to the annotation data set to obtain an iteratively updated target semi-supervised detection model until the model reaches the expected performance or the annotation quantity reaches the budget.
[0111] In an embodiment of the present application, further, it further includes:
[0112] Design the initial semi-supervised detection model in the following manner:
[0113] Extract the multi-scale feature map of the image, use the center point estimation as the first branch, and the weakly-supervised global average pooling as the second branch. The first branch and the second branch share the parameters of part of the multi-scale feature map;
[0114] In the first branch, the multi-scale feature map is convolved to obtain predicted location information;
[0115] In the second branch, the multi-scale feature map is convolved to obtain a response map that can be supervised by the image-level label.
[0116] In an embodiment of the present application, further, it further includes:
[0117] A detection module that predicts the target object through an initial semi-supervised detection model to obtain the localization information and classification information corresponding to the target object.
[0118] In an embodiment of the present application, further, it further includes:
[0119] The first preset condition includes target objects with classification information amount higher than the first specific threshold, and the second preset condition includes target objects with localization information amount higher than the second specific threshold.
[0120] In an embodiment of the present application, the overall detection model structure is as Figure 3 shown.
[0121] The technical effect of the present application: fully utilizes the characteristics of decoupling of classification and localization and separability of multiple targets in the detection task, respectively estimates the classification information amount and localization information amount of each target in the image, adaptively adds category labels or bounding box annotations to valuable targets, and at the same time designs a detection model that can jointly train full-supervised and weakly-supervised data, which can not only significantly save the annotation cost, but also specifically improve the judgment of the detection algorithm on the target category and location.
[0122] To achieve the above object, an embodiment of the third aspect of the present application proposes a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the method for actively learning image target detection with adaptive annotation type described in the first aspect embodiment of the present application.
[0123] To achieve the above object, an embodiment of the fourth aspect of the present application proposes a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, it implements the method for actively learning image target detection with adaptive annotation type described in the first aspect embodiment of the present application.
[0124] Although the present application is disclosed in detail with reference to the accompanying drawings, it should be understood that these descriptions are merely exemplary and are not used to limit the application of the present application. The protection scope of the present application is defined by the appended claims and may include various modifications, improvements, and equivalent solutions made to the invention without departing from the protection scope and spirit of the present application.
[0125] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0126] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0127] Any process or method description in a flowchart or described in other ways herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0128] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0129] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0130] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0131] In addition, in each embodiment of the present invention, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0132] The above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. An active learning image object detection method with adaptive annotation types, characterized in that, Including the following steps: Using an initial semi-supervised detection model to detect a target detection object, and obtaining the corresponding positioning information and classification information of the target detection object; For the target detection object whose classification information meets the first preset condition, select the most valuable object for annotation according to the quantization index quota to obtain the corresponding class label of the target detection object. For the target detection object whose positioning information meets the second preset condition, select the most valuable detection object for annotation according to the quantization index quota to obtain the corresponding supplementary bounding box label of the target detection object; Generating the first annotation data of the target detection object according to the class label and the supplementary bounding box label, and adding the first annotation data to the annotation dataset, wherein the second annotation data is pre-stored in the annotation dataset; Retraining the semi-supervised detection model according to the annotation dataset to obtain an iteratively updated target semi-supervised detection model until the model reaches the expected performance or the annotation quantity reaches the budget; Designing the initial semi-supervised detection model in the following manner: Extracting the multi-scale feature map of the image, taking the center point estimation as the first branch and the weakly supervised global average pooling as the second branch, and the first branch and the second branch share part of the parameters of the multi-scale feature map; In the first branch, performing convolution on the multi-scale feature map to obtain the predicted position information; In the second branch, performing convolution on the multi-scale feature map to obtain a response map that can be supervised by the image-level label; The first preset condition includes a target object with a classification information amount higher than a first specific threshold, and the second preset condition includes a target object with a positioning information amount higher than a second specific threshold; Annotating the target object whose classification information meets the first preset condition to obtain the corresponding class label of the target object, and annotating the target object whose positioning information meets the second preset condition to obtain the corresponding supplementary bounding box label of the target object, including: Measuring the classification information amount of the target with entropy: Among them, is the classification information quantity of the target measured by entropy, represents the central point coordinates is the class prediction probability distribution at, and c is the total number of candidate classes; To calculate the coordinate of the center point of the point For the positioning information amount at, first calculate the scale compensation matrix The expected value of the local probability distribution of: where r defines the local neighborhood radius; Secondly, the difference between the entropy of the local average prediction value and the mean of the entropy of the prediction value is used to measure the mutual information between the data distribution and the model prediction distribution, and is used as an estimate of the positioning information volume: Among them, calculate the information entropy, defined here as: Similarly, obtain the center point coordinates the dimensional information at Use to represent the total positioning information: Set thresholds ∈ for classification and localization respectively c , ∈ l , and respectively screen out the targets whose information amount exceeds the corresponding threshold, quantitatively select the target with the largest information amount, and provide the annotation of the corresponding type.
2. An active learning image object detection device with adaptive annotation types, characterized in that, Including: A detection module, which is used to detect a target detection object with an initial semi-supervised detection model to obtain the corresponding positioning information and classification information of the target detection object; An evaluation module, which is used to select the most valuable object for annotation according to the quantization index quota for the target detection object whose classification information meets the first preset condition to obtain the corresponding class label of the target detection object, and select the most valuable detection object for annotation according to the quantization index quota for the target detection object whose positioning information meets the second preset condition to obtain the corresponding supplementary bounding box label of the target detection object. The first preset condition includes a target object with a classification information amount higher than a first specific threshold, and the second preset condition includes a target object with a positioning information amount higher than a second specific threshold; A labeling module, which is used to generate the first annotation data of the target detection object according to the class label and the supplementary bounding box label, and add the first annotation data to the annotation dataset, wherein the second annotation data is pre-stored in the annotation dataset; A training module for retraining the semi-supervised detection model according to the labeled data set to obtain an iteratively updated target semi-supervised detection model until the model reaches the expected performance or the number of labels reaches the budget; Design the initial semi-supervised detection model in the following way: Extract the multi-scale feature maps of the image, use the center point estimation as the first branch, and the weakly supervised global average pooling as the second branch, where the first branch and the second branch share the parameters of part of the multi-scale feature maps; In the first branch, perform convolution on the multi-scale feature maps to obtain the predicted position information; In the second branch, perform convolution on the multi-scale feature maps to obtain a response map that can be supervised by the image-level label; The evaluation module is also used to measure the classification information of the target with entropy: Among them, is the classification information amount of the target measured by entropy, represents the center point coordinates is the class prediction probability distribution at, and c is the total number of candidate classes; To calculate the coordinate of the center point of the point For the positioning information amount at, first calculate the scale compensation matrix The expected value of the local probability distribution of where r defines the local neighborhood radius; Secondly, the difference between the entropy of the local average prediction value and the mean of the prediction value entropy is used to measure the mutual information between the data distribution and the model prediction distribution, and is used as an estimate of the positioning information amount: Among them, Calculate the information entropy, defined here as: Similarly, obtain the center point coordinates The dimensional information at Use To represent the total positioning information: Set thresholds ∈ for classification and positioning respectively c , ∈ l , and respectively screen out the targets whose information volume exceeds the corresponding threshold, quantitatively select the target with the largest information volume, and provide the annotation of the corresponding type.
3. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the annotation type adaptive active learning image target detection method according to claim 1.
Citation Information
Patent Citations
Target detection model building method
CN107038448A
An automatic image annotation method for weakly supervised semantic segmentation
CN109255790A