Object Detection and Model Training Methods, Devices, Computer Equipment, and Storage Media
By using the label assign strategy of adaptive thresholds during the training of the object detection model, the first threshold is automatically updated, and the missed detection and false detection problems caused by artificial setting of thresholds in the prior art are solved, and the detection capability and accuracy of the model are improved.
Patent Information
- Application Number
- CN202210293222.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-23
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-03-23
AI Technical Summary
During the training process of existing target detection models, due to human-set threshold problems, missed or misdetected, affecting the detection ability of the model.
The label assign strategy of adaptive threshold is adopted to automatically update the first threshold during model training. By obtaining the lowest overlap of the prediction boxes of the sample image, the threshold is dynamically adjusted to improve the detection capability of the model.
It effectively improves the detection capability of the target detection model, reduces missed and missed detection, and improves the accuracy of the model.
Smart Images

Figure CN114694218B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the technical field of image processing, and particularly to object detection and model training methods, devices, computer devices, and storage media. Background Art
[0002] Object detection is a research hotspot in the fields of computer vision and machine learning, which refers to identifying objects in an image and marking the positions of the objects. The idea of some object detection tasks is that a machine learning model generates multiple windows anchor (which can be called anchor boxes or candidate boxes) for pre-anchoring possible positions of the object according to sample images, and the sample images are pre-calibrated with ground truth boxes (gtbox or gt) that surround the object and represent the true positions of the object in the image. Then, after determining the overlap degrees of each anchor of each sample image with the gtbox, by comparing with a preset threshold, the anchors as positive samples and the anchors as negative samples can be determined, and then the model is trained. Among them, since the threshold for comparing the overlap degree of the anchor with the gtbox is set artificially when determining positive samples, if the setting is too high, some anchors may not be assigned as positive samples, resulting in missed detections by the model; if the setting is too low, low-quality anchors will be determined as positive samples, affecting the accuracy of the model. Based on this, it is necessary to improve the detection ability of the object detection model. Summary of the Invention
[0003] To overcome the problems existing in the related art, the embodiments of this specification provide object detection and model training methods, devices, computer devices, and storage media.
[0004] According to the first aspect of the embodiments of this specification, a training method for an object detection model is provided, including:
[0005] Obtain a sample set, where the sample images in the sample set are calibrated with ground truth boxes that surround the object;
[0006] After obtaining multiple candidate boxes of the sample image and determining the overlap degree of each candidate box of the sample image with the ground truth box of the sample image, loop through the following steps until the object detection model meets the preset iteration termination condition:
[0007] After determining a sample box as a sample by using the multiple candidate boxes of the sample image, the object detection model learns; among them, the sample box as a positive sample is determined by comparing the overlap degree of the candidate box with the ground truth box with a first threshold;
[0008] Obtain the predicted boxes with objects predicted by the object detection model from the multiple candidate boxes of the sample image after learning;
[0009] Based on the prediction bounding boxes of multiple said sample images, obtain a set of prediction bounding boxes with the lowest overlap degree, and update the first threshold according to the overlap degree of this set of prediction bounding boxes.
[0010] In some examples, the initial value of the first threshold is greater than a first preset value.
[0011] In some examples, the updating the first threshold according to the overlap degree of this set of prediction bounding boxes includes:
[0012] Determine the mean value of the overlap degree of this set of prediction bounding boxes, and update the first threshold according to the difference between the mean value and the first threshold.
[0013] In some examples, there are at least two prediction bounding boxes of the sample images; the set of prediction bounding boxes with the lowest overlap degree includes: the prediction bounding box with the lowest overlap degree among at least two prediction bounding boxes of the sample images.
[0014] In some examples, the preset iteration termination condition includes: the difference between the mean value of the set of prediction bounding boxes with the lowest overlap degree and the first threshold is less than a second preset value.
[0015] According to the second aspect of the embodiments of the present specification, there is provided an object detection method, and the method includes:
[0016] Obtain an image of the object to be recognized;
[0017] Input the image into an object detection model, and obtain the position information of the object in the image predicted by the object detection model; wherein, the object detection model is trained by using the method embodiments of the first aspect.
[0018] According to the third aspect of the embodiments of the present specification, there is provided a training device for an object detection model, including:
[0019] A sample acquisition module, configured to: acquire a sample set, and the sample images in the sample set are labeled with ground truth bounding boxes surrounding the object;
[0020] A determination module, configured to: acquire multiple candidate bounding boxes of the sample image, and after determining the overlap degree between each candidate bounding box of the sample image and the ground truth bounding box of this sample image, loop to execute the following steps until the object detection model meets the preset iteration termination condition:
[0021] An iterative module is used to: after determining a sample box as a sample by using multiple candidate boxes of the sample image, the target detection model is used for learning; among them, the sample box as a positive sample is determined by comparing the overlap degree between the candidate box and the ground truth box with a first threshold; obtaining a predicted box with a target predicted by the target detection model from multiple candidate boxes of the sample image after learning; according to the predicted boxes of multiple sample images, obtaining a group of predicted boxes with the lowest overlap degree, and updating the first threshold according to the overlap degree of this group of predicted boxes.
[0022] In some examples, the initial value of the first threshold is greater than a first preset value.
[0023] In some examples, the iterative module is further used to:
[0024] Determine the mean value of the overlap degree of this group of predicted boxes, and update the first threshold according to the difference between the mean value and the first threshold.
[0025] In some examples, there are at least two predicted boxes of the sample image; the group of predicted boxes with the lowest overlap degree includes: the predicted box with the lowest overlap degree among at least two predicted boxes of the sample image.
[0026] In some examples, the preset iteration termination condition includes: the difference between the mean value of the group of predicted boxes with the lowest overlap degree and the first threshold is less than a second preset value.
[0027] According to the fourth aspect of the embodiments of the present specification, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the method described in the first aspect or the second aspect are implemented.
[0028] According to the fifth aspect of the embodiments of the present specification, a computer storage medium is provided, which stores a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect or the second aspect are implemented.
[0029] According to the sixth aspect of the embodiments of the present specification, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps of the method described in the second aspect are implemented.
[0030] The technical solutions provided by the embodiments of the present specification may include the following beneficial effects:
[0031] In the embodiments of this specification, after each round of learning of the model, for the prediction boxes with targets predicted by the target detection model for multiple candidate boxes of the sample image, a group of prediction boxes with the lowest overlap degree is obtained according to the prediction boxes of each sample image, and the first threshold is updated according to the overlap degree of this group of prediction boxes; therefore, an adaptive threshold label assign strategy is adopted in the training process of the target detection model, that is, the first threshold used to determine positive samples is not fixed, but is automatically determined during the model training process; therefore, if the overlap degree of this group of prediction boxes is relatively high, the first threshold can be appropriately lowered to add more positive samples; if the overlap degree of this group of prediction boxes is relatively low, the first threshold can be appropriately raised to reduce low-quality positive samples. Therefore, this embodiment can adaptively select an appropriate first threshold during the model training process, effectively improving the ability of the target detection model.
[0032] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. Brief Description of the Drawings
[0033] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this specification, and are used together with the specification to explain the principles of this specification.
[0034] Figure 1A It is a flowchart of a method for training a target detection model shown according to an exemplary embodiment of this specification.
[0035] Figure 1B It is a schematic diagram of a sample image shown according to an exemplary embodiment of this specification.
[0036] Figure 2 It is a flowchart of a target detection method shown according to an exemplary embodiment of this specification.
[0037] Figure 3 It is a hardware structure diagram of a computer device where a training device of a target detection model is located shown according to an exemplary embodiment of this specification.
[0038] Figure 4 It is a block diagram of a training device of a target detection model shown according to an exemplary embodiment of this specification. Detailed Embodiments
[0039] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0040] The terms used in this specification are for the purpose of describing particular embodiments only and are not intended to limit this specification. The singular forms "a", "the", and "said" used in this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0041] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "upon" or "in response to determining".
[0042] Currently, in the field of image processing, there is a need for object detection in many scenarios; for example, segmenting the foreground region as the object from an image, such as in the fields of security, intelligent transportation, or autonomous driving, where pedestrians or vehicles are the objects, and locating the positions of objects such as pedestrians or vehicles from the image; or, in image face detection, where the face is the object and locating the face position from the image, and so on.
[0043] In some object detection solutions, during the training process of an object detection model, an anchor-based label assignment strategy is involved; among them, an anchor is a position window where an anchored target may exist, segmented from a sample image based on the set conditions of the anchor. In this embodiment, it is referred to as a candidate box. Many anchors can be segmented from each sample image, and these anchors may or may not include a target. In the label assignment strategy, the overlap (Intersection over Union, IoU) between the anchors of the sample image and the gtbox is calculated. If the overlap is greater than a manually set threshold, a label representing a target is assigned to the anchor, otherwise, a label representing a non-target is assigned. This process is the aforementioned label assignment. As mentioned above, the threshold for comparing the overlap between the anchor and the gtbox is manually set. If the setting is too high, some anchors may not be assigned as positive samples, resulting in missed detections by the model; if the setting is too low, low-quality anchors will be determined as positive samples, affecting the accuracy of the model.
[0044] Based on this, as Figure 1A shown, this is a flowchart of a training method for an object detection model provided by an embodiment of this specification, including the following steps:
[0045] In step 102, a sample set is obtained, and the sample images in the sample set are calibrated with ground truth boxes surrounding the target;
[0046] In step 104, after obtaining multiple candidate boxes of the sample image and determining the overlap between each candidate box of the sample image and the ground truth box of the sample image, the following steps are repeatedly executed until the object detection model meets the preset iteration termination condition:
[0047] In step 106, after determining the candidate boxes as samples using the multiple candidate boxes of the sample image, the object detection model is used for learning; among them, the candidate boxes as positive samples are determined by comparing the overlap between the candidate box and the ground truth box with a first threshold;
[0048] In step 108, after the object detection model learns, prediction boxes with targets are predicted from the multiple candidate boxes of the sample image;
[0049] In step 110, according to the prediction boxes of each sample image, a group of prediction boxes with the lowest overlap is obtained, and the first threshold is updated according to the overlap of this group of prediction boxes.
[0050] In this embodiment, the task of the object detection model is to detect the position of the object in the input image. The training process of the model can be as follows: first, represent a model through modeling, then evaluate the model by constructing an evaluation function (also called an objective function), and finally optimize the evaluation function according to the training data and optimization method to adjust the model to meet the set conditions. The object detection model in this embodiment can be selected according to needs in actual applications, such as a model based on the Region Proposal Network (RPN) or a model based on the Regions with CNN features (RCNN) network, etc. This embodiment does not limit this.
[0051] In some examples, a training set can be prepared in advance for training the object detection model; in this embodiment, the sample image is calibrated with a ground truth box that encloses the object, and this ground truth box encloses the object and represents the true position of the object in the image.
[0052] In this embodiment, the training process of the object detection model is carried out in an anchor-based manner, that is, the positive and negative labels in the sample image are automatically determined based on the anchor. In some examples, the anchoring method of the candidate box and the method of determining positive and negative samples can be set in advance.
[0053] Regarding the anchoring method of the candidate box, as an example, the size (scale) and aspect ratio of the anchor can be defined; for example, assuming that n different scales and m different aspect ratios (the aspect ratios of different wide, medium, and narrow anchors) are defined, through combination, k (k = n * m) types of anchors can be obtained. For each pixel point on each feature map of each sample image, k anchors can be generated. Therefore, each sample image includes multiple anchors. Optionally, the above process of determining the candidate box anchor can be to input the training set into the object detection model, and the object detection model determines it based on the set anchoring method of the candidate box; in other examples, it can also be to use other neural networks independent of the model to determine multiple anchors for each sample image. This embodiment does not limit this.
[0054] In this embodiment, the Intersection over Union (IoU) can be used to evaluate the matching degree between the anchor and the gtbox. The IoU defines the overlap degree of two rectangular boxes as the ratio of the overlapping area of the two rectangular boxes to the area of the union of the two rectangular boxes.
[0055] Regarding the method of determining positive and negative samples, as an example, positive samples can be defined as anchors that meet one of the following conditions: among the multiple anchors of each sample image, the anchor with the highest overlap with the gtbox, and / or, the anchor with an overlap with the gtbox greater than or equal to the first threshold. For negative samples, they can be defined as anchors with an overlap lower than the second threshold. Except for the above positive and negative samples, other anchors (i.e., those that do not meet the conditions for positive or negative samples) can not participate in the training and do not contribute to the optimization of the loss function. In this embodiment, the anchors used for training are called sample boxes. Among them, the first threshold and the second threshold are actually set by technicians according to needs in practical applications. The determined candidate boxes as positive samples and candidate boxes as negative samples can be used to train the model. During the training process, the model will learn the features of each candidate box as a positive sample and the features of the candidate box as a negative sample; since the anchor as a positive sample does not exactly match the gtbox, during the training process, border regression (anchor regression) will also be performed on the anchor as a positive sample, that is, the process of approximating the anchor as a positive sample with the labeled true box as the target. The model learns from this how to correctly frame the target from the image and learns the ability to find the location where the target is located.
[0056] For the object detection model trained with the artificially set first threshold, after putting the model online, the inventor analyzed the prediction results of the object detection model in actual applications to seek solutions to improve the model performance. Through analysis, it was found that there was a situation of missed detection of targets after the object detection model was put online, that is, the model failed to predict the presence of targets in the image. Through the analysis of these images missed by the model, it was found that since the model failed to find the anchor that could represent the location of the target from the various anchors of the missed detection images, the missed detection phenomenon occurred; and the reason why the model failed to find the anchor representing the location of the target was that the location and size of the target represented in the image did not match the anchor that could be used as a positive sample, resulting in the face not being assigned an anchor and thus missed detection.
[0057] One way to solve this problem is to set a smaller first threshold during training, that is, as long as the anchor has a certain overlap with the true box gtbox, it can be label assigned as a positive sample, but this will cause many low-quality anchors to be label assigned as positive samples, resulting in a large number of false detections of the model.
[0058] The training of the object detection model in this embodiment adopts a label assign strategy with an adaptive threshold, which can automatically determine an appropriate first threshold during the model training process, effectively improving the ability of the object detection model. Among them, in steps 208 and 210 of this embodiment, it is possible to obtain the predicted bounding boxes with objects predicted by the object detection model from multiple candidate bounding boxes of each sample image after learning; according to the predicted bounding boxes of each sample image, obtain a group of predicted bounding boxes with the lowest overlap degree, and update the first threshold according to the overlap degree of this group of predicted bounding boxes.
[0059] In this embodiment, multiple rounds of iteration are required during the training process of the object detection model; one round of iteration in this embodiment refers to the process in which, after determining the positive sample anchors and negative sample anchors of the sample images in the dataset with the first threshold of this round of iteration, the object detection model uses all the samples to go through the learning process; after the object detection model completes one round of iteration using some or all of the images in the sample set, the object detection model has gone through one round of learning. In this embodiment, it is possible to obtain the predicted bounding boxes with objects predicted by the object detection model from multiple candidate bounding boxes of the sample image after learning, that is, use the prediction results of the object detection model after learning to update the first threshold again. In this embodiment, for the convenience of distinction, the trained object detection model will predict all the anchors of the sample image to find the anchors with objects. In this embodiment, the anchors with objects predicted by the model for the sample image are called predicted bounding boxes.
[0060] In some examples, the sample images used for prediction by the object detection model after learning are exactly the same as the sample images used when determining the positive and negative samples; that is, for all the sample images used when determining the positive and negative samples, the object detection model also makes predictions for all the sample images during prediction; in other examples, the two may not be exactly the same. For example, it is optional that there are certain differences in quantity, etc. This embodiment does not limit this.
[0061] Among them, for the predicted bounding boxes with objects predicted by the object detection model, the overlap degrees of these predicted bounding boxes with the ground truth boxes are high and low. If the overlap degree of the predicted bounding box with the ground truth box is very high, it means that the prediction accuracy of the object detection model for this image is very high; and this embodiment uses a group of predicted bounding boxes with the lowest overlap degree among the predicted bounding boxes of multiple sample images. If the overlap degree of this group of predicted bounding boxes is also relatively high, the first threshold can be appropriately lowered to add more positive samples; if the overlap degree of this group of predicted bounding boxes is relatively low, the first threshold can be appropriately raised to reduce low-quality positive samples.
[0062] In some examples, the initial value of the first threshold can be relatively small, meaning that the number of positive samples is relatively large in the initial stage; subsequently, during the training process, the first threshold is gradually increased to gradually improve the quality of the positive samples.
[0063] In other examples, the initial value of the first threshold in this embodiment is greater than the first preset value, that is, a relatively high first threshold is adopted in the initial stage of model training. Since the ability of the object detection model is relatively poor during the initial training, by setting a relatively high initial value for the first threshold, the prediction ability of the model can be quickly improved using fewer and higher-quality positive samples in the initial stage; and through the above embodiment, in subsequent iterations, a group of prediction boxes with the lowest overlap degree can be determined from the prediction boxes of multiple sample images, and the first threshold can be slowly reduced using the overlap degree of this group of prediction boxes, so that the model can gradually obtain more positive samples, thereby reducing the missed detection of the model and also improving the performance of the model. That is to say, during the model training process, the first threshold in the first round of iteration is the largest and is greater than the first thresholds used in subsequent other rounds of iteration. Since the initial first threshold is greater than the first thresholds in subsequent iterations, the number of initial positive samples is less than the first thresholds in subsequent iterations; that is, during multiple rounds of iteration, the first threshold gradually decreases and the positive samples gradually increase. Among them, the fact that the initial value of the first threshold is greater than the first preset value means that the first threshold is relatively large initially, and it can be flexibly determined according to needs in practical applications, and this embodiment does not limit this.
[0064] In some examples, updating the first threshold according to the overlap degree of this group of prediction boxes includes:
[0065] Determine the mean value of the overlap degree of this group of prediction boxes, and update the first threshold according to the difference between the mean value and the first threshold.
[0066] In some examples, after each round of training is completed, the object detection model can predict each sample image, score all anchors of each sample image, and the object detection model can output: the anchor(s) predicted as the target for each sample image. For the anchor(s) predicted as the target for each sample image, the overlap degree between each anchor and the gtbox can be determined. When selecting a group of prediction boxes with the lowest overlap degree, it can be flexibly determined according to needs in practical applications. For example, an overlap degree ratio s can be set, and the prediction boxes with an overlap degree lower than this overlap degree ratio s among all anchors are used as the group of prediction boxes with the lowest overlap degree.
[0067] In some other examples, there are at least two predicted bounding boxes for the sample image; the set of predicted bounding boxes with the lowest overlap includes: the predicted bounding box with the lowest overlap among at least two predicted bounding boxes of the sample image. That is, one or more anchors with the lowest overlap among the anchors predicted as the target for each sample image can be determined. For example, in the first round of training, there are multiple anchors for the first sample image, and among them, 5 anchors are predicted as the target. One or more anchors with the lowest overlap with the gtbox are determined from these 5 anchors, and the overlap of the one or more anchors with the lowest overlap is stored; the same applies to other sample images. Optionally, if the overlaps of all the anchors predicted as the target for the sample image are the same, any one can be selected; if there is only one anchor predicted as the target for the sample image, the overlap of this anchor is stored.
[0068] Therefore, the anchor for which the overlap is stored is the set of predicted bounding boxes with the lowest overlap. Finally, the mean value of the stored overlaps is compared with the first threshold, and the first threshold is updated according to the difference between the two; for example, if the mean value is greater than the first threshold, the first threshold can be increased; if the mean value is less than the first threshold, the first threshold can be decreased.
[0069] In some examples, the preset iteration termination condition may include: the difference between the mean value of the set of predicted bounding boxes with the lowest overlap and the first threshold is less than a second preset value, and this second preset value can be flexibly set as needed, and this embodiment does not limit this; the fact that the difference between the mean value and the first threshold is less than the second preset value indicates that after multiple rounds of iteration, the update granularity of the first threshold has become smaller, indicating that the label assign is in a good state; in practical applications, the preset iteration termination condition can also adopt other conditions as needed, or a combination of one or more other conditions. For example, it can be that the number of iterations is greater than or equal to a set number threshold; or, conditions set based on the evaluation function of the object detection model, etc.
[0070] As can be seen from the above embodiments, after each round of learning of the model, for the prediction boxes with targets predicted by the object detection model for multiple candidate boxes of the sample image, a group of prediction boxes with the lowest overlap degree is obtained according to the prediction boxes of each sample image, and the first threshold is updated according to the overlap degree of this group of prediction boxes; therefore, an adaptive threshold label assign strategy is adopted in the training process of the object detection model, that is, the first threshold for determining positive samples is not fixed, but is automatically determined during the model training process; therefore, if the overlap degree of this group of prediction boxes is relatively high, the first threshold can be appropriately lowered to add more positive samples; if the overlap degree of this group of prediction boxes is relatively low, the first threshold can be appropriately raised to reduce low-quality positive samples. Therefore, this embodiment can adaptively select an appropriate first threshold during the model training process, effectively improving the ability of the object detection model.
[0071] Next, it will be further described through another embodiment. Taking the intelligent identity verification scenario as an example, in an intelligent identity verification product, after the face detector detects the position of the face in the image, subsequent live body algorithms, etc. can be executed; in the online scenario, the inventor found that in addition to some cases where the face in the image is missed by the face detector, the face detector fails to detect the face in the image.
[0072] Based on this, this embodiment designs a dynamic threshold label assign strategy. At the beginning of the face detector training, a relatively high threshold is set in advance. As the training iteration progresses, there will be some sample images where the overlap degree between the anchor and the gtbox is slightly less than the first threshold, but the face detector still scores it as a face. Therefore, this embodiment will readjust to lower the first threshold, readjust the positive samples through label assign and retrain. Such a dynamic adjustment process iterates until the iteration termination condition is met. Therefore, this embodiment can effectively improve the ability of the face detector.
[0073] In this embodiment, the initial value of the first threshold σ1 is relatively high. In practical applications, this initial value can be set as needed, and this embodiment does not limit this.
[0074] This embodiment also sets the condition θ for dynamically updating the threshold. This condition can represent the iteration termination condition of label assign, that is, when this condition is met, it means that the first threshold no longer needs to be updated and the iteration can be terminated. As an example, in this embodiment, the initial relatively large first threshold is taken as an example. The first threshold will gradually decrease as the number of iterations increases. This condition can represent the lower limit of the first threshold, that is, when it decreases to this condition, the iteration is terminated. It also means that the difference between the first threshold updated in the last iteration and the first threshold in the previous iteration is very small, that is, the current face detector has iterated to a better quality and the first threshold no longer needs to be updated. In practical applications, it can be flexibly set according to needs, and this embodiment does not limit it.
[0075] In this embodiment, a rule for obtaining anchors is also set. In practical applications, it can be flexibly set according to needs, and this embodiment does not limit it.
[0076] After the settings are completed, the training of the face detector can be started. During the training process, according to the strategy of label assign, the face detector obtains multiple anchors for the sample images, determines the positive sample anchors according to the first threshold, and determines the negative sample anchors according to the second threshold. The face detector performs one round of training based on the positive and negative samples.
[0077] After one round of training on the sample set, the trained face detector makes predictions on each sample image, finds all the anchors of each sample image for face scoring, and the face detector can output: the anchors (one or more) predicted to be faces for each sample image.
[0078] For the anchors predicted to be faces for each sample image, determine the overlap degree between each anchor and the gtbox; determine the anchor with the lowest overlap degree among the anchors predicted to be faces for each sample image.
[0079] Therefore, for n sample images, the set of anchors with the lowest overlap degree obtained in the first round of training is: σ 11 , σ 12 , …, σ 1n ; where represents the overlap degree of the anchor with the lowest overlap degree in the j-th sample image in the i-th round of training. For example, in the first round of training, the first sample image has multiple anchors, and among the anchors predicted to be faces, there are 5 anchors. Determine the anchor with the lowest overlap degree with the gtbox from these 5 anchors, and the overlap degree of this anchor is σ 11 . The same applies to other sample images.
[0080] Calculate the average of the minimum overlap degrees for the current round of training If |σ2 - σ1| > θ, update the first threshold to σ2; after obtaining the new first threshold, repeat the above training process until the first threshold is no longer updated and the model training is completed.
[0081] As Figure 1B shown, it is a schematic diagram of a sample image shown in this specification according to an exemplary embodiment. On the left side of the schematic diagram, there are three rectangular boxes in the sample image, and on the right side, the image content of the three rectangular boxes is shown. From top to bottom, they are gtbox, anchor1, and anchor2; among the two anchors, the overlap degree between anchor1 and gtbox is larger, and the overlap degree between anchor2 and gtbox is lower. In the first round of iteration, due to the relatively large first threshold set, only anchor1 is assigned as a positive sample; as the number of iterations deepens, the first threshold decreases, and anchor2 will also be adaptively assigned as a positive sample to participate in training, thereby increasing the effective face labels, ensuring both a high detection rate and no false detections.
[0082] By testing the trained face detector, taking a public test set as an example for testing, this embodiment can effectively increase the recall rate by 2%; when applied to intelligent identity verification products, it operates well with almost no missed detections. The label assign with the adaptive threshold in this embodiment effectively improves the effect of the face detection algorithm.
[0083] As Figure 2 shown, it is a target detection method shown in this specification according to an exemplary embodiment. The method includes:
[0084] In step 202, obtain an image of the target to be recognized;
[0085] In step 204, input the image into the target detection model to obtain the position information of the target in the image predicted by the target detection model; wherein, the target detection model is trained using the training method of the target detection model in the foregoing embodiment.
[0086] Among them, the embodiment of the training of the target detection model can refer to the foregoing embodiment; when inputting the image of the target to be recognized into the target detection model, the model can identify the position information of the target in the image and output it, where the position information represents the position of the target in the image.
[0087] Since the target detection model adaptively selects a suitable first threshold during the training process, effectively improving the ability of the target detection model, this model can accurately identify the target in the image, and the possibility of missed detections or false detections is relatively low.
[0088] Corresponding to the embodiments of the training method of the foregoing object detection model, this specification also provides embodiments of a training apparatus for an object detection model and a computer device to which it is applied.
[0089] The embodiments of the training apparatus for the object detection model in this specification can be applied to a computer device, such as a server or a terminal device. The apparatus embodiments can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful apparatus, it is formed by the corresponding computer program instructions in the non-volatile memory being read into the memory and run by the processor where it is located. From a hardware level, as Figure 3 shown, it is a hardware structure diagram of the computer device where the training apparatus for the object detection model in this specification is located. In addition to Figure 3 the processor 310, memory 330, network interface 320, and non-volatile memory 340 shown, the computer device where the training apparatus 331 for the object detection model in the embodiment is located usually includes other hardware according to the actual functions of the computer device, which will not be elaborated here.
[0090] As Figure 4 shown, Figure 4 is a block diagram of a training apparatus for an object detection model shown according to an exemplary embodiment of this specification. The apparatus includes:
[0091] A sample acquisition module 41, configured to: acquire a sample set, where the sample images in the sample set are labeled with ground truth boxes surrounding the objects;
[0092] A determination module 42, configured to: after acquiring a plurality of candidate boxes of the sample image and determining the overlap degree between each candidate box of the sample image and the ground truth box of the sample image, repeatedly execute the following steps until the object detection model meets a preset iteration termination condition:
[0093] An iteration module 43, configured to: after determining a sample box as a sample by using the plurality of candidate boxes of the sample image, perform learning by the object detection model; among them, the sample box as a positive sample is determined by comparing the overlap degree between the candidate box and the ground truth box with a first threshold; acquire the predicted boxes with objects predicted by the object detection model from the plurality of candidate boxes of the sample image after learning; according to the predicted boxes of multiple sample images, acquire a group of predicted boxes with the lowest overlap degree, and update the first threshold according to the overlap degree of this group of predicted boxes.
[0094] In some examples, the initial value of the first threshold is greater than a first preset value.
[0095] In some examples, the updating the first threshold according to the overlap degree of this group of predicted boxes includes:
[0096] Determine the average overlap degree of this set of predicted bounding boxes, and update the first threshold according to the difference between the average value and the first threshold.
[0097] In some examples, there are at least two predicted bounding boxes in the sample image; the set of predicted bounding boxes with the lowest overlap degree includes: the predicted bounding boxes with the lowest overlap degree among at least two predicted bounding boxes of the sample image.
[0098] In some examples, the preset iteration termination condition includes: the difference between the average value of the set of predicted bounding boxes with the lowest overlap degree and the first threshold is less than a second preset value.
[0099] For the implementation processes of the functions and roles of each module in the above training device of the object detection model, please refer to the implementation processes of the corresponding steps in the above training method of the object detection model for details, which will not be elaborated here.
[0100] Correspondingly, an embodiment of this specification further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the embodiment of the training method of the foregoing object detection model, or the steps of the embodiment of the foregoing object detection method.
[0101] Correspondingly, an embodiment of this specification further provides a computer storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the embodiment of the training method of the foregoing object detection model, or the steps of the embodiment of the foregoing object detection method.
[0102] Correspondingly, an embodiment of this specification further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the embodiment of the foregoing object detection method.
[0103] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this specification. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0104] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0105] The step divisions of the above various methods are only for clear description. In implementation, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of this patent. Making insignificant modifications to the algorithm or process or introducing insignificant designs, but without changing the core design of the algorithm and process, are all within the protection scope of this application.
[0106] Among them, the description of "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of this specification. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiments or examples. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0107] Those skilled in the art will readily conceive of other embodiments of this specification after considering the specification and practicing the invention claimed herein. This specification is intended to cover any variations, uses, or adaptations of this specification, which follow the general principles of this specification and include the known common knowledge or conventional technical means in the technical field not claimed in this application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of this specification are pointed out by the following claims.
[0108] It should be understood that this specification is not limited to the exact structures already described and shown in the figures, and various modifications and changes can be made without departing from its scope. The scope of this specification is only limited by the appended claims.
[0109] The above are only the preferred embodiments of this specification and are not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included within the protection scope of this specification.
Claims
1. A training method for a target detection model, comprising: Obtaining a sample set, where the sample images in the sample set are labeled with ground truth boxes surrounding the target; After obtaining multiple candidate boxes of the sample image and determining the overlap degree between each candidate box of the sample image and the ground truth box of the sample image, the following steps are repeatedly executed until the target detection model meets a preset iteration termination condition: After determining a sample box as a sample by using the multiple candidate boxes of the sample image, the target detection model performs learning; among them, the sample box as a positive sample is determined by comparing the overlap degree between the candidate box and the ground truth box with a first threshold; Obtaining the predicted boxes with targets predicted by the target detection model from the multiple candidate boxes of the sample image after learning; According to the predicted boxes of multiple sample images, obtaining a group of predicted boxes with the lowest overlap degree, and updating the first threshold according to the overlap degree of this group of predicted boxes.
2. The method according to claim 1, wherein the initial value of the first threshold is greater than a first preset value.
3. The method according to claim 1, wherein the updating the first threshold according to the overlap degree of this group of predicted boxes comprises: Determining the mean value of the overlap degree of this group of predicted boxes, and updating the first threshold according to the difference between the mean value and the first threshold.
4. According to the method described in claim 1, there are at least two prediction boxes for the sample image; the group of prediction boxes with the lowest overlap degree includes: The predicted box with the lowest overlap degree among at least two predicted boxes of the sample image.
5. The method according to claim 1, wherein the preset iteration termination condition includes: The difference between the mean value of the group of predicted boxes with the lowest overlap degree and the first threshold is less than a second preset value.
6. A target detection method, the method comprising: Obtaining an image of the target to be recognized; Inputting the image into a target detection model to obtain the position information of the target in the image predicted by the target detection model; wherein, the target detection model is trained by using the method according to any one of claims 1 to 5.
7. A training device for a target detection model, comprising: A sample acquisition module, configured to: obtain a sample set, where the sample images in the sample set are labeled with ground truth boxes surrounding the target; A determination module, configured to: after obtaining multiple candidate boxes of the sample image and determining the overlap degree between each candidate box of the sample image and the ground truth box of the sample image, repeatedly execute the following steps until the target detection model meets a preset iteration termination condition: An iteration module, configured to: after determining a sample box as a sample by using the multiple candidate boxes of the sample image, the target detection model performs learning; among them, the sample box as a positive sample is determined by comparing the overlap degree between the candidate box and the ground truth box with a first threshold; obtaining the predicted boxes with targets predicted by the target detection model from the multiple candidate boxes of the sample image after learning; according to the predicted boxes of multiple sample images, obtaining a group of predicted boxes with the lowest overlap degree, and updating the first threshold according to the overlap degree of this group of predicted boxes.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, The processor implements the steps of the method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, on which a computer program is stored, and the computer program implements the steps of the method according to any one of claims 1 to 6 when being executed by a processor.
10. A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method according to claim 6.
Citation Information
Patent Citations
Model training method, target detection method and device and storage medium
CN111444828A
Target detection model training method and device, computer equipment and storage medium
CN112749726A