An object detection method and device in which anchor boxes participate in training
By optimizing the anchor frame size in the target detection system and updating the redundant anchor frame by calculating the interleaving ratio and gradient descent method, the detection accuracy problem caused by anchor frame redundancy is solved, and higher detection accuracy is achieved.
Patent Information
- Application Number
- CN202210724738.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-06-23
AI Technical Summary
In the existing target detection system, redundancy is caused by unreasonable anchor frame size setting, resulting in excessive negative samples, which reduces detection accuracy.
During the training process, by calculating the intersection ratio of the anchor box and the marked box to be detected, counting the anchor box utilization rate, resetting and updating the size of the redundant anchor box, and optimizing the anchor box size by using the gradient descent method to avoid the occurrence of redundant anchor box.
By optimizing the anchor frame size, reducing redundant anchor frames, reducing the number of negative samples, and improving the target detection accuracy.
Smart Images

Figure CN115205519B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection, and in particular to an object detection method and device in which anchor boxes participate in training. Background Art
[0002] The task of object detection is to find all the objects of interest in an image and determine their positions and sizes, which is one of the core problems in the field of machine vision. In the object detection task, the effect of the algorithm is often affected by various factors, and the multi-scale problem, that is, identifying objects of different regions and different sizes in the image, is a difficult problem that we usually encounter in object detection. In current popular object detection systems, the method of using the anchor box size as a preset size and making corrections is often used to improve the detection accuracy of the object bounding box.
[0003] The setting of the anchor box size is crucial for improving the detection accuracy. The mainstream method is to obtain it by clustering the actual bounding box sizes, or to manually set several sizes and set these anchor boxes at each anchor point in the image, so the anchor boxes set at each anchor point are the same. But in fact, this causes redundancy in the anchor box setting. For example, if an anchor box for detecting extremely large and wide objects is set at the leftmost edge of the image, then this anchor box will never be actually utilized because the center point of an extremely large and wide object will not appear at the leftmost side of the image. The existence of a large number of redundant anchor boxes will lead to an excessive number of negative samples, resulting in a decrease in detection accuracy. Summary of the Invention
[0004] Aiming at the problems proposed in the above background art, the present invention provides an object detection method in which anchor boxes participate in training to avoid a large number of redundant anchor boxes in each region of the image.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] An object detection method in which anchor boxes participate in training, including the following steps:
[0007] S1. At the start of iterative training of the object detection model, perform region determination on the input image to determine the positions where anchor points are generated in the image, and assign anchor boxes with the same size to each anchor point;
[0008] S2. In the process of each iterative training of the object detection model, calculate the intersection over union (IoU) between each anchor box and the annotation bounding box of the object to be detected, and the anchor boxes with an IoU greater than the set threshold are designated for detecting the object;
[0009] S3. After one round of iterative training is completed, the utilization rate of each anchor box is statistically calculated. Anchor boxes with a utilization rate lower than the set threshold are defined as redundant anchor boxes, and the sizes of the redundant anchor boxes are reset; for anchor boxes with a utilization rate not lower than the set threshold, calculate the gradient of the anchor box size with respect to the overall loss, and update the anchor box size based on the gradient.
[0010] S4. Use the updated anchor boxes to repeat steps S2 - S4 to iteratively train the object detection model until the overall loss converges, and obtain the trained object detection model.
[0011] S5. Use the trained object detection model to perform object detection on the image to be detected.
[0012] Furthermore, the overall loss includes confidence loss, class loss, and bounding box loss. In step S3, the gradient descent method is used to calculate the gradient of the anchor box size with respect to the overall loss.
[0013] Furthermore, the method for updating the anchor box size based on the gradient in step S3 is: manually set the update step size, and according to the partial derivative calculated by the gradient descent method, subtract the product of the update step size and the partial derivative from the size of the anchor box to obtain the updated anchor box size.
[0014] Furthermore, the utilization rate of the anchor box is the number of times the anchor box participates in training in one round of iterative training.
[0015] Furthermore, the method for resetting the size of the redundant anchor box is: statistically calculate the actual bounding box size of the object to be detected with the largest intersection over union ratio with this anchor box in each iterative training, put the statistically calculated actual bounding box sizes into a set for K - means clustering, and the clustering result is the size of this anchor box in the next round of training.
[0016] Furthermore, in the object detection process of step S5, the object detection model predicts an offset for each anchor box, including a size offset and a position offset. The coordinates and size of each anchor box plus the offset obtain a predicted bounding box; when obtaining the predicted bounding box, a class vector and a probability of the existence of an object are also predicted for each predicted bounding box. The class vector is used to predict which class of object exists in the predicted bounding box, and the probability of the existence of an object is used to predict the probability that a detected object exists in each predicted bounding box; select the anchor boxes with a probability of the existence of an object greater than a certain threshold, and use the NMS algorithm to merge the predicted bounding boxes belonging to the same object.
[0017] The present invention also provides an object detection device in which anchor boxes participate in training, including a processor and a memory. A computer program is stored on the memory, and when the computer program is executed by the processor, the object detection method in which anchor boxes participate in training is implemented.
[0018] The present invention further provides a computer storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the object detection method in which the anchor boxes participate in training is implemented.
[0019] Due to the above technical solution, the present invention has the following beneficial effects:
[0020] For the above object detection method and device in which the anchor boxes participate in training, at the start of training, the same anchor boxes are assigned to each anchor point. During the training process, the gradient of the anchor box size with respect to the loss is calculated for each iteration of training, and the anchor box size is updated, so as to achieve the setting of different anchor box groups at each anchor point position, making the anchor boxes at each anchor point better reflect the bounding box features of the objects in the corresponding area range, avoiding a large number of redundant anchor boxes in each area of the image, reducing the number of negative samples in the object detection process, and improving the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flowchart of the object detection method in which the anchor boxes participate in training according to a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0023] It should be noted that when a component is referred to as being "fixed to" another component, it can be directly on the other component or there may also be an intermediate component. When a component is considered to be "connected to" another component, it can be directly connected to the other component or there may be an intermediate component at the same time. When a component is considered to be "disposed on" another component, it can be directly disposed on the other component or there may be an intermediate component at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are only for the purpose of illustration.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0025] Please refer to Figure 1, a preferred embodiment of the present invention provides an object detection method involving anchor boxes in training, including the following steps:
[0026] S1. At the beginning of iterative training of the object detection model, perform region determination on the input image, determine the positions where anchor points are generated in the image, and assign anchor boxes with the same size to each anchor point.
[0027] In this embodiment, the object detection model used in step S1 is an object detection model based on a convolutional neural network in the prior art, which can be divided into two categories: single-step detectors and two-step detectors. Single-step detectors such as YOLO (You only look once, an object detection system) and SSD (single shot multibox detector) and other object detection models have higher speed but lower accuracy; two-step detectors such as Faster R-CNN (Faster Regions with Convolution Neural Network) and other object detection models have higher accuracy but slower speed.
[0028] An anchor box is a prior rectangular box in the two-dimensional space for detection. Its size includes the height and width of the anchor box, and the height and width of the anchor box are parallel to the y-axis and x-axis respectively. The generation of anchor boxes belongs to the prior art. To save space, it will not be elaborated here.
[0029] S2. During each iterative training of the object detection model, calculate the intersection over union (IoU) between each anchor box and the annotation box of the target to be detected. Anchor boxes with an IoU greater than a set threshold are designated for detecting the target.
[0030] In step S2, the annotation box, that is, the GT box, is the smallest bounding rectangle containing the target object. These boxes are usually pre-annotated manually and belong to the supervision information.
[0031] The intersection over union (IoU) is used to measure the overlap degree of two rectangular boxes in the two-dimensional space. Suppose A and B are two rectangular boxes, then their intersection over union (IoU) is defined as:
[0032] where, A∩B is the intersection of A and B, and |A∩B| is the area of this intersection region; A∪B is the union of A and B, and |A∪B| is the area of this union region.
[0033] The set threshold of the intersection over union in step S2 can be set according to actual needs and will not be elaborated in the present invention. If the intersection over union between the anchor box and the annotation box of the target to be detected is greater than the set threshold, mark this anchor box as a positive sample and be responsible for detecting the corresponding target; otherwise, mark this anchor box as a negative sample.
[0034] In S3, after one round of iterative training, the utilization rate of each anchor box is counted, and the anchor boxes with utilization rates lower than the set threshold are defined as redundant anchor boxes, and the sizes of the redundant anchor boxes are reset; for the anchor boxes with utilization rates not lower than the set threshold, the gradient of the anchor box size with respect to the overall loss is calculated, and the anchor box size is updated based on the gradient.
[0035] In step S3, the utilization rate of the anchor box is the number of times the anchor box participates in training in one round of iterative training. The set threshold of the utilization rate can be set according to actual needs and will not be elaborated in this invention.
[0036] The method for resetting the size of the redundant anchor box is as follows: count the actual bounding box sizes of the detected objects with the largest intersection over union (IoU) with this anchor box in each iterative training, put the counted actual bounding box sizes into a set for K-means clustering, and the clustering result is the size of this anchor box in the next round of training.
[0037] In step S3, the overall loss includes confidence loss, class loss, and bounding box loss, where:
[0038] Confidence loss loss obj The calculation is shown in the following formula (1):
[0039]
[0040] In formula (1), S×N represents dividing the image into S×N cells, with n anchor boxes set in each cell, and S×N×n is the total number of anchor boxes; γ ij represents whether the j-th anchor box in the i-th cell contains an object. If it does, its value is 1, otherwise 0; obj ij represents the predicted probability value that there is an object in the j-th anchor box in the i-th cell.
[0041] Bounding box loss loss bbox The calculation is shown in the following formula (2):
[0042]
[0043] In formula (2), S×N represents dividing the image into S×N cells, with n anchor boxes set in each cell, and S×N×n is the total number of anchor boxes; γ ijIndicates whether the j-th anchor box in the i-th cell contains a target. If it does, its value is 1; otherwise, it is 0. (x', y', w', h') represents a predicted bounding box predicted by the object detection model based on the anchor box, where (x', y') represents the center coordinates of the predicted bounding box, and (w', h') represents the width and height of the predicted bounding box. (x, y, w, h) represents an actual bounding box or label, where (x, y) represents the center coordinates of the actual bounding box, and (w, h) represents the width and height of the actual bounding box.
[0044] Class loss class The calculation is shown in the following formula (3):
[0045]
[0046] In formula (3), S×N represents dividing the image into S×N cells, with n anchor boxes set in each cell, and S×N×n is the total number of anchor boxes; γ ij Indicates whether the j-th anchor box in the i-th cell contains a target. If it does, its value is 1; otherwise, it is 0. K is the total number of categories of the target to be detected in the predicted bounding box predicted by the object detection model, and p1, p2, …, p k represent the probabilities of each category existing in the predicted bounding box predicted by the object detection model; c1, c2, …, c k is the category label vector within the predicted bounding box.
[0047] Overall loss all The calculation method is shown in the following formula (4):
[0048] loss all = loss class + loss obj + loss bbox (4)
[0049] It can be understood that the confidence loss, class loss, and bounding box loss can also use other algorithms, such as the cross-entropy algorithm for calculation.
[0050] In step S3, the gradient descent method is used to calculate the gradient of the anchor box size with respect to the overall loss, that is, to calculate the partial derivative of the overall loss with respect to the anchor box size. The method for updating the anchor box size based on the gradient in step S3 is: artificially set an update step size, and according to the partial derivative calculated by the gradient descent method, subtract the product of the update step size and the partial derivative from the size of the anchor box to obtain the updated anchor box size.
[0051] S4. Repeat steps S2 - S4 using the updated anchor boxes to iteratively train the object detection model until the overall loss converges, obtaining a trained object detection model. It can be understood that the iterative training of the object detection model may also include other steps, which belong to the prior art. For the sake of brevity, they are not elaborated here.
[0052] S5. Use the trained object detection model to perform object detection on the image to be detected.
[0053] During the object detection process in step S5, the object detection model predicts an offset for each anchor box, including a size offset and a position offset. The coordinates and size of each anchor box plus the offset result in a predicted bounding box. At the same time as obtaining the predicted bounding box, a class vector and a probability of the existence of an object are also predicted for each predicted bounding box. The class vector is used to predict which class of object exists in the predicted bounding box, and the probability of the existence of an object is used to predict the probability that a target to be detected exists in each predicted bounding box. The anchor boxes with a probability of the existence of an object greater than a certain threshold are selected, and the NMS algorithm is used to select the anchor box with the highest probability of the existence of an object, and the remaining anchor boxes are deleted as redundant anchor boxes to achieve the purpose of simplifying the output, and finally a predicted bounding box is predicted for each object. Step S5 belongs to the prior art. For the sake of brevity, it is not elaborated here.
[0054] In the above object detection method and device in which the anchor boxes participate in training, at the beginning of training, the same anchor boxes are assigned to each anchor point. During the training process, the gradient of the anchor box size with respect to the loss is calculated for each iterative training, and the anchor box size is updated, so as to set different groups of anchor boxes at each anchor point position, making the anchor boxes at each anchor point better reflect the characteristics of the bounding box of the target in the corresponding area range, avoiding a large number of redundant anchor boxes in each area of the image, reducing the number of negative samples in the object detection process, and improving the detection accuracy.
[0055] In the above object detection method and device in which the anchor boxes participate in training, during iterative training, the utilization rate of each anchor box participating in training is statistically calculated. The anchor boxes with a high utilization rate have their sizes updated using the gradient descent method during each iterative training process; for the anchor boxes with a low utilization rate, it indicates that the anchor box settings are unreasonable, and due to insufficient gradient information, their sizes cannot be effectively updated using the gradient descent method. In this embodiment, such anchor boxes are defined as redundant anchor boxes, and the method in step S3 is used to change the anchor box sizes of the redundant anchor boxes to improve their utilization rate in the next round of training, thereby reducing the number of redundant anchor boxes.
[0056] An embodiment of the present invention also provides an object detection device in which the anchor boxes participate in training, including a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, the object detection method in which the anchor boxes participate in training is implemented.
[0057] An embodiment of the present invention also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, the object detection method in which the anchor box participates in training is implemented.
[0058] The memory may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory. The memory may also be a memory array. The memory may also be partitioned, and the partitions may be combined into virtual volumes according to certain rules. The processor may be a central processing unit CPU, or a GPU, or may be an application specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0059] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, an apparatus, or a computer program product. Therefore, the present invention may be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may be implemented in the form of a computer program product of an embodiment on one or more computer-usable non-transitory storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0060] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses, and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the specified function in one flow Figure 1 or multiple flows and / or one block or multiple blocks in the block diagram.
[0061] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the specified function in one flow Figure 1 or multiple flows and / or one block or multiple blocks in the block diagram.
[0062] The above description is a detailed description of the preferred and feasible embodiments of the present invention. However, the embodiments are not intended to limit the scope of the patent application of the present invention. Any equivalent changes or modifications made under the technical spirit disclosed by the present invention shall fall within the scope of the patent covered by the present invention.
Claims
1. A target detection method in which anchor boxes participate in training, characterized in that, It includes the following steps: S1. At the start of the iterative training of the target detection model, perform region determination on the input image, determine the positions where anchor points are generated in the image, and assign anchor boxes with the same size to each anchor point; S2. During each iterative training of the target detection model, calculate the intersection over union (IoU) between each anchor box and the annotation box of the target to be detected. The anchor boxes with an IoU greater than the set threshold are designated for detecting the target; S3. After one round of iterative training is completed, count the utilization rate of each anchor box. Define the anchor boxes with a utilization rate lower than the set threshold as redundant anchor boxes, and reset the sizes of the redundant anchor boxes; for the anchor boxes with a utilization rate not lower than the set threshold, calculate the gradient of the anchor box size with respect to the overall loss, and update the anchor box size based on the gradient; S4. Repeat steps S2 - S4 using the updated anchor boxes to perform iterative training on the target detection model until the overall loss converges, and obtain the trained target detection model; S5. Use the trained target detection model to perform target detection on the image to be detected.
2. The object detection method with anchor boxes participating in training according to claim 1, wherein The overall loss includes confidence loss, class loss, and bounding box loss. In step S3, the gradient of the anchor box size with respect to the overall loss is calculated using the gradient descent method.
3. The object detection method with anchor boxes participating in training according to claim 1, wherein The method for updating the anchor box size based on the gradient in step S3 is: manually set an update step size, and according to the partial derivative calculated by the gradient descent method, subtract the product of the update step size and the partial derivative from the size of the anchor box to obtain the updated anchor box size.
4. The object detection method with anchor boxes participating in training according to claim 1, wherein, The utilization rate of the anchor box is the number of times the anchor box participates in training in one round of iterative training.
5. The object detection method with anchor boxes participating in training according to claim 1, wherein The method for resetting the size of the redundant anchor box is: count the actual bounding box sizes of the targets to be detected with the largest IoU with this anchor box in each iterative training, put the counted actual bounding box sizes into a set for K - means clustering, and the clustering result is the size of this anchor box in the next round of training.
6. The object detection method with anchor boxes participating in training according to claim 1, wherein During the target detection process in step S5, the target detection model predicts an offset for each anchor box, including a size offset and a position offset. The coordinates and size of each anchor box plus the offset obtain a predicted bounding box; at the same time as obtaining the predicted bounding box, a class vector and a probability of the existence of a target are also predicted for each predicted bounding box. The class vector is used to predict which class of target exists in the predicted bounding box, and the probability of the existence of a target is used to predict the probability that a target to be detected exists in each predicted bounding box; select the anchor boxes with a probability of the existence of a target greater than a certain threshold, and use the non - maximum suppression (NMS) algorithm to merge the predicted bounding boxes belonging to the same target.
7. An object detection device in which anchor boxes participate in training, characterized in that It includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, it implements the target detection method in which the anchor box participates in training as described in any one of claims 1 - 6.
8. A computer storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the target detection method in which the anchor box participates in training as described in any one of claims 1 - 6.
Citation Information
Patent Citations
Target detection method for rotating object
CN111524095A
Target object detection method and device, electronic equipment and storage medium
CN114581652A