Object Detection Method, Device, Electronic Device and Storage Medium
By combining the target's minimum external rectangular frame and mask information to generate a compact frame, the object detection model is trained, and the problem of low target detection accuracy and uncommon representation in the prior art is solved, and a higher precision and more general object detection method is achieved.
Patent Information
- Application Number
- CN202210178676.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-25
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-02-25
AI Technical Summary
The existing arbitrary target detection method cannot fully utilize the target details and cannot represent the target in a tight manner, resulting in low detection accuracy and uncommon representation form.
By combining the minimum external rectangular frame of the sample target in the sample image and the target mask to generate a sample compaction frame, the object detection model is trained based on the sample compaction frame, so that the model can generate a detection frame that tightens the target and accurately describes the detailed information of the target.
The precise description of detailed information such as target position and direction is achieved, and the accuracy of target detection is improved. The target representation method is more versatile than the existing technology, which broadens the application scenarios of target detection.
Smart Images

Figure CN114549825B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to an object detection method, apparatus, electronic device, and storage medium. Background Art
[0002] As one of the expanding branches in the field of object detection, arbitrary direction object detection has been widely applied in fields such as intelligent transportation, remote sensing image object detection, scene text detection, and fisheye image pedestrian detection. In some scenarios, the objects have problems such as dense arrangement, arbitrary directions, cluttered backgrounds, and large aspect ratios. At this time, using the traditional horizontal bounding box to represent the position of the object will have problems of including too much background information or foreground-background ambiguity.
[0003] Existing arbitrary direction object detection methods usually use the five-parameter representation method of a rotated rectangle box or the eight-parameter representation method of an arbitrary quadrilateral to represent the position of the object. Although these two object representation methods can alleviate the problems existing in the horizontal bounding box representation method to a certain extent, they still cannot accurately depict the detailed information of the object. Summary of the Invention
[0004] The present invention provides an object detection method, apparatus, electronic device, and storage medium, which are used to solve the defects that the prior art cannot make full use of the detailed information of the object and cannot represent the object in a compact manner, and realizes the accurate depiction of detailed information such as the position and direction of the object.
[0005] The present invention provides an object detection method, including:
[0006] Determine the image to be detected;
[0007] Based on an object detection model, perform object detection on the image to be detected to obtain a compact box in the image to be detected, where the compact box is circumscribed to the object in the image to be detected, and the compact box is within the minimum circumscribed rectangle box of the object;
[0008] The object detection model is trained based on a sample image and a sample compact box in the sample image, and the sample compact box is determined based on the minimum circumscribed rectangle box and object mask of the sample object in the sample image.
[0009] According to an object detection method provided by the present invention, the performing object detection on the image to be detected based on the object detection model to obtain a compact box in the image to be detected includes:
[0010] Based on the rectangle box detection network in the object detection model, perform object detection on the image to be detected to obtain a rectangle box in the image to be detected;
[0011] Based on the tight box detection network in the object detection model, applying the image features within the rectangular box, perform object detection within the rectangular box to obtain the tight box.
[0012] According to an object detection method provided by the present invention, the step of based on the tight box detection network in the object detection model, applying the image features within the rectangular box, performing object detection within the rectangular box to obtain the tight box includes:
[0013] Based on the tight box detection network in the object detection model, applying the image features within the rectangular box, perform object detection within the rectangular box to obtain the sliding offsets of the vertices of the rectangular box, and based on the sliding offsets of the vertices of the rectangular box, determine the vertices of the tight box, and based on the vertices of the tight box, determine the tight box.
[0014] According to an object detection method provided by the present invention, before the step of determining the vertices of the tight box based on the sliding offsets of the vertices of the rectangular box, further includes:
[0015] If the sliding offset of any vertex of the rectangular box is less than a preset threshold, update the sliding offset of the any vertex to zero.
[0016] According to an object detection method provided by the present invention, the sliding offsets of the vertices of the rectangular box include the offsets of the vertices on their corresponding multiple edges.
[0017] According to an object detection method provided by the present invention, the loss function of the object detection model is determined based on the difference between the predicted sliding offset and the true sliding offset, the predicted sliding offset is determined by the object detection model based on the sample image, and the true sliding offset is determined based on the minimum bounding rectangle of the sample object and the sample tight box.
[0018] According to an object detection method provided by the present invention, the sample tight box is determined based on the following steps:
[0019] Construct contour auxiliary lines in the sample image;
[0020] Obtain the intersection points of the contour auxiliary lines and the minimum bounding rectangle of the sample object when the contour auxiliary lines are tangent to the object mask;
[0021] Based on the intersection points of the contour auxiliary lines and the minimum bounding rectangle of the sample object, determine the sample tight box.
[0022] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the target detection method as described in any one of the above is implemented.
[0023] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the target detection method as described in any one of the above is implemented.
[0024] The target detection method, device, electronic device, and storage medium provided by the present invention generate a sample compact box by combining the minimum bounding rectangle of the sample target in the sample image and the target mask, and train a target detection model based on the sample compact box, so that the trained target detection model can generate a compact box of the target in the image based on the input image to be detected, realizing the accurate characterization of the detailed information of the target, improving the accuracy of target detection, and this target representation method is more general than the prior art, broadening the application scenarios of target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 is one of the flow diagrams of the target detection method provided by the present invention;
[0027] Figure 2 is an example diagram of the target compact box representation method provided by the present invention;
[0028] Figure 3 is another flow diagram of the target detection method provided by the present invention;
[0029] Figure 4 is the flow diagram of the method for determining the sample compact box provided by the present invention;
[0030] Figure 5 is an example diagram of the method for determining the sample compact box provided by the present invention;
[0031] Figure 6 is the structural diagram of the target detection model provided by the present invention;
[0032] Figure 7 is the training flow chart of the target detection model provided by the present invention;
[0033] Figure 8It is the test flow chart of the object detection model provided by the present invention;
[0034] Figure 9 It is the structural schematic diagram of the object detection device provided by the present invention;
[0035] Figure 10 It is the structural schematic diagram of the electronic device provided by the present invention. Detailed implementation manners
[0036] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0037] With the rapid rise of deep learning related technologies, the field of object detection has been greatly developed. Traditional 2D (Two-Dimensional) object detection algorithms aim to represent the position of an object with a horizontal bounding box and give the corresponding category of the object, mainly for axis-aligned images in natural scenes. In recent years, two-stage object detectors based on R-CNN (Regions with Convolutional Network) and single-stage object detectors based on YOLO (You Only Look Once) and SSD (Single Shot MultiBox Detector) have achieved excellent detection performance on axis-aligned images and have been successfully applied to related fields such as security, transportation, and life, and more and more fields have put forward new application requirements for this technology.
[0038] As one of the extended branches in the field of object detection, object detection in arbitrary directions has been widely applied in fields such as intelligent transportation, remote sensing image object detection, scene text detection, and fisheye image pedestrian detection. There are problems such as dense arrangement, arbitrary directions, cluttered backgrounds, and large aspect ratios of objects in some scenes. If horizontal bounding boxes are still used to represent the objects, there will be problems of including too much background information or foreground-background ambiguity. Especially in scenes where objects are densely arranged, in arbitrary directions and with cluttered backgrounds, using horizontal bounding boxes to represent objects will not only cause problems of foreground-background confusion, that is, information that may be foreground for one object may be background information for another object, and the overlapping of objects of the same category will further affect the training of discriminators and detectors. In addition, even if the miss detection rate of objects is very low and the position regression of horizontal bounding boxes is very accurate, the visual experience effect of using horizontal bounding boxes to represent objects is also very poor.
[0039] Existing arbitrary direction object detection methods are mainly divided into arbitrary direction object detection based on rotated rectangular boxes and arbitrary direction object detection based on arbitrary quadrilaterals. Arbitrary direction object detection represented by a rotated rectangular box adds an angle information θ on the basis of traditional 2D object detection to represent the direction of the object, and finally uses five parameters (x, y, w, h, θ) to represent the position information of the object; for arbitrary direction object detection represented by an arbitrary quadrilateral, the four corner coordinates of the detection box are directly regressed by the network, and the position information of the object is represented by an arbitrary quadrilateral (x 1 , y 1 , x 2 , y 2 , x 3 , y 3 , x 4 , y 4 ).
[0040] To a certain extent, both of these two representation methods alleviate the problems existing in the horizontal bounding box representation method, improve the discrimination ability of the classifier and the positioning accuracy of the detector, and have been widely applied in fields such as remote sensing image object detection and scene text detection. Nevertheless, the annotation information adopted by the object in the above two methods is still a rotated rectangular box or an arbitrary quadrilateral, without combining finer-grained mask information as a supervision signal, unable to accurately depict the detailed information of the object, and its representation form is a rotated rectangular box or an arbitrary quadrilateral, lacking generality.
[0041] In view of this, the present invention provides an object detection method. Figure 1 It is one of the schematic flowcharts of the object detection method provided by the present invention. As Figure 1 shown, the method includes:
[0042] Step 110, determining an image to be detected;
[0043] Step 120, performing object detection on the image to be detected based on an object detection model to obtain a tight box in the image to be detected. The tight box is circumscribed to the object in the image to be detected, and the tight box is within the minimum circumscribed rectangle of the object;
[0044] The object detection model is trained based on a sample image and sample tight boxes in the sample image. The sample tight boxes are determined based on the minimum circumscribed rectangle of the sample object in the sample image and the object mask.
[0045] Specifically, the image to be detected, i.e., the image for which object detection needs to be performed, can be, for example, an image captured by a camera, a video frame extracted from a video captured by a camera, etc. The embodiments of the present invention do not make specific limitations thereto. The image to be detected is input into an object detection model, and the object detection model can identify the objects in the image to be detected and mark them on the image to be detected with detection boxes.
[0046] Considering that in the prior art, the annotation information used for objects is an arbitrary quadrilateral or a rotated rectangle box, without combining finer-grained mask information as the supervision signal of the model, resulting in the detection boxes generated by the model being unable to accurately depict the detailed information of the objects. To address this problem, in the training process of the object detection model in the embodiments of the present invention, taking the sample image as a sample, and using the sample tight box determined by the minimum bounding rectangle of the sample object and the object mask in the sample image as the sample label, the object detection model is supervised by applying the object mask, so that the trained object detection model can determine a detection box that more tightly surrounds the object for the input image to be detected, i.e., a tight box. The tight box is circumscribed to the object in the image to be detected and is within the minimum bounding rectangle of the object.
[0047] Here, the tight box can be any polygon such as a pentagon, a hexagon, etc. For example, Figure 2 is an example diagram of the target tight box representation provided by the present invention. As Figure 2 shown, the tight box is an octagon composed of vertices left_top, top_left, top_right, right_top, right_bottom, bottom_right, bottom_left, and left_bottom. It can be seen that the tight box is circumscribed to the object in the image to be detected and is within the minimum bounding rectangle ABCD of the object, and surrounds the object more tightly than the minimum bounding rectangle. Correspondingly, the sample tight box is a polygon that is circumscribed to the object mask of the sample object and is within the minimum bounding rectangle of the sample object.
[0048] It should be noted that in response to challenges in object detection such as arbitrary object directions, dense arrangements, large aspect ratios, and cluttered backgrounds, the embodiments of the present invention use the object mask to generate a polygon sample tight box as the supervision signal of the object detection model, and finally enable the object detection model to generate a polygon tight box to represent the objects in the input image, so as to depict the detailed information such as the position, pose, and scale of the objects in a more compact manner. Moreover, the tight boxes generated by the object detection model are not limited to quadrilaterals, and are more general compared to the quadrilateral representation method in the prior art.
[0049] The method provided by the embodiments of the present invention generates a sample compact box by combining the minimum bounding rectangle of the sample target in the sample image and the target mask, and trains a target detection model based on the sample compact box, so that the trained target detection model can generate a compact box of the target in the input image to be detected, realizing the accurate characterization of the detailed information of the target, improving the accuracy of target detection, and this target representation method is more general than the prior art, broadening the application scenarios of target detection.
[0050] Based on the above embodiments, Figure 3 is the second flowchart of the target detection method provided by the present invention. As Figure 3 shown, step 120 includes:
[0051] Step 121, perform target detection on the image to be detected based on the rectangle detection network in the target detection model to obtain the rectangles in the image to be detected;
[0052] Step 122, perform target detection within the rectangle based on the compact box detection network in the target detection model by applying the image features within the rectangle to obtain the compact box.
[0053] Specifically, in order to further improve the accuracy of target detection, the target detection model in the embodiments of the present invention may include a rectangle detection network and a compact box detection network. When the image to be detected is input into the target detection model, the rectangle detection network can perform target detection on the image to be detected to obtain the candidate detection boxes of the target in the image to be detected, that is, the rectangles, and output them to the compact box detection network. Then, the compact box detection network can apply the image features within the rectangle to perform target detection within the rectangle, so as to obtain a detection box that more tightly surrounds the target, that is, the compact box.
[0054] Here, the determination method of the image features within the rectangle may be to extract features according to the regional image within the rectangle in the image to be detected, or to first extract the features of the image to be detected and then map the rectangle to the features to obtain the image features corresponding to the rectangle. The embodiments of the present invention do not make specific limitations on this. The compact box detection network can directly output the coordinates of each vertex of the compact box to directly obtain the compact box, or output the coordinate offsets of each vertex of the rectangle relative to the compact box, so as to migrate the rectangle to the compact box. The embodiments of the present invention also do not make specific limitations on this.
[0055] Based on any of the above embodiments, step 122 includes:
[0056] Based on the tight bounding box detection network in the object detection model, apply the image features within the rectangular bounding box to perform object detection within the rectangular bounding box, obtain the sliding offsets of the vertices of the rectangular bounding box, and based on the sliding offsets of the vertices of the rectangular bounding box, determine the vertices of the tight bounding box, and determine the tight bounding box based on the vertices of the tight bounding box.
[0057] Specifically, after obtaining the minimum bounding rectangle of the object in the image to be detected, the tight bounding box detection network can apply the image features within the rectangular bounding box to perform more fine-grained object detection within the rectangular bounding box, that is, further determine the tight bounding box of the object from within the rectangular bounding box. The specific process can be, first, determine the sliding offsets of the vertices of the rectangular bounding box, then, according to the sliding offsets of the vertices of the rectangular bounding box, determine the vertices of the tight bounding box, and finally, connect the vertices of the tight bounding box to obtain the tight bounding box.
[0058] Here, the vertices of the tight bounding box can be directly determined according to the sliding offsets of the vertices of the rectangular bounding box, or the sliding offsets of the vertices of the rectangular bounding box can be first updated and adjusted, and then the vertices of the tight bounding box are determined according to the updated values. The embodiments of the present invention do not make specific limitations on this. The sliding offset refers to the coordinate offset of each vertex of the rectangular bounding box relative to the tight bounding box in the sliding direction. In order to obtain a tighter tight bounding box surrounding the object, the sliding directions corresponding to each vertex are multiple directions sliding towards the object direction, which can be the directions of the multiple sides where each vertex is located, or other directions. The embodiments of the present invention do not make specific limitations on this either.
[0059] For example, as Figure 2 shown, if ABCD is a rectangular bounding box, for vertex A, the sliding offset of vertex A includes the coordinate offset in the AB direction and the coordinate offset in the AD direction. According to the sliding offset of vertex A, the vertices top_left and left_top of the tight bounding box can be directly obtained; for another example, if there are any two points b and c on the line segment left_top, top_left, the sliding offset of vertex A includes the coordinate offset in the Ab direction and the coordinate offset in the Ac direction. According to the sliding offset of vertex A, the coordinates of b and c can be obtained. According to the coordinates of these two points, a straight line can be determined, and the vertices top_left and left_top of the tight bounding box can be determined according to the intersection of this straight line and the rectangular bounding box.
[0060] Furthermore, the tight bounding box detection network includes a classification branch and a regression branch. Among them, the classification branch is used for predicting the category of the object in the image to be detected. In addition to obtaining the sliding offsets of the vertices of the rectangular bounding box relative to the tight bounding box, the regression branch can also obtain the coordinate offsets of the vertices of the rectangular bounding box itself to achieve position refinement of the rectangular bounding box, and finally obtain the minimum bounding rectangle of the object.
[0061] Based on any of the above embodiments, before determining the vertices of the tight frame based on the sliding offsets of the vertices of the rectangular frame, it further includes:
[0062] If the sliding offset of any vertex of the rectangular frame is less than the preset threshold, update the sliding offset of this vertex to zero.
[0063] Specifically, in order to avoid the situation that the model does not converge and there are redundant edges in the tight frame generated by the model, the embodiments of the present invention preset a threshold corresponding to the sliding offset, that is, the preset threshold. When the sliding offset of any vertex of the rectangular frame is less than the preset threshold, the sliding offset of this vertex is updated to zero, which means that this vertex can directly be used as one of the vertices of the tight frame. Here, the preset threshold can be set according to the empirical values during the test process or obtained by intelligent calculation, and the embodiments of the present invention do not make specific limitations on this.
[0064] For example, in the above example, if the distance between point A and top_left is extremely small, it means that the two sides at the corner of point A are already relatively close to the target, and the included redundant background area is relatively small, and there is no need to further shrink here. At this time, the vertex A of the rectangular frame can directly be used as one of the vertices of the tight frame, and the finally obtained tight frame is a heptagon composed of vertices A, top_right, right_top, right_bottom, bottom_right, bottom_left, and left_bottom.
[0065] Specially, if the sliding offset of each vertex of the rectangular frame is less than the preset threshold, it means that the original rectangular frame already fits the target well. For this situation, the tight frame can be the original quadrilateral rectangular frame.
[0066] It should be noted that by introducing the preset threshold of the sliding offset, the embodiments of the present invention can adaptively select the most suitable representation form of the target through the setting of the threshold for targets with approximate horizontal and arbitrary directions, improving the generality and detection accuracy of target detection.
[0067] Based on any of the above embodiments, the sliding offsets of the vertices of the rectangular frame include the offsets of the vertices of the rectangular frame on its corresponding multiple sides.
[0068] Specifically, in order not to introduce redundant direction information and reduce the calculation amount, the sliding offsets of the vertices of the rectangular frame in the embodiments of the present invention include the offsets of the vertices of the rectangular frame on its corresponding multiple sides. For example, as Figure 2As shown in the figure, if ABCD is a rectangular box, for the upper left corner point A of the rectangular box, the sliding offset of A includes the coordinate offset in the AB direction and the coordinate offset in the AD direction. For the upper right corner point B of the rectangular box, the sliding offset of B includes the coordinate offset in the BA direction and the coordinate offset in the BC direction.
[0069] It should be noted that in the embodiment of the present invention, the target detection model only increases the output dimension compared with the original target detection network. Since the sliding offsets of the vertices of the rectangular box and the coordinate offsets of the vertices of the rectangular box itself are output simultaneously, the time complexity is the same and no angle information is involved. Therefore, while the introduced additional computational amount can be ignored, the detailed information of the target can be described in a more compact manner. Moreover, in the embodiment of the present invention, the compact box is generated by migrating on the basis of the rectangular box, which can ensure that the regression order of the vertices of the compact box is consistent with the true order, and well solves the problems of angle regression sensitivity and sequential label points in the current arbitrary direction target detection technology.
[0070] Based on any of the above embodiments, the loss function of the target detection model is determined based on the difference between the predicted sliding offset and the true sliding offset. The predicted sliding offset is determined by the target detection model based on the sample image, and the true sliding offset is determined based on the minimum bounding rectangle of the sample target and the sample compact box.
[0071] Specifically, since the compact box is determined according to the sliding offsets of the vertices of the rectangular box, in order to further improve the prediction accuracy of the compact box, in the training process of the target detection model in the embodiment of the present invention, the loss function of the target detection model is determined according to the difference between the predicted sliding offset and the true sliding offset. Here, the predicted sliding offset is predicted by the target detection model according to the input sample image, and specifically, it can be the predicted sliding offsets of the vertices of the sample rectangular box obtained by the compact box detection network in the target detection model. The true sliding offset, that is, the actual value of the sliding offset, can be determined according to the coordinates of the vertices of the minimum bounding rectangle of the sample target and the coordinates of the vertices of the sample compact box.
[0072] Furthermore, the target detection model can adopt Faster RCNN (Faster Regions with Convolutional Neural Network). The loss function of the target detection model can further include the position regression loss of the rectangular box in the original Faster RCNN in addition to the loss between the predicted sliding offset and the true sliding offset.
[0073] Based on any of the above embodiments, Figure 4It is a schematic flowchart of the method for determining the sample compact frame provided by the present invention. As Figure 4 shown, the sample compact frame is determined based on the following steps:
[0074] Step 410: Construct contour auxiliary lines in the sample image;
[0075] Step 420: Obtain the intersection points of the contour auxiliary lines and the minimum bounding rectangle of the sample target when the contour auxiliary lines are tangent to the target mask;
[0076] Step 430: Determine the sample compact frame based on the intersection points of the contour auxiliary lines and the minimum bounding rectangle of the sample target.
[0077] Specifically, after determining the training samples of the target detection model, that is, the sample images, contour auxiliary lines can be first constructed in the sample images. Here, the contour auxiliary lines are the auxiliary lines constructed for determining the contour of the sample compact frame. Translate the contour auxiliary lines to obtain the intersection points of the contour auxiliary lines and the minimum bounding rectangle of the sample target in the sample image when the contour auxiliary lines are tangent to the target mask of the sample target in the sample image. Then, according to the intersection points of the contour auxiliary lines and the minimum bounding rectangle of the sample target, the sample label, that is, the sample compact frame, can be obtained.
[0078] For example, the contour auxiliary lines can be the straight lines drawn through the vertices of the minimum bounding rectangle of the sample target. The inclination angles of each straight line can be preset. For example, the inclination angles of the straight lines corresponding to the upper left corner point and the lower right corner point can be set to 45°, and the inclination angles of the straight lines corresponding to the upper right corner point and the lower left corner point can be set to -45°. Translate the contour auxiliary lines until they are tangent to the target mask of the sample target, and the intersection points of the contour auxiliary lines and the minimum bounding rectangle of the sample target at the tangent time can be obtained. Then, according to the intersection points of the contour auxiliary lines and the minimum bounding rectangle of the sample target, the sample compact frame can be obtained.
[0079] Another example, Figure 5 is an example diagram of the method for determining the sample compact frame provided by the present invention. As Figure 5As shown in the figure, considering the symmetry of the circumscribed rectangle, the contour auxiliary lines can be the straight lines L1 and L4 drawn through the upper left corner point A and the upper right corner point B of the minimum circumscribed rectangle of the sample target respectively. The inclination angles of L1 and L4 can be preset in advance. Then, traverse all the pixel points in the target mask, calculate the shortest distance from each of the above lines and the pixel points corresponding to the farthest distance, so as to obtain two pixel points a and b corresponding to the line L1, and two pixel points c and d corresponding to the line L4. Immediately afterwards, translate L1 until it passes through points a and b respectively, and the contour auxiliary lines L2 and L3 tangent to the target mask can be obtained. Among them, L2 intersects the upper side top and the left side Left of the minimum circumscribed rectangle at two points top_left and left_top respectively, and L3 intersects the right side right and the lower side bottom of the minimum circumscribed rectangle at two points right_bottom and bottom_right respectively. In the same way, translate L4 until it passes through points c and d respectively, and the contour auxiliary lines L5 and L6 tangent to the target mask can be obtained. Among them, L5 intersects the upper side top and the right side right of the minimum circumscribed rectangle at two points top_right and right_top respectively, and L6 intersects the left side left and the lower side bottom of the minimum circumscribed rectangle at two points left_bottom and bottom_left respectively. Finally, connect the above determined points in sequence to obtain the sample compact frame.
[0080] Based on any of the above embodiments, the existing object detection methods in any direction design anchor boxes based on the semantic features of images, and at the same time predict the position and size of the anchor boxes through the center point prediction branch and the shape prediction branch, and have good performance on the remote sensing image data set. However, there are still some problems. For example, adding an additional loss function may lead to the problem of non-convergence. At the same time, the introduction of the anchor box center point prediction branch, the anchor box shape prediction branch and various shapes of anchor boxes increases the additional computational complexity and other problems. And for some densely arranged targets with large aspect ratios, the effect of this method is not good.
[0081] In addition, existing arbitrary-direction object detection methods are mainly divided into arbitrary-direction object detection based on rotated rectangular boxes and arbitrary-direction object detection based on arbitrary quadrilaterals. To a certain extent, both of these representation methods alleviate the problems existing in the horizontal bounding box representation method, improve the discriminative ability of the classifier and the positioning accuracy of the detector, and are widely applied in fields such as remote sensing image object detection and scene text detection. However, the above two methods still have certain limitations. The five-parameter representation method based on rotated rectangular boxes has strict requirements for the prediction accuracy of angles. A slight angle deviation may lead to a significant decrease in the intersection over union (IoU) of the object, resulting in a decline in detection performance, especially for objects with a large aspect ratio. Moreover, a large amount of computation is added in both the box generation stage and the post-processing stage. The eight-parameter representation method based on arbitrary quadrilaterals has the problem of sequential labeled points, that is, how to define the regression order of the four corner points so that it is consistent with the order of the true values.
[0082] The existing technology predicts the center points of anchor boxes and the shapes of anchor boxes respectively based on the multi-scale features extracted from the original image. On the one hand, the introduction of the anchor box center point prediction branch, the anchor box shape prediction branch, and anchor boxes of various shapes increases the computational complexity, and excessive loss calculation may lead to non-convergence of network training. On the other hand, the annotation information used for the object is still a rotated rectangular box or an arbitrary quadrilateral, without combining finer-grained mask information as the supervision signal of the model. The detection boxes generated by the model cannot accurately depict the detailed information of the object, and its representation form is not universal.
[0083] In response to this, the present invention improves the problems existing in the previous arbitrary-direction object detection, such as computational complexity, sensitivity to the accuracy of angle prediction, and sequential labeled points, and provides a precise object detection method based on polygon compact box representation. The method includes the following steps:
[0084] Step S1, data preparation:
[0085] First, preprocess the data set. Optionally, the MS COCO2017 data set can be used as the sample image. According to the annotation information of the COCO data set, the minimum bounding rectangle of the sample object and the object mask in the sample image can be obtained. Subsequently, the minimum bounding rectangle and the object mask are converted through a certain formula to obtain the corresponding polygon annotation box, that is, the sample compact box, so as to realize obtaining the corresponding polygon annotation box by combining annotation information of different granularities.
[0086] Step S2, construction of the object detection model:
[0087] Figure 6 is the structural schematic diagram of the object detection model provided by the present invention, as Figure 6As shown, the model is modified based on the two-stage object detector Faster RCNN, including a rectangular box detection network and a tight box detection network. Among them, the rectangular box detection network can include a feature extraction module and an RPN (Region Proposal Network) module, and the tight box detection network can adopt a modified ROI Head (Region of Interest Head). Specifically, eight regression parameters are added to the tail of the ROI Head, which are respectively used to represent the sliding offsets of the vertices of the rectangular box generated in the first stage of the model on their corresponding sides.
[0088] Furthermore, the backbone network of the object detection model (i.e., Figure 6 the Backbone in Figure 6 ), that is, the feature extraction module, can adopt a pre-trained ResNet101 (Residual Networks) plus FPN (Feature Pyramid Network) structure, where the C1-C5 layers of the FPN are used. The feature map extracted by the backbone network (i.e.,
[0089] the Feature map in Figure 6Flattened in FC*2) in the middle, and input into two parallel fully-connected layers, which are respectively used for the classification prediction of target categories, and the regression prediction of the rectangular box offset and eight sliding offsets. The output dimension of the classification branch is N * the number of categories, and the output dimension of the regression branch is N * 12, where N is the number of targets in the image, the number of categories includes the number of target categories M and the background, and the output of the regression branch includes the coordinate offset of the rectangular box itself and the sliding offsets of each vertex on its corresponding two sides.
[0090] Step S3, Training of the object detection model:
[0091] Define the calculation method of the intersection over union and the calculation method of the loss function of the object detection model, Figure 7 is the training flowchart of the object detection model provided by the present invention. As Figure 7 shown, use the preprocessed sample images, the minimum bounding rectangles of the corresponding sample targets, and the sample compact boxes to train the parameters of the model, and obtain the trained object detection model based on the polygon compact box representation. The loss function of the object detection model is Loss = Loss rpn + loss m-Fasterrcnn where Loss rpn is the same as the setting in Faster RCNN, including the foreground score prediction loss of the anchor box and the offset prediction loss of the anchor box; the loss loss Figure 7 of the modified ROI Head (i.e., the detection head in m-Fasterrcnn is different from the original Faster RCNN detection head network. It not only includes the regression parameters of the rectangular box, but also introduces the regression parameters of eight sliding offsets. However, the calculation cost of these eight sliding offsets can be ignored.
[0092] Step S4, Testing of the object detection model:
[0093] Figure 8 is the test flowchart of the object detection model provided by the present invention. As Figure 8 shown, preprocess the image to be detected and input it into the trained object detection model for detection. During the detection process, for the predicted value of the sliding offset, when the predicted value of the sliding offset is less than a preset threshold, such as 0.05, directly set it to 0. Using this method, a good processing can also be obtained for the case where the target corners are approximately horizontal. Finally, the output is the compact box representation of the target in the image to be detected and its corresponding target category.
[0094] It should be noted that, in order to avoid excessive number of sides of the compact box, which may lead to similarity with contour point detection and introduce excessive computational complexity, in the embodiments of the present invention, the maximum number of sides of the compact box is limited to an octagon. Therefore, the compact boxes finally detected by the object detection model can be from any quadrilateral to any octagon, and the model can adaptively select any quadrilateral to any octagon for representation according to the shape of the object.
[0095] The method provided by the embodiments of the present invention generates a polygon supervision signal by using the mask information of the sample object, designs a polygon compact box representation method for the object, combines visual tasks of different granularities, and effectively solves the problems of the dependence on the angle regression accuracy and the problem of sequential label points in the previous arbitrary direction object detection based on regression. On the premise that the introduced additional computational complexity is negligible, the position, pose, scale and other detailed information of the object can be depicted in a more compact manner. Moreover, the compact box is not limited to a quadrilateral, which is more general and broadens the application scenarios of object detection and improves the performance of object detection.
[0096] Meanwhile, a sliding offset threshold is set. For approximately horizontal objects and objects in any direction, through the setting of the threshold, the network can adaptively select the most suitable representation form of the object, improving the generality and detection accuracy of object detection. And finally, experiments prove that the proposed method is applicable to both axis-aligned objects in natural scenes and objects in any direction in remote sensing images, traffic scenes, etc.
[0097] The object detection device provided by the present invention will be described below. The object detection device described below can be mutually referred to the object detection method described above.
[0098] Based on any of the above embodiments, the present invention provides an object detection device. Figure 9 is a schematic structural diagram of the object detection device provided by the present invention, as Figure 9 shown, the device includes:
[0099] A determination unit 910, configured to determine an image to be detected;
[0100] A detection unit 920, configured to perform object detection on the image to be detected based on an object detection model, and obtain a compact box in the image to be detected. The compact box is circumscribed to the object in the image to be detected, and the compact box is within the minimum circumscribed rectangle of the object.
[0101] The object detection model is trained based on a sample image and a sample compact box in the sample image. The sample compact box is determined based on the minimum circumscribed rectangle of the sample object in the sample image and the object mask.
[0102] The device provided by the embodiment of the present invention generates a sample compact box by combining the minimum bounding rectangle of the sample target in the sample image and the target mask, and trains a target detection model based on the sample compact box, so that the trained target detection model can generate a compact box of the target in the image based on the input image to be detected, realizing the accurate characterization of the detailed information of the target, improving the accuracy of target detection, and this target representation method is more general than the prior art, broadening the application scenarios of target detection.
[0103] Based on any of the above embodiments, the detection unit 920 includes:
[0104] A rectangle box detection subunit, configured to perform target detection on the image to be detected based on the rectangle box detection network in the target detection model, and obtain the rectangle box in the image to be detected;
[0105] A compact box detection subunit, configured to perform target detection within the rectangle box based on the compact box detection network in the target detection model and apply the image features within the rectangle box to obtain a compact box.
[0106] Based on any of the above embodiments, the compact box detection subunit is configured to:
[0107] Based on the compact box detection network in the target detection model, apply the image features within the rectangle box, perform target detection within the rectangle box, obtain the sliding offsets of the vertices of the rectangle box, and based on the sliding offsets of the vertices of the rectangle box, determine the vertices of the compact box, and determine the compact box based on the vertices of the compact box.
[0108] Based on any of the above embodiments, before determining the vertices of the compact box based on the sliding offsets of the vertices of the rectangle box, it further includes:
[0109] If the sliding offset of any vertex of the rectangle box is less than a preset threshold, update the sliding offset of this vertex to zero.
[0110] Based on any of the above embodiments, the sliding offsets of the vertices of the rectangle box include the offsets of the vertices of the rectangle box on their corresponding multiple sides.
[0111] Based on any of the above embodiments, the loss function of the target detection model is determined based on the difference between the predicted sliding offset and the true sliding offset. The predicted sliding offset is determined by the target detection model based on the sample image, and the true sliding offset is determined based on the minimum bounding rectangle of the sample target and the sample compact box.
[0112] Based on any of the above embodiments, the sample compact box is determined based on the following steps:
[0113] Construct contour auxiliary lines in the sample image;
[0114] Obtain the intersection points of the contour auxiliary line and the minimum bounding rectangle of the sample target when the contour auxiliary line is tangent to the target mask;
[0115] Determine the sample compact box based on the intersection points of the contour auxiliary line and the minimum bounding rectangle of the sample target.
[0116] Figure 10 An example of the physical structure diagram of an electronic device is shown as Figure 10 shown. The electronic device may include: a processor 1010, a communication interface 1020, a memory 1030, and a communication bus 1040. Among them, the processor 1010, the communication interface 1020, and the memory 1030 complete mutual communication through the communication bus 1040. The processor 1010 can call the logical instructions in the memory 1030 to execute the object detection method, and the method includes: determining the image to be detected; based on the object detection model, performing object detection on the image to be detected to obtain the compact box in the image to be detected, the compact box is circumscribed to the object in the image to be detected, and the compact box is within the minimum bounding rectangle of the object; the object detection model is trained based on the sample image and the sample compact box in the sample image, and the sample compact box is determined based on the minimum bounding rectangle of the sample target and the target mask in the sample image.
[0117] In addition, when the logical instructions in the above-mentioned memory 1030 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0118] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the object detection method provided by the above-mentioned various methods. The method includes: determining an image to be detected; based on an object detection model, performing object detection on the image to be detected to obtain a tight box in the image to be detected, where the tight box is circumscribed to the object in the image to be detected and the tight box is within the minimum circumscribed rectangle of the object; the object detection model is trained based on a sample image and a sample tight box in the sample image, and the sample tight box is determined based on the minimum circumscribed rectangle of the sample object in the sample image and an object mask.
[0119] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the object detection method provided by the above-mentioned various methods. The method includes: determining an image to be detected; based on an object detection model, performing object detection on the image to be detected to obtain a tight box in the image to be detected, where the tight box is circumscribed to the object in the image to be detected and the tight box is within the minimum circumscribed rectangle of the object; the object detection model is trained based on a sample image and a sample tight box in the sample image, and the sample tight box is determined based on the minimum circumscribed rectangle of the sample object in the sample image and an object mask.
[0120] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0121] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A target detection method, characterized in that, it includes: Determine the image to be detected; Based on the target detection model, perform target detection on the image to be detected to obtain the tight box in the image to be detected. The tight box is circumscribed to the target in the image to be detected, and the tight box is within the minimum circumscribed rectangle of the target; The target detection model is trained based on the sample image and the sample tight box in the sample image. The sample tight box is determined based on the minimum circumscribed rectangle of the sample target and the target mask in the sample image; the sample tight box is a polygon that is circumscribed to the target mask of the sample target and within the minimum circumscribed rectangle of the sample target.
2. The target detection method according to claim 1, characterized in that, The performing target detection on the image to be detected based on the target detection model to obtain the tight box in the image to be detected includes: Based on the rectangle detection network in the target detection model, perform target detection on the image to be detected to obtain the rectangle in the image to be detected; Based on the tight box detection network in the target detection model, apply the image features within the rectangle and perform target detection within the rectangle to obtain the tight box.
3. The target detection method according to claim 2, characterized in that, The performing target detection within the rectangle based on the tight box detection network in the target detection model, applying the image features within the rectangle, and obtaining the tight box includes: Based on the tight box detection network in the target detection model, apply the image features within the rectangle, perform target detection within the rectangle to obtain the sliding offsets of the vertices of the rectangle, and based on the sliding offsets of the vertices of the rectangle, determine the vertices of the tight box, and based on the vertices of the tight box, determine the tight box.
4. The target detection method according to claim 3, characterized in that, Before the determining the vertices of the tight box based on the sliding offsets of the vertices of the rectangle, it further includes: If the sliding offset of any vertex of the rectangle is less than the preset threshold, update the sliding offset of the any vertex to zero.
5. The target detection method according to claim 3, characterized in that, The sliding offsets of the vertices of the rectangle include the offsets of the vertices of the rectangle on their corresponding multiple sides.
6. The target detection method according to claim 3, characterized in that, The loss function of the target detection model is determined based on the difference between the predicted sliding offset and the true sliding offset. The predicted sliding offset is determined by the target detection model based on the sample image, and the true sliding offset is determined based on the minimum circumscribed rectangle of the sample target and the sample tight box.
7. The target detection method according to any one of claims 1 to 6, characterized in that, The sample tight box is determined based on the following steps: Construct contour auxiliary lines in the sample image; Obtain the intersection points of the contour auxiliary line and the minimum bounding rectangle of the sample target when the contour auxiliary line is tangent to the target mask; Determine the sample compact box based on the intersection points of the contour auxiliary line and the minimum bounding rectangle of the sample target.
8. A target detection device, Characterized in that, Comprising: A determination unit for determining an image to be detected; A detection unit for performing target detection on the image to be detected based on a target detection model to obtain a compact box in the image to be detected, the compact box being circumscribed to the target in the image to be detected, and the compact box being within the minimum bounding rectangle of the target; The target detection model is trained based on a sample image and a sample compact box in the sample image, and the sample compact box is determined based on the minimum bounding rectangle of the sample target and a target mask in the sample image; the sample compact box is a polygon that is circumscribed to the target mask of the sample target and is within the minimum bounding rectangle of the sample target.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, Characterized in that, When the processor executes the program, the target detection method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium, on which a computer program is stored, Characterized in that, When the computer program is executed by a processor, the target detection method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Building target detection method based on compact quadrilateral representation
CN112084869A