Target detection method and device, computer readable storage medium, and electronic device
By combining the object detection method of instance segmentation network and edge extraction network, the problem of poor universality of traditional computer vision algorithms relying on prior information and parameters is solved, and the object detection effect with high accuracy and universality is achieved.
Patent Information
- Application Number
- CN202110007102.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-05
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-01-05
AI Technical Summary
In the prior art, object detection based on traditional computer vision algorithms requires relying on prior information of objects, and the parameters of traditional vision algorithms are poorly universal and are greatly affected by the environment, resulting in low accuracy of object detection.
A target detection method combining an instance segmentation network and an edge extraction network is adopted, and RGBD data and depth data are generated by acquiring RGB data and point cloud data, input instance segmentation network to obtain the enclosure box and the first pixel data, input edge extraction network to obtain edge pixel data, and correct the first pixel data based on the edge pixel data to determine the coordinate information of the target object.
It improves the accuracy of target detection, avoids the problem of excessive time-consuming and excessive collection of target object information, and enhances the universality of target detection methods.
Smart Images

Figure CN113793349B_ABST
Abstract
Description
Background Art
[0002] With the development of artificial intelligence technology, robots have replaced human labor in many aspects. For example, in warehouse automation applications, robots can complete box picking. Robot box picking refers to the technology that robots, based on visual guidance and according to the picking tasks issued by the system, take out the corresponding number of goods from the turnover box according to the task requirements and put them into the designated delivery box.
[0003] In the prior art, target detection is generally achieved by using traditional computer vision algorithms. Most target detection based on traditional computer vision algorithms needs to rely on prior information of objects and build models based on prior information. However, objects are frequently updated, and it takes a long time to collect prior information of objects. In addition, the parameters of traditional vision algorithms are less universal and are greatly affected by the environment.
[0004] In view of this, there is an urgent need in this field to develop a new target detection method and device.
[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0006] The purpose of the present disclosure is to provide a target detection method, a target detection device, a computer-readable storage medium and an electronic device, thereby improving the accuracy of target detection at least to a certain extent.
[0007] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.
[0008] According to one aspect of the present disclosure, a target detection method is provided, the method comprising: acquiring RGB data and point cloud data containing a plurality of target objects, generating RGBD data and depth data according to the RGB data and the point cloud data; inputting the RGBD data into an instance segmentation network to obtain a plurality of bounding boxes and first pixel data corresponding to each of the bounding boxes, and inputting the depth data into an edge extraction network to obtain edge pixel data corresponding to the depth data; performing data correction on the first pixel data according to the edge pixel data to obtain target pixel data corresponding to each of the bounding boxes, and determining coordinate information of each of the target objects according to the target pixel data.
[0009] In some exemplary embodiments of the present disclosure, the edge pixel data and the first pixel data are binary data, and the binary data includes a first numerical value and a second numerical value; the first pixel data is corrected according to the edge pixel data to obtain target pixel data corresponding to each of the bounding boxes, including: traversing the first pixel data in each of the bounding boxes respectively, obtaining the target coordinate point where the first pixel data is the first numerical value, and obtaining the target edge data corresponding to the target coordinate point in the edge pixel data; judging whether the target edge data is the first numerical value, and determining the target pixel data according to the judgment result.
[0010] In some exemplary embodiments of the present disclosure, the target pixel data is determined based on the judgment result, including: if the target edge data is the first value, the first pixel data corresponding to the target coordinate point is corrected to a second value, and the corrected first pixel data is used as the target pixel data.
[0011] In some exemplary embodiments of the present disclosure, the coordinate information of each of the target objects is determined based on the target pixel data, including: performing an open operation on the target pixel data to obtain second pixel data corresponding to each of the bounding boxes, and determining the coordinate information of each of the target objects based on the second pixel data.
[0012] In some exemplary embodiments of the present disclosure, the coordinate information of each of the target objects is determined based on the second pixel data, including: performing connected domain analysis on the second pixel data corresponding to each of the bounding boxes to obtain multiple connected domains corresponding to each of the bounding boxes; obtaining the coordinate information of the maximum connected domain corresponding to each of the bounding boxes, and determining the coordinate information of each of the target objects based on the coordinate information of each of the maximum connected domains.
[0013] In some exemplary embodiments of the present disclosure, RGB data and point cloud data containing multiple target objects are obtained, including: obtaining the RGB data containing the multiple target objects based on shooting with a two-dimensional camera, and obtaining the point cloud data containing the multiple target objects based on shooting with a three-dimensional camera, wherein a resolution of the two-dimensional camera is greater than a resolution of the three-dimensional camera.
[0014] In some exemplary embodiments of the present disclosure, the point cloud data includes first coordinate data, second coordinate data, and third coordinate data; generating RGBD data and depth data based on the RGB data and the point cloud data includes: configuring the third coordinate data as the depth data; performing coordinate transformation on the RGB data according to a transformation matrix to obtain three-channel data having the same number as the third coordinate data, and generating the RGBD data based on the third coordinate data and the three-channel data.
[0015] According to one aspect of the present disclosure, a target detection device is provided, comprising: a data acquisition module, used to acquire RGB data and point cloud data containing multiple target objects, and generate RGBD data and depth data according to the RGB data and the point cloud data; a data analysis module, used to input the RGBD data into an instance segmentation network to obtain multiple bounding boxes and first pixel data corresponding to each of the bounding boxes, and input the depth data into an edge extraction network to obtain edge pixel data corresponding to the depth data; a coordinate determination module, used to perform data correction on the first pixel data according to the edge pixel data to obtain target pixel data corresponding to each of the bounding boxes, and determine the coordinate information of each of the target objects according to the target pixel data.
[0016] According to one aspect of the present disclosure, a computer-readable medium is provided, on which a computer program is stored. When the program is executed by a processor, the target detection method as described in the above embodiments is implemented.
[0017] According to one aspect of the present disclosure, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the target detection method as described in the above embodiments.
[0018] It can be seen from the above technical solutions that the target detection method and device, computer-readable storage medium, and electronic device in the exemplary embodiments of the present disclosure have at least the following advantages and positive effects:
[0019] The target detection method disclosed in the present invention first obtains RGB data and point cloud data containing multiple target objects, and generates RGBD data and depth data based on the RGB data and point cloud data; then the RGBD data is input into the instance segmentation network to obtain multiple bounding boxes and the first pixel data corresponding to each bounding box, and the depth data is input into the edge extraction network to obtain edge pixel data corresponding to the depth data; finally, the first pixel data is corrected according to the edge pixel data to obtain the target pixel data corresponding to each bounding box, and the coordinate information of each target object is determined according to the target pixel data. On the one hand, the target detection method disclosed in the present invention can use the edge pixel data obtained by the edge extraction network to correct the first pixel data obtained by the instance segmentation network, combining the advantages of the instance segmentation network and the edge extraction network, and improving the accuracy of target detection; on the other hand, the target detection method does not need to obtain the target object information in advance, avoiding the problem of too long time caused by collecting the target object information, and improving the universality of the target detection method.
[0020] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0022] Figure 1 A schematic diagram of a process flow of a target detection method according to an embodiment of the present disclosure is schematically shown;
[0023] Figure 2 The schematic diagram of the process of generating depth data and RGBD data according to an embodiment of the present disclosure is schematically shown;
[0024] Figure 3 A schematic diagram of a process for determining coordinate information of each target object according to an embodiment of the present disclosure is schematically shown;
[0025] Figure 4 A schematic diagram of a process of correcting first pixel data according to an embodiment of the present disclosure is schematically shown;
[0026] Figure 5 A schematic diagram schematically shows a structure of a specific application scenario according to an embodiment of the present disclosure;
[0027] Figure 6 A schematic diagram of a process flow of a target detection method according to a specific embodiment of the present disclosure is schematically shown;
[0028] Figure 7 A schematic diagram schematically shows a structure of an image composed of first pixel data according to an embodiment of the present disclosure;
[0029] Figure 8 A schematic diagram showing the structure of an image composed of edge pixel data according to an embodiment of the present disclosure is shown;
[0030] Fig. 9 A schematic diagram schematically shows the structure of an image composed of second pixel data according to an embodiment of the present disclosure;
[0031] Fig.10 A schematic diagram showing the structure of an image formed by a minimum bounding box according to an embodiment of the present disclosure is shown;
[0032] Fig.11A block diagram of a target detection device according to an embodiment of the present disclosure is schematically shown;
[0033] Fig.12 The module diagram of the electronic device according to an embodiment of the present disclosure is schematically shown;
[0034] Fig.13 A schematic diagram of a program product according to an embodiment of the present disclosure is shown schematically. DETAILED DESCRIPTION
[0035] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more comprehensive and complete and will fully convey the concept of the example embodiments to those skilled in the art.
[0036] In addition, the described features, structures or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the present disclosure.
[0037] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0038] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0039] In the related technologies in this field, on the one hand, target detection is achieved based on traditional computer vision algorithms, such as scale-invariant feature transform (SIFT), shape-based feature matching (shapebase), point pair feature matching (PPF), template matching, etc. On the other hand, target detection is achieved based on deep learning, such as bounding box prediction plus mask instance segmentation, or deep learning edge extraction. However, most target detection based on traditional computer vision algorithms needs to rely on prior information of the target and model the prior information in advance, such as extracting SIFT, shapebase, PPF features, saving image templates, recording target size, etc. It takes a long time to collect prior information of the target, and due to the frequent update of the target, the collection work needs to be carried out frequently, and the parameters of traditional computer vision algorithms have poor universality and do not take advantage of the seamless migration of projects in different environments. In addition, the method based on deep learning adopts the method of bounding box prediction plus mask instance segmentation, which has low segmentation accuracy and leads to low target detection accuracy. Moreover, based on the deep learning edge extraction method, if the edge is extracted on the color image, the target texture and shadow will bring a large number of false edges without further segmentation. If the edge is extracted on the depth map, its accuracy is seriously dependent on the point cloud quality.
[0040] Based on the problems existing in the related art, a target detection method is proposed in one embodiment of the present disclosure. Figure 1 A schematic diagram of the process of target detection method is shown, such as Figure 1 As shown, the target detection method at least includes the following steps:
[0041] Step S110: acquiring RGB data and point cloud data containing multiple target objects, and generating RGBD data and depth data according to the RGB data and point cloud data;
[0042] Step S120: inputting the RGBD data into an instance segmentation network to obtain a plurality of bounding boxes and first pixel data corresponding to each bounding box, and inputting the depth data into an edge extraction network to obtain edge pixel data corresponding to the depth data;
[0043] Step S130: performing data correction on the first pixel data according to the edge pixel data to obtain target pixel data corresponding to each bounding box, and determining coordinate information of each target object according to the target pixel data.
[0044] The information recommendation method in the disclosed embodiment, on the one hand, can use the edge pixel data obtained by the edge extraction network to correct the first pixel data obtained by the instance segmentation network, combining the advantages of the instance segmentation network and the edge extraction network, and improving the accuracy of target detection; on the other hand, the target detection method does not need to obtain the target object information in advance, avoiding the problem of too long time caused by collecting the target object information, and improving the universality of the target detection method.
[0045] In order to make the technical solution of the present disclosure clearer, each step of the target detection method is described below.
[0046] In step S110, RGB data and point cloud data containing multiple target objects are acquired, and RGBD data and depth data are generated according to the RGB data and the point cloud data.
[0047] In an exemplary embodiment of the present disclosure, RGB data and point cloud data are obtained by photographing the same scene of the multiple target objects. The RGB data includes pixel values of an RGB image containing multiple target objects, and the RGB data specifically includes pixel values of each pixel point on the RGB image in three channels of R (red), G (green), and B (blue). The point cloud data includes XYZ coordinate data containing multiple target objects, specifically including the X-axis coordinate value, Y-axis coordinate value, and Z-axis coordinate value of the three channels of X-axis, Y-axis, and Z-axis of each pixel point in the point cloud image. Among them, the coordinate origin of the point cloud data is the position of the three-dimensional camera, the plane formed by the X-axis and the Y-axis is a plane parallel to the camera imaging plane, and the Z-axis coordinate value represents the distance between the target object and the three-dimensional camera.
[0048] In an exemplary embodiment of the present disclosure, RGBD data is two-dimensional image data including four-channel information of R (red), G (green), B (blue), and D (depth value), and depth data includes single-channel two-dimensional image data of D (depth value).
[0049] In an exemplary embodiment of the present disclosure, an RGB image containing multiple target objects can be obtained by photographing with a two-dimensional camera, and point cloud data containing multiple target objects can be obtained by photographing with a three-dimensional camera, wherein the resolution of the two-dimensional camera is greater than the resolution of the three-dimensional camera.
[0050] In addition, the RGB image can also be obtained by photographing a three-dimensional camera, and the present disclosure does not specifically limit the photographing parameters of the two-dimensional camera and the three-dimensional camera.
[0051] In an exemplary embodiment of the present disclosure, the point cloud data includes first coordinate data, second coordinate data and third coordinate data, the first coordinate data can represent the X-axis coordinate value, the second coordinate data can represent the Y-axis coordinate value; or the first coordinate data can represent the Y-axis coordinate value, the second coordinate data can represent the X-axis coordinate value; the third coordinate data represents the Z-axis coordinate value.
[0052] In an exemplary embodiment of the present disclosure, RGBD data and depth data under a three-dimensional camera perspective can be generated based on RGB data and point cloud data, and RGBD data and depth data under a two-dimensional camera perspective can also be generated, and the disclosure does not make specific limitations on this.
[0053] The present disclosure embodiment takes the generation of RGBD data and depth data under the perspective of a three-dimensional camera as an example for explanation. Figure 2 A schematic diagram of the process of generating depth data and RGBD data is shown, such as Figure 2 As shown, the process may at least include steps S210 to S220, which are described in detail as follows:
[0054] In step S210 , the third coordinate data is configured as depth data.
[0055] In an exemplary embodiment of the present disclosure, the Z-axis coordinate value of the Z-axis channel in the point cloud data is acquired, and each Z-axis coordinate value is configured as depth data.
[0056] For example, the RGB data output by a two-dimensional camera is marked as rgb, the point cloud data output by a three-dimensional camera is marked as depth (x, y, z), the RGBD data generated from the three-dimensional camera perspective is marked as rgbd, and the depth data generated from the three-dimensional camera perspective is marked as d.
[0057] Specifically, to obtain the coordinate value of the z channel of the point cloud data depth(x,y,z), the pseudo code can be expressed as: d=depth.splitChannels[2] (separate the channels of the point cloud data and extract the z channel). The point cloud data includes three channel data of x, y, and z respectively, and the array subscripts of the three channel data of x, y, and z can be 0, 1, and 2 respectively.
[0058] In step S220, coordinate transformation is performed on the RGB data according to the transformation matrix to obtain three-channel data having the same number as the third coordinate data, and RGBD data is generated according to the third coordinate data and the three-channel data.
[0059] In an exemplary embodiment of the present disclosure, a transformation matrix T of a two-dimensional camera coordinate system to a three-dimensional camera coordinate system is obtained. The transformation matrix T can be obtained by an external parameter calibration technology of a camera, and the present disclosure does not specifically limit this. The transformation matrix is shown in formula (1):
[0060]
[0061] Among them, R 33 Represents the rotation part, which is a 3*3 matrix, t 31 Represents the translation part, which is a 3*1 matrix.
[0062] In addition, the intrinsic parameter matrix M of the two-dimensional camera is obtained through the intrinsic parameter calibration technology, where the intrinsic parameter matrix is shown by formula (2):
[0063]
[0064] Among them, f x 、f y is the focal length of the 2D camera, c x 、c y is the center pixel coordinate of the image.
[0065] In an exemplary embodiment of the present disclosure, each coordinate point in the point cloud data depth (x, y, z) is traversed one by one. First, all the coordinate points are obtained, as shown in formula (3):
[0066]
[0067] Where 0≤i <depth.rows,0≤j<depth.cols。
[0068] Next, the coordinate points in the point cloud data are transformed into the two-dimensional camera coordinate system, as shown in formula (4):
[0069]
[0070] Then, the coordinate points of the point cloud data are projected onto the two-dimensional camera imaging plane, as shown in formula (5):
[0071]
[0072] Then, the pixel coordinates in the RGB data corresponding to the point cloud coordinate point are calculated, as shown in formula (6):
[0073]
[0074] Finally, obtain rgb(v,u) corresponding to the coordinate point (v,u) in the RGB data, configure rgbd(i,j) according to rgb(v,u) and d(i,j), and the generated RGBD data is shown in formula (7):
[0075] rgbd(i,j)={rgb(v,u),d(i,j)}, (7)
[0076] Among them, rgb(v,u) contains three-channel data of R (red), G (green), and B (blue), and d(i,j) contains depth data.
[0077] In step S120, the RGBD data is input into an instance segmentation network to obtain a plurality of bounding boxes and first pixel data corresponding to each bounding box, and the depth data is input into an edge extraction network to obtain edge pixel data corresponding to the depth data.
[0078] In an exemplary embodiment of the present disclosure, each bounding box is a rectangle that is parallel to the length and width of the RGB image, and the instance segmentation network outputs the position coordinate information of each bounding box, which may specifically include the two-dimensional image coordinate information of the upper left vertex and the lower right vertex of each bounding box.
[0079] In addition, the instance segmentation network also outputs the first pixel data corresponding to each bounding box, and the first pixel data is the pixel value of all pixels contained in each bounding box. The first pixel data includes binary data, specifically a binary data set including a first value and a second value, the first value is any positive integer, and the second value is zero. For example, the first value can be 1, or 255 or other positive integers, and the present disclosure does not specifically limit the value of the first value.
[0080] It should be noted that, assuming that the first value is 1 and the second value is 0, if the first pixel data corresponding to a pixel in the bounding box is 1, it means that the pixel is the pixel where the target object is located; if the first pixel data of a pixel is 0, it means that the pixel does not contain the target object. The first pixel data corresponding to each pixel in the bounding box can be judged, and the smaller bounding box where the target object is located can be calculated based on the first pixel data. The cv::minAreaRect() function of OpenCV can be used to calculate the smaller bounding box.
[0081] In an exemplary embodiment of the present disclosure, RGBD data is input into an instance segmentation network, and features are extracted from the RGBD data through the instance segmentation network to obtain multiple bounding boxes and first pixel data corresponding to each bounding box. The instance segmentation network can be any network model with bounding box prediction function and mask instance segmentation function, for example, the instance segmentation network can be a MaskRCNN network model, or a DeepMask, MultipathNet, FCIS network model, etc. The present disclosure does not specifically limit the type of instance segmentation network.
[0082] In an exemplary embodiment of the present disclosure, the image formed by the edge pixel data is a binary edge image of the same size as the image formed by the depth data, and the edge pixel data is binary data including a first value and a second value, and the binary edge image represents edge features of multiple target objects. The value of the first value in the first pixel data and the value of the first value in the edge pixel data may be the same or different, and the present disclosure does not specifically limit this.
[0083] It should be noted that, assuming that the first value is 1 and the second value is 0, if the edge pixel data corresponding to a pixel point is 1, it means that the pixel point is the edge of the target object; if the edge pixel data corresponding to the pixel point is 0, it means that the pixel point is a non-edge point of the target object, and the non-edge point can be a pixel point where the other part of the target object except the edge is located, or it can be a pixel point that is not related to the target object.
[0084] In an exemplary embodiment of the present disclosure, the edge extraction network extracts features from depth data based on a deep learning edge extraction method to obtain edge pixel data including multiple target objects. The edge extraction network can be any neural network model with edge extraction function, such as an RCF network model, or a DeepContour, DeepEdge network model, etc., which is not specifically limited in the present disclosure.
[0085] In step S130, the first pixel data is corrected according to the edge pixel data to obtain target pixel data corresponding to each bounding box, and the coordinate information of each target object is determined according to the target pixel data.
[0086] In an exemplary embodiment of the present disclosure, data correction is performed on the first pixel data according to the edge pixel data to obtain target pixel data corresponding to each enclosing frame. Figure 3 A schematic diagram of the process of determining the coordinate information of each target object is shown in FIG. Figure 3 As shown, the process at least includes steps S310 to S320, which are described in detail as follows:
[0087] In step S310, the first pixel data in each bounding box is traversed respectively, the target coordinate point whose first pixel data is a first value is obtained, and the target edge data corresponding to the target coordinate point is obtained from the edge pixel data.
[0088] In an exemplary embodiment of the present disclosure, the first pixel data in each bounding box is traversed one by one, and if the first pixel data is a first value, the target coordinate point corresponding to the first pixel data is obtained, and the target edge data corresponding to the target coordinate point is obtained in the edge pixel data according to the target coordinate point.
[0089] In an exemplary embodiment of the present disclosure, if the first pixel data within the bounding box is a second value, the first pixel data is maintained unchanged.
[0090] In step S320, it is determined whether the target edge data is a first value, and the target pixel data is determined according to the determination result.
[0091] In an exemplary embodiment of the present disclosure, if the target edge data is a first value, the first pixel data corresponding to the target coordinate point is corrected to a second value, and the corrected first pixel data is used as the target pixel data.
[0092] In an exemplary embodiment of the present disclosure, if the target edge data is a second value, the first pixel data at the target coordinate point is maintained unchanged.
[0093] For example, Figure 4 FIG. 4 is a schematic diagram showing a flow chart of correcting first pixel data according to a specific embodiment of the present disclosure. Figure 4 As shown, in step S410, a first bounding box and first pixel data corresponding to the first bounding box are obtained from multiple bounding boxes; in step S420, the first pixel data corresponding to the first bounding box is traversed to determine whether the first pixel data is a first value; in step S430, if the first pixel data is a second value, the first pixel data is maintained unchanged; in step S440, if the first pixel data is a first value, a target coordinate point corresponding to the first pixel value is obtained, and target edge data of the edge pixel data at the target coordinate point is obtained; in step S450, it is determined whether the target edge data is a first value; in step S460, if the target edge data is a second value, the first pixel data is maintained unchanged; in step S470; if the target edge data is a first value, the first pixel data corresponding to the target coordinate point is corrected to the second value to obtain the corrected target pixel data; in step S480, the next bounding box and the corresponding first pixel data are taken out, and steps S420 to S470 are repeatedly executed until all bounding boxes are processed as above.
[0094] That is to say, at the same coordinate point, if both the first pixel data and the edge pixel data are not zero, the first pixel data of the coordinate point is corrected to zero. 0 ,y 0 ), the first pixel data is 1 and the edge pixel data is 255, then the first pixel data is corrected to 0 to obtain the corrected first pixel data.
[0095] The target pixel data is generated according to the corrected first pixel data and the uncorrected first pixel data, and the generated target pixel data is binary data having the same size as the first pixel data.
[0096] In an exemplary embodiment of the present disclosure, determining coordinate information of each target object according to target pixel data includes: acquiring a first coordinate point with a first value in the target pixel data, wherein the coordinate position of the first coordinate point is the coordinate information of multiple target objects.
[0097] In an exemplary embodiment of the present disclosure, an opening operation is performed on the target pixel data to obtain second pixel data corresponding to each bounding box, and coordinate information of each target object is determined according to the second pixel data.
[0098] The opening operation is an operation of first corroding and then dilating in the image morphological processing technology. The target pixel data in each bounding box is opened according to the preset detection box. In the process of opening each bounding box, the pixel size of the preset detection box can be dynamically changed according to the pixel size of the bounding box. For example, if the pixel size of the first bounding box is 50*50, the preset detection box with a pixel size of 5*5 is used to open the target pixel data in the first bounding box; if the pixel size of the second bounding box is 30*30, the preset detection box with a pixel size of 3*3 is used to open the target pixel data in the second bounding box. The present disclosure does not specifically limit the pixel size of the preset detection box.
[0099] Specifically, a connected domain analysis is performed on the second pixel data corresponding to each bounding box to obtain multiple connected domains corresponding to each bounding box; the coordinate information of the maximum connected domain corresponding to each bounding box is obtained, and the coordinate information of each target object is determined according to the coordinate information of each maximum connected domain.
[0100] A connected domain analysis is performed on the second pixel data corresponding to each bounding box whose pixel value is the first value, to obtain multiple connected domains corresponding to each bounding box, and the maximum connected domain corresponding to each bounding box is retained.
[0101] In addition, the second pixel values corresponding to the remaining connected domains are corrected to 0.
[0102] In an exemplary embodiment of the present disclosure, the pixel coordinates of the maximum connected domain corresponding to each bounding box are obtained, and the minimum bounding box of multiple target objects can be calculated using the cv::minAreaRect() function of OpenCV. The coordinate information of each target object is the coordinate position corresponding to the minimum bounding box.
[0103] In a specific embodiment of the present disclosure, the target detection method is applied to a scenario in which a robot picks items in a turnover box. Figure 5 A structural diagram showing a specific application scenario of the present disclosure is shown in FIG. Figure 5As shown, the item picking scene includes a picking station 501, an item turnover box 502, a plurality of items 503 placed in the item turnover box 502, a two-dimensional camera 504 and a three-dimensional camera 505 arranged above the item turnover box 502, and a robot 506.
[0104] Figure 6 A schematic diagram of a target detection method according to a specific embodiment is shown. Figure 6 As shown, the process at least includes steps S610 to S670, which are specifically described as follows:
[0105] In step S610, the two-dimensional camera 504 is started to take photos of the multiple items 503 in the item turnover box 502 to obtain RGB image data, and the three-dimensional camera 505 is started to take photos of the multiple items 503 to obtain point cloud data (xyz coordinate data);
[0106] In step S620, the coordinate transformation matrix of the two-dimensional camera 504 and the three-dimensional camera 505 and the internal parameter matrix of the two-dimensional camera 504 are used to generate RGBD data and depth data under the perspective of the three-dimensional camera 504 according to the RGB image data and the point cloud data;
[0107] In step S630, the RGBD data is input into an instance segmentation network to obtain a plurality of bounding boxes and first pixel data corresponding to the bounding boxes, and the depth data is input into an edge extraction network to obtain edge pixel data; wherein the image composed of the first pixel data output by the instance segmentation network is as shown in FIG. Figure 7 As shown, the image composed of edge pixel data output by the edge extraction network is as follows Figure 8 shown.
[0108] In step S640, the first pixel data corresponding to each enclosing frame is corrected according to the edge pixel data to obtain the target pixel data corresponding to each enclosing frame;
[0109] In step S650, the target pixel data corresponding to each bounding box is opened according to the preset detection frame to obtain the second pixel data corresponding to each bounding box; wherein the image composed of the second pixel data is as follows: Fig. 9 As shown, each bounding box is broken at the edge of the object 503.
[0110] In step S660, a connected domain analysis is performed on the second pixel data corresponding to each bounding box to obtain the maximum connected domain corresponding to each bounding box, and a minimum bounding box is determined according to the pixel coordinates of the maximum connected domain. The coordinate position of the minimum bounding box is the coordinate information of each object 503. The image formed by the minimum bounding box is as shown in FIG. Fig.10 shown.
[0111] In step S670 , the coordinate information of each object 503 is sent to the robot 506 , so that the robot 506 picks the object 503 according to the coordinate information.
[0112] Those skilled in the art will appreciate that all or part of the steps for implementing the above embodiments are implemented as a computer program executed by a CPU. When the computer program is executed by the CPU, the above functions defined by the above method provided by the present invention are performed. The program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk.
[0113] In addition, it should be noted that the above-mentioned figures are only schematic illustrations of the processes included in the method according to an exemplary embodiment of the present invention, and are not intended to be limiting. It is easy to understand that the processes shown in the above-mentioned figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.
[0114] The following describes an apparatus embodiment of the present disclosure, which can be used to perform the target detection method described above. For details not disclosed in the apparatus embodiment of the present disclosure, please refer to the embodiment of the target detection method described above.
[0115] Fig.11 A block diagram of a target detection device according to an embodiment of the present disclosure is schematically shown.
[0116] Reference Fig.11 As shown, according to an embodiment of the present disclosure, the target detection device 1100 includes: a data acquisition module 1101, a data analysis module 1102 and a coordinate determination module 1103. Specifically:
[0117] The data acquisition module 1101 is used to acquire RGB data and point cloud data containing multiple target objects, and generate RGBD data and depth data according to the RGB data and point cloud data;
[0118] The data analysis module 1102 is used to input the RGBD data into an instance segmentation network to obtain a plurality of bounding boxes and first pixel data corresponding to each bounding box, and input the depth data into an edge extraction network to obtain edge pixel data corresponding to the depth data;
[0119] The coordinate determination module 1103 is used to perform data correction on the first pixel data according to the edge pixel data to obtain target pixel data corresponding to each bounding box, and determine the coordinate information of each target object according to the target pixel data.
[0120] In an exemplary embodiment of the present disclosure, the coordinate determination module 1103 can also be used to preset the detection frame edge pixel data and the preset detection frame first pixel data as binary data, the preset detection frame binary data including a first value and a second value; perform data correction on the preset detection frame first pixel data according to the preset detection frame edge pixel data to obtain target pixel data corresponding to each preset detection frame enclosing frame, including: traversing the preset detection frame first pixel data in each preset detection frame enclosing frame respectively, obtaining the target coordinate point where the preset detection frame first pixel data is the first value, and obtaining the target edge data corresponding to the preset detection frame target coordinate point in the preset detection frame edge pixel data; judging whether the preset detection frame target edge data is the preset detection frame first value, and determining the preset detection frame target pixel data according to the judgment result.
[0121] In an exemplary embodiment of the present disclosure, the coordinate determination module 1103 can also be used to determine the preset detection frame target pixel data based on the preset detection frame judgment result, including: if the preset detection frame target edge data is the preset detection frame first value, then the preset detection frame first pixel data corresponding to the preset detection frame target coordinate point is corrected to the second value, and the corrected first pixel data is used as the preset detection frame target pixel data.
[0122] In an exemplary embodiment of the present disclosure, the coordinate determination module 1103 can also be used to determine the coordinate information of each preset detection frame target object based on the preset detection frame target pixel data, including: performing an open operation on the preset detection frame target pixel data to obtain second pixel data corresponding to each preset detection frame enclosing frame, and determining the coordinate information of each preset detection frame target object based on the preset detection frame second pixel data.
[0123] In an exemplary embodiment of the present disclosure, the coordinate determination module 1103 can also be used to determine the coordinate information of the target object of each preset detection frame based on the second pixel data of the preset detection frame, including: performing connected domain analysis on the second pixel data of the preset detection frame corresponding to each preset detection frame enclosing frame, to obtain multiple connected domains corresponding to each preset detection frame enclosing frame; obtaining the coordinate information of the maximum connected domain corresponding to each preset detection frame enclosing frame, and determining the coordinate information of the target object of each preset detection frame based on the coordinate information of the maximum connected domain of each preset detection frame.
[0124] In an exemplary embodiment of the present disclosure, the data acquisition module 1101 can also be used to acquire RGB data and point cloud data containing multiple target objects, including: preset detection frame RGB data containing multiple target objects in a preset detection frame obtained by shooting with a two-dimensional camera, and preset detection frame point cloud data containing multiple target objects in a preset detection frame obtained by shooting with a three-dimensional camera of the preset detection frame, wherein the resolution of the preset detection frame two-dimensional camera is greater than the resolution of the preset detection frame three-dimensional camera.
[0125] In an exemplary embodiment of the present disclosure, the data acquisition module 1101 can also be used to preset detection frame point cloud data including first coordinate data, second coordinate data and third coordinate data; generate RGBD data and depth data according to the preset detection frame RGB data and the preset detection frame point cloud data, including: configuring the preset detection frame third coordinate data as preset detection frame depth data; performing coordinate transformation on the preset detection frame RGB data according to the transformation matrix to obtain three-channel data with the same number as the preset detection frame third coordinate data, and generate the preset detection frame RGBD data according to the preset detection frame third coordinate data and the preset detection frame three-channel data.
[0126] The specific details of each of the above target detection devices have been described in detail in the corresponding target detection method, so they will not be repeated here.
[0127] It should be noted that, although several modules or units of the device for execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0128] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.
[0129] It will be appreciated by those skilled in the art that various aspects of the present invention may be implemented as a system, method or program product. Therefore, various aspects of the present invention may be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to herein as a "circuit", "module" or "system".
[0130] Refer to the following Fig.12 The electronic device 1200 according to this embodiment of the present invention is described. Fig.12 The electronic device 1200 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0131] like Fig.12 As shown, the electronic device 1200 is in the form of a general computing device. The components of the electronic device 1200 may include, but are not limited to: the at least one processing unit 1210, the at least one storage unit 1220, a bus 1230 connecting different system components (including the storage unit 1220 and the processing unit 1210), and a display unit 1240.
[0132] The storage unit stores program codes, which can be executed by the processing unit 1210, so that the processing unit 1210 performs the steps according to various exemplary embodiments of the present invention described in the above “Exemplary Method” section of this specification. For example, the processing unit 1210 can perform the following steps: Figure 1 In step S110 shown in the figure, RGB data and point cloud data containing multiple target objects are obtained, and RGBD data and depth data are generated according to the RGB data and the point cloud data; in step S120, the RGBD data is input into an instance segmentation network to obtain multiple bounding boxes and first pixel data corresponding to each bounding box, and the depth data is input into an edge extraction network to obtain edge pixel data corresponding to the depth data; in step S130, data correction is performed on the first pixel data according to the edge pixel data to obtain target pixel data corresponding to each bounding box, and coordinate information of each target object is determined according to the target pixel data.
[0133] The storage unit 1220 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 12201 and / or a cache storage unit 12202 , and may further include a read-only storage unit (ROM) 12203 .
[0134] The storage unit 1220 may also include a program / utility 12204 having a set (at least one) of program modules 12205, such program modules 10205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0135] Bus 1230 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0136] The electronic device 1200 may also communicate with one or more external devices 1400 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable viewers to interact with the electronic device 1200, and / or communicate with any device that enables the electronic device 1200 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 1250. In addition, the electronic device 1200 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 1260. As shown, the network adapter 1260 communicates with other modules of the electronic device 1200 via a bus 1230. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0137] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0138] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the above method of the present specification is stored. In some possible implementations, various aspects of the present invention can also be implemented in the form of a program product, which includes a program code, and when the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of the present specification.
[0139] refer to Fig.13 As shown, a program product 1300 for implementing the above method according to an embodiment of the present invention is described, which can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.
[0140] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0141] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0142] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0143] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0144] In addition, the above-mentioned figures are only schematic illustrations of the processes included in the method according to an exemplary embodiment of the present invention, and are not intended to be limiting. It is easy to understand that the processes shown in the above-mentioned figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.
[0145] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and examples are to be considered as exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.
[0146] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A target detection method, It is characterized in that include: Acquire RGB data and point cloud data containing multiple target objects, and generate RGBD data and depth data according to the RGB data and the point cloud data; Inputting the RGBD data into an instance segmentation network to obtain a plurality of bounding boxes and first pixel data corresponding to each of the bounding boxes, and inputting the depth data into an edge extraction network to obtain edge pixel data corresponding to the depth data; wherein the edge pixel data and the first pixel data are binary data, and the binary data includes a first value and a second value; Performing data correction on the first pixel data according to the edge pixel data to obtain target pixel data corresponding to each of the bounding boxes, and determining coordinate information of each of the target objects according to the target pixel data; Among them, the first pixel data is corrected according to the edge pixel data to obtain the target pixel data corresponding to each of the bounding boxes, including: traversing the first pixel data in each of the bounding boxes respectively, obtaining the target coordinate point where the first pixel data is a first value, and obtaining the target edge data corresponding to the target coordinate point in the edge pixel data; judging whether the target edge data is the first value, and determining the target pixel data according to the judgment result.
2. The target detection method according to claim 1, It is characterized in that Determining the target pixel data according to the judgment result includes: If the target edge data is the first value, the first pixel data corresponding to the target coordinate point is corrected to a second value, and the corrected first pixel data is used as the target pixel data.
3. The target detection method according to claim 1, It is characterized in that Determining coordinate information of each target object according to the target pixel data includes: An opening operation is performed on the target pixel data to obtain second pixel data corresponding to each of the bounding boxes, and coordinate information of each of the target objects is determined according to the second pixel data.
4. The target detection method according to claim 3, It is characterized in that Determining coordinate information of each of the target objects according to the second pixel data includes: Performing connected domain analysis on the second pixel data corresponding to each of the bounding boxes respectively to obtain a plurality of connected domains corresponding to each of the bounding boxes; The coordinate information of the maximum connected domain corresponding to each of the bounding boxes is obtained, and the coordinate information of each of the target objects is determined according to the coordinate information of each of the maximum connected domains.
5. The target detection method according to claim 1, It is characterized in that Get RGB data and point cloud data containing multiple target objects, including: The RGB data containing the multiple target objects is obtained by shooting with a two-dimensional camera, and the point cloud data containing the multiple target objects is obtained by shooting with a three-dimensional camera, wherein the resolution of the two-dimensional camera is greater than the resolution of the three-dimensional camera.
6. The target detection method according to claim 5, It is characterized in that The point cloud data includes first coordinate data, second coordinate data and third coordinate data; Generating RGBD data and depth data according to the RGB data and the point cloud data, including: configuring the third coordinate data as the depth data; The RGB data is coordinate transformed according to a transformation matrix to obtain three-channel data having the same number as the third coordinate data, and the RGBD data is generated according to the third coordinate data and the three-channel data.
7. A target detection device, It is characterized in that include: A data acquisition module, used to acquire RGB data and point cloud data containing multiple target objects, and generate RGBD data and depth data according to the RGB data and the point cloud data; A data analysis module, configured to input the RGBD data into an instance segmentation network to obtain a plurality of bounding boxes and first pixel data corresponding to each of the bounding boxes, and input the depth data into an edge extraction network to obtain edge pixel data corresponding to the depth data; wherein the edge pixel data and the first pixel data are binary data, and the binary data includes a first value and a second value; A coordinate determination module, configured to perform data correction on the first pixel data according to the edge pixel data to obtain target pixel data corresponding to each of the bounding boxes, and determine coordinate information of each of the target objects according to the target pixel data; Among them, the first pixel data is corrected according to the edge pixel data to obtain the target pixel data corresponding to each of the bounding boxes, including: traversing the first pixel data in each of the bounding boxes respectively, obtaining the target coordinate point where the first pixel data is a first value, and obtaining the target edge data corresponding to the target coordinate point in the edge pixel data; judging whether the target edge data is the first value, and determining the target pixel data according to the judgment result.
8. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the target detection method according to any one of claims 1 to 6 is implemented.
9. An electronic device, It is characterized in that include: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enables the one or more processors to implement the target detection method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Indoor scene outline detection method fusing color and depth information
CN107578418A
Human skeleton key point extraction method and computer readable storage medium
CN110619285A