A palletizing and depalletizing method, device, electronic device and system based on vision guidance
By extracting the combined segmentation method of point cloud area of the top-level box and visible light image of the box stack, the problem of low box positioning accuracy in the prior art is solved, and higher box positioning accuracy and destacking efficiency are achieved.
Patent Information
- Application Number
- CN202211204153.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-09-29
AI Technical Summary
In the existing visual guidance-based depalletization method, box position recognition heavily relies on the recognition accuracy of two-dimensional images, resulting in low positioning accuracy.
By obtaining the depth map and visible light image of the box stack, the top-level box point cloud area of the box stack is extracted, its corresponding area in the visible light image is determined, and the box instance segmentation is performed to obtain the divided area of each box, and destack is performed based on these divided areas.
The positioning accuracy of the box during unstacking is improved, the interference of other information in the visible light image is reduced, and the positioning accuracy of the box segmentation area is enhanced.
Smart Images

Figure CN115533902B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular, to a depalletizing method, device, electronic device and system based on visual guidance. Background Art
[0002] The technology of using visual guidance to control a robotic arm to perform the depalletizing task of boxes is widely used in the fields of logistics, warehousing, etc. The so-called visual guidance means using a vision device to obtain the image and three-dimensional point cloud information of the box stack, determining the spatial pose (X, Y, Z, Rx, Ry, Rz) of six degrees of freedom (6 Degree of Freedom, 6DOF) of each box to be depalletized through methods such as image processing and point cloud processing, and finally controlling the robotic arm to perform grasping.
[0003] In the existing depalletizing method based on visual guidance, a two-dimensional image and three-dimensional point cloud data of the box stack are collected, the boxes in the two-dimensional image are directly recognized through computer vision technology to obtain the positions of the boxes in the two-dimensional image; based on the coordinate conversion relationship between the two-dimensional image and the point cloud data, the mapping area of each box in the two-dimensional image in the point cloud data is determined to obtain the positions of the boxes in three-dimensional space; then the positions of the boxes in three-dimensional space are converted into the corresponding 6DOF spatial poses of the robotic arm, so as to use the robotic arm to perform depalletizing. However, using the above method, the recognition of the box positions seriously depends on the recognition accuracy of the boxes in the two-dimensional image, and there is a technical problem of low positioning accuracy. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a depalletizing method, device, electronic device and system based on visual guidance to improve the positioning accuracy of the boxes during depalletizing. The specific technical solutions are as follows:
[0005] According to the first aspect of the embodiments of the present application, a depalletizing method based on visual guidance is provided, including:
[0006] Obtain a depth map and a visible light image of the box stack, where the box stack includes at least one box;
[0007] Based on the depth map, extract the point cloud region of the top-layer box of the box stack;
[0008] Determine the region corresponding to the point cloud region of the top-layer box in the visible light image to obtain the box mask region;
[0009] Perform box instance segmentation on the box mask region to obtain the respective segmentation regions of each box;
[0010] According to the segmentation region of each box, perform depalletizing on the box stack.
[0011] In a possible implementation manner, extracting the top-layer box body point cloud region of the box stack based on the depth map includes:
[0012] In the depth image coordinate system, convert the depth map into first point cloud data;
[0013] Map the first point cloud data to the coordinate system of the box stack to obtain second point cloud data;
[0014] Segment the second point cloud data to obtain at least one point cloud classification, where the plane features of the point clouds in the same point cloud classification are the same;
[0015] In the at least one point cloud classification, select the point cloud classification with the highest height to obtain the top-layer box body point cloud region of the box stack.
[0016] In a possible implementation manner, performing box instance segmentation on the box body mask region to obtain the respective segmentation regions of each box body includes:
[0017] Use a pre-trained deep learning model to perform box instance segmentation on the box body mask region to obtain the respective segmentation regions of each box body, where the deep learning model is trained by sample box body mask images, and the sample box body mask images are images of the top-layer box bodies in the sample box stack.
[0018] In a possible implementation manner, unstacking the box stack according to the segmentation region of each box body includes:
[0019] For each box body, extract the image edge feature of the segmentation region of the box body, or extract the image edge feature of the grayscale image of the segmentation region of the box body;
[0020] Map the segmentation region of each box body to the point cloud data converted from the depth map to obtain the point cloud mapping region of each box body;
[0021] For each box body, extract the point cloud edge feature of the point cloud mapping region of the box body;
[0022] For each box body, determine the edge straight line feature of the box body according to the image edge feature and the point cloud edge feature of the box body;
[0023] For each box body, perform rectangular splicing on the edge straight line feature of the box body to obtain at least one candidate rectangular frame of the box body;
[0024] For each box body, select the box body rectangular frame of the box body from the candidate rectangular frames of the box body;
[0025] Unstack the box stack according to the box body rectangular frame of each box body.
[0026] In a possible implementation manner, for each box body, determining the edge straight line feature of the box body according to the image edge feature and the point cloud edge feature of the box body includes:
[0027] For each box body, determining the edge pixel region of the box body in the depth map according to the image edge feature and the point cloud edge feature of the box body;
[0028] Performing straight line feature extraction on the edge pixel region of the box body to obtain the edge straight line feature of the box body.
[0029] In a possible implementation manner, for each box body, performing rectangular splicing on the edge straight line feature of the box body to obtain at least one candidate rectangular frame of the box body, including:
[0030] Calculating a first conversion relationship between the visible light image coordinate system of the visible light image and the top box body coordinate system of the top box body point cloud region;
[0031] For each box body, converting the segmentation region of the box body into the top box body coordinate system according to the first conversion relationship to obtain the size range of the box body;
[0032] Performing rectangular splicing on the edge straight line feature of the box body according to the size range of the box body and a preset rectangular angle threshold to obtain at least one candidate rectangular frame of the box body.
[0033] In a possible implementation manner, for each box body, selecting the box body rectangular frame of the box body from the candidate rectangular frames of the box body, including:
[0034] For each box body, respectively calculating the projection point cloud occupancy ratio of each candidate rectangular frame of the box body according to the point cloud mapping region of the box body; respectively calculating the image edge intensity and rectangularity of each candidate rectangular frame of the box body;
[0035] For each candidate rectangular frame, calculating the score of the candidate rectangular frame according to the projection point cloud occupancy ratio, image edge intensity and rectangularity of the candidate rectangular frame;
[0036] Calculating the interference degree between adjacent candidate rectangular frames according to the positional relationship between adjacent candidate rectangular frames;
[0037] Determining the box body rectangular frame of each box body according to the scores of each candidate rectangular frame and the interference degree between adjacent candidate rectangular frames.
[0038] In a possible implementation manner, the determining the box body rectangular frame of each box body according to the candidate scores of each candidate rectangular frame and the interference degree between adjacent candidate rectangular frames includes:
[0039] Among the candidate rectangular frames, select the top N candidate rectangular frames with the highest scores as seed rectangular frames, where N is a preset integer;
[0040] For each candidate rectangular frame among the N seed rectangular frames, use this candidate rectangular frame as a reference, and determine a group of candidate rectangular frames including this candidate rectangular frame with scores, interference degrees, and the number of boxes as constraints;
[0041] Among the N groups of candidate rectangular frames, select the group with the highest score to obtain the box rectangular frame of each box.
[0042] In a possible implementation manner, before the step of unstacking the box stack according to the box rectangular frame of each box, the method further includes:
[0043] For each box rectangular frame, determine the depth pixel region corresponding to this box rectangular frame in the depth map, and expand the depth pixel region of this box rectangular frame to obtain the depth pixel expansion region of this box rectangular frame;
[0044] Determine the visible light pixel region corresponding to this box rectangular frame in the visible light image, and expand the visible light pixel region of this box rectangular frame to obtain the visible light pixel expansion region of this box rectangular frame;
[0045] Determine the intersection region of the depth pixel expansion region and the visible light pixel expansion region of this box rectangular frame in the same coordinate system;
[0046] Use the image gradient of the edge of this box rectangular frame, the straightness of the gray - level edge points, and the distance from the edge of the top - layer box point cloud region as constraints to correct this box rectangular frame in the intersection region of this box rectangular frame;
[0047] The step of unstacking the box stack according to the box rectangular frame of each box includes:
[0048] Unstack the box stack according to the corrected box rectangular frame of each box.
[0049] According to the second aspect of the embodiments of the present application, a vision - guided unstacking device is provided, including:
[0050] An acquisition module, configured to acquire the depth map and visible light image of the box stack, where the box stack includes at least one box;
[0051] An extraction module, configured to extract the top - layer box point cloud region of the box stack based on the depth map;
[0052] An acquisition module, configured to determine the region corresponding to the top-layer box body point cloud region in the visible light image, and acquire a box body mask region;
[0053] A segmentation module, configured to perform box body instance segmentation on the box body mask region to obtain the segmentation region of each box body;
[0054] A pallet-unstacking module, configured to unstack the pallet according to the segmentation region of each box body.
[0055] In a possible implementation manner, the extraction module includes:
[0056] A conversion sub-module, configured to convert the depth map into first point cloud data in a depth image coordinate system;
[0057] A first mapping sub-module, configured to map the first point cloud data to the coordinate system of the pallet to obtain second point cloud data;
[0058] A segmentation sub-module, configured to segment the second point cloud data to obtain at least one point cloud classification, where the plane features of the point clouds in the same point cloud classification are the same;
[0059] A first selection sub-module, configured to select the point cloud classification with the highest height from the at least one point cloud classification to obtain the top-layer box body point cloud region of the pallet.
[0060] In a possible implementation manner, the segmentation module is specifically configured to:
[0061] Use a pre-trained deep learning model to perform box body instance segmentation on the box body mask region to obtain the segmentation region of each box body, where the deep learning model is trained by a sample box body mask image, and the sample box body mask image is an image of the top-layer box body in a sample pallet.
[0062] In a possible implementation manner, the pallet-unstacking module includes:
[0063] A first extraction sub-module, configured to extract the image edge feature of the segmentation region of each box body, or extract the image edge feature of the grayscale image of the segmentation region of each box body;
[0064] A second mapping sub-module, configured to map the segmentation region of each box body to the point cloud data converted from the depth map to obtain the point cloud mapping region of each box body;
[0065] A second extraction sub-module, configured to extract the point cloud edge feature of the point cloud mapping region of each box body;
[0066] A determination sub-module, configured to determine, for each box body, an edge straight-line feature of the box body according to the image edge feature and the point cloud edge feature of the box body;
[0067] A splicing sub-module, configured to perform rectangular splicing on the edge straight-line feature of each box body to obtain at least one candidate rectangular frame of the box body;
[0068] A second selection sub-module, configured to select, for each box body, a box body rectangular frame of the box body from the candidate rectangular frames of the box body;
[0069] A palletizing removal sub-module, configured to remove the pallet of boxes according to the box body rectangular frame of each box body.
[0070] In a possible implementation manner, the determination sub-module is specifically configured to:
[0071] For each box body, determine an edge pixel region of the box body in the depth map according to the image edge feature and the point cloud edge feature of the box body;
[0072] Perform straight-line feature extraction on the edge pixel region of the box body to obtain the edge straight-line feature of the box body.
[0073] In a possible implementation manner, the splicing sub-module includes:
[0074] A first calculation unit, configured to calculate a first conversion relationship between the visible light image coordinate system of the visible light image and the top box body coordinate system of the top box body point cloud region;
[0075] A conversion unit, configured to, for each box body, convert the segmentation region of the box body into the top box body coordinate system according to the first conversion relationship to obtain the size range of the box body;
[0076] A splicing unit, configured to perform rectangular splicing on the edge straight-line feature of the box body according to the size range of the box body and a preset rectangular angle threshold to obtain at least one candidate rectangular frame of the box body.
[0077] In a possible implementation manner, the second selection sub-module includes:
[0078] A second calculation unit, configured to, for each box body, calculate the projection point cloud occupancy ratio of each candidate rectangular frame of the box body according to the point cloud mapping region of the box body; calculate the image edge intensity and rectangularity of each candidate rectangular frame of the box body respectively;
[0079] A third calculation unit, configured to calculate a score of each candidate rectangular frame according to the projection point cloud occupancy ratio, the image edge intensity, and the rectangularity of the candidate rectangular frame;
[0080] A fourth computing unit, configured to calculate the interference degree between each pair of adjacent candidate rectangular frames according to the positional relationship between adjacent candidate rectangular frames;
[0081] A first determining unit, configured to determine the rectangular frame of each box according to the scores of each of the candidate rectangular frames and the interference degree between adjacent candidate rectangular frames.
[0082] In a possible implementation manner, the determining unit is specifically configured to:
[0083] Among each of the candidate rectangular frames, select the top N candidate rectangular frames with the highest scores as seed rectangular frames, where N is a preset integer;
[0084] For each candidate rectangular frame among the N seed rectangular frames, use this candidate rectangular frame as a reference, and use the score, interference degree, and number of boxes as constraint conditions to determine a set of candidate rectangular frames including this candidate rectangular frame;
[0085] Among the N sets of candidate rectangular frames, select the set with the highest score to obtain the rectangular frame of each box.
[0086] In a possible implementation manner, the apparatus further includes:
[0087] A first expansion unit, configured to, before performing the step of unstacking the box stack according to the rectangular frame of each box, for each rectangular frame of the box, determine the depth pixel region corresponding to this rectangular frame of the box in the depth map, and expand the depth pixel region of this rectangular frame of the box to obtain the depth pixel expansion region of this rectangular frame of the box;
[0088] A second expansion unit, configured to determine the visible light pixel region corresponding to this rectangular frame of the box in the visible light image, and expand the visible light pixel region of this rectangular frame of the box to obtain the visible light pixel expansion region of this rectangular frame of the box;
[0089] A second determining unit, configured to determine the intersection region of the depth pixel expansion region and the visible light pixel expansion region of this rectangular frame of the box in the same coordinate system;
[0090] A correction unit, configured to use the image gradient of the edge of this rectangular frame of the box, the straightness of the gray edge points, and the distance from the edge of the point cloud region of the top box as constraint conditions to correct this rectangular frame of the box in the intersection region of this rectangular frame of the box;
[0091] The unstacking sub-module is specifically configured to: unstack the box stack according to the corrected rectangular frame of each box.
[0092] According to a third aspect of the embodiments of the present application, there is provided an electronic device, including: a processor and a memory;
[0093] The memory stores instructions executable by the at least one processor.
[0094] The instructions are executed by the at least one processor to implement any of the above-mentioned vision-guided palletizing methods.
[0095] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any of the above-mentioned vision-guided palletizing methods is implemented.
[0096] According to a fifth aspect of the embodiments of the present application, a vision-guided palletizing system is provided, including:
[0097] A depth camera, a palletizing robotic arm, and a control device;
[0098] The depth camera is configured to collect a depth map and a visible light image of the pallet.
[0099] The palletizing robotic arm is configured to palletize the pallet in response to an instruction of the control device.
[0100] The control device is configured to implement any of the above-mentioned vision-guided palletizing methods during operation.
[0101] Advantages of the embodiments of the present application:
[0102] In the technical solution provided by the embodiments of the present application, by extracting the point cloud region of the top layer box of the pallet and determining the corresponding region of the point cloud region of the top layer box in the visible light image, the interference regions of non-top layer boxes in the visible light image can be filtered out, and the box instance segmentation is performed on the box mask region of the top layer box, which is equivalent to combining the information of the depth map and the visible light image for box instance segmentation, and can improve the accuracy of the positioning of the box segmentation region; in addition, only the box mask region of the top layer box is subjected to box instance segmentation, compared with performing instance segmentation on the entire visible light image, it can reduce the interference of other information in the visible light image, thereby improving the accuracy of the positioning of the box segmentation region, and finally can effectively improve the positioning accuracy of the box during palletizing. Of course, implementing any product or method of the present application does not necessarily need to achieve all the above-mentioned advantages at the same time. Description of the Drawings
[0103] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments according to these drawings.
[0104] Figure 1 This is the first flowchart of the palletizing and depalletizing method based on vision guidance provided by the embodiments of the present application;
[0105] Figure 2 This is the flowchart of extracting the point cloud region of the top layer box of the box stack provided by the embodiments of the present application;
[0106] Figure 3 This is the flowchart of instance segmentation training of the deep learning model provided by the embodiments of the present application;
[0107] Figure 4 This is the second flowchart of the palletizing and depalletizing method based on vision guidance provided by the embodiments of the present application;
[0108] Figure 5 This is the flowchart of stitching candidate rectangular boxes provided by the embodiments of the present application;
[0109] Figure 6 This is the flowchart of determining the rectangular box of the box body provided by the embodiments of the present application;
[0110] Figure 7 This is one of the flowcharts of screening the optimal rectangular box provided by the embodiments of the present application;
[0111] Figure 8 This is the flowchart of refining the contour of the rectangular box of the box body provided by the embodiments of the present application;
[0112] Figure 9 This is the flowchart of calculating the grasping pose provided by the embodiments of the present application;
[0113] Figure 10 This is the first structural schematic diagram of the palletizing and depalletizing device based on vision guidance provided by the embodiments of the present application;
[0114] Figure 11 This is the second structural schematic diagram of the palletizing and depalletizing device based on vision guidance provided by the embodiments of the present application;
[0115] Figure 12 This is the third structural schematic diagram of the palletizing and depalletizing device based on vision guidance provided by the embodiments of the present application;
[0116] Figure 13 This is the fourth structural schematic diagram of the palletizing and depalletizing device based on vision guidance provided by the embodiments of the present application;
[0117] Figure 14 This is the schematic diagram of an electronic device provided by the embodiments of the present application;
[0118] Figure 15 This is the schematic diagram of another electronic device provided by the embodiments of the present application;
[0119] Figure 16 A schematic diagram of the calibration situation provided by the embodiment of the present application. Detailed implementation manners
[0120] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0121] In order to improve the positioning accuracy of the box body during palletizing, the embodiment of the present application provides a palletizing method, device, electronic device and system based on visual guidance, which will be described in detail below.
[0122] Figure 1 The first flowchart of the palletizing method based on visual guidance provided by the embodiment of the present application. As Figure 1 shown, it includes the following steps:
[0123] Step S110, obtain the depth map and visible light image of the pallet, where the pallet includes at least one box body.
[0124] The pallet is stacked by box bodies, and the pallet includes at least one box body. The depth map and visible light image of the pallet can be collected by a depth camera, such as an RGBD camera.
[0125] Step S120, extract the point cloud region of the top box body of the pallet based on the depth map.
[0126] The depth map includes information in three dimensions of the length, width, and height of the pallet. Considering that palletizing must start from the top box body of the pallet, the point cloud region of the top box body of the pallet can be extracted using the depth map. For example, the depth map of the pallet can be converted into the point cloud data of the pallet, and the point cloud of the top box body can be extracted to obtain the point cloud region of the top box body.
[0127] Step S130, determine the corresponding region of the top box body point cloud region in the visible light image to obtain the box body mask region.
[0128] Convert the top box body point cloud region into the image coordinate system of the visible light image to obtain the region of the top box body in the visible light image, which is called the box body mask region.
[0129] Step S140, perform box body instance segmentation on the box body mask region to obtain the respective segmentation regions of each box body.
[0130] Using computer vision technology, the box instance is segmented from the box mask area, so as to obtain the respective segmentation areas of each box.
[0131] Step S150, according to the segmentation area of each box, the box stack is unstacked.
[0132] After obtaining the box segmentation area based on the two-dimensional visible light image, the manipulator can be controlled to unstack the box stack according to the box segmentation area.
[0133] In the technical solution provided by the embodiment of the present application, by extracting the point cloud area of the top box of the box stack and determining the area corresponding to the point cloud area of the top box in the visible light image, the interference area of the non-top box in the visible light image can be filtered out, and the box instance segmentation is performed on the box mask area of the top box. It is equivalent to combining the information of the depth map and the visible light image to perform box instance segmentation, which can improve the accuracy of the positioning of the box segmentation area; in addition, only the box mask area of the top box is subjected to box instance segmentation, compared with performing instance segmentation on the entire visible light image, it can reduce the interference of other information in the visible light image, thereby improving the accuracy of the positioning of the box segmentation area, and finally effectively improving the positioning accuracy of the box during unstacking.
[0134] When extracting the point cloud area of the top box of the box stack, the depth image needs to be converted into point cloud data. In a possible implementation, see Figure 2 The extraction of the point cloud area of the top box of the box stack may include the following steps:
[0135] Step S121, in the depth image coordinate system, convert the depth map into the first point cloud data;
[0136] The depth map includes information in the three dimensions of the length, width, and height of the box stack. According to the information in these three dimensions, the point cloud data of the box stack in the depth image coordinate system, that is, the first point cloud data, can be obtained.
[0137] Step S122, map the first point cloud data to the coordinate system of the box stack to obtain the second point cloud data;
[0138] According to the pre-calibrated correspondence between the depth image coordinate system and the box stack coordinate system, the first point cloud data is converted into the coordinate system of the box stack to obtain the second point cloud data.
[0139] In the unstacking scenario, such as Figure 16As shown, the box stack 1063 is stacked on the pallet 1064, and the RGBD camera 1061 is arranged above the box stack 1063 through the support rod 1062. Among them, the RGBD camera 1061 includes a left-eye camera Left, a right-eye camera Right, and a RGB camera. First, the pose transformation matrix Rw, Tw from the depth image coordinate system to the pallet plane coordinate system can be calibrated. By adding a height dimension on the basis of the pallet plane coordinate system, the box stack coordinate system is obtained. In one example, the calibration of the transformation matrix from the depth image coordinate system to the pallet plane coordinate system can be achieved through the following two schemes:
[0140] 1) Calibration board calibration: The RGBD camera includes a left-eye camera, a RGB camera, and a right-eye camera. Depth maps can be acquired through the left-eye camera and the right-eye camera, and visible light images can be obtained through the RGB camera. According to the factory parameters of the RGBD camera, the transformation matrix (Rw, Tw) between the depth image coordinate system and the visible light image can be obtained. Place a calibration board flat on the pallet, control the RGB camera to acquire a 2D visible light image, extract the feature points of the calibration image and calculate the external parameters from the RGB camera to the pallet. Finally, using the transformation matrix (Rc, Tc) from the RGB camera to the depth camera, calculate the transformation matrix (Rw, Tw) from the depth image coordinate system to the pallet plane coordinate system. Among them, the calibration board can be a checkerboard, a circle, etc., and the external parameter calculation can adopt the Zhang Zhengyou plane calibration method.
[0141] 2) Depth map calibration: Control the depth camera to acquire a depth map of an empty pallet, extract the point cloud on the pallet. The collected point cloud data is in the depth image coordinate system. Perform plane fitting on it, and construct the pallet coordinate system (Xw, Yw, Zw) with the normal vector of the fitting plane and the centroid of the point cloud. Then, calculate the transformation matrix (Rw, Tw) from the depth image coordinate system to the pallet plane coordinate system according to the method of aligning the coordinate axes with the origin.
[0142] Step S123, segment the second point cloud data to obtain at least one point cloud classification, where the plane features of the point clouds in the same point cloud classification are the same;
[0143] In one example, the second point cloud data can be segmented using a point cloud segmentation method, and the point clouds with the same plane features are divided into one category. Among them, point cloud segmentation can adopt methods such as concave-convex clustering, normal clustering, and iterative RANSAC (Random Sample Consensus) plane search.
[0144] Step S124, in the at least one point cloud classification, select the point cloud classification with the highest height to obtain the top box body point cloud area of the box stack.
[0145] In one example, the point clouds obtained by point cloud segmentation can be sorted by height for each point cloud classification, and the point clouds of the highest layer can be retained to obtain the point cloud region of the top layer box of the pallet.
[0146] In the subsequent process of obtaining the box mask region, the point cloud data of the top layer box of the pallet can be converted to the RGB camera coordinate system by using the calibrated transformation matrices Rw, Tw and Rc, Tc, projected into the RGB image coordinate system through the RGB camera internal parameters to generate a mask image, and then the region corresponding to the point cloud region of the top layer box in the visible light image can be determined, and the RGB information corresponding to the invalid pixel points in the mask image can be removed, and only the RGB information of the visible light image corresponding to the point cloud region of the top layer box is retained.
[0147] The interference background in the RGB image can be removed through the point cloud of the top layer box, and instance segmentation is performed on the RGB image that only retains the top layer box part in the visible light image. During the process of training the model, it is not necessary to collect a large number of samples, the generalization of the instance segmentation model is improved, and the false detection ratio can be effectively reduced.
[0148] The segmentation of the box mask region can be realized by a pre-trained deep learning model. In a possible implementation manner, the instance segmentation of the box mask region to obtain the respective segmentation regions of each box includes: using the pre-trained deep learning model to perform instance segmentation on the box mask region to obtain the respective segmentation regions of each box, wherein the deep learning model is trained by a sample box mask image, and the sample box mask image is an image of the top layer box in the sample pallet.
[0149] In one example, refer to Figure 3 , Figure 3 which is the flowchart of the instance segmentation training of the deep learning model provided by the embodiment of the present application. As Figure 3 shown, the instance segmentation training of the deep learning model may include the following steps:
[0150] Step S301, build a framework for the instance segmentation deep learning model;
[0151] For example, the TensorFlow deep learning framework can be used to establish a deep learning model library with the help of the keras application.
[0152] Step S302, obtain the box mask image retaining the RGB information of the top layer box in the sample pallet as the sample image data;
[0153] Step S303, use the sample image data set of the box as the training set to train the deep learning model;
[0154] Step S304: When the loss of the deep learning model converges or reaches the preset number of training times, obtain the pre-trained deep learning model.
[0155] Calculate the loss of the deep learning model based on the prediction result of the deep learning model and the ground truth calibration of the sample image data. When the loss of the deep learning model converges or reaches the preset number of training times, complete the training process to obtain the pre-trained deep learning model.
[0156] In this embodiment, the deep learning model is trained using the sample box mask image that only includes the top-level box in the sample box stack. Compared with training using the sample images of the entire box, since it can effectively reduce the interference of other information outside the top-level box, its segmentation result is more accurate and the model training speed is also faster.
[0157] Use the captured visible light image as the input of the deep learning model. The trained deep learning model will output the mask for discriminating and segmenting the target box, which can quickly segment the boxes in the box stack and effectively improve the detection rate and accuracy of the deep learning segmentation result.
[0158] After obtaining the segmentation region of the box, the segmentation region can be further transformed to obtain the box rectangle frame of each box. In one possible implementation, as Figure 4 shown, the steps for depalletizing according to the segmentation region of each box may include the following steps:
[0159] Step S151: For each box, extract the image edge feature of the segmentation region of the box, or extract the image edge feature of the grayscale image of the segmentation region of the box;
[0160] In one example, according to the rough localization result of the instance segmentation mask obtained by the deep learning model, in the RGB image coordinate system, traverse the instance segmentation mask to extract the image edge feature of the segmentation region of the box, or according to the pixel point coordinates in the mask image, obtain the corresponding grayscale data, and then extract the image edge feature of the grayscale image; among them, the extraction of the grayscale edge feature can be implemented using the Canny operator.
[0161] Step S152: Map the segmentation region of each box to the point cloud data converted from the depth map to obtain the point cloud mapping region of each box;
[0162] In one example, the point cloud data converted from the depth map can be the first point cloud data mentioned above. The segmentation region of each box can be mapped into the first point cloud data in the depth image coordinate system to obtain the point cloud mapping region of each box. In one example, the point cloud data converted from the depth map can be the second point cloud data mentioned above. The segmentation region of each box can be mapped into the second point cloud data in the box stack coordinate system to obtain the point cloud mapping region of each box.
[0163] Step S153: For each box, extract the point cloud edge feature of the point cloud mapping region of this box.
[0164] For example, the extraction of the point cloud edge feature can be realized by using an integral image.
[0165] Step S154: For each box, determine the edge straight line feature of this box according to the image edge feature and the point cloud edge feature of this box.
[0166] For each box, according to the image edge feature and the point cloud edge feature of this box, determine the edge pixel region of this box in the depth map; perform straight line feature extraction on the edge pixel region of this box to obtain the edge straight line feature of this box. The extraction of the straight line feature can be realized by methods such as Hough transform, short straight line clustering or chain search, all of which are within the protection scope of this application.
[0167] In this embodiment, based on the characteristic that the box grasping surface is rectangular, the edge straight line feature of this box can be extracted according to the texture, edge points of the grayscale image and the point cloud edge. Among them, the straight line feature extraction can be realized by methods such as Hough transform, short straight line clustering and chain search.
[0168] Step S155: For each box, perform rectangular splicing on the edge straight line feature of this box to obtain at least one candidate rectangular frame of this box.
[0169] According to the rough positioning of the segmentation region of the box and the constraint of the box plane, calculate the box size range and perform edge feature splicing, and candidate rectangular frames with sizes and angles meeting the conditions can be obtained.
[0170] Step S156: For each box, select the box rectangular frame of this box from the candidate rectangular frames of this box.
[0171] Traverse and screen the candidate rectangular frames, calculate the score for each spliced candidate rectangular frame. The score calculation can be a weighted sum calculation of image attributes; according to the intersection over union of the rectangles and the interference relationship between two adjacent rectangles, the group of candidate rectangular frames with the best result can be used as the box rectangular frame.
[0172] Step S157: Unstack the box stack according to the box rectangular frame of each box.
[0173] By separately extracting the image edge features and the point cloud data features in the instance segmentation mask, combining the two to determine the edge pixel region, and then extracting the straight edge features, and combining the rough positioning of the segmentation region of the box body and the plane constraint of the box body plane, a candidate rectangular box can be obtained, which can effectively improve the positioning accuracy of the grasping surface of the box body.
[0174] The candidate rectangular box can be obtained based on the box body size and the rectangular box constraint conditions. In a possible implementation manner, as Figure 5 shown, for each box body, rectangular splicing is performed on the edge straight line features of the box body to obtain at least one candidate rectangular box of the box body, including:
[0175] Step S501, calculate the first conversion relationship between the visible light image coordinate system of the visible light image and the top box body coordinate system of the top box body point cloud region;
[0176] Plane fitting can be performed on the box body point cloud region, the point cloud region of the top box body is saved, and the conversion matrix Rb,Tb from the visible light image coordinate system to the top box body coordinate system of the top box body point cloud region is calculated as the first conversion relationship, where the plane fitting can be implemented by a consistency algorithm.
[0177] Step S502, for each box body, according to the first conversion relationship, convert the segmentation region of the box body into the top box body coordinate system to obtain the size range of the box body;
[0178] The image edge features in the depth image coordinate system can be converted into the length edge features in the top box body coordinate system through the conversion matrices Rb,Tb and Rc,Tc.
[0179] Step S503, according to the size range of the box body and the preset rectangular angle threshold, perform rectangular splicing on the edge straight line features of the box body to obtain at least one candidate rectangular box of the box body.
[0180] The size range of the box body can be calculated based on the instance segmentation region and the box body plane of the box body in the top box body coordinate system. Rectangular splicing is performed on the straight line features of the physical plane based on the size range and the rectangular angle threshold, and the candidate rectangular boxes that meet the size and angle conditions are retained. Among them, the rectangular angle threshold is the geometric condition of the preset rectangular box. When performing rectangular splicing, the 4 angle values need to meet the preset geometric conditions.
[0181] By calculating the size range of the box body and combining the preset rectangular angle threshold for splicing the candidate rectangular box, the splicing accuracy of the candidate rectangular box of the box body can be effectively improved.
[0182] The box body may include multiple candidate rectangular frames, and further selection needs to be carried out among the candidate rectangular frames to obtain the box body rectangular frame of the box body. In a possible implementation manner, as Figure 6 shown, for each box body, selecting the box body rectangular frame of the box body from the candidate rectangular frames of the box body includes:
[0183] Step S601, for each box body, calculate the projection point cloud occupancy ratio of each candidate rectangular frame of the box body according to the point cloud mapping area of the box body; calculate the image edge intensity and rectangularity of each candidate rectangular frame of the box body respectively;
[0184] Among them, the projection point cloud occupancy ratio is the area ratio of the point cloud projected onto a plane model after projecting the point cloud of the rectangular frame onto the model; the larger the projection point cloud occupancy ratio, the more regular the shape of the rectangular frame; the image edge intensity is the amplitude of the edge point gradient, which can be calculated by a gradient operator; the rectangularity reflects the degree to which an object fills its circumscribed rectangle and is a parameter reflecting the similarity degree of an object to a rectangle.
[0185] Step S602, for each candidate rectangular frame, calculate the score of the candidate rectangular frame according to the projection point cloud occupancy ratio, image edge intensity and rectangularity of the candidate rectangular frame;
[0186] In an example, for each candidate rectangular frame, the rectangularity, projection point cloud occupancy ratio, and image edge intensity of the candidate rectangular frame can be weighted and summed to calculate the score of the candidate rectangular frame.
[0187] Step S603, calculate the interference degree between adjacent candidate rectangular frames according to the positional relationship between adjacent candidate rectangular frames;
[0188] The interference degree between candidate rectangular frames can be the ratio of the overlapping area between candidate rectangular frames. In an example, the interference relationship between two adjacent rectangles can be calculated according to the intersection over union of the rectangles, and an adjacency list can be constructed. The options in the adjacency list are "adjacent, non-interfering, interfering, and interfering area". For example, a local coordinate system can be established on one of the rectangles, and the coordinates of 8 vertices on the non-interfering boundary can be calculated in the local coordinate system. Then, the coordinates of the centroid of another rectangle are converted to the local coordinate system, and it is judged whether the centroid is within the non-interfering boundary. If the centroid falls within the non-interfering boundary, the two planar rectangles interfere, otherwise they do not interfere.
[0189] Step S604, determine the box body rectangular frame of each box body according to the scores of each candidate rectangular frame and the interference degree between adjacent candidate rectangular frames.
[0190] The higher the score of the candidate rectangle, the greater the probability that the candidate rectangle is determined as the box rectangle; the smaller the interference degree between adjacent candidate rectangles, the greater the probability that the adjacent candidate rectangle is determined as the box rectangle. The score of the candidate rectangle and the interference degree between adjacent candidate rectangles can be combined to limit the constraint conditions, so as to obtain the box rectangle of each box.
[0191] For each candidate rectangle, calculate the weighted sum of image attributes and calculate the interference relationship between its two adjacent rectangles as the screening condition for determining the box rectangle, which can effectively improve the stability of the box rectangle.
[0192] In a possible implementation manner, as Figure 7 shown, determining the box rectangle of each box according to the candidate scores of the candidate rectangles and the interference degree between adjacent candidate rectangles includes:
[0193] Step S701, among the candidate rectangles, select the top N candidate rectangles with the highest scores as the seed rectangles; where N is a preset integer;
[0194] In an example, to prevent excessive calculation, take N ≤ 5. When N = 5, that is, among the candidate rectangles, select the top 5 candidate rectangles with the highest scores.
[0195] Step S702, for each candidate rectangle among the N seed rectangles, this candidate rectangle can be used as a reference, and with the score, interference degree, and number of boxes as constraint conditions, determine a group of candidate rectangles including this candidate rectangle;
[0196] Specifically, when N = 5, take the top 5 candidate rectangles with the highest scores as the seed rectangles. Respectively take each seed rectangle as a reference, and the depth - first traversal algorithm can be used to search with the score, interference degree, and number of boxes as constraint conditions to determine the candidate rectangles of the current reference rectangle, obtain a group of candidate rectangles including this seed rectangle, and select 5 groups of candidate rectangles; among them, the larger the rectangle score, the better, the smaller the interference degree, the better, and the more the number of boxes, the better, that is, the number of candidate rectangles searched with the current reference rectangle should be as large as possible.
[0197] The following takes the calculation of a group of candidate rectangles as an example to introduce the selection process of each group of candidate rectangles. It can be understood that the following process is only for example and does not limit the protection scope of this application.
[0198] In an example, assume that there are m rectangles in a group of candidate rectangles. A feasible step is as follows:
[0199] Step 1: Obtain a set of rectangular boxes in the manner of depth - first traversal;
[0200] In one example, taking a seed rectangular box as a reference, traverse each candidate rectangular box and obtain it from the candidate rectangular boxes adjacent to and outside the seed rectangular box to get a set of rectangular boxes.
[0201] Step 2: Calculate the sum of the rectangular scores of this set of candidate rectangular boxes, that is, add up the scores of each rectangular box;
[0202] In one example, the scores of each candidate rectangle can be calculated by weighted summation according to attributes such as rectangularity, projection point cloud occupancy ratio, and image edge intensity.
[0203] Step 3: Calculate the interference degree of this set of candidate rectangular boxes;
[0204] In one example, the overlapping area of each rectangle in the candidate rectangular boxes can be calculated to obtain the average overlapping rate of a set of candidate rectangular boxes;
[0205] Specifically, for each rectangle, count the overlapping rate with the largest overlapping area in the candidate rectangular boxes; then add up the maximum overlapping rates corresponding to each rectangle, and divide by the number m of candidate rectangular boxes to obtain the average overlapping rate; finally, use the average overlapping rate plus 1 as the interference degree. Here, adding 1 is to prevent the average overlapping rate from being 0. The smaller the overlapping area between any two rectangles in the candidate rectangles, the smaller the interference degree and the higher the interference score.
[0206] Step 4: Combine the rectangular score, interference degree, and the number of rectangles to calculate the weighted score.
[0207] In one example, combining the number m of candidate rectangular boxes searched with the current seed rectangular box as a reference, perform weighted scoring on the rectangular box score, interference degree, and the number m of rectangular boxes obtained above. Among them, the weight of the rectangular box score is positively correlated with the score value, the weight of the interference degree score is negatively correlated with the score value, and the value of the number m of boxes is positively correlated with the score value.
[0208] Step S703: Select the group with the highest score among N groups of candidate rectangular boxes to obtain the rectangular box of each box.
[0209] Based on the above steps, N groups of candidate rectangular boxes with the seed rectangular box as a reference are obtained. For each group of candidate rectangular boxes, calculate the weighted score, and select the group with the largest weighted score value as the optimal rectangular box screening result of each box, which is used as the rectangular box of the box.
[0210] For each candidate rectangular box that has been calculated, respectively, using the seed rectangular box as a reference, and taking the rectangular score, the degree of interference, and the number of boxes of the candidate rectangular box as constraint conditions, a set of candidate rectangular boxes including the candidate rectangular box is obtained as the optimal box rectangular box, which can effectively improve the accuracy and stability of box positioning.
[0211] To further improve the accuracy of box positioning, in a possible implementation, as Figure 8 shown, before the step of unstacking the pallet according to the box rectangular box of each box, the method further includes:
[0212] Step S801, for each box rectangular box, determine the depth pixel region corresponding to the box rectangular box in the depth map, and expand the depth pixel region of the box rectangular box to obtain the depth pixel expansion region of the box rectangular box;
[0213] The depth pixel region with equal density within a certain width neighborhood of the depth image edge of the box rectangular box can be extracted as the extraction candidate points for depth edge correction;
[0214] Step S802, determine the visible light pixel region corresponding to the box rectangular box in the visible light image, and expand the visible light pixel region of the box rectangular box to obtain the visible light pixel expansion region of the box rectangular box;
[0215] The visible light pixel region with equal density within a certain width neighborhood of the visible light image edge of the box rectangular box can be extracted as the extraction candidate points for visible light edge correction;
[0216] Step S803, determine the intersection region of the depth pixel expansion region and the visible light pixel expansion region of the box rectangular box in the same coordinate system;
[0217] Step S804, taking the image gradient of the edge of the box rectangular box, the straightness of the gray edge points, and the distance from the edge of the top box point cloud region as constraint conditions, correct the box rectangular box in the intersection region of the box rectangular box.
[0218] Specifically, the correction process is to correct the four sides of the rectangular box respectively. The visible light image gradient and the intersection region of the depth point cloud edge can be combined, and the RANSAC idea is used to extract a straight line in the intersection region, and the correction constraint conditions are constructed to determine the refined straight line. Among them, the edge correction is carried out with the image gradient, the straightness of the gray edge points within the neighborhood of the line segment, and the distance between the corrected rectangular edge line segment and the real depth edge of the top box as constraint conditions.
[0219] In this embodiment, the image gradient on the corrected rectangular edge line segment should be as large as possible; the straightness (linear fitting error) of the grayscale image edge point closest to the corrected rectangular edge line segment along the normal direction should be as small as possible; the distance between the corrected rectangular edge line segment and the true depth edge of the top box body should be as small as possible.
[0220] The correction process corrects the four sides of the rectangular frame respectively. Taking one side of the rectangular frame as an example below, the edge correction process is introduced, and the correction steps for the other three sides are the same. It can be understood that the following process is only for example and does not limit the protection scope of this application.
[0221] In one example, the visible light image gradient and the intersection region of the depth point cloud edge can be combined. The RANSAC idea can be used to extract a straight line in the intersection region, and the correction constraint conditions are constructed to determine the refined straight line. Among them, the correction constraint target can be expressed as the highest constraint score. For example, the constraint score can be the weighted sum of the image gradient score of the edge, the straightness score of the grayscale edge point, and the distance score from the edge of the top box body point cloud region. Among them, the stronger the image gradient on the rectangular edge line segment, the higher the image gradient score of the edge; the better the straightness of the grayscale edge point, the higher the straightness score of the grayscale edge point; the closer the distance to the edge of the top box body point cloud region, the higher the distance score from the edge of the top box body point cloud region.
[0222] Multiple candidate straight lines can be randomly extracted according to a preset number or at equal intervals according to the pixel interval in the intersection region of the box rectangular frame, and the weighted scores of the multiple candidate straight lines are calculated, and the group with the highest weighted score is selected as the corrected straight line. In this embodiment, there is no specific limitation on the extraction of candidate straight lines, and the extraction method can be selected according to the actual situation.
[0223] Taking the score calculation of one candidate straight line as an example, a feasible step is provided as follows:
[0224] Step 1: Uniformly sample on the candidate straight line to obtain N points;
[0225] In one example, the RANSAC idea can be used to extract candidate points with equal density in a neighborhood with a certain width W in the intersection region, and the sampling density can be set to 2 to 5 pixels.
[0226] Step 2: Calculate the total gradient on the candidate straight line, that is, add the gradients of the N sampling points.
[0227] Step 3: Calculate the grayscale edge straightness of the candidate straight line;
[0228] In one example, it is possible to first search along the straight-line normal direction to obtain the grayscale edge points closest to each sampling point, perform least-squares linear fitting on these grayscale edge points, and add 1 to the calculated linear fitting error value as the straightness. Here, adding 1 is to prevent the fitting error from being 0.
[0229] Step 4: Calculate the distance from the candidate straight line to the depth edge of the top-level box corresponding to the candidate straight line, and then add 1 as the distance score. Here, adding 1 is to prevent the distance score from being 0.
[0230] Step 5: Combine the image gradient, the straightness of the grayscale edge straight line, and the distance from the candidate straight line to the depth edge of the top-level box corresponding to the candidate straight line to calculate the weighted score of the candidate straight line.
[0231] In one example, the total gradient, the straightness of the grayscale edge straight line, and the distance from the candidate straight line to the depth edge of the top-level box corresponding to the candidate straight line calculated for the candidate straight line are used for weighted scoring. Among them, the weight of the image gradient value is positively correlated with the score value, the straightness of the grayscale edge points, that is, the weight of the linear fitting error value, is negatively correlated with the score value, and the weight of the distance from the candidate straight line to the depth edge of the top-level box corresponding to the candidate straight line is negatively correlated with the score value.
[0232] Based on the above steps, each candidate straight line is calculated to obtain a weighted score, and the group with the largest weighted score value is selected and determined as the straight line after edge refinement as the correction result.
[0233] The depalletizing of the palletized boxes according to the rectangular frame of each box includes:
[0234] Depalletize the palletized boxes according to the corrected rectangular frame of each box.
[0235] In this embodiment, the edge correction can be performed in the pallet coordinate system or in the palletized box reference plane coordinate system. Through the edge correction, the accuracy and stability of the rectangular frame of the box can be effectively improved.
[0236] In one example, Figure 9 is a flowchart of the grasping pose calculation provided by the embodiment of the present application. As Figure 9 shown, the grasping pose calculation may include the following steps:
[0237] Step S901, based on the rectangular frame of the box, re-extract the point cloud on the surface of the box in the depth camera coordinate system, and perform point cloud plane fitting;
[0238] Specifically, the point cloud plane fitting can adopt methods such as global least-squares fitting and RANSAC fitting.
[0239] Step S902: Calculate the grasping points (X, Y, Z, Rx, Ry, Rz) based on the centroid and normal of the point cloud.
[0240] Step S903: Calculate the minimum bounding rectangle based on the contour of the point cloud.
[0241] In this embodiment, the minimum bounding rectangle refers to the maximum range of a two-dimensional rectangle represented by two-dimensional coordinates, that is, a rectangle defined by the maximum abscissa, minimum abscissa, maximum ordinate, and minimum ordinate among the vertices of the given two-dimensional rectangle.
[0242] Step S904: Sort all the grasping poses according to the scores and output them.
[0243] Specifically, the pose score can be calculated by weighted summation according to attributes such as height and point cloud area.
[0244] This solution is applicable to scenarios such as single-piece unpacking of single-category boxes, single-piece unpacking of multi-category boxes, and mixed unpacking of multi-category boxes, and has the advantages of good scene applicability and high deployment efficiency.
[0245] Based on the same inventive concept, a vision-guided palletizing and depalletizing device is provided corresponding to the vision-guided depalletizing method. Figure 10 FIG. 1 is a first structural schematic diagram of the vision-guided depalletizing device provided in the embodiment of the present application. The device includes:
[0246] An acquisition module 110, configured to acquire a depth map and a visible light image of a pallet of boxes, where the pallet of boxes includes at least one box.
[0247] An extraction module 120, configured to extract the point cloud region of the top-layer box of the pallet of boxes based on the depth map.
[0248] An obtaining module 130, configured to determine the region corresponding to the point cloud region of the top-layer box in the visible light image, and obtain a box mask region.
[0249] A segmentation module 140, configured to perform instance segmentation on the box mask region to obtain the respective segmentation regions of each box.
[0250] A depalletizing module 150, configured to depalletize the pallet of boxes according to the segmentation region of each box.
[0251] In the technical solution provided by the embodiment of the present application, by extracting the point cloud region of the top layer of the box stack, determining the corresponding region of the point cloud region of the top layer of the box in the visible light image, filtering out interfering background information, obtaining the box mask region, combining the visible light image features and the depth image features, performing box instance segmentation on the box mask region, and obtaining the respective segmentation regions of each box, the box stack is thus unstacked, thereby achieving the purpose of effectively improving the positioning accuracy of the target box during unstacking.
[0252] In a possible implementation manner, referring to Figure 11 , the extraction module 120 may include:
[0253] A conversion sub-module 121, configured to convert the depth map into first point cloud data in the depth image coordinate system;
[0254] A first mapping sub-module 122, configured to map the first point cloud data to the coordinate system of the box stack to obtain second point cloud data;
[0255] A segmentation sub-module 123, configured to segment the second point cloud data to obtain at least one point cloud classification, where the plane features of the point clouds in the same point cloud classification are the same;
[0256] A first selection sub-module 124, configured to select the point cloud classification with the highest height from the at least one point cloud classification to obtain the point cloud region of the top layer of the box stack.
[0257] The interfering background in the RGB image can be removed through the point cloud of the top layer of the box, and instance segmentation is performed on the RGB image that only retains the top layer of the box in the visible light image. During the process of training the model, it is not necessary to collect a large number of samples, the generalization of the instance segmentation model is improved, and the false detection ratio can be effectively reduced.
[0258] In a possible implementation manner, as Figure 12 shown, the unstacking module 150 may include:
[0259] A first extraction sub-module 151, configured to extract the image edge feature of the segmentation region of each box, or extract the image edge feature of the grayscale image of the segmentation region of each box;
[0260] A second mapping sub-module 152, configured to map the segmentation region of each box to the point cloud data converted from the depth map to obtain the point cloud mapping region of each box;
[0261] A second extraction sub-module 153, configured to extract the point cloud edge feature of the point cloud mapping region of each box;
[0262] The determination sub-module 154 is configured to determine the edge straight-line feature of each box according to the image edge feature and the point cloud edge feature of the box.
[0263] The splicing sub-module 155 is configured to perform rectangular splicing on the edge straight-line feature of each box to obtain at least one candidate rectangular frame of the box.
[0264] Specifically, the first calculation unit 1551 is configured to calculate the first conversion relationship between the visible light image coordinate system of the visible light image and the top box coordinate system of the top box point cloud region.
[0265] The conversion unit 1552 is configured to convert the segmentation region of each box into the top box coordinate system according to the first conversion relationship to obtain the size range of the box.
[0266] The splicing unit 1553 is configured to perform rectangular splicing on the edge straight-line feature of the box according to the size range of the box and the preset rectangular angle threshold to obtain at least one candidate rectangular frame of the box.
[0267] The second selection sub-module 156 is configured to select the box rectangular frame of each box from the candidate rectangular frames of the box.
[0268] Specifically, the second calculation unit 1561 is configured to calculate the projection point cloud occupancy ratio of each candidate rectangular frame of each box according to the point cloud mapping region of the box; calculate the image edge intensity and rectangularity of each candidate rectangular frame of the box respectively.
[0269] The third calculation unit 1562 is configured to calculate the score of each candidate rectangular frame according to the projection point cloud occupancy ratio, image edge intensity and rectangularity of the candidate rectangular frame.
[0270] The fourth calculation unit 1563 is configured to calculate the interference degree between adjacent candidate rectangular frames according to the positional relationship between adjacent candidate rectangular frames.
[0271] The first determination unit 1564 is configured to determine the box rectangular frame of each box according to the scores of the candidate rectangular frames and the interference degree between adjacent candidate rectangular frames.
[0272] The unstacking sub-module 157 is configured to unstack the box stack according to the box rectangular frame of each box.
[0273] By separately extracting the image edge features and the point cloud data features in the instance segmentation mask, combining the two to determine the edge pixel region, and then extracting the straight edge features, and combining the rough positioning of the segmentation region of the box body and the plane constraint of the box body, a candidate rectangular frame can be obtained, which can effectively improve the positioning accuracy of the grasping surface of the box body.
[0274] In a possible implementation, as Figure 13 shown, the device may further include:
[0275] A first expansion unit 1401, configured to, before performing the unstacking step on the pallet according to the box body rectangular frame of each box body, for each box body rectangular frame, determine the depth pixel region corresponding to the box body rectangular frame in the depth map, and expand the depth pixel region of the box body rectangular frame to obtain the depth pixel expansion region of the box body rectangular frame;
[0276] A second expansion unit 1402, configured to determine the visible light pixel region corresponding to the box body rectangular frame in the visible light image, and expand the visible light pixel region of the box body rectangular frame to obtain the visible light pixel expansion region of the box body rectangular frame;
[0277] A second determination unit 1403, configured to determine the intersection region of the depth pixel expansion region and the visible light pixel expansion region of the box body rectangular frame in the same coordinate system;
[0278] A correction unit 1404, configured to use the image gradient of the edge of the box body rectangular frame, the straightness of the gray edge points, and the distance from the edge of the top box body point cloud region as constraint conditions to correct the box body rectangular frame in the intersection region of the box body rectangular frame.
[0279] In this embodiment, the edge correction can be performed in the pallet coordinate system or in the box pallet reference plane coordinate system. Through the edge correction, the accuracy and stability of the box body rectangular frame can be effectively improved.
[0280] The embodiment of the present application also provides an electronic device, as Figure 14 shown, including a processor 151 and a memory 152;
[0281] The memory 152 stores instructions executable by the at least one processor 151;
[0282] The instructions are executed by the at least one processor 151, so that the at least one processor 151 can execute any one of the vision-guided unstacking methods in the present application.
[0283] In a possible implementation, as Figure 15 shown, the electronic device further includes a communication bus 154 and a communication interface 153.
[0284] The communication bus mentioned in the above-mentioned electronic device may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in illustration, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.
[0285] The communication interface is used for communication between the above-mentioned electronic device and other devices.
[0286] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. In a possible implementation manner, the memory may also be at least one storage device located far from the aforementioned processor.
[0287] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0288] In another embodiment provided by this application, a computer-readable storage medium is also provided. The computer-readable storage medium stores a computer program, and the computer program is executed by the processor to perform any one of the above-mentioned vision-guided palletizing and depalletizing methods.
[0289] An embodiment of this application also provides a vision-guided depalletizing system, including a depth camera, a depalletizing robotic arm, and a control device;
[0290] The depth camera is used to collect the depth map and visible light image of the pallet;
[0291] The depalletizing robotic arm is used to depalletize the pallet in response to the instruction of the control device;
[0292] The control device is used to implement the above-mentioned vision-guided palletizing method during operation.
[0293] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state disk (SSD)).
[0294] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0295] Each embodiment in this specification is described in a related manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0296] The foregoing are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A pallet - unstacking method based on visual guidance, characterized in that, the method includes: Obtain a depth map and a visible - light image of the pallet stack, where the pallet stack includes at least one box body; Based on the depth map, extract the point - cloud region of the top - layer box body of the pallet stack; Determine the corresponding region of the point - cloud region of the top - layer box body in the visible - light image to obtain the box - body mask region; Perform box - body instance segmentation on the box - body mask region to obtain the respective segmentation regions of each box body; For each box body, extract the image - edge features of the segmentation region of this box body, or extract the image - edge features of the grayscale image of the segmentation region of this box body; Map the segmentation region of each box body to the point - cloud data converted from the depth map to obtain the point - cloud mapping region of each box body; For each box body, extract the point - cloud edge features of the point - cloud mapping region of this box body; For each box body, determine the edge - straight - line features of this box body according to the image - edge features and point - cloud edge features of this box body; For each box body, perform rectangular splicing on the edge - straight - line features of this box body to obtain at least one candidate rectangular frame of this box body; For each box body, calculate the projected - point - cloud occupancy ratio of each candidate rectangular frame of this box body according to the point - cloud mapping region of this box body; calculate the image - edge intensity and rectangularity of each candidate rectangular frame of this box body respectively; For each candidate rectangular frame, calculate the score of this candidate rectangular frame according to the projected - point - cloud occupancy ratio, image - edge intensity and rectangularity of this candidate rectangular frame; According to the positional relationship between adjacent candidate rectangular frames, calculate the interference degree between each adjacent candidate rectangular frame; According to the scores of each candidate rectangular frame and the interference degree between adjacent candidate rectangular frames, determine the box - body rectangular frame of each box body; Unstack the pallet stack according to the box - body rectangular frame of each box body.
2. The method according to claim 1, characterized in that, the extracting the point - cloud region of the top - layer box body of the pallet stack based on the depth map includes: In the depth - image coordinate system, convert the depth map into first - point - cloud data; Map the first - point - cloud data to the coordinate system of the pallet stack to obtain second - point - cloud data; Segment the second - point - cloud data to obtain at least one point - cloud classification, where the plane features of the point clouds in the same point - cloud classification are the same; Among the at least one point - cloud classification, select the point - cloud classification with the highest height to obtain the point - cloud region of the top - layer box body of the pallet stack.
3. The method according to claim 1, characterized in that, the performing box - body instance segmentation on the box - body mask region to obtain the respective segmentation regions of each box body includes: Use a pre - trained deep - learning model to perform box - body instance segmentation on the box - body mask region to obtain the respective segmentation regions of each box body, where the deep - learning model is trained by sample box - body mask images, and the sample box - body mask images are images of the top - layer box bodies in the sample pallet stack.
4. The method according to claim 1, characterized in that, the determining the edge - straight - line features of each box body according to the image - edge features and point - cloud edge features of this box body includes: For each box, according to the image edge features and point cloud edge features of the box, determine the edge pixel region of the box in the depth map; Extract the straight line features of the edge pixel region of the box to obtain the edge straight line features of the box.
5. The method according to claim 1, wherein, for each box, performing rectangular splicing on the edge straight line features of the box to obtain at least one candidate rectangular frame for the box, including: Calculate the first conversion relationship between the visible light image coordinate system of the visible light image and the top box coordinate system of the top box point cloud region of the box; For each box, according to the first conversion relationship, convert the segmentation region of the box into the top box coordinate system to obtain the size range of the box; According to the size range of the box and the preset rectangular angle threshold, perform rectangular splicing on the edge straight line features of the box to obtain at least one candidate rectangular frame for the box.
6. The method according to claim 1, wherein, determining the box rectangular frame of each box according to the candidate scores of the candidate rectangular frames and the interference degree between adjacent candidate rectangular frames, including: Among the candidate rectangular frames, select the top N candidate rectangular frames with the highest scores as the seed rectangular frames; where N is a preset integer; For each candidate rectangular frame among the N seed rectangular frames, taking the candidate rectangular frame as a reference, and using the score, interference degree, and number of boxes as constraint conditions, determine a group of candidate rectangular frames including the candidate rectangular frame; Among the N groups of candidate rectangular frames, select the group with the highest score to obtain the box rectangular frame of each box.
7. The method according to claim 1, wherein, before the step of depalletizing the pallet according to the box rectangular frame of each box, the method further includes: For each box rectangular frame, determine the depth pixel region corresponding to the box rectangular frame in the depth map, and expand the depth pixel region of the box rectangular frame to obtain the depth pixel expansion region of the box rectangular frame; Determine the visible light pixel region corresponding to the box rectangular frame in the visible light image, and expand the visible light pixel region of the box rectangular frame to obtain the visible light pixel expansion region of the box rectangular frame; Determine the intersection region of the depth pixel expansion region and the visible light pixel expansion region of the box rectangular frame in the same coordinate system; Using the image gradient of the edge of the box rectangular frame, the straightness of the gray edge points, and the distance from the edge of the top box point cloud region as constraint conditions, correct the box rectangular frame in the intersection region of the box rectangular frame; The step of depalletizing the pallet according to the box rectangular frame of each box includes: Depalletize the pallet according to the corrected box rectangular frame of each box.
8. A depalletizing device based on vision guidance, wherein, the device includes: An acquisition module for acquiring a depth map and a visible light image of a pallet, where the pallet includes at least one box; An extraction module for extracting the top box point cloud region of the pallet based on the depth map; An acquisition module, configured to determine a region corresponding to the top-layer box point cloud region in the visible light image, and acquire a box mask region; A segmentation module, configured to perform box instance segmentation on the box mask region to obtain a segmentation region for each box; A depalletizing module, comprising: A first extraction sub-module, configured to, for each box, extract the image edge feature of the segmentation region of the box, or extract the image edge feature of the grayscale image of the segmentation region of the box; A second mapping sub-module, configured to map the segmentation region of each box to the point cloud data converted from the depth map to obtain a point cloud mapping region for each box; A second extraction sub-module, configured to, for each box, extract the point cloud edge feature of the point cloud mapping region of the box; A determination sub-module, configured to, for each box, determine the edge straight line feature of the box according to the image edge feature and the point cloud edge feature of the box; A splicing sub-module, configured to, for each box, perform rectangular splicing on the edge straight line feature of the box to obtain at least one candidate rectangular frame for the box; A second selection sub-module, configured to, for each box, select the box rectangular frame of the box from the candidate rectangular frames of the box; A depalletizing sub-module, configured to depalletize the pallet according to the box rectangular frame of each box; Wherein, the second selection sub-module is specifically configured to: for each box, calculate the projected point cloud occupancy ratio of each candidate rectangular frame of the box according to the point cloud mapping region of the box; calculate the image edge intensity and rectangularity of each candidate rectangular frame of the box respectively; for each candidate rectangular frame, calculate the score of the candidate rectangular frame according to the projected point cloud occupancy ratio, image edge intensity and rectangularity of the candidate rectangular frame; calculate the interference degree between adjacent candidate rectangular frames according to the positional relationship between adjacent candidate rectangular frames; determine the box rectangular frame of each box according to the scores of the candidate rectangular frames and the interference degree between adjacent candidate rectangular frames.
9. The apparatus according to claim 8, wherein, the extraction module comprises: A conversion sub-module, configured to convert the depth map into first point cloud data in the depth image coordinate system; A first mapping sub-module, configured to map the first point cloud data to the coordinate system of the pallet to obtain second point cloud data; A segmentation sub-module, configured to segment the second point cloud data to obtain at least one point cloud classification, wherein the plane features of the point clouds in the same point cloud classification are the same; A first selection sub-module, configured to select the point cloud classification with the highest height from the at least one point cloud classification to obtain the top-layer box point cloud region of the pallet; The segmentation module is specifically configured to: Use a pre-trained deep learning model to perform box instance segmentation on the box mask region to obtain a segmentation region for each box, wherein the deep learning model is trained by a sample box mask image, and the sample box mask image is an image of the top-layer box in a sample pallet; The determination sub-module is specifically configured to: For each box, according to the image edge features and point cloud edge features of the box, determine the edge pixel region of the box in the depth map; Extract the straight line features from the edge pixel region of the box to obtain the edge straight line features of the box; The splicing sub-module includes: A first calculation unit for calculating a first conversion relationship between the visible light image coordinate system of the visible light image and the top box coordinate system of the top box point cloud region; A conversion unit for, for each box, according to the first conversion relationship, convert the segmentation region of the box into the top box coordinate system to obtain the size range of the box; A splicing unit for, according to the size range of the box and a preset rectangular angle threshold, perform rectangular splicing on the edge straight line features of the box to obtain at least one candidate rectangular frame of the box; The second selection sub-module is specifically used for: Among the candidate rectangular frames, select the top N candidate rectangular frames with the highest scores as the seed rectangular frames, where N is a preset integer; For each candidate rectangular frame among the N seed rectangular frames, take the candidate rectangular frame as a reference, and use the score, interference degree, and number of boxes as constraint conditions to determine a group of candidate rectangular frames including the candidate rectangular frame; Among the N groups of candidate rectangular frames, select the group with the highest score to obtain the box rectangular frame of each box; The device further includes: A first expansion unit for, before performing the unstacking step on the box stack according to the box rectangular frame of each box, for each box rectangular frame, determine the depth pixel region corresponding to the box rectangular frame in the depth map, and expand the depth pixel region of the box rectangular frame to obtain the depth pixel expansion region of the box rectangular frame; A second expansion unit for determining the visible light pixel region corresponding to the box rectangular frame in the visible light image, and expanding the visible light pixel region of the box rectangular frame to obtain the visible light pixel expansion region of the box rectangular frame; A second determination unit for determining the intersection region of the depth pixel expansion region and the visible light pixel expansion region of the box rectangular frame in the same coordinate system; A correction unit for, using the image gradient of the edge of the box rectangular frame, the straightness of the gray edge points, and the distance from the edge of the top box point cloud region as constraint conditions, correct the box rectangular frame in the intersection region of the box rectangular frame; The unstacking sub-module is specifically used for: Unstack the box stack according to the corrected box rectangular frame of each box.
10. An electronic device, Characterized in that, It includes a processor and a memory; The memory stores instructions executable by the at least one processor; The instructions are executed by the at least one processor, so that the at least one processor can execute the method according to any one of claims 1-7.
11. A computer-readable storage medium, Characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method steps according to any one of claims 1-7.
12. A vision-guided unstacking system, Characterized in that, It includes: Depth camera, depalletizing robotic arm and control device; The depth camera is used to collect the depth map and visible light image of the pallet of boxes; The depalletizing robotic arm is used to depalletize the pallet of boxes in response to the instruction of the control device; The control device is used to implement the method steps described in any one of claims 1-7 during operation.
Citation Information
Patent Citations
Warehouse box body identification and positioning method based on contour features
CN111507390A
Method and device for recognizing stacked box bodies in stack type and robot
CN112907668A
Method and device for determining space grabbing point of robot
CN114170442A