Object grasping method, system, device and storage medium based on robotic arm
By combining image and point cloud data, the positioning accuracy and grasping accuracy of the robot arm in object grabbing are improved, and the problem of low positioning accuracy in the prior art is solved.
Patent Information
- Application Number
- CN202210511704.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-05-10
AI Technical Summary
In the prior art, the robotic arm is not very accurate in grasping objects, resulting in a low gripping accuracy.
By receiving the target image and target point cloud data sent by the acquisition device, the coordinates of the target object are determined using the target image and the pre-stored target segmentation model, the point cloud data of the object is determined in combination with the point cloud data, and finally the target position is determined based on the point cloud data, and sent to the robot arm for grabbing.
The object positioning accuracy and grabbing accuracy are improved. Compared with the direct use of point cloud data, the image data has better continuity, resulting in higher accuracy of point cloud data.
Smart Images

Figure CN115213896B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robot technology, and in particular to a method, system, device and storage medium for grasping an object based on a robot arm. Background Art
[0002] Traditional industrial production processes generally rely on manual labor to grasp, transport, and install workpieces, which has low production efficiency, high work risks, high labor costs, high work intensity, and high operator turnover rates, such as object sorting in the logistics industry and assembly of parts on industrial production lines. With the improvement of industrial automation and intelligence, there is a huge demand for the intelligent grasping of objects by robotic arms.
[0003] In the prior art, the point cloud of the target object is usually used to determine the position of the target object, and then the robot arm is controlled to grasp the object according to the position of the target object. However, the positioning accuracy of the target object determined by this method is not high, resulting in a low accuracy rate of the robot arm grasping the object. Summary of the invention
[0004] The present application provides a method, system, device and storage medium for grasping an object based on a robotic arm, which solves the problem of low accuracy in grasping an object by a robotic arm due to low object positioning accuracy.
[0005] In order to achieve the above objectives, this application adopts the following technical solutions:
[0006] In a first aspect of an embodiment of the present application, a method for grasping an object based on a robot arm is provided, the method comprising: receiving a target image and target point cloud data of a target area sent by a collection device, the target area including a target object to be grasped;
[0007] Determine the target region coordinates according to the target image, target identification information corresponding to the target object and a pre-stored target segmentation model, where the target region coordinates include the coordinates of each pixel point in each pixel point of the target object;
[0008] Determine the object point cloud data of the target object according to the target point cloud data and the area coordinates;
[0009] Determine the target pose of the target object according to the object point cloud data;
[0010] Send the target pose to the robot arm, which is used by the robot arm to grasp the target object.
[0011] In one embodiment, the target region coordinates are determined according to the target image, target identification information corresponding to the target object, and a pre-stored target segmentation model, including:
[0012] Input the target image into the target segmentation model to obtain label information corresponding to each pixel area in multiple pixel areas of the target image, wherein the label information includes identification information and area coordinates corresponding to the pixel area, and the area coordinates include multiple pixel points;
[0013] Determine the target pixel area corresponding to the target identification information according to the correspondence between the pixel area and the identification information;
[0014] The region coordinates corresponding to the target pixel region are determined as the target region coordinates.
[0015] In one embodiment, determining object point cloud data of a target object according to the target point cloud data and the region coordinates includes:
[0016] Acquire the mapping relationship between the coordinates of each pixel point included in the target image and the coordinates of each target point cloud data;
[0017] The object point cloud data corresponding to the area coordinates is determined according to the mapping relationship.
[0018] In one implementation, before determining the region coordinates according to the target image, target identification information corresponding to the target object, and a pre-stored target segmentation model, the method further includes:
[0019] Obtain sample images of multiple objects;
[0020] Determine a sample pixel region of each sample image, and determine label information corresponding to each sample pixel region of each sample image, wherein the label information includes region coordinates and identification information corresponding to the sample pixel region;
[0021] The preset positioning segmentation model is trained using sample images of multiple objects and label information corresponding to each pixel area in each sample image to obtain a target segmentation model.
[0022] In one implementation, determining a sample pixel area of each sample image includes:
[0023] Perform edge segmentation processing on each sample image to obtain the target contour of the object included in each sample image;
[0024] According to the target contour corresponding to each sample image, each sample image is divided into regions to obtain a first sample pixel region and a second sample pixel region of each sample image;
[0025] The pixels within the target contour form a first sample pixel region, and the second sample pixel region is a blank region in the sample image.
[0026] In one implementation, determining label information corresponding to each sample pixel region of each sample image includes:
[0027] When the target contour corresponding to the first sample pixel region successfully matches the pre-stored pixel contour, the identification information corresponding to the pixel contour is used as the identification information corresponding to the first sample pixel region;
[0028] Obtaining preset identification information corresponding to the second sample pixel area;
[0029] The region coordinates corresponding to the pixel region are determined according to the coordinates of each pixel point in the pixel region.
[0030] In one implementation, the object point cloud data is point cloud data of the target object in a target coordinate system, and the target coordinate system is a coordinate system used by the acquisition device;
[0031] Determine the target pose of the target object based on the object point cloud data, including:
[0032] Obtaining point cloud template data corresponding to the target identification information, where the point cloud template data is point cloud data of the target object in a preset coordinate system;
[0033] Determine the target pose based on the object point cloud data and point cloud template data.
[0034] In one implementation, determining a target pose according to object point cloud data and point cloud template data includes:
[0035] The initial position and posture of the target object is obtained according to the object point cloud data, the point cloud template data, the preset point feature histogram and the preset feature matching algorithm based on sampling matching consistency. The initial position and posture is the position and posture of the target object based on the acquisition device;
[0036] Iteratively calculating the initial pose and the object point cloud to obtain the optimized pose of the target object;
[0037] Obtaining the target coordinate transformation relationship between the acquisition device and the base of the robotic arm;
[0038] The target pose is determined according to the optimized pose and the target coordinate transformation relationship, and the target pose is the pose of the target object based on the base of the robot arm.
[0039] In one implementation, obtaining a target coordinate transformation relationship between a collection device and a base of a robotic arm includes:
[0040] Obtaining a first coordinate transformation relationship and a corresponding second coordinate transformation relationship of the object in different positions and postures, the first coordinate transformation relationship being the coordinate transformation relationship between the acquisition device and the gripper of the robotic arm, and the second coordinate transformation relationship being the coordinate transformation relationship between the base and the gripper;
[0041] According to each first coordinate transformation relationship and the corresponding second coordinate transformation relationship, a target coordinate transformation relationship is obtained.
[0042] In one implementation, obtaining a target coordinate transformation relationship according to each first coordinate transformation relationship and the corresponding second coordinate transformation relationship includes:
[0043] According to each first coordinate transformation relationship and the corresponding second coordinate transformation relationship, a third coordinate transformation relationship corresponding to each first coordinate transformation relationship is obtained;
[0044] The least squares fitting calculation is performed on multiple third coordinate transformation relations to obtain the target coordinate transformation relation.
[0045] In one implementation, before obtaining the point cloud template corresponding to the target object, the method further includes:
[0046] Select at least two point cloud data from the point cloud data of the target object to establish a preset coordinate system;
[0047] Determine the point cloud template data according to the preset coordinate system.
[0048] In a second aspect of the embodiment of the present application, there is also provided an object grasping system based on a robotic arm, the system comprising: a collection device, an electronic device and a robotic arm;
[0049] A collection device, used to collect a target image and target point cloud data of a target area, the target area including a target object to be captured, and send the collected target image and target point cloud data to an electronic device;
[0050] An electronic device, used for receiving a target image and target point cloud data sent by a collection device, wherein the image content of the target image includes a target object to be captured, and the target point cloud data includes object point cloud data of the target object;
[0051] The electronic device is further used to perform image processing on the target image using a pre-stored target segmentation model to obtain region coordinates, where the region coordinates include each pixel of the target object and the coordinates of each pixel;
[0052] The electronic device is further used to determine the object point cloud data based on the target point cloud data and the area coordinates;
[0053] The electronic device is further used to determine a target pose of the target object based on the object point cloud data and send the target pose to the robotic arm;
[0054] A robotic arm is used to grasp the target object according to the target posture.
[0055] According to a third aspect of an embodiment of the present application, there is also provided an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the robot-based object grasping method according to the first aspect of an embodiment of the present application is implemented.
[0056] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the robot-based object grasping method of the first aspect of the embodiments of the present application is implemented.
[0057] The beneficial effects brought by the technical solution provided by the embodiment of the present application include at least:
[0058] The object grasping method based on a robotic arm provided in the present application receives the target image and target point cloud data of the area including the object to be grasped sent by the acquisition device, and determines the coordinates of each pixel point of the target object according to the target identification information corresponding to the target object and the preset target segmentation model, and then determines the object point cloud data of the target object according to the target point cloud data and the area coordinates, and finally determines the target posture of the target object according to the object point cloud data, and sends the target to the robotic arm, so that the robotic arm grasps the target object according to the target posture. The object grasping method based on a robotic arm provided in the embodiment of the present application uses the area coordinates determined by the image, and then obtains the point cloud data of the target object according to the area coordinates and the point cloud data. Since the image data has better continuity than the discrete point cloud data, the point cloud of the target object determined by the method of the present application is more accurate than the point cloud of the target object obtained by directly using the point cloud data in the prior art, so the positioning accuracy of the object can be improved.
[0059] Furthermore, since the present application segments the image to obtain the regional coordinates of the target object, compared with the prior art which directly segments the point cloud data, the data processing efficiency is higher, thereby making the object positioning efficiency higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 A schematic diagram of the internal structure of a computer device provided in an embodiment of the present application;
[0061] Figure 2 A flowchart of a method for grasping an object based on a robotic arm provided in an embodiment of the present application;
[0062] Figure 3 A schematic diagram of an object grasping principle based on a robotic arm provided in an embodiment of the present application;
[0063] Figure 4 A structural diagram of an object grasping system based on a robotic arm provided in an embodiment of the present application. DETAILED DESCRIPTION
[0064] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0065] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, unless otherwise specified, "plurality" means two or more.
[0066] Additionally, the use of “based on” or “according to” is meant to be open and inclusive, as a process, step, calculation, or other action “based on” or “according to” one or more conditions or values may, in practice, be based on additional conditions or beyond values.
[0067] Traditional industrial production processes generally rely on manual labor to grasp, transport, and install workpieces, which has low production efficiency, high work risks, high labor costs, high work intensity, and high operator turnover rates, such as object sorting in the logistics industry and assembly of parts on industrial production lines. With the improvement of industrial automation and intelligence, there is a huge application demand for intelligent grasping of objects by robotic arms, but current industrial robots are not flexible enough and can only complete single grasping and installation according to tutorials, and cannot make corresponding judgments based on different postures of objects. In the actual production process, a large number of robots are often required to work together, occupying a large amount of space.
[0068] In industry, most robot arm grasping operations use traditional teaching methods. However, for a new operating object or a new operating environment, the robot arm needs to be manually taught again. In addition, the teaching grasping method can only grasp a single object and cannot adapt to the different postures of objects in complex scenes. At the same time, as the number of sensors increases, the cost also increases. With the development and application of machine vision, more and more vision-based intelligent robot arm grasping posture calculation methods have been proposed. These methods can be roughly divided into two categories. The first type of method is based on machine learning, and the second type of method is based on template matching.
[0069] Computational methods based on machine learning process features in visual images in a learning manner to estimate the grasping posture. This type of method relies on the surface texture information of the grasped object and has better grasping posture calculation results for objects with rich texture information. However, this method is obviously not ideal when encountering grasping objects with surface lacking texture information. The template matching-based method matches the contour information of the grasped object with the template contour in the template library, and estimates the grasping posture of the grasped object based on the grasping posture of the best matching template. This type of method is no longer based on the texture information of the object surface, but only the contour of the object. Therefore, this type of method can improve the grasping of objects with missing texture.
[0070] In the prior art, the point cloud of the target object is usually used to determine the position of the target object, and then the robotic arm is controlled to grasp the object according to the position of the target object. However, the positioning accuracy of the target object determined by this method is not high, resulting in a low accuracy rate of the robotic arm grasping the object. In addition, in the process of posture determination, the traditional machine vision-based robotic arm grasping method often only uses two-dimensional information and ignores the three-dimensional structure information. The two-dimensional target detection method cannot determine the three-dimensional posture of the target. Therefore, it is difficult to plan the best grasping method for randomly placed targets according to their different postures.
[0071] In order to solve the above problems, the embodiment of the present application provides an object grasping method based on a robotic arm, which receives a target image and target point cloud data of the area including the object to be grasped sent by an acquisition device, and determines the coordinates of each pixel of the target object according to the target identification information corresponding to the target object and the preset target segmentation model, and then determines the object point cloud data of the target object according to the target point cloud data and the area coordinates, and finally determines the target posture of the target object according to the object point cloud data, and sends the target to the robotic arm, so that the robotic arm grasps the target object according to the target posture. The object grasping method based on a robotic arm provided in the embodiment of the present application is to use the area coordinates determined by the image, and then obtain the point cloud data of the target object according to the area coordinates and the point cloud data. Since the image data has better continuity than the discrete point cloud data, the point cloud of the target object determined by the method of the present application is more accurate than the point cloud of the target object obtained by directly using the point cloud data in the prior art, so the positioning accuracy of the object can be improved.
[0072] Furthermore, since the present application segments the image to obtain the regional coordinates of the target object, compared with the prior art which directly segments the point cloud data, the data processing efficiency is higher, thereby making the object positioning efficiency higher.
[0073] The executor of the robot-based object grasping method provided in the embodiment of the present application may be an electronic device, which may be a computing device, a terminal device or a server, wherein the terminal device may be various personal computers, laptops, smart phones, tablet computers and portable wearable devices, etc., and the present application does not make any specific limitation.
[0074] Optionally, the electronic device may also be a processor or a processing chip. When the electronic device is a processor or a processing chip, the electronic device may be integrated into the robotic arm.
[0075] Figure 1 The internal structure diagram of a computer device provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the computer device includes a processor and a memory connected via a system bus. The processor is used to provide computing and control capabilities. The memory may include a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The computer program can be executed by the processor to implement the steps of a method for determining gas diffusion layer parameters provided in each of the above embodiments. The internal memory provides a cache operating environment for the operating system and the computer program in the non-volatile storage medium.
[0076] Those skilled in the art will understand that Figure 1 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0077] Based on the above execution subject, the embodiment of the present application provides an object grasping method based on a robotic arm. Figure 2 As shown, the method comprises the following steps:
[0078] Step 201: Receive a target image and target point cloud data of a target area sent by a collection device.
[0079] The target area includes the target object to be captured. The target object is the object to be captured, and the target area is the area photographed or captured by the acquisition device when capturing the target object.
[0080] It should be noted that the acquisition device may be one device or an integration of multiple devices, and the acquisition device may realize the acquisition of the target area image and the acquisition of the target area point cloud data.
[0081] Step 202: Determine the target region coordinates according to the target image, the target identification information corresponding to the target object and the pre-stored target segmentation model. The target region coordinates include the coordinates of each pixel point of each pixel point of the target object.
[0082] The preset target segmentation model is a model trained using sample images. The target object in the target image can be extracted by the trained segmentation model to obtain the coordinates of each pixel in each pixel of the target object.
[0083] Step 203: Determine the object point cloud data of the target object according to the target point cloud data and the target area coordinates.
[0084] After obtaining each pixel point of the target object and the coordinates of each pixel point, the object point cloud data corresponding to the target area coordinates can be obtained according to the target point cloud data and the target area coordinates.
[0085] Step 204: Determine the target pose of the target object according to the object point cloud data.
[0086] Step 205: Send the target posture to the robotic arm, and the target posture is used for the robotic arm to grasp the target object.
[0087] The object grasping method based on a robotic arm provided in the present application receives the target image and target point cloud data of the area including the object to be grasped sent by the acquisition device, and determines the coordinates of each pixel point of the target object according to the target identification information corresponding to the target object and the preset target segmentation model, and then determines the object point cloud data of the target object according to the target point cloud data and the area coordinates, and finally determines the target posture of the target object according to the object point cloud data, and sends the target to the robotic arm, so that the robotic arm grasps the target object according to the target posture. The object grasping method based on a robotic arm provided in the embodiment of the present application uses the area coordinates determined by the image, and then obtains the point cloud data of the target object according to the area coordinates and the point cloud data. Since the image data has better continuity than the discrete point cloud data, the point cloud of the target object determined by the method of the present application is more accurate than the point cloud of the target object obtained by directly using the point cloud data in the prior art, so the positioning accuracy of the object can be improved.
[0088] Furthermore, since the present application segments the image to obtain the regional coordinates of the target object, compared with the prior art which directly segments the point cloud data, the data processing efficiency is higher, thereby making the object positioning efficiency higher.
[0089] Optionally, the specific implementation process of the above step 202 can be: input the target image into the target segmentation model to obtain the label information corresponding to each pixel area in the multiple pixel areas of the target image, the label information includes the identification information and area coordinates corresponding to the pixel area, the area coordinates include multiple pixel points, and then determine the target pixel area corresponding to the target identification information based on the correspondence between the pixel area and the identification information; finally, determine the area coordinates corresponding to the target pixel area as the target area coordinates.
[0090] The region coordinates are the multiple pixels constituting the region and the coordinates of each pixel. The multiple pixel regions obtained are the pixel regions of each object obtained by the segmentation model for segmenting the image, and the identification information corresponding to each pixel region, which can be the name of the object composed of the pixels in each region.
[0091] For example, if a picture includes parts, people and other blank areas, the picture is input into a trained picture segmentation model, and the model can output the segmentation marks as parts, people and blank areas.
[0092] It should be noted that the target segmentation model is a model trained using sample images. Therefore, before inputting the target image into the target segmentation model to obtain label information corresponding to each pixel area in multiple pixel areas of the target image, the segmentation model needs to be trained. The specific training method can be:
[0093] Sample images of multiple objects are obtained, and a sample pixel area of each sample image is determined, as well as label information corresponding to each sample pixel area of each sample image, wherein the label information includes area coordinates and identification information corresponding to the sample pixel area. Finally, the preset positioning segmentation model is trained using the sample images of multiple objects and the label information corresponding to each pixel area in each sample image to obtain a target segmentation model.
[0094] Among them, the segmentation model can be a semantic segmentation model based on deep learning. The model can adopt a LinkNet network structure and adopt a deep model fine-tuning training method for training during the training process. This training method can reduce the time and resource consumption of repeated training due to the addition of new object categories.
[0095] Among them, determining the sample pixel area of each sample image and determining the label information corresponding to each sample pixel area of each sample image is to determine the label information of each sample image, and then use the label information and the sample image to train the preset segmentation model.
[0096] Optionally, the process of determining the sample pixel area of each sample image may be:
[0097] Perform edge segmentation processing on each sample image to obtain the target contour of the object included in each sample image; perform region division on each sample image according to the target contour corresponding to each sample image to obtain the first sample pixel region and the second sample pixel region of each sample image; wherein the pixel points within the target contour form the first sample pixel region, and the second sample pixel region is the blank region in the sample image.
[0098] Specifically, the specific process of determining the label information corresponding to each sample pixel area of each sample image in the above process can be: when the target contour corresponding to the first sample pixel area successfully matches the pre-stored pixel contour, the identification information corresponding to the pixel contour is used as the identification information corresponding to the first sample pixel area; the preset identification information corresponding to the second sample pixel area is obtained; and the area coordinates corresponding to the pixel area are determined according to the coordinates of each pixel point in the pixel area.
[0099] It should be noted that the above process of determining the sample pixel area of each sample image is actually the process of how to generate the label of the sample image. When using deep learning for automatic positioning and detection, a large number of labeled samples are usually required. In the prior art, it is usually necessary to manually label the pictures using labeling tools. Manual labeling is time-consuming and laborious, and requires a lot of manpower and time costs. This application uses image processing algorithms to extract the area and contour of the target, and then outputs the label information of the area coordinates and category. This can improve the efficiency of label generation.
[0100] In one embodiment, the specific implementation process of the above step 203 may be: obtaining a mapping relationship between the coordinates of each pixel point included in the target image and the coordinates of each target point cloud data, and determining the object point cloud data corresponding to the region coordinates according to the mapping relationship.
[0101] It should be noted that there is a preset mapping relationship between the pixel points of the target image captured by the acquisition device and the target point cloud data. Therefore, after obtaining the target area coordinates of the target object, the point cloud data of the target object corresponding to the pixel points in the target area coordinates can be obtained according to the preset mapping relationship.
[0102] In one embodiment, the object point cloud data is point cloud data of the target object in a target coordinate system, and the target coordinate system is a coordinate system used by the acquisition device;
[0103] The specific implementation process of the above step 204 may be: obtaining the point cloud template data corresponding to the target identification information, and determining the target posture according to the object point cloud data and the point cloud template data.
[0104] The point cloud template data is the point cloud data of the target object in a preset coordinate system, and the target pose is the pose based on the base of the robotic arm.
[0105] Since the point cloud template data is the point cloud data of the target object in a preset coordinate system, before obtaining the point cloud template, it is necessary to obtain the point cloud template data corresponding to each object in advance.
[0106] Specifically, a preset coordinate system may be established by selecting at least two point cloud data from the point cloud data of the target object, and the point cloud template data may be determined according to the preset coordinate system.
[0107] In the actual implementation process, according to the two points selected on the point cloud data of the object acquired in advance, the origin of the coordinate system and the point cloud normal vector n at this point are determined according to the first point, and it is used as the Z axis of the grabbing coordinate system, and then the tangent plane is obtained according to the normal vector, recorded as XOY, and then the vector composed of the projection point of the second point on the tangent plane and the origin is used as the X axis, and the normal vector of the X axis is obtained on the formed XOY plane and recorded as the Y axis. Then the equation of the XOY plane is: A*(x-x0)+B*(y-y0)+C*(z–z0)=0, where the normal vector n=(A, B, C), and the two vectors are perpendicular: X*n=0, Y*n=0, so that the preset coordinate system can be established through two points.
[0108] It should be noted that the object point cloud data obtained in step 203 is the object point cloud data based on the acquisition device. Therefore, it is necessary to first determine the position and posture of the object under the acquisition device based on the object point cloud data and the point cloud template data. In the actual process of object grasping by the robotic arm, it is necessary to convert the object point cloud data based on the acquisition device into point cloud data based on the robotic arm base, so that the target position and posture can be obtained.
[0109] Optionally, the specific process of determining the target pose based on the object point cloud data and the point cloud template data may be:
[0110] The initial pose of the target object is obtained according to the object point cloud data, the point cloud template data, the preset point feature histogram and the preset feature matching algorithm based on sampling matching consistency. The initial pose is the pose of the target object based on the acquisition device. The initial pose and the object point cloud are iteratively calculated to optimize the target object, and the target coordinate transformation relationship between the acquisition device and the base of the robotic arm is obtained. The target pose is determined according to the optimized pose and the target coordinate transformation relationship. The target pose is the pose of the target object based on the base of the robotic arm.
[0111] Among them, the initial pose and the optimized pose are both the poses of the target object based on the acquisition device, and the target pose is the pose of the target object based on the base of the robotic arm.
[0112] Specifically, the initial pose and the object point cloud are iteratively calculated to obtain the optimization of the target object. The specific implementation process can be: the point cloud data of the initial pose and the point cloud data of the target object are subjected to iterative error optimization calculations based on the nearest neighbor points. When the error reaches the set standard, the optimal pose is output and obtained.
[0113] Optionally, an iterative nearest neighbor algorithm may be used to iteratively calculate the initial pose and the object point cloud.
[0114] In the actual implementation process, the preset point feature histogram and the preset feature matching algorithm based on sampling matching consistency are used to match the object point cloud data and the point cloud template data to obtain the initial pose of the target object, and then the initial pose is optimized to obtain the optimized pose of the target object. Finally, according to the target coordinate transformation relationship between the acquisition device and the base of the robotic arm, the optimized pose is converted into the target pose, and the target pose is sent to the robotic arm to grasp the target object.
[0115] Optionally, the specific process of obtaining the target coordinate transformation relationship between the acquisition device and the base of the robotic arm can be: obtaining the first coordinate transformation relationship and the corresponding second coordinate transformation relationship of the object in each different posture, the first coordinate transformation relationship is the coordinate transformation relationship between the acquisition device and the gripper of the robotic arm, and the second coordinate transformation relationship is the coordinate transformation relationship between the base and the gripper; according to each first coordinate transformation relationship and the corresponding second coordinate transformation relationship, the target coordinate transformation relationship is obtained.
[0116] Specifically, a target coordinate transformation relationship is obtained according to each first coordinate transformation relationship and the corresponding second coordinate transformation relationship, including: according to each first coordinate transformation relationship and the corresponding second coordinate transformation relationship, a third coordinate transformation relationship corresponding to each first coordinate transformation relationship is obtained; and a least squares fitting calculation is performed on multiple third coordinate transformation relationships to obtain a target coordinate transformation relationship.
[0117] Furthermore, the above-mentioned first coordinate transformation relationship and second coordinate transformation relationship can be obtained through actual calibration. The specific calibration process can be: fix the calibration plate on the robotic arm, rotate the robotic arm to place the calibration plate under the left eye of the camera, rotate the robotic arm to make the calibration plate change different postures under the camera, and take pictures with the camera to record the posture of the calibration plate relative to the camera under different postures of the robotic arm, and at the same time record the posture information of the gripper relative to the base. In this way, the postures of multiple sets of calibration plates relative to the camera and the postures of the gripper relative to the base under different postures of the robotic arm can be obtained, and the first coordinate transformation relationship can be obtained by constructing a group of equations based on the spatial relationship.
[0118] like Figure 3As shown, it is a schematic diagram of the object grasping process based on the robot arm provided in the embodiment of the present application. The entire object grasping process can be divided into two stages: the offline model building process and the online actual grasping. Specifically, the grasped object is used as a part for explanation. First, three modeling processes are performed in the offline model building stage. Specifically, it includes: two-dimensional camera parts acquisition to build AI model training samples to obtain a part segmentation model, two-dimensional calibration plate image acquisition, and obtaining a spatial pose conversion model between the camera coordinate system and the robot arm base coordinate system. Three-dimensional point cloud acquisition to obtain a point cloud template for each part.
[0119] In the actual grasping process, when the system receives the start work instruction, the camera starts to collect part images and point clouds, and sends the collected data to the deep semantic segmentation model, and obtains the target part point cloud through the point cloud positioning and segmentation module of this application; then the target part point cloud is feature matched with the point cloud template in the template library to obtain the grasping posture of the target part in the real environment scene, and the grasping posture is obtained under the robot arm base through the posture conversion model, and finally the actual grasping posture is transmitted to the remote robot arm execution system through the network to complete the final robot arm grasping operation.
[0120] The object grasping method based on a robotic arm provided in the present application aims at the deficiencies and existing problems of the existing part point cloud positioning and segmentation method, and provides a three-dimensional spatial positioning and segmentation method for parts by integrating the deep learning model of part semantic segmentation with the part point cloud. Compared with the traditional distance-based point cloud segmentation method, the point cloud segmentation method proposed in the present application has higher accuracy and faster efficiency. At the same time, a label data set for part semantic segmentation is produced by an automatic annotation method. This method replaces the traditional manual annotation and greatly improves the efficiency of part segmentation model training. And accurately estimate the 6-degree-of-freedom position and posture of the part space by matching the part point cloud features and building a two-point system. The point cloud feature matching method can accurately estimate the grasping point position of irregular parts and the angle of the three-dimensional x, y, and z axes in space. For any part placed, its 6-degree-of-freedom position and posture in space can be accurately calculated, which is far more applicable than the vertical grasping method of the two-dimensional plane angle. From the perspective of three-dimensional posture, the part posture obtained by the method of the present application is more accurate than the three-dimensional posture obtained by the fusion of the plane posture and the depth of field map. At the same time, the grasping point position and grasping coordinate system of the part are quickly determined by marking 2 points. Furthermore, an external parameter calibration method of a laser point cloud 3D camera based on a 2D image is provided. A traditional 3D point cloud camera calculates the pose of the 3D camera on the base of the robotic arm through the point cloud data of the target object. This application calculates the pose of the target object through the 2D image data obtained by shooting the target object with the left camera. This method has the advantages of convenient robotic arm operation, convenient data acquisition, and high efficiency in solving 2D images.
[0121] like Figure 4As shown, an embodiment of the present application provides an object grasping system based on a robotic arm, the system comprising: a collection device 10, an electronic device 20 and a robotic arm 30;
[0122] The acquisition device 10 is used to acquire a target image and target point cloud data of a target area, the target area including a target object to be captured, and send the acquired target image and target point cloud data to an electronic device;
[0123] The electronic device 20 is used to receive the target image and target point cloud data sent by the acquisition device, wherein the image content of the target image includes the target object to be captured, and the target point cloud data includes the object point cloud data of the target object;
[0124] The electronic device 20 is further used to perform image processing on the target image using a pre-stored target segmentation model to obtain region coordinates, where the region coordinates include each pixel of the target object and the coordinates of each pixel;
[0125] The electronic device 20 is further used to determine the object point cloud data according to the target point cloud data and the area coordinates;
[0126] The electronic device 20 is further used to determine the target pose of the target object according to the object point cloud data, and send the target pose to the robot arm;
[0127] The robot arm 30 is used to grasp the target object according to the target posture.
[0128] In one embodiment, the electronic device 20 is specifically used to: input the target image into the target segmentation model to obtain label information corresponding to each pixel area in multiple pixel areas of the target image, wherein the label information includes identification information and area coordinates corresponding to the pixel area, and the area coordinates include multiple pixel points;
[0129] Determine the target pixel area corresponding to the target identification information according to the correspondence between the pixel area and the identification information;
[0130] The region coordinates corresponding to the target pixel region are determined as the target region coordinates.
[0131] In one embodiment, the electronic device 20 is specifically used for:
[0132] Acquire the mapping relationship between the coordinates of each pixel point included in the target image and the coordinates of each target point cloud data;
[0133] The object point cloud data corresponding to the area coordinates is determined according to the mapping relationship.
[0134] In one embodiment, the electronic device 20 is further configured to:
[0135] Obtain sample images of multiple objects;
[0136] Determine a sample pixel region of each sample image, and determine label information corresponding to each sample pixel region of each sample image, wherein the label information includes region coordinates and identification information corresponding to the sample pixel region;
[0137] The preset positioning segmentation model is trained using sample images of multiple objects and label information corresponding to each pixel area in each sample image to obtain a target segmentation model.
[0138] In one embodiment, the electronic device 20 is specifically used to: perform edge segmentation processing on each sample image to obtain a target contour of an object included in each sample image;
[0139] According to the target contour corresponding to each sample image, each sample image is divided into regions to obtain a first sample pixel region and a second sample pixel region of each sample image;
[0140] The pixels within the target contour form a first sample pixel region, and the second sample pixel region is a blank region in the sample image.
[0141] In one embodiment, the electronic device 20 is specifically configured to: when the target contour corresponding to the first sample pixel region successfully matches the pre-stored pixel contour, use the identification information corresponding to the pixel contour as the identification information corresponding to the first sample pixel region;
[0142] Obtaining preset identification information corresponding to the second sample pixel area;
[0143] The region coordinates corresponding to the pixel region are determined according to the coordinates of each pixel point in the pixel region.
[0144] In one embodiment, the object point cloud data is point cloud data of the target object in a target coordinate system, and the target coordinate system is a coordinate system used by the acquisition device;
[0145] Electronic equipment is specifically used for:
[0146] Obtaining point cloud template data corresponding to the target identification information, where the point cloud template data is point cloud data of the target object in a preset coordinate system;
[0147] Determine the target pose based on the object point cloud data and point cloud template data.
[0148] In one embodiment, the electronic device 20 is specifically used to: obtain the initial posture of the target object according to the object point cloud data, the point cloud template data, the preset point feature histogram and the preset feature matching algorithm based on sampling matching consistency, where the initial posture is the posture of the target object based on the acquisition device;
[0149] Iterate the initial pose and object point cloud to obtain the optimized pose of the target object;
[0150] Obtaining the target coordinate transformation relationship between the acquisition device and the base of the robotic arm;
[0151] The target pose is determined according to the optimized pose and the target coordinate transformation relationship, and the target pose is the pose of the target object based on the base of the robot arm.
[0152] In one embodiment, the electronic device 20 is specifically used to: obtain a first coordinate transformation relationship and a corresponding second coordinate transformation relationship of the object in different positions and postures, the first coordinate transformation relationship is a coordinate transformation relationship between the acquisition device and the gripper of the robotic arm, and the second coordinate transformation relationship is a coordinate transformation relationship between the base and the gripper;
[0153] According to each first coordinate transformation relationship and the corresponding second coordinate transformation relationship, a target coordinate transformation relationship is obtained.
[0154] In one embodiment, the electronic device 20 is specifically used to: obtain a third coordinate transformation relationship corresponding to each first coordinate transformation relationship according to each first coordinate transformation relationship and the corresponding second coordinate transformation relationship;
[0155] The least squares fitting calculation is performed on multiple third coordinate transformation relations to obtain the target coordinate transformation relation.
[0156] In one embodiment, the electronic device 20 is further used to: select at least two point cloud data from the point cloud data of the target object to establish a preset coordinate system;
[0157] Determine the point cloud template data according to the preset coordinate system.
[0158] The embodiment of the present application provides an object grasping system based on a robotic arm. The electronic device receives the target image and target point cloud data of the area including the object to be grasped sent by the acquisition device, and determines the coordinates of each pixel of the target object according to the target identification information corresponding to the target object and the preset target segmentation model, and then determines the object point cloud data of the target object according to the target point cloud data and the area coordinates, and finally determines the target posture of the target object according to the object point cloud data, and sends the target to the robotic arm, so that the robotic arm grasps the target object according to the target posture. Since the present application uses the area coordinates determined by the image, and then obtains the point cloud data of the target object according to the area coordinates and the point cloud data, and the image data has better continuity than the discrete point cloud data, the point cloud of the target object determined by the present application is more accurate than the prior art in which the point cloud of the target object is obtained directly using the point cloud data, so the positioning accuracy of the object can be improved.
[0159] Furthermore, since the present application segments the image to obtain the regional coordinates of the target object, compared with the prior art which directly segments the point cloud data, the data processing efficiency is higher, thereby making the object positioning efficiency higher.
[0160] The robot-based object grasping system provided in this embodiment can execute the above method embodiments, and its implementation principles and technical effects are similar, which will not be elaborated here.
[0161] For the specific limitations of the robot-based object grasping system, please refer to the limitations of the robot-based object grasping method above, which will not be repeated here.
[0162] In another embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the steps of the object grasping method based on a robotic arm as in the embodiment of the present application are implemented.
[0163] In another embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the object grasping method based on a robotic arm as in the embodiment of the present application are implemented.
[0164] In another embodiment of the present application, a computer program product is also provided. The computer program product includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes each step of the object grasping method based on a robotic arm in the method flow shown in the above method embodiment.
[0165] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When loading and executing computer execution instructions on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. Computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, data center, etc. that contains one or more servers that can be integrated with a medium. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state drive (SSD)).
[0166] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0167] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for grasping an object based on a robotic arm, characterized in that: The method comprises: Receiving a target image and target point cloud data of a target area sent by a collection device, wherein the target area includes a target object to be captured; Determine the target region coordinates according to the target image, the target identification information corresponding to the target object and the pre-stored target segmentation model, wherein the target region coordinates include the coordinates of each pixel point of each pixel point of the target object; Determining object point cloud data of the target object according to the target point cloud data and the target area coordinates; Determine the target pose of the target object according to the object point cloud data; Sending the target posture to the robotic arm, where the target posture is used for the robotic arm to grasp the target object; Wherein, the object point cloud data is the point cloud data of the target object in the target coordinate system, and the target coordinate system is the coordinate system used by the acquisition device; The step of determining the target pose of the target object according to the object point cloud data comprises: Acquire point cloud template data corresponding to the target identification information, wherein the point cloud template data is point cloud data of the target object in a preset coordinate system; Obtaining an initial posture of the target object according to the object point cloud data, the point cloud template data, a preset point feature histogram, and a preset feature matching algorithm based on sampling matching consistency, wherein the initial posture is the posture of the target object based on the acquisition device; Iteratively calculating the initial pose and the object point cloud to obtain an optimized pose of the target object; Acquire a target coordinate transformation relationship between the acquisition device and the base of the robotic arm; Determine the target posture according to the optimized posture and the target coordinate transformation relationship, wherein the target posture is the posture of the target object based on the base of the robotic arm; Before determining the target area coordinates according to the target image, the target identification information corresponding to the target object and the pre-stored target segmentation model, the method further includes: Obtain sample images of multiple objects; Determine a sample pixel region of each sample image, and determine label information corresponding to each sample pixel region of each sample image, wherein the label information includes region coordinates and identification information corresponding to the sample pixel region; The preset positioning segmentation model is trained using the sample images of the multiple objects and the label information corresponding to each pixel area in each sample image to obtain the target segmentation model.
2. The method according to claim 1, characterized in that The step of determining the target area coordinates according to the target image, the target identification information corresponding to the target object and the pre-stored target segmentation model includes: Inputting the target image into the target segmentation model to obtain label information corresponding to each pixel area in multiple pixel areas of the target image, wherein the label information includes identification information and area coordinates corresponding to the pixel area, and the area coordinates include multiple pixel points; Determine the target pixel area corresponding to the target identification information according to the correspondence between the pixel area and the identification information; The region coordinates corresponding to the target pixel region are determined as the target region coordinates.
3. The method according to claim 1 or 2, characterized in that: Determining the object point cloud data of the target object according to the target point cloud data and the target area coordinates includes: Acquire a mapping relationship between the coordinates of each pixel point included in the target image and the coordinates of each target point cloud data; The object point cloud data corresponding to the area coordinates is determined according to the mapping relationship.
4. The method according to claim 1, characterized in that: The step of determining a sample pixel area of each sample image comprises: Perform edge segmentation processing on each sample image to obtain the target contour of the object included in each sample image; According to the target contour corresponding to each sample image, each sample image is divided into regions to obtain a first sample pixel region and a second sample pixel region of each sample image; The pixels within the target contour form the first sample pixel area, and the second sample pixel area is a blank area in the sample image.
5. The method according to claim 4, characterized in that The determining of label information corresponding to each sample pixel region of each sample image includes: When the target contour corresponding to the first sample pixel region successfully matches the pre-stored pixel contour, using the identification information corresponding to the pixel contour as the identification information corresponding to the first sample pixel region; Acquire preset identification information corresponding to the second sample pixel area; The region coordinates corresponding to the pixel region are determined according to the coordinates of each pixel point in the pixel region.
6. The method according to claim 1, characterized in that The step of obtaining a target coordinate conversion relationship between the acquisition device and the base of the robotic arm includes: Acquire a first coordinate transformation relationship and a corresponding second coordinate transformation relationship of the object in different positions and postures, wherein the first coordinate transformation relationship is a coordinate transformation relationship between the acquisition device and the gripper of the robotic arm, and the second coordinate transformation relationship is a coordinate transformation relationship between the base and the gripper; The target coordinate transformation relationship is obtained according to each first coordinate transformation relationship and the corresponding second coordinate transformation relationship.
7. The method according to claim 6, characterized in that The step of obtaining the target coordinate transformation relationship according to each first coordinate transformation relationship and the corresponding second coordinate transformation relationship includes: According to each first coordinate transformation relationship and the corresponding second coordinate transformation relationship, a third coordinate transformation relationship corresponding to each first coordinate transformation relationship is obtained; The plurality of third coordinate transformation relationships are subjected to least square fitting calculation to obtain the target coordinate transformation relationship.
8. The method according to claim 1, characterized in that Before obtaining the point cloud template corresponding to the target object, the method further includes: Selecting at least two point cloud data from the point cloud data of the target object to establish the preset coordinate system; The point cloud template data is determined according to the preset coordinate system.
9. An object grasping system based on a robotic arm, characterized in that: The system comprises: a collection device, an electronic device and a mechanical arm; The acquisition device is used to acquire a target image and target point cloud data of a target area, wherein the target area includes a target object to be captured, and send the acquired target image and target point cloud data to the electronic device; The electronic device is used to receive the target image and target point cloud data sent by the acquisition device, wherein the image content of the target image includes the target object to be captured, and the target point cloud data includes the object point cloud data of the target object; The electronic device is further used to determine the target area coordinates according to the target image, the target identification information corresponding to the target object and the pre-stored target segmentation model, wherein the target area coordinates include the coordinates of each pixel point of each pixel point of the target object; The electronic device is further used to determine the object point cloud data of the target object according to the target point cloud data and the target area coordinates; The electronic device is further used to determine the target posture of the target object according to the object point cloud data, and send the target posture to the robotic arm; The robotic arm is used to grasp the target object according to the target posture; Wherein, the object point cloud data is the point cloud data of the target object in the target coordinate system, and the target coordinate system is the coordinate system used by the acquisition device; The electronic device is specifically used for: Acquire point cloud template data corresponding to the target identification information, wherein the point cloud template data is point cloud data of the target object in a preset coordinate system; obtain an initial pose of the target object according to the object point cloud data, the point cloud template data, a preset point feature histogram and a preset feature matching algorithm based on sampling matching consistency, wherein the initial pose is the pose of the target object based on the acquisition device; iteratively calculate the initial pose and the object point cloud to obtain an optimized pose of the target object; acquire a target coordinate transformation relationship between the acquisition device and the base of the robotic arm; determine the target pose according to the optimized pose and the target coordinate transformation relationship, wherein the target pose is the pose of the target object based on the base of the robotic arm; The electronic device is further used for: Acquire sample images of multiple objects; determine a sample pixel area of each sample image, and determine label information corresponding to each sample pixel area of each sample image, wherein the label information includes area coordinates and identification information corresponding to the sample pixel area; train a preset positioning segmentation model using the sample images of the multiple objects and the label information corresponding to each pixel area in each sample image to obtain the target segmentation model.
10. An electronic device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the object grasping method based on a robot arm according to any one of claims 1 to 8 is implemented.
11. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the object grasping method based on a robotic arm according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Detection method, system and device for fusing image and point cloud information and storage medium
CN112861653A
Industrial robot disordered grabbing method based on double-clamp real-time switching
CN112873205A