Object grabbing method and device, terminal equipment and storage medium
By generating and filtering candidate placement poses, using point cloud image difference calculations to optimize the grab and placement operation of the robot, the problem that the robot cannot place objects to the target pose at one time is solved, and accurate object placement under incomplete images is achieved.
Patent Information
- Application Number
- CN202510538542.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-27
AI Technical Summary
When a robot moves an object, it is often impossible to grab and place the object to the target posture at one time due to the limitations of the workspace and robot structure, and it requires multiple grab and placement. The existing technology relies on the complete image of the object to accurately determine the position position and cannot handle the situation of incomplete images.
By obtaining the point cloud image of the target object, multiple candidate placement positions are generated, and the optimal placement positions are filtered according to the difference in point cloud information, and the robot controls the grab and placement to achieve the accurate placement of the target object.
In the case where the complete image of the target object is not obtained, the optimal placement position of the object can be accurately determined and realized, reducing grabbing errors, and improving the accuracy of object movement.
Smart Images

Figure CN120244970A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of robot control, and particularly relates to an object grasping method, device, terminal device, and storage medium. Background Art
[0002] When a robot performs the task of moving an object, it needs to first grasp the object and then place the object at the target position in the target pose. During the process of the robot moving the object, often due to limitations such as the working space and the robot structure, the robot cannot place the object at the target position in the target pose through a single grasping and placing operation. The robot often needs to perform multiple grasping and placing operations to complete the movement of the object. Therefore, how to more accurately determine the placement pose of the object after the robot grasps the object is a problem that needs to be solved currently. Summary of the Invention
[0003] Embodiments of this application provide an object grasping method, device, terminal device, and storage medium, which can more accurately determine the placement pose of the object after the robot grasps the object.
[0004] In a first aspect, embodiments of this application provide an object grasping method, including:
[0005] During the process of the manipulator moving the target object to the target pose, if the i-th grasping of the target object is required, obtain the i-th point cloud image of the target object collected at the current moment, where the target object is in the i-th placement pose at the current moment, and i≥1;
[0006] If the object structure in the i-th point cloud image does not match the preset object structure of the target object, generate multiple candidate placement poses according to the i-th point cloud image, where the candidate placement pose is the pose that the manipulator can place the target object after the i-th grasping of the target object;
[0007] According to the difference between the first point cloud information corresponding to each candidate placement pose and the second point cloud information of the target object in the i-th point cloud image, screen out the optimal placement pose from the multiple candidate placement poses, where the first point cloud information is the point cloud information of the target object that can be collected when the target object is predicted to be in the candidate placement pose;
[0008] Based on the optimal placement pose, control the manipulator to perform the i-th grasping and placing operation, and place the target object in the optimal placement pose, where the optimal placement position is the (i + 1)-th placement pose of the target object.
[0009] In a possible implementation of the first aspect, after obtaining the i-th point cloud image of the target object collected at the current moment, the method further includes:
[0010] If the object structure in the i-th point cloud image matches the preset object structure of the target object, determine the optimal placement pose according to the i-th placement pose and the preset target pose of the target object.
[0011] In a possible implementation of the first aspect, the obtaining the i-th point cloud image of the target object collected at the current moment includes:
[0012] Control the imaging device to collect the i-th image information of the target object, where the imaging device is fixed in a preset area, and the image information is a single-view RGB-D image;
[0013] Generate the i-th point cloud image of the target object according to the i-th image information.
[0014] In a possible implementation of the first aspect, the generating the i-th point cloud image of the target object according to the i-th image information includes:
[0015] Generate the current point cloud image of the target object according to the i-th image information;
[0016] Obtain the (i - 1)-th point cloud image of the target object, where the (i - 1)-th point cloud image is obtained based on the collected (i - 1)-th image information, and the (i - 1)-th image information is the image information of the target object collected when the target object is in the (i - 1)-th placement pose;
[0017] Perform image fusion on the (i - 1)-th point cloud image and the current point cloud image to obtain the i-th point cloud image of the target object.
[0018] In a possible implementation of the first aspect, the screening out the optimal placement pose from multiple candidate placement poses according to the difference between the first point cloud information corresponding to each candidate placement pose and the second point cloud information of the target object in the i-th point cloud image includes:
[0019] For each candidate placement pose, predict the first point cloud information of the target object collected when the target object is in the candidate placement pose;
[0020] Extract the second point cloud information of the target object in the i-th point cloud image;
[0021] Screen out the point cloud information with differences in the first feature information and the second feature information to obtain the differential feature information;
[0022] Screen the same point cloud information in the first point cloud information and the second point cloud information to obtain coincidence feature information;
[0023] Based on the difference feature information and the coincidence feature information, calculate the score of each candidate placement pose;
[0024] According to the scores of each candidate placement pose, screen out the optimal placement pose from multiple candidate placement poses.
[0025] In a possible implementation manner of the first aspect, the calculating the score of each candidate placement pose based on the difference feature information and the coincidence feature information includes:
[0026] Based on the difference feature information, calculate the information gain of the first point cloud information relative to the second point cloud information;
[0027] Based on the coincidence feature information, calculate the information coincidence degree between the first point cloud information and the second point cloud information;
[0028] Based on the information gain and the information coincidence degree, calculate the score of each candidate placement pose.
[0029] In a possible implementation manner of the first aspect, the predicting the first point cloud information of the target object collected when the target object is in the candidate placement pose includes:
[0030] According to the candidate placement pose, determine the virtual placement position of the virtual camera device relative to the target object, so that the target object in the i-th placement pose observed by the virtual camera device is the same as the target object in the candidate placement pose observed by the camera device, where the camera device is a device fixed in a preset area for photographing the target object at the current moment;
[0031] Construct a virtual space, where the target object is in the i-th placement pose in the virtual space, and the virtual camera device is set at the virtual placement position;
[0032] In the virtual space, simulate the virtual camera device to emit ray signals to the target object;
[0033] According to the return signal of the ray signal, obtain the first point cloud information of the target object collected when the target object is in the candidate placement pose.
[0034] In a second aspect, an embodiment of the present application provides an object grasping device, including:
[0035] An image acquisition module, configured to, when the manipulator moves the target object to the target pose and needs to perform the i-th grasping of the target object, acquire the i-th point cloud image of the target object collected at the current moment, where the target object is in the i-th placement pose at the current moment, and i≥1;
[0036] A position prediction module, configured to, if the i-th point cloud image does not include a complete image of the target object, generate a plurality of candidate placement poses according to the i-th point cloud image, where the candidate placement poses are the poses in which the manipulator can place the target object after the i-th grasping of the target object by the manipulator;
[0037] An optimal pose selection module, configured to screen out the optimal placement pose from the plurality of candidate placement poses according to the difference between the first point cloud information corresponding to each candidate placement pose and the second point cloud information of the target object in the i-th point cloud image, where the first point cloud information is the point cloud information of the target object that can be acquired when the target object is predicted to be in the candidate placement pose;
[0038] A control module, configured to control the manipulator to perform the i-th grasping and placement operations based on the optimal placement pose, and place the target object in the optimal placement pose, where the optimal placement position is the (i + 1)-th placement pose of the target object.
[0039] In a third aspect, an embodiment of the present application provides a terminal device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the object grasping method described in any one of the first aspects above is implemented.
[0040] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the object grasping method described in any one of the first aspects above is implemented.
[0041] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when running on a terminal device, causes the terminal device to execute the object grasping method described in any one of the first aspects above.
[0042] The beneficial effects of the first aspect of this application compared with the prior art are as follows: When the manipulator performs re-grasping, if it is necessary to perform the i-th grasping of the target object, first obtain the i-th point cloud image of the target object collected at the current moment. If the object structure in the i-th point cloud image does not match the preset object structure of the target object, it means that there are missing parts in the point cloud image of the currently collected target object. Therefore, multiple candidate placement poses that the target object can present next can be predicted based on the i-th point cloud image first; according to the differences between the first point cloud information of the target object that can be collected at each candidate placement pose and the second point cloud information in the i-th point cloud image, the optimal placement pose is selected from multiple candidate placement poses, and finally the manipulator is controlled to grasp and place the target object so that the target object presents the optimal placement pose. The grasping of the target object in this application does not depend on the complete image of the target object. When the complete image of the target object is not obtained, this application can first predict the candidate placement poses where the target object can be placed, and then obtain the optimal placement pose through screening. It can not only grasp the target object based on the incomplete image of the target object, but also place the target object in the optimal pose after grasping the target object, making the grasping of the target object more accurate.
[0043] It can be understood that the beneficial effects of the second to fifth aspects above can refer to the relevant descriptions in the first aspect above and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0045] Figure 1 It is a schematic diagram of a scenario where an object cannot be placed in a target pose with a single grasping provided by an embodiment of this application;
[0046] Figure 2 It is a schematic diagram of a scenario where an object can be placed in a target pose with multiple graspings provided by an embodiment of this application;
[0047] Figure 3 It is a schematic flowchart of an object grasping method provided by an embodiment of this application;
[0048] Figure 4 It is a comparison schematic diagram of an incomplete image and a complete image of an object provided by an embodiment of this application;
[0049] Figure 5It is a schematic flowchart of a method for obtaining a point cloud image provided by an embodiment of the present application;
[0050] Figure 6 It is a schematic flowchart of a method for obtaining a point cloud image provided by another embodiment of the present application;
[0051] Figure 7 It is a schematic flowchart of a method for selecting an optimal placement pose provided by an embodiment of the present application;
[0052] Figure 8 It is a schematic diagram showing the display of a virtual space constructed by an embodiment of the present application;
[0053] Figure 9 It is a schematic structural diagram of an object grasping device provided by an embodiment of the present application;
[0054] Figure 10 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0055] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0056] As used in the specification and appended claims of the present application, the term "if" may be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" may be interpreted as meaning "once determined" or "in response to determining" or "once detecting [the described condition or event]" or "in response to detecting [the described condition or event]" according to the context.
[0057] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0058] Referring to "an embodiment" or "some embodiments" etc. described in the specification of the present application means that specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Thus, statements such as "in an embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.
[0059] In robotic operations, object redirection is a common task. Object redirection refers to the process by which a robot adjusts an object from an initial pose to a target pose.
[0060] Specifically, when a robotic arm needs to grasp an object and move it, often due to limitations in space and the performance of the robotic arm itself, after the robotic arm grasps an object once, it cannot place the object in the target position in the target pose. For example Figure 1 As shown, due to the limitations of the robotic arm itself, the object cannot be laid flat through a single grasp. At this time, the robotic arm needs to perform an object redirection operation and place the object in the target position in the target pose after multiple grasps and placements. For example Figure 2 As shown, the robotic arm lays the object flat through two grasps.
[0061] Currently, when a robotic arm performs multiple grasps on an object, it needs to rely on a complete image of the object. If the image of the object is incomplete, the pose of the object after grasping cannot be calculated. Therefore, when grasping an object, it is necessary to first collect images of multiple perspectives of the object in order to obtain a complete image of all angles of the object. This method is relatively complex, and when dealing with objects with complex structures, the object cannot be accurately moved either. In the case where a complete image of the object cannot be obtained, grasping and moving cannot be performed.
[0062] Therefore, this application proposes an object grasping method. Through the image of the object taken from a single perspective, when the image cannot completely reproduce the structure of the object, the object can be grasped and placed more accurately. Specifically, when the structure of the collected object is incomplete, multiple placement poses where the object can be placed after grasping can be predicted first, and then the optimal placement pose can be selected from multiple poses. Finally, the grasped object is placed in the optimal placement pose. After multiple grasps and placements, the effect of grasping and more accurately placing the object without a complete image of the object is achieved.
[0063] The following will Figure 3 describe in detail the object grasping method of the embodiments of this application.
[0064] Figure 3 shows a schematic flowchart of the object grasping method provided by this application. Referring to Figure 3 the following is a detailed description of this method:
[0065] S101, during the process of the robotic arm moving the target object to the target pose, if the i-th grasp of the target object is required, obtain the i-th point cloud image of the target object collected at the current moment, where the target object is in the i-th placement pose at the current moment, i≥1.
[0066] In this embodiment, the i-th point cloud image is the point cloud image of the target object at the current moment collected by a camera device fixed in a preset area. The point cloud image of an object obtained by a camera device fixed in a preset area is a point cloud image in a single view.
[0067] The placement pose includes position and orientation.
[0068] When i = 1, the i-th placement pose is the initial pose of the target object. At this time, the target object needs to be grasped for the first time. Correspondingly, the currently collected point cloud image of the target object is the first point cloud image.
[0069] When i ≠ 1, the current placement pose of the target object is the pose where the target object is placed after the previous grasping of the object. For example, if i = 3, at the current moment, the target object is in the 3rd placement pose, and the 3rd placement pose is the pose where the target object is placed after the second grasping. At this time, the target object needs to be grasped for the third time, and the point cloud image of the target object collected at the current moment is the third point cloud image.
[0070] S102. If the object structure in the i-th point cloud image does not match the preset object structure of the target object, generate multiple candidate placement poses according to the i-th point cloud image, where the candidate placement pose is the pose that the manipulator can place the target object after the i-th grasping of the target object.
[0071] In this embodiment, the object structure of the target object is stored in advance, and the stored object structure of the target object is the complete structure of the object. If the object structure in the i-th point cloud image does not match the preset object structure of the target object, it means that the point cloud image of the object collected at this time cannot completely present the object structure. For example Figure 4 as shown Figure 4 only the last image in Figure 4 is the complete image of the object, and Figure 4 the first 3 images in
[0072] cannot completely present the object structure of the target object;
[0073] In one way, input the i-th point cloud image into a trained convolutional neural network to obtain multiple candidate placement poses.
[0074] In another way, according to the i-th point cloud image, the performance parameters of the manipulator, etc., find the poses that the manipulator can reach and the target object can be stably held after being placed, and obtain multiple candidate placement poses.
[0075] S103. According to the difference between the first point cloud information corresponding to each candidate placement pose and the second point cloud information of the target object in the i-th point cloud image, screen out the optimal placement pose from the multiple candidate placement poses, where the first point cloud information is the point cloud information of the target object that can be collected when the target object is in the candidate placement pose.
[0076] In this embodiment, first predict that if the target object is placed according to the candidate placement pose. In this case, the point cloud information of the target object collected by the camera device fixed in the preset area is recorded as the first point cloud information. In order to collect a more complete image of the target object next time, therefore, according to the difference between the first point cloud information corresponding to the candidate placement pose and the second point cloud information of the target object in the i-th point cloud image, screen out the optimal placement pose.
[0077] In one way, calculate the missing point cloud ratio, the newly added point cloud ratio, and the overlapping point cloud ratio of the first point cloud information relative to the second point cloud information. Based on the missing point cloud ratio, the newly added point cloud ratio, and the overlapping point cloud ratio, screen out the optimal placement pose. Missing point cloud ratio: The ratio of the number of point clouds in the second point cloud information that have no corresponding point clouds in the first point cloud information to the number of point clouds in the second point cloud information. Newly added point cloud ratio: The ratio of the number of point clouds in the first point cloud information that have no corresponding point clouds in the second point cloud information to the number of point clouds in the first point cloud information. Overlapping point cloud ratio: The ratio of the number of point clouds in the first point cloud information that have corresponding point clouds in the second point cloud information to the number of point clouds in the first point cloud information.
[0078] Specifically, obtain the first weight corresponding to the missing point cloud ratio, the second weight of the newly added point cloud ratio, and the third weight of the overlapping point cloud ratio, and use the weighted algorithm to obtain the score of the candidate placement pose. Select the candidate placement pose with the highest score as the optimal placement pose.
[0079] S104. Based on the optimal placement pose, control the manipulator to perform the i-th grasping and placing operation, and place the target object in the optimal placement pose, where the optimal placement position is the (i + 1)-th placement pose of the target object.
[0080] In this embodiment, the candidate placement poses may further include the distance that the manipulator should translate and the angle of rotation when converting from the current position of the target object to the candidate placement pose. Therefore, the optimal placement position includes the distance that the manipulator should translate and the angle of rotation during this grasping. Control the manipulator to act according to the translation distance and rotation angle of the manipulator, so that the target object reaches the optimal placement pose.
[0081] Alternatively, according to the optimal placement pose, the i-th placement pose of the target object, and the performance parameters of the manipulator, determine the distance that the manipulator should translate and the angle of rotation during this grasping. Then control the manipulator to act according to the determined translation distance and rotation angle.
[0082] In this application, when the manipulator performs re-grasping, if it is necessary to grasp the target object for the i-th time, first obtain the i-th point cloud image of the target object collected at the current moment. If the object structure in the i-th point cloud image does not match the preset object structure of the target object, it means that there are missing parts in the currently collected point cloud image of the target object. Therefore, multiple candidate placement poses that the target object can present next can be predicted according to the i-th point cloud image first; then, according to the difference between the first point cloud information of the target object that can be collected at each candidate placement pose and the second point cloud information in the i-th point cloud image, screen the optimal placement pose from multiple candidate placement poses. Finally, control the manipulator to grasp and place the target object so that the target object presents the optimal placement pose. This application does not depend on the complete image of the target object for grasping the target object. When the complete image of the target object is not obtained, this application can first predict the candidate placement poses where the target object can be placed, and then obtain the optimal placement pose through screening. It can not only grasp the target object according to the incomplete image of the target object, but also place the target object in the optimal pose after grasping the target object, making the grasping of the target object more accurate.
[0083] Since this application is all executed based on the point cloud image of the target object, and the RGB-D image includes the depth and color of the target object, etc., the point cloud image of the target object can be determined based on the RGB-D image.
[0084] Thus, as Figure 5 shown, the implementation process of step S101 may include:
[0085] S1011, control the imaging device to collect the i-th image information of the target object, where the imaging device is fixed in a preset area, and the image information is a single-view RGB-D image.
[0086] In this embodiment, the imaging device can be a depth camera, a radar system, or the like. Since the imaging device is fixed in a preset area and its position remains unchanged, the target object captured by the imaging device is a single-view image.
[0087] S1012. Generate the i-th point cloud image of the target object according to the i-th image information.
[0088] In this embodiment, the i-th image information is the i-th RGB-D image, which includes color visual information and depth information. Therefore, the point cloud image of the target object can be generated according to the i-th image information.
[0089] Since the image information of the target object obtained each time is the image information under a single view, in order to make the point cloud image more accurate, after obtaining the point cloud image of the target object at the current moment, the point cloud image at the current moment and the previous point cloud image can be fused to obtain a point cloud image containing more features of the target object, and then a complete point cloud image of the target object can be obtained.
[0090] Specifically, as Figure 6 shown, the implementation process of the above step S1012 may further include:
[0091] S201. Generate the current point cloud image of the target object according to the i-th image information.
[0092] S202. Obtain the (i - 1)-th point cloud image of the target object, where the (i - 1)-th point cloud image is obtained based on the collected (i - 1)-th image information, and the (i - 1)-th image information is the image information of the target object collected when the target object is in the (i - 1)-th placement pose.
[0093] In this embodiment, when i = 1, there is no (i - 1)-th point cloud image. Therefore, when i = 1, the obtained (i - 1)-th point cloud image is 0.
[0094] S203. Perform image fusion on the (i - 1)-th point cloud image and the current point cloud image to obtain the i-th point cloud image of the target object.
[0095] In this embodiment, when i = 2, obtain the 1st point cloud image of the target object, and fuse the 1st point cloud image with the currently collected point cloud image to obtain the 2nd point cloud image. The 1st point cloud image is the point cloud image of the target object collected when the target object is in the initial pose.
[0096] When i = 3, obtain the 2nd point cloud image of the target object, and fuse the 2nd point cloud image with the currently collected point cloud image to obtain the 3rd point cloud image.
[0097] By analogy, after several captures, a complete point cloud image of the target object can be obtained. For example, Figure 4 the last picture in
[0098] is the complete point cloud image of the target object. In this application, through continuous fusion of point cloud images, a complete point cloud image of the target object can finally be obtained. Grabbing the target object based on the increasingly complete image of the target object can enable the manipulator to grab the target object more accurately, reducing the error of the manipulator in grabbing the target object.
[0099] For example, Figure 7 as shown in
[0100] S1031, for each of the candidate placement poses, predict the first point cloud information of the target object collected when the target object is in the candidate placement pose.
[0101] In one way, simulate the scenario where the target object is placed in the candidate placement pose, and then simulate using a camera device to photograph the target object to obtain the first point cloud information of the target object.
[0102] In another way, use the principle of Truncated Signed Distance Function (TSDF) to determine the first point cloud information.
[0103] Specifically, S11, according to the candidate placement pose, determine the virtual placement position of the virtual camera device relative to the target object, so that the target object observed by the virtual camera device in the i-th placement pose is the same as the target object observed by the camera device in the candidate placement pose. The camera device is a device fixed in a preset area for photographing the target object at the current moment.
[0104] In this embodiment, without changing the current pose of the target object, that is, the target object is still in the i-th placement pose. Change the position of the virtual camera device to find the position where the virtual camera device photographs the target object, which is equivalent to the image photographed by the camera device when the target object is set to the candidate placement pose, to obtain the virtual placement position of the virtual camera device.
[0105] S12, construct a virtual space, in which the target object is in the i-th placement pose and the virtual camera device is set at the virtual placement position.
[0106] As an example, for example, Figure 8As shown, place the target object in the i-th placement pose. A is a camera device fixed in a preset area. B is the virtual placement position where the virtual camera device is located according to the first candidate placement position. C is the virtual placement position where the virtual camera device is located according to the second candidate placement position. D is the virtual placement position where the virtual camera device is located according to the third candidate placement position. E is the virtual placement position where the virtual camera device is located according to the fourth candidate placement position.
[0107] S13. In the virtual space, simulate the virtual camera device emitting a ray signal towards the target object.
[0108] S14. According to the return signal of the ray signal, obtain the first point cloud information of the target object collected when the target object is in the candidate placement pose.
[0109] In this embodiment, if the ray signal can hit the target object, there is a return signal for the ray signal. If the ray signal cannot hit the target object, there is no return signal for the ray signal.
[0110] S1032. Extract the second point cloud information of the target object in the i-th point cloud image.
[0111] In this embodiment, the point cloud information includes the data of each point cloud in the point cloud image, and the point cloud information can include geometric information, intensity information, color information, etc. of each point cloud.
[0112] S1033. Screen the point cloud information with differences between the first point cloud information and the second point cloud information to obtain difference feature information.
[0113] In this embodiment, the point cloud information with differences between the first point cloud information and the second point cloud information is the difference feature information.
[0114] Specifically, compare each point cloud in the first point cloud information with each point cloud in the second information to obtain the point cloud information with differences between the first point cloud information and the second point cloud information.
[0115] Specifically, if point cloud A in the first point cloud information does not exist in the second point cloud information, then point cloud A is a differential point cloud. If point cloud B in the second point cloud information does not exist in the first point cloud information, then point cloud B is also a differential point cloud.
[0116] S1034. Screen the point cloud information that is the same between the first point cloud information and the second point cloud information to obtain coincidence feature information.
[0117] In this embodiment, the point cloud information that is the same between the first point cloud information and the second point cloud information is the coincidence feature information.
[0118] S1035, calculate the score of each of the candidate placement poses based on the difference feature information and the coincidence feature information.
[0119] In this embodiment, calculate the score of the candidate placement pose according to the number of the differential point clouds in the difference feature information and the number of the coincident point clouds in the coincidence feature information.
[0120] In one way, according to the number of the differential point clouds, query the preset interval where the number of the differential point clouds is located. Each preset interval corresponds to a score, and determine the score corresponding to the preset interval where the number of the differential point clouds is located as the first score. According to the number of the coincident point clouds, query the preset interval where the number of the coincident point clouds is located, and determine the score corresponding to the preset interval where the number of the coincident point clouds is located as the second score. The sum of the first score and the second score is the score of the candidate placement pose.
[0121] In another way, the calculation method of the score may further include:
[0122] Based on the difference feature information, calculate the information gain of the first point cloud information relative to the second point cloud information; based on the coincidence feature information, calculate the information coincidence degree between the first point cloud information and the second point cloud information; based on the information gain and the information coincidence degree, calculate the score of each of the candidate placement poses.
[0123] In this embodiment, the information gain represents the newly added information of the first point cloud information relative to the second point cloud information, that is, the features of the target object newly collected. The information coincidence degree represents the degree to which the point clouds in the second point cloud information are retained in the first point cloud information, that is, the retention degree of the features of the target object collected previously. Determine the optimal placement position jointly according to the information gain and the information coincidence degree, which takes into account that after grasping and placing the target object this time, when collecting the image of the target object next time, as much point cloud information of the target object as possible is collected, so that the constructed point cloud image of the target object is more complete.
[0124] Specifically, query the number of the first point clouds that are the newly added point clouds in the first point cloud information in the difference feature information. Divide the number of the first point clouds by the total number of the point clouds in the first point cloud information to obtain the information gain.
[0125] Query the number of the second point clouds that are the repeated point clouds in the first point cloud information and the second point cloud information in the coincidence feature information; multiply the number of the second point clouds by a preset parameter to obtain the information coincidence degree. Alternatively, calculate the total number of the point clouds in the first point cloud information and the number of the point clouds in the second point cloud information; divide the number of the second point clouds by the total number to obtain the information coincidence degree.
[0126] Obtain the weight of the information gain and the weight of the information coincidence degree, and use the weighted algorithm to obtain the score of the candidate placement pose.
[0127] S1036. Screen out the optimal placement pose from multiple candidate placement poses according to the scores of each candidate placement pose.
[0128] In this embodiment, find the highest score from all the scores, and determine the candidate placement pose corresponding to the highest score as the optimal placement pose.
[0129] In this application, by comparing the first point cloud information and the second point cloud information, calculate the scores of each candidate placement pose, and then screen out the optimal placement pose according to the scores of each candidate placement pose; this application provides a method for screening the optimal placement pose, making the finally determined optimal placement pose more in line with the requirements.
[0130] The above introduces a method for screening out the optimal placement pose according to the difference between the first point cloud information corresponding to each candidate placement pose and the second point cloud information of the target object in the i-th point cloud image when the complete point cloud image of the target object is not obtained. If the complete point cloud image of the target object is obtained, the pose where the target object should be placed after grasping the target object can be determined according to the current pose and the target pose of the target object.
[0131] Specifically, if the object structure in the i-th point cloud image matches the preset object structure of the target object, determine the (i + 1)-th placement pose of the target object according to the i-th placement pose and the preset target pose of the target object.
[0132] In this embodiment, if the object structure in the i-th point cloud image matches the preset object structure of the target object, it means that there is a complete point cloud image of the target object in the i-th point cloud image.
[0133] According to the i-th point cloud image and the performance parameters of the manipulator, it can be predicted whether the target object can be placed in the target pose after this grasping. If the target object can be placed in the target pose after this grasping, the target pose can be determined as the pose that the target object should present after grasping the target object this time. If the target object cannot be placed in the target pose after this grasping, a virtual model of the target object and the manipulator can be established, and the process of grasping the target object can be simulated through the virtual model to determine how many times of grasping are required to place the target object in the target pose, and then determine the placement pose of the target object after each grasping.
[0134] Alternatively, input the i-th point cloud image and the target pose of the target object into the trained convolutional neural network to obtain the (i + 1)-th placement pose.
[0135] The following introduces another implementation manner of the method of this application.
[0136] S21, Collect the initial image (the first image information) of the target object, and generate the first point cloud image based on the initial image.
[0137] S22, If the structure of the target object in the first point cloud image does not match the preset object structure (that is, the structure of the target object in the first point cloud image is incomplete), generate multiple candidate placement poses based on the first image information.
[0138] S23, According to the difference between the first point cloud information of the target object corresponding to the candidate placement pose and the second point cloud information in the first point cloud image, screen out the optimal placement pose from multiple candidate placement poses.
[0139] S24, Control the manipulator to perform the first grasping and placement, so that the target object is in the optimal placement pose (the second placement pose).
[0140] S25, Collect the second image information of the target object when it is in the second placement pose, and perform image fusion on the point cloud image generated from the second image information and the first point cloud image to generate the second point cloud image.
[0141] S26, If the structure of the target object in the second point cloud image does not match the preset object structure, repeat the operations of S22 to S25 above.
[0142] S27, If the structure of the target object in the collected Nth point cloud image matches the preset object structure, determine the (N + 1)th placement pose of the target object according to the Nth placement pose of the target object and the preset target pose of the target object. Control the manipulator to perform the Nth grasping and placement, so that the target object is in the (N + 1)th placement pose until the target object is moved to the target pose.
[0143] In this application, through the fusion of the point cloud data of the target object and the screening of the optimal placement pose, the purpose of reconstructing the structure of the target object is achieved, and finally a complete image of the target object is obtained, reducing the error in the object grasping process and making the object grasping more accurate.
[0144] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.
[0145] Corresponding to the object grasping method described in the above embodiments, Figure 9 The structural block diagram of the object grasping device provided by the embodiment of this application is shown. For the sake of convenience of description, only the parts related to the embodiment of this application are shown.
[0146] Refer to Figure 9, the device 300 may include: an image acquisition module 310, a position prediction module 320, an optimal pose selection module 330, and a control module 340.
[0147] Among them, the image acquisition module 310 is configured to, when the manipulator moves the target object to the target pose and needs to perform the i-th grasping of the target object, acquire the i-th point cloud image of the target object collected at the current moment, where the target object is in the i-th placement pose at the current moment, and i≥1;
[0148] The position prediction module 320 is configured to, if the i-th point cloud image does not include a complete image of the target object, generate a plurality of candidate placement poses according to the i-th point cloud image, where the candidate placement pose is the pose at which the manipulator can place the target object after the i-th grasping of the target object;
[0149] The optimal pose selection module 330 is configured to screen out the optimal placement pose from the plurality of candidate placement poses according to the difference between the first point cloud information corresponding to each candidate placement pose and the second point cloud information of the target object in the i-th point cloud image, where the first point cloud information is the point cloud information of the target object that can be acquired when the target object is in the candidate placement pose;
[0150] The control module 340 is configured to, based on the optimal placement pose, control the manipulator to perform the i-th grasping and placement operations, and place the target object in the optimal placement pose, where the optimal placement position is the (i + 1)-th placement pose of the target object.
[0151] In a possible implementation manner, the control module 340 may specifically further be configured to:
[0152] If the object structure in the i-th point cloud image matches the preset object structure of the target object, determine the (i + 1)-th placement pose of the target object according to the i-th placement pose and the preset target pose of the target object.
[0153] In a possible implementation manner, the image acquisition module 310 may specifically be configured to:
[0154] Control the imaging device to acquire the i-th image information of the target object, where the imaging device is fixed in a preset area, and the image information is a single-view RGB-D image;
[0155] Generate the i-th point cloud image of the target object according to the i-th image information.
[0156] In a possible implementation manner, the image acquisition module 310 may specifically be configured to:
[0157] Generate the current point cloud image of the target object according to the i-th image information.
[0158] Obtain the (i - 1)-th point cloud image of the target object, where the (i - 1)-th point cloud image is obtained based on the collected (i - 1)-th image information, and the (i - 1)-th image information is the image information of the target object collected when the target object is in the (i - 1)-th placement pose.
[0159] Fuse the (i - 1)-th point cloud image with the current point cloud image to obtain the i-th point cloud image of the target object.
[0160] In a possible implementation, the optimal pose selection module 330 may specifically be configured to:
[0161] For each of the candidate placement poses, predict the first point cloud information of the target object collected when the target object is in the candidate placement pose.
[0162] Extract the second point cloud information of the target object in the i-th point cloud image.
[0163] Filter the point cloud information with differences in the first point cloud information and the second point cloud information to obtain the difference feature information.
[0164] Filter the same point cloud information in the first point cloud information and the second point cloud information to obtain the coincidence feature information.
[0165] Based on the difference feature information and the coincidence feature information, calculate the score of each candidate placement pose.
[0166] According to the scores of each candidate placement pose, screen out the optimal placement pose from multiple candidate placement poses.
[0167] In a possible implementation, the optimal pose selection module 330 may specifically be configured to:
[0168] Based on the difference feature information, calculate the information gain of the first point cloud information relative to the second point cloud information.
[0169] Based on the coincidence feature information, calculate the information coincidence degree between the first point cloud information and the second point cloud information.
[0170] Based on the information gain and the information coincidence degree, calculate the score of each candidate placement pose.
[0171] In a possible implementation, the optimal pose selection module 330 may specifically be configured to:
[0172] According to the candidate placement pose, determine the virtual placement position of the virtual camera device relative to the target object, so that the target object in the i-th placement pose observed by the virtual camera device is the same as the target object in the candidate placement pose observed by the camera device, where the camera device is a device fixed in a preset area for photographing the target object at the current moment;
[0173] Construct a virtual space, in which the target object is in the i-th placement pose, and the virtual camera device is set at the virtual placement position;
[0174] In the virtual space, simulate the virtual camera device to emit ray signals to the target object;
[0175] According to the return signal of the ray signal, obtain the first point cloud information of the target object collected when the target object is in the candidate placement pose.
[0176] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / modules, due to being based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, please refer to the method embodiment part specifically, and will not be elaborated here.
[0177] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments and will not be elaborated here.
[0178] The embodiment of the present application also provides a terminal device. Refer to Figure 10 , the terminal device 400 may include: at least one processor 410, a memory 420, and a computer program stored in the memory 420 and executable on the at least one processor 410. When the processor 410 executes the computer program, it implements the steps in any of the above method embodiments, for example Figure 3Steps S101 to S104 in the illustrated embodiments. Alternatively, when the processor 410 executes the computer program, it implements the functions of each module / unit in the above-described device embodiments, such as Figure 9 the functions of the illustrated image acquisition module 310 to the control module 340.
[0179] Exemplarily, the computer program can be segmented into one or more modules / units. One or more modules / units are stored in the memory 420 and executed by the processor 410 to complete the present application. The one or more modules / units can be a series of computer program segments capable of performing specific functions, and these program segments are used to describe the execution process of the computer program in the terminal device 400.
[0180] Those skilled in the art can understand that Figure 10 merely examples of the terminal device, which do not constitute a limitation to the terminal device, and may include more or fewer components than those shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, buses, etc.
[0181] The processor 410 can be a central processing unit (CPU), or can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0182] The memory 420 can be an internal storage unit of the terminal device, or can also be an external storage device of the terminal device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. The memory 420 is used to store the computer program and other programs and data required by the terminal device. The memory 420 can also be used to temporarily store data that has been output or is to be output.
[0183] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.
[0184] The object grasping method, device, terminal device, and storage medium provided in the embodiments of this application can be applied to terminal devices such as computers, tablet computers, laptop computers, netbooks, and personal digital assistants (PDAs). The embodiments of this application do not impose any restrictions on the specific types of terminal devices.
[0185] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0186] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0187] In the embodiments provided in this application, it should be understood that the disclosed terminal devices, devices, and methods can be implemented in other ways. For example, the terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in an electrical, mechanical, or other form.
[0188] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0189] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0190] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-described embodiment methods of the present application can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by one or more processors, the steps of the above-described method embodiments can be implemented.
[0191] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-described embodiment methods of the present application can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by one or more processors, the steps of the above-described method embodiments can be implemented.
[0192] Similarly, as a computer program product, when the computer program product runs on a terminal device, it enables the terminal device to execute the steps in the above-described method embodiments.
[0193] Among them, the computer program includes computer program code, which may be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0194] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An object grasping method, characterized in that, Including: During the process of the manipulator moving the target object to the target pose, if the target object needs to be grasped for the $i$-th time, acquire the $i$-th point cloud image of the target object collected at the current moment, where the target object is in the $i$-th placement pose at the current moment, and $i\geq1$; If the object structure in the $i$-th point cloud image does not match the preset object structure of the target object, generate multiple candidate placement poses according to the $i$-th point cloud image, where the candidate placement pose is the pose that the manipulator can place the target object after grasping the target object for the $i$-th time; According to the difference between the first point cloud information corresponding to each candidate placement pose and the second point cloud information of the target object in the $i$-th point cloud image, screen out the optimal placement pose from multiple candidate placement poses, where the first point cloud information is the point cloud information of the target object that can be collected when the target object is predicted to be in the candidate placement pose; Based on the optimal placement pose, control the manipulator to perform the $i$-th grasping and placement operation, and place the target object in the optimal placement pose, where the optimal placement position is the $(i + 1)$-th placement pose of the target object.
2. The object grasping method according to claim 1, wherein After acquiring the $i$-th point cloud image of the target object collected at the current moment, the method further includes: If the object structure in the $i$-th point cloud image matches the preset object structure of the target object, determine the $(i + 1)$-th placement pose of the target object according to the $i$-th placement pose and the preset target pose of the target object.
3. The object grasping method according to claim 1, wherein The acquiring the $i$-th point cloud image of the target object collected at the current moment includes: Control the imaging device to collect the $i$-th image information of the target object, where the imaging device is fixed in a preset area, and the image information is a single-view RGB-D image; Generate the $i$-th point cloud image of the target object according to the $i$-th image information.
4. The object grasping method according to claim 3, characterized in that, The generating the $i$-th point cloud image of the target object according to the $i$-th image information includes: Generate the current point cloud image of the target object according to the $i$-th image information; Acquire the $(i - 1)$-th point cloud image of the target object, where the $(i - 1)$-th point cloud image is obtained based on the collected $(i - 1)$-th image information, and the $(i - 1)$-th image information is the image information of the target object collected when the target object is in the $(i - 1)$-th placement pose; Fuse the $(i - 1)$-th point cloud image with the current point cloud image to obtain the $i$-th point cloud image of the target object.
5. The object grasping method according to any one of claims 1 to 4, characterized in that, The screening out the optimal placement pose from multiple candidate placement poses according to the difference between the first point cloud information corresponding to each candidate placement pose and the second point cloud information of the target object in the $i$-th point cloud image includes: For each candidate placement pose, predict the first point cloud information of the target object that can be collected when the target object is in the candidate placement pose; Extract the second point cloud information of the target object in the $i$-th point cloud image; Screen the point cloud information with differences in the first point cloud information and the second point cloud information to obtain differential feature information; Screen the identical point cloud information in the first point cloud information and the second point cloud information to obtain coincidence feature information; Based on the differential feature information and the coincidence feature information, calculate the score of each candidate placement pose; According to the scores of each candidate placement pose, screen out the optimal placement pose from multiple candidate placement poses.
6. The object grasping method according to claim 5, characterized in that, The calculating the score of each candidate placement pose based on the differential feature information and the coincidence feature information includes: Based on the differential feature information, calculate the information gain of the first point cloud information relative to the second point cloud information; Based on the coincidence feature information, calculate the information coincidence degree between the first point cloud information and the second point cloud information; Based on the information gain and the information coincidence degree, calculate the score of each candidate placement pose.
7. The object grasping method according to claim 5, characterized in that, The acquiring the first point cloud information of the target object when predicting that the target object is in the candidate placement pose includes: According to the candidate placement pose, determine the virtual placement position of the virtual camera device relative to the target object, so that the target object observed by the virtual camera device in the i-th placement pose is the same as the target object observed by the camera device in the candidate placement pose, where the camera device is a device fixed in a preset area for photographing the target object at the current moment; Construct a virtual space, where the target object is in the i-th placement pose in the virtual space, and the virtual camera device is set at the virtual placement position; In the virtual space, simulate the virtual camera device to emit a ray signal to the target object; According to the return signal of the ray signal, obtain the first point cloud information of the target object when the target object is in the candidate placement pose.
8. An object grasping device, characterized in that, including: An image acquisition module, configured to, when the manipulator moves the target object to the target pose and needs to grasp the target object for the i-th time, acquire the i-th point cloud image of the target object acquired at the current moment, where the target object is in the i-th placement pose at the current moment, and i≥1; A position prediction module, configured to, if the i-th point cloud image does not include a complete image of the target object, generate multiple candidate placement poses according to the i-th point cloud image, where the candidate placement pose is the pose where the manipulator can place the target object after the i-th grasp of the target object by the manipulator; An optimal pose selection module, configured to screen out the optimal placement pose from multiple candidate placement poses according to the difference between the first point cloud information corresponding to each candidate placement pose and the second point cloud information of the target object in the i-th point cloud image, where the first point cloud information is the point cloud information of the target object that can be acquired when predicting that the target object is in the candidate placement pose; The control module is configured to control the manipulator to perform the i-th grasping and placing operation based on the optimal placement pose, and place the target object in the optimal placement pose, where the optimal placement position is the (i + 1)-th placement pose of the target object.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the object grasping method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the object grasping method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Vision calibration system for robotic carton unloading
CN111618842A
Method for detecting object grabbing poses in three-dimensional point clouds
CN111652928A
Robot grabbing pose detection method based on domain migration under single-view-angle point cloud
CN112489117A
Unknown object six-degree-of-freedom grabbing method considering point cloud skeleton characteristics
CN116460845A
Autonomous pickup and placement pose acquisition method for robot in disordered scene
CN118081758A
Cited By
Cargo loading method, device and equipment based on deep reinforcement learning
CN121340319A