A Fine Grasping Method for Manipulator in Hybrid Scenarios Based on Deep Reinforcement Learning

Through deep reinforcement learning, the robotic arm learns the plantar grasping in mixed scenarios, and uses the enhanced feedback signal to the plantar degree to solve the problem of insufficient generalization ability and reliability in the existing technology, and achieves efficient and stable grasping in mixed environments.

CN114986519BActive Publication Date: 2025-06-13SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210859754.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-21
Publication Date
2025-06-13
Estimated Expiration
2042-07-21

AI Technical Summary

Technical Problem

It is difficult for existing robot grasping technology to have both generalization capabilities and reliability, especially in mixed object scenarios.

Method used

The deep reinforcement learning method is adopted to enable the robotic arm to learn how to grasp objects on the plantar in mixed scenarios through continuous trial and error, and introduce feedback signals on the plantar degree as a self-supervised grab learning to enhance the stability and reliability of the grab.

Benefits of technology

It significantly improves the robot's grasping ability in mixed environments, improves the tolerance and stability of grabbing, and is suitable for scenarios such as garbage sorting, mixed workpiece selection and home finishing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114986519B_ABST
    Figure CN114986519B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for fine grasping of a robotic arm in a hybrid scenario based on deep reinforcement learning, including the following steps: Step S1: Based on deep reinforcement learning, the robotic arm continuously attempts to grasp in the grasping environment to train the fine grasping network, where the antipodal degree and the determination of whether the grasping is successful constitute the reward of deep reinforcement learning; Step S2: Use a camera to collect image information in the working scenario and convert the collected color image I c and depth information I d into a color top view I hc and a depth top view I hd ; Step S3: Input the depth top view I hc and the color top view I hd into the fine grasping network obtained in Step S1, output a multi-channel grasping indication diagram, and select the optimal grasping action G according to the generated multi-channel grasping indication diagram; Step S4: Send the optimal grasping action G from the operation server to the robot controller, and the robotic arm plans a trajectory to execute the optimal grasping action G. The present invention can perform antipodal grasping in a hybrid environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a fine grasping method for a robotic arm in a hybrid scenario based on deep reinforcement learning, belonging to the technical field of robot applications. Background Art

[0002] Robot grasping is one of the long-term challenges in the field of robot research. In recent years, as robots have penetrated into all walks of life in production and life, the application scenarios of robot grasping have become more and more extensive. For example, robot grasping is applied to scenarios such as garbage sorting, workpieces stacked in a mixed manner, and service robots in the home environment. The increasing number of application scenarios means that the corresponding grasping ability of the robot needs to be improved, such as robustness, stability, and flexibility.

[0003] Traditional methods can plan 6-DoF grasps in the point cloud space according to the force / form closure criterion. However, such methods highly depend on the models of the object and the end effector, and rely on some strong assumptions, such as simplified contact models, Coulomb friction, and rigid body model modeling methods. The actual effects of these methods are related to the matching degree between the assumptions in the physical world and the real physical world. However, in the real environment, there are inevitably errors in the models of the object and the robot, and the vision sensor is also inevitably noisy. Therefore, such methods basically do not have grasping generalization ability, that is, they cannot grasp objects with unknown models, or can only grasp objects similar to the template object, and are particularly poor in adaptability to the mixed aggregation scenario. The analytical method is difficult to achieve dexterous grasping of a variety of objects in a mixed environment like humans.

[0004] In recent years, the emerging data-driven grasping methods do not depend on the target model and can learn how to correctly grasp from the grasping actions and the observed image data, and have strong ability to grasp various objects in complex environments. As a representative method among them, the self-supervised grasping learning method learns the best grasping strategy through repeated grasping attempts, without the object model and without explicitly estimating the shape and pose of the object. Such methods can avoid a large amount of manual annotation efforts for samples, have strong model generalization ability, and have obvious advantages. Nevertheless, such grasping actions are rough, with only success or failure, and cannot guarantee the antipodality of the grasp, often resulting in unexpected movements of the grasping target and changes in the surrounding environment, thus damaging the stability and reliability of the grasp.

[0005] Generally speaking, the existing robot grasping technologies are difficult to meet the realistic requirements in various scenarios. Therefore, it is necessary to further improve the reliability of grasping on the premise of meeting a certain grasping generalization ability. Summary of the Invention

[0006] In view of the fact that existing robot grasping cannot possess both generalization ability and reliability, and it is difficult to adapt to various grasping scenarios, especially the mixed object scenario. The present invention provides a fine grasping method for a robotic arm in a mixed scenario based on deep reinforcement learning, which uses deep reinforcement learning to enable the robot to learn how to grasp objects in an antipodal manner from continuous trial and error, and introduces the antipodality of the grasping scenario to enhance the feedback signal of self-supervised grasping learning. The antipodality can finely reflect the strength of the grasping antipodality. - The antipodality is obtained by transforming the destructiveness through a non-increasing function. The destructiveness is reflected by the difference between the scene image before grasping the target object and the scene image after grasping and then putting the object back to the original position along the original path. Based on deep reinforcement learning and enhanced grasping feedback, we can significantly improve the grasping ability of the robot in a mixed environment, and it is expected to be applied to sorting tasks in scenarios such as garbage sorting, picking of mixed workpieces, and home organization.

[0007] The present invention is realized through the following technical solutions:

[0008] A fine grasping method for a robotic arm in a mixed scenario based on deep reinforcement learning of the present invention specifically includes the following steps:

[0009] Step S1: Based on deep reinforcement learning, the robotic arm continuously attempts to grasp in the grasping environment to train the fine grasping network, where the antipodality and the determination of whether the grasping is successful constitute the reward value in deep reinforcement learning;

[0010] Step S2: Use a camera to collect image information in the working scene and convert the collected color image I c and depth information I d into a color top view I hc and a depth top view I hd ;

[0011] Step S3: Input the depth top view I hc and the color top view I hd into the fine grasping network obtained in Step S1, output a multi-channel grasping demonstration diagram, and select the optimal grasping action G according to the generated multi-channel grasping demonstration diagram;

[0012] Step S4: Send the optimal grasping action G from the server to the robot controller, and the robotic arm plans a trajectory to execute the optimal grasping action G.

[0013] Furthermore, the training of the fine grasping network based on deep reinforcement learning in Step S1 specifically includes the following steps:

[0014] Step 1-1: Build a fine grasping network model Q θ, where θ represents the model parameters of the network. The fine grasping network has the same size as the input image and outputs a multi-channel grasping indication diagram. The multi-channel grasping indication diagram consists of N heatmaps with the same size as the input image, representing the highest grasping success rate. Different heatmaps indicate different rotation angles of the gripper around the Z-axis. Then, the N heatmaps can represent N grasping directions in total, and the rotation angle of the gripper around the Z-axis between two adjacent ones is 360° / N;

[0015] Step 1-2: Capture the image information in the grasping working scene and convert the collected image information I = (I c , I d ), where I c represents the color image, and I d represents the depth information. Then convert I c and I d into a color top view I hc and a depth top view I hd ;

[0016] Step 1-3: Input the depth top view I hc and the color top view I hd into the fine grasping network to output a multi-channel grasping indication diagram Q;

[0017] Step 1-4: Select the grasping action corresponding to the point with the maximum pixel value in Q as the optimal grasping action G = (T, φ), where T is the three-dimensional position (x, y, z) of the gripper, and φ is the angle of rotation of the three-dimensional grasping position of the gripper around the Z-axis;

[0018] Step 1-5: The robot executes the repeated grasping sampling operation M according to the optimal grasping action G. The process of the repeated grasping sampling operation M is as follows: The gripper first moves to a position l above the target position T perpendicular to it and rotates by an angle φ, then moves to the target position T along a straight-line trajectory. Determine whether the object is successfully grasped according to the closing condition of the gripper. If the object is successfully grasped, pick up the object and still move to a position l above T at an angle φ. If the object does not fall at this time, it is determined that this grasping is successful, and record the flag bit g = 1, and place the object back to the position T along the same trajectory; Capture the image information of the working space after the action Convert the color image and the depth information and convert them into a depth top view and a color top view If this grasping is successful, execute the grasping action G again to move the placed object out of the working space; If the grasping fails or the object falls when moving to a position l above T, it is determined that the grasping fails, and record the flag bit g = 0;

[0019] Step 1-6: Calculate I and I +The image difference degree P between them is obtained, including but not limited to calculating by formula (1):

[0020]

[0021] where B is the binarization operation of the image, and H and W are the length and width of I respectively;

[0022] Step 1-7: Input the image difference degree P into a non-increasing monotonic function to obtain the antipodal degree A, including but not limited to calculating A by formula (2):

[0023] A = -P + 1 (2)

[0024] Step 1-8: Combine the antipodal degree A and the flag bit g indicating whether the grasping is successful to obtain the grasping reward r, including but not limited to calculating r by formula (3):

[0025]

[0026] where is the first-order reward, is the baseline reward, and the two satisfy e < r 0 ;

[0027] Step 1-9: Update the fine grasping network based on the action value function iteration formula in deep reinforcement learning, that is

[0028]

[0029] where is the learning rate, is the loss function, including but not limited to using the root mean square error function;

[0030]

[0031] where γ is the discount factor, is the target network of Q θ and Q θ have the same network structure, with the network parameters θ - delayed;

[0032] Step 1-10: Determine whether the set maximum number of iteration steps is reached. If not, return to Step 1-2. If so, output the trained fine antipodal grasping network Q θ .

[0033] Furthermore, in Step S2, the collected color image I c and the depth information I d are converted into a color top view I hcand the depth top view I hd including the following steps:

[0034] Step 2-1: Obtain the internal and external parameter matrices of the fixed-position camera through hand-eye calibration;

[0035] Step 2-2: When the camera captures images, the depth image should be registered to the color image;

[0036] Step 2-3: First, use the internal and external parameter matrices of the camera to convert the depth image into a 3D point cloud image, and obtain the depth top view and color top view inside the working space through the projection method.

[0037] Furthermore, the step of selecting the optimal grasping action G according to the generated multi-channel grasping view map described in step S3 includes the following steps:

[0038] Step 3-1: Search for the index of the point with the largest pixel value in the multi-channel grasping view map Q,

[0039]

[0040] where c is the channel index where the maximum pixel value appears, and are respectively the positions of the maximum points in the heat map of this channel;

[0041] Step 3-2: According to c, and the depth map of the grasping space, obtain the three-dimensional coordinates T and the rotation angle φ around the Z-axis in the grasping space through the image and robotic arm coordinate transformation matrix

[0042]

[0043] Then, comprehensively obtain the optimal grasping action G=(T, φ) based on the three-dimensional coordinates T and the rotation angle φ around the Z-axis in the grasping space.

[0044] Furthermore, the step of the robotic arm planning a trajectory to execute the optimal grasping action G described in step S4 specifically includes the following steps:

[0045] Step 4-1: Establish communication between the operation server and the robotic arm controller;

[0046] Step 4-2: Plan a trajectory in the joint space to generate a grasping object trajectory based on the current robotic arm pose and the optimal grasping action G. If the grasping is successful, plan a trajectory and place the object at the target position; if the grasping fails, the robotic arm plans a trajectory and returns to the initial position.

[0047] Beneficial effects: The present invention can perform antipodal grasping in a mixed environment. Without the need for a target model and an artificial annotation dataset, a deep reinforcement learning method is used to train a robot to generate reliable antipodal grasping in a scene. When facing a grasping scene with mixed stacking, it can avoid large-scale damage to the environment caused by grasping actions and improve the reliability of grasping actions. Description of the Drawings

[0048] Figure 1 is the schematic diagram of repeated sampling in the present invention.

[0049] Figure 2 is the schematic diagram of the grasping damage degree and antipodal degree of the robot, where (a) and (b) respectively represent different grasping actions G to be performed in the same scene 1 and G 2 . (c) and (d) respectively show the comparison of the scene images before and after grasping, where the dashed-line object block represents the position of the object block before grasping, and the solid-line object block represents the position of the object block after grasping. The object in (c) has a larger range of unexpected movement compared to the object in (d). Detailed Implementation Manner

[0050] The present invention will be further described in detail below with reference to the accompanying drawings.

[0051] This embodiment provides a method for antipodal grasping of a robotic arm in a mixed scene based on deep reinforcement learning. The method includes the following steps:

[0052] Step S1: Based on deep reinforcement learning, the robotic arm continuously attempts to grasp in the grasping environment to train a fine grasping network, where the antipodal degree and the determination of whether the grasping is successful constitute the reward value in deep reinforcement learning.

[0053] Step S2: Use a camera to collect image information in the grasping working scene and convert the collected color image I c and depth information I d into a color top view I hc and a depth top view I hd ;

[0054] Step S3: Input the depth top view I hc and the color top view I hd into the fine grasping network obtained in Step S1 to output a multi-channel grasping indication diagram, and select the optimal grasping action G according to the generated multi-channel grasping indication diagram;

[0055] Step S4: Send the optimal grasping action G from the server to the robot controller, and the robotic arm plans a trajectory to execute the optimal grasping action G;

[0056] In this embodiment, the fine grasping network trained based on deep reinforcement learning in step S1 specifically includes the following steps:

[0057] Step 1-1: Build a fine grasping network model Q θ , where θ represents the model parameters of the network. The fine grasping network has the same size as the input image and outputs a multi-channel grasping indication diagram. The multi-channel grasping indication diagram consists of N heatmaps with the same size as the input image, representing the highest grasping success rate. Different heatmaps indicate different rotation angles of the gripper around the Z-axis. Then, the N heatmaps can represent N grasping directions in total, and the rotation angle of the gripper around the Z-axis between two adjacent heatmaps is 360° / N; Step 1-2: Capture the image information in the grasping working scene and convert the collected image information I=(I c ,I d ) according to the camera extrinsic parameter matrix, where I c represents the color image, and I d represents the depth information, and convert I c and I d into a color top view I hc and a depth top view I hd ;

[0058] Step 1-3: Input the depth top view I hc and the color top view I hd into the fine grasping network to output a multi-channel grasping indication diagram Q;

[0059] Step 1-4: Select the grasping action corresponding to the point with the largest pixel value in Q as the optimal grasping action G=(T,φ), where T is the three-dimensional position (x,y,z) of the gripper, and φ are the three-dimensional grasping position of the gripper and the rotation angle around the Z-axis respectively;

[0060] Step 1-5: The robot performs a repeated grasping sampling operation M according to the optimal grasping action G. The process of the repeated grasping sampling operation M is as follows: The gripper first moves to a position l above the target position T perpendicular to it and rotates to the angle φ, then moves to the target position T along a straight-line trajectory. Determine whether the object is successfully grasped according to the closing situation of the gripper. If the object is successfully grasped, pick up the object and still move to a position l above T at the angle φ. If the object does not fall at this time, it is determined that this grasping is successful, record the flag bit g = 1, and put the object back to T along the same trajectory; Capture the image information of the working space after the action, color image and depth information and convert them into a depth top view and a color top view respectively. If this grasping is successful, perform the grasping action G again to move the replaced object out of the working space; If the grasping fails or the object falls when moving to a position l above T, it is determined that the grasping fails, and record the flag bit g = 0.

[0061] Step 1-6: Calculate the image difference between I and I + to obtain the image difference degree P, including but not limited to calculating using formula (1):

[0062]

[0063] where B is the binarization operation of the image, and H and W are the length and width of I respectively

[0064] Step 1-7: Input the image difference degree P into a non-increasing monotonic function U to obtain the antipodal degree A, including but not limited to calculating A using formula (2):

[0065] A = -P + 1 (2)

[0066] Step 1-8: Combine the antipodal degree A and the grasping success flag g to obtain the grasping reward r, including but not limited to calculating r using formula (3):

[0067]

[0068] where is the first-order reward, is the benchmark reward, and both satisfy e < r 0 .

[0069] Step 1-9: Update the fine grasping network based on the action value function iteration formula in deep reinforcement learning, that is

[0070]

[0071] where is the learning rate, is the loss function, including but not limited to using the root mean square error function.

[0072]

[0073] where is the target network of Q θ , and Q θ has the same network structure, with the network parameters θ - . Copy the network parameters of Q θ every τ steps

[0074] Step 1-10: Determine whether the set maximum number of iteration steps is reached. If not, return to Step 1-2. If so, output the trained fine antipodal grasping network Q θ .

[0075] The fine grasping network in this embodiment can be composed of an encoding module, a fusion module, and a decoding module. The encoding module takes an image as input and outputs shallow color features and shallow depth information features. The fusion module takes the shallow color features and shallow depth information features as input and inputs the fused latent features. The decoding module takes the latent features as input and outputs a multi-channel grasping view map. The encoding module uses ResNet-50 as the backbone network. The fusion module contains two convolutional blocks, and each convolutional block includes two convolutional layers and a ReLU activation function. The decoding module contains three transposed convolutional blocks, and each transposed convolutional block contains two transposed convolutional layers and a ReLU activation function. The fine grasping network can be implemented by the PyTorch deep learning library.

[0076] The number of rotation angles of the discrete gripper around the Z-axis in this embodiment can be set to 16, that is, the rotation angle between two adjacent rotations around the Z-axis is 360° / 16 = 22.5°.

[0077] The repeated grasping sampling operation m in this embodiment is as Figure 1 shown. First, execute trajectory (1), and the gripper moves from the initial position to After that, the gripper executes trajectory (2) and moves to T along a straight-line trajectory i . The gripper closes to attempt to grasp the target object, and it is judged whether the object is successfully grasped according to whether the gripper closes. If the object is successfully grasped, then execute trajectory (3), and the gripper moves to Subsequently, the gripper executes trajectory (4) again and runs to T along a straight-line trajectory i . At this time, open the gripper and put down the object. Finally, the gripper returns to the initial position along trajectories (5) and (6).

[0078] The image binarization operation B in this example can adopt the OTSU binarization method, which divides the image into a background and a target according to the gray characteristics of the image. The typical size of the image I in this example: H = 640, W = 640.

[0079] The non-increasing monotonic function adopted in this example is not limited to being expressed by a linear formula, as long as the non-increasing monotonic function satisfies the following characteristics:

[0080]

[0081] Typically, the following formula can also be used to transform P to A:

[0082]

[0083] where Typically, ∈ = 0.1.

[0084] This example combines Figure 2Further elaborate on the damage degree and the antipodal degree. Figure 2 (a) and (b) respectively represent different grasping actions G to be executed in the same scenario 1 and G 2 . (c) and (d) respectively show the comparison of the scene images before and after grasping, where the dashed-line object block represents the position of the object block before grasping, and the solid-line object block represents the position of the object block after grasping. The object in (c) has a larger range of unexpected movement compared to the object in (d). This means that G 1 has a greater damage degree than G 2 , that is, G 1 has a better antipodal degree than G 2 .

[0085] In this example, the benchmark reward r 0 = 1, and the first-order reward e = 0.5.

[0086] In this example, copy the network parameters of Q θ every τ = 3 steps

[0087] In this embodiment, in step S2, the collected color image I c and the depth information I d are converted into a color top view I hc and a depth top view I hd :

[0088] Step 2-1: Obtain the internal and external parameter matrices of the fixed-position camera through hand-eye calibration;

[0089] Step 2-2: When the camera collects images, the depth image should be registered to the color image;

[0090] Step 2-3: First, use the internal and external parameter matrices of the camera to convert the depth image into a 3D point cloud image, and obtain the depth top view and color top view inside the working space through the projection method.

[0091] The camera adopted in this embodiment should be able to obtain both the color image information and the depth information of the scene at the same time. The fixed camera position should be set so that the field of view of the camera can completely cover the working space of the robot. The distance between the camera and the working space should be within the optimal distance range for the camera to capture depth information. In this example, a Realsense D435 camera is used to capture the image information = (I c , I d ), and the coordinates of the target object in the robot coordinate system are calculated through coordinate transformation.

[0092] The step of selecting the optimal grasping action G according to the generated multi-channel grasping indication diagram in step S3 of this embodiment includes the following steps:

[0093] Step 3-1: Search and find the index of the point with the largest pixel value in the multi-channel grasping view map Q.

[0094]

[0095] Among them, c is the channel index where the maximum pixel value appears. And Are respectively the positions of the maximum points in the heat map of this channel.

[0096] Step 3-2: According to c, And the depth map of the grasping space, obtain the three-dimensional coordinates T=(x, y, z) in the grasping space and the rotation angle φ around the Z-axis through the image and robotic arm coordinate transformation matrix.

[0097]

[0098] In this example, the z coordinate of the gripper is obtained by reading the depth information of the optimal grasping position in the plane. This example uses a UR3 collaborative robot and a two-finger gripper to perform the grasping task.

[0099] The specific steps for the robotic arm to plan the trajectory and execute the optimal grasping action G described in step S4 of this embodiment are as follows:

[0100] Step 4-1, establish communication between the operation server and the robotic arm controller.

[0101] Step 4-2, plan a trajectory in the joint space based on the current robotic arm pose and the optimal grasping action G to generate a grasping object trajectory. If the grasping is successful, plan the trajectory and place the object at the target position; if the grasping fails, the robotic arm plans the trajectory and returns to the initial position.

[0102] In this example, the system used is Ubuntu 18.04, and the server is equipped with an Intel Xeon CPU E5-2620v3@2.4Ghz and an NVIDIA RTX 3090. This example uses the forward and inverse kinematic solvers inside the robot control to plan the actions of the robotic arm. The initial position of the robotic arm in this example should ensure that the robot does not appear within the field of view of the camera.

[0103] A robot sorting system was built to prove the effectiveness of the present invention. Extensive experiments were conducted on common metal workpieces in the industry in a real scenario to test the method proposed in this paper. In an automated production line, the grasping task of stacked metal workpieces is very common. Current data-driven research usually uses daily objects as experimental subjects. However, compared with daily objects, metal workpieces have more complex and irregular shapes and cannot be approximated as simple geometric bodies, which requires higher requirements for grasping postures. In addition, stacked metal workpieces are denser than daily objects, and the occlusion between objects is more serious. This requires the grasping method to have a stronger ability to grasp objects from a mixed stacking scenario. Metal workpieces were selected as experimental subjects to demonstrate the advantages of the method in this paper when grasping mixed stacked objects. Twenty kinds of metal workpieces were randomly selected as experimental subjects.

[0104] Three scenarios were designed: 1) the scenario of isolated objects, 2) the scenario of multiple objects, and 3) the scenario of mixed objects. In the isolated scenario, a single workpiece was randomly placed, and each time the workpiece was placed with a random position and angle. In the scenario of multiple workpieces, there was no overlap and occlusion between the workpieces. In the scenario of mixed aggregation, there was serious overlap and occlusion between the workpieces. At this time, improper grasping would seriously damage the environment. The effects of the present invention were evaluated in the above three scenarios. Each method was tested for 30 grasps, and the experimental results were statistically analyzed to obtain Table 1. Method 1 in Table 1 is similar to the method of the present invention, except that the antipodal degree is not used to enhance the grasping feedback. The grasping feedback of Method 1 is obtained according to whether the grasping is successful, that is, if the grasping is successful, the grasping feedback is 1, and if the grasping fails, the grasping feedback is 0.

[0105] The grasping success rate in Table 1 is obtained according to the following formula:

[0106]

[0107] where K is the number of successful grasps in M grasps, and M is the number of attempts in each round of experiments. Here, M = 30.

[0108] The average antipodal degree in Table 1 is obtained according to the following formula:

[0109]

[0110] where A i is the antipodal degree of each grasp.

[0111] Experimental results of grasping in different scenarios in Table 1

[0112]

[0113] When grasping in a single-scene, scattered-object, and mixed-object scenario, the grasping success rate and average antipodality of the method proposed in this paper are higher than those of Method 1. In particular, the proposed invention has a higher average antipodality in each grasping scenario compared to Method 1. This indicates that the proposed invention has good antipodality and the ability to avoid collisions. This shows that the method proposed in this paper is more reliable, stable, and causes less damage to the environment. This indicates that although these two methods can perform high-antipodality grasps in an isolated scenario, in a scenario with multiple objects, it is also possible to cause environmental damage due to collisions between the gripper and the surrounding environment.

[0114] In addition, the generalization ability of the proposed invention in different scenarios was tested on 6 novel objects. The experimental results are shown in Table 2. We can see that the method proposed in this paper achieved a grasping success rate similar to that of grasping known objects in each scenario when grasping novel objects. Although the average antipodality decreased slightly when grasping novel objects, especially in the mixed scenario, it was still better than Method 1. Our method demonstrated good generalization ability when grasping novel objects. The generalization ability here not only includes the ability to successfully grasp objects. More importantly, it includes the ability to perform fine-grained antipodal grasps for different scenarios.

[0115] Table 2 Grasping unseen objects

[0116]

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and not to limit the scope of its protection. Although this application has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that after reading this application, various changes, modifications, or equivalent replacements can still be made to the specific implementation manners of the application. However, these changes, modifications, or equivalent replacements are all within the scope of the protection of the pending claims of the application.

Claims

1. A fine grasping method for a robotic arm in a hybrid scenario based on deep reinforcement learning, characterized in that, it includes the following steps: Step S1: Based on deep reinforcement learning, the robotic arm continuously attempts to grasp in the grasping environment to train the fine grasping network, where the antipodality and the determination of whether the grasping is successful constitute the reward value in deep reinforcement learning; The training of the fine grasping network based on deep reinforcement learning in Step S1 specifically includes the following steps: Step 1-1: Build a fine-grained grasping network model Q θ , where θ represents the model parameters of the network. The fine-grained grasping network has the same size as the input image and outputs a multi-channel grasping indication diagram. The multi-channel grasping indication diagram consists of N heatmaps with the same size as the input image, representing the highest grasping success rate. Different heatmaps indicate different rotation angles of the gripper around the Z-axis. Then, the N heatmaps can represent N grasping directions in total, and the rotation angle of the gripper around the Z-axis between two adjacent ones is 360° / N; Step 1-2: Capture the image information in the grasping work scene and convert the collected image information I=(I c , I d ), where I c represents the color image, I d represents the depth information, and convert I c and I d into a color top view image I hc and a depth top view image I hd ; Step 1-3: Input the depth top view I hc and the color top view I hd into the fine grasping network to output the multi-channel grasping view map Q; Step 1-4: Select the grasping action corresponding to the point with the largest pixel value in Q as the optimal grasping action G=(T,φ), where T is the three-dimensional position of the gripper (x,y,z), and φ is the angle of rotation of the three-dimensional grasping position of the gripper around the Z-axis; Step 1-5: The robot performs repeated grasping sampling operation M according to the optimal grasping action G. The process of the repeated grasping sampling operation M is as follows: The gripper first moves to a position l above the target position T perpendicular to it and rotates by an angle φ, then moves to the target position T along a straight-line trajectory. It determines whether the object is successfully grasped based on the closing condition of the gripper. If the object is successfully grasped, it picks up the object and still moves to a position l above T at an angle φ. If the object does not fall at this time, it is determined that this grasping is successful, and the flag bit g is set to 1, and the object is placed back to T along the same trajectory; the workspace image information after the capture action The color image and the depth information are respectively converted into a depth top view and a color top view If this grasping is successful, the grasping action G is performed again to move the placed object out of the workspace; if the grasping fails or the object falls when moving to a position l above T, it is determined that the grasping fails, and the flag bit g is set to 0; Step 1-6: Calculate the image difference between I and I + to obtain the image difference degree P, including but not limited to calculating using formula (1): where B is the binarization operation of the image, and H and W are the length and width of I respectively; Step 1-7: Input the image difference degree P into a non-increasing monotonic function to obtain the antipodality A, including but not limited to calculating A using formula (2): A = -P + 1 (2) Step 1-8: Combine the antipodality A and the flag bit g indicating whether the grasping is successful to obtain the grasping reward r, including but not limited to calculating r using formula (3): Among them is the first-order reward, is the benchmark reward, and the two satisfy e < r 0 ; Step 1-9: Update the fine grasping network based on the action value function iteration formula in deep reinforcement learning, that is Among them, is the learning rate, is the loss function, including but not limited to using the root mean square error function; where γ is the discount factor, is the target network of Q θ and has the same network structure as Q θ with the network parameters θ delayed - ; Step 1-10: Determine whether the set maximum number of iteration steps is reached. If not, return to Step 1-2. If so, output the trained fine anti-podal grasping network Q θ ; Step S2: Use a camera to collect image information in the working scene and convert the collected color image I c and depth information I d into a color top view image I hc and a depth top view image I hd ; Step S3: Input the depth top view I hc and the color top view I hd into the fine grasping network obtained in Step S1, output a multi-channel grasping indication diagram, and select the optimal grasping action G according to the generated multi-channel grasping indication diagram; The selection of the optimal grasping action G according to the generated multi-channel grasping demonstration diagram in Step S3 includes the following steps: Step 3-1: Search for the index of the point with the largest pixel value in the multi-channel grasping demonstration diagram Q, where c is the channel index where the maximum pixel value appears, and are respectively the positions of the maximum points in the channel heat map where the maximum pixel value appears; Step 3-2: According to c, and the depth map of the grasping space, the three-dimensional coordinates T in the grasping space and the rotation angle φ around the Z-axis are obtained through the image and the robotic arm coordinate transformation matrix and then synthesize the three-dimensional coordinates T in the grasping space and the rotation angle φ around the Z-axis to obtain the optimal grasping action G=(T,φ); Step S4: Send the optimal grasping action G from the server to the robot controller, and the robotic arm plans a trajectory to execute the optimal grasping action G.

2. A fine grasping method for a robotic arm in a hybrid scenario based on deep reinforcement learning according to claim 1, characterized in that, In step S2, the acquired color image I c and depth information I d are converted into a color top view I hc and a depth top view I hd , including the following steps: Step 2-1: Obtain the internal and external parameter matrices of the fixed-position camera through hand-eye calibration; Step 2-2: When the camera collects images, the depth image should be registered to the color image; Step 2-3: First, use the internal and external parameter matrices of the camera to convert the depth image into a 3D point cloud image, and obtain the depth top view and color top view inside the working space through the projection method.

3. A fine grasping method for a robotic arm in a hybrid scenario based on deep reinforcement learning according to claim 1, characterized in that, The specific steps for the robotic arm to plan a trajectory to execute the optimal grasping action G in Step S4 include the following: Step 4-1, establish communication between the computing server and the robotic arm controller; Step 4-2, plan a trajectory in the joint space based on the current pose of the robotic arm and the optimal grasping action G to generate a trajectory for grasping the object. If the grasping is successful, plan a trajectory and place the object at the target position; if the grasping fails, the robotic arm plans a trajectory and returns to the initial position.