Dynamic object grabbing method based on 3D visual reinforcement learning
By combining depth camera and robot body information, a deep reinforcement learning network with a multi-stage reward function is designed to solve the problems of low dynamic grasping success rate and high algorithm complexity, and achieve efficient grasping in industrial assembly and human-machine collaboration scenarios.
Patent Information
- Application Number
- CN202510721113.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-05
AI Technical Summary
How to improve the success rate of dynamic grasping and simplify the complexity of the algorithm, especially when applying deep reinforcement learning algorithms in industrial assembly and human-machine collaboration scenarios, to solve the challenge of simulation-to-reality migration.
By collecting depth camera visual information and converting it into point cloud data, combined with the robot body information, a deep reinforcement learning network with a multi-stage reward function is designed, and the improved SAC algorithm and grasping point feature compression coding network are used to realize the grasping of dynamic objects.
It improves the success rate of dynamic capture, reduces the difficulty of migrating from simulation to reality, simplifies the algorithm complexity, and has certain robustness and engineering implementation advantages.
Smart Images

Figure CN120599595A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object grasping, and in particular to a dynamic object grasping method based on 3D vision reinforcement learning. Background Art
[0002] Reinforcement Learning (RL), a key branch of machine learning, has revolutionized robotic grasping technology with its unique trial-and-error-reward mechanism. This approach, by constructing a closed-loop learning system in which an intelligent agent continuously interacts with its environment, has pioneered a new paradigm for robots to autonomously explore optimal grasping strategies. Compared to traditional methods, RL exhibits three significant advantages in robotic grasping tasks: First, it abandons the traditional rule-based control framework and adopts an end-to-end learning model to achieve intelligent mapping from raw visual input to grasping actions. Second, it possesses a powerful ability to process high-dimensional continuous state / action spaces, effectively addressing positional changes and scene interference of target objects in unstructured environments. Third, it significantly reduces the reliance on customized specialized fixtures and manual parameter adjustment, significantly improving the system's versatility and adaptability.
[0003] In recent years, with the rapid development of deep reinforcement learning algorithms and the widespread application of high-fidelity physical simulation platforms (MuJoCo, PyBullet, Isaac Sim, etc.), researchers have built a complete simulation environment training system. Through the physics engine, it accurately simulates the kinematic parameters of the joints (including angles, angular velocities, etc.), the interaction state of the gripper-object (covering contact force, slip detection, etc.), and environmental perception information such as the target object's posture. It is particularly noteworthy that although simulation training effectively avoids the high cost and high risk of ontological training in real scenarios, the challenges of migration from simulation to reality (Sim2Real), such as sensor noise modeling deviations and differences in the dynamic characteristics of the transmission system, remain the main bottlenecks restricting the large-scale application of this technology in industrial scenarios.
[0004] Dynamic object grasping technology has important application value in industrial assembly and human-machine collaboration. Although grasping methods based on the decoupling of 6DoF grasping detection and motion planning in static scenes have made significant progress, dynamic object grasping tasks based on deep reinforcement learning still face challenges. In addition, when implementing robot grasping skill training based on deep reinforcement learning algorithms, scene visual information and the robot's body state are usually used as algorithm inputs, and reward function design, neural network optimization and other technologies are combined to accelerate model convergence. Overall, the effectiveness of deep reinforcement learning algorithms depends not only on the degree of adaptation of the algorithm itself to specific tasks, but is also closely related to factors such as experimental environment construction, test plan design, task difficulty division and training strategy selection. These factors will have different degrees of influence on the skill learning effect.
[0005] How to improve the success rate of dynamic grasping and simplify the complexity of the algorithm is an urgent problem to be solved. To this end, a dynamic object grasping method based on 3D visual reinforcement learning is proposed. Summary of the Invention
[0006] The technical problem to be solved by the present invention is how to improve the success rate of dynamic grasping and simplify the complexity of the algorithm, and provides a dynamic object grasping method based on 3D visual reinforcement learning.
[0007] The present invention solves the above technical problems through the following technical solutions, which include the following steps:
[0008] S1: Collection and processing of observation information
[0009] The depth camera collects visual information, converts the visual information into point cloud data, performs grasping point detection based on the point cloud data, obtains the grasping point information of the target object, and then extracts features from the grasping point information to obtain visual grasping features. At the same time, the robot body related information is obtained and processed to obtain the robot body features. The visual grasping features and the robot body features are observation information, which are spliced together as the state information of the deep reinforcement learning network.
[0010] S2: Deep reinforcement learning network design and processing
[0011] The state information in step S1 is input into the deep reinforcement learning network for processing to obtain the decision action; a multi-stage reward function is designed in the deep reinforcement learning network, which includes a proximity reward function, a grasping contact reward function, and a target reward function;
[0012] S3: Robot body and gripper motion output
[0013] Based on the decision-making actions obtained by the deep reinforcement learning network, the action output of the robot body and the action output of the gripper are obtained to complete the task of grasping the target object.
[0014] Furthermore, in step S1, the process of obtaining the grasping point information is as follows:
[0015] S101: Acquire visual information through a depth camera, and convert the visual information into point cloud data with three-dimensional spatial information based on the transformation relationship between pixel coordinates and world coordinates, that is, obtain local point cloud data of the scene in the camera coordinate system;
[0016] S102: transforming the scene local point cloud data in the camera coordinate system into the base coordinate system to obtain the scene local point cloud data in the base coordinate system, wherein the base coordinate system is the coordinate system where the robot base is located;
[0017] S103: Before starting to grasp, the relative relationship between the target object and the robot remains unchanged, and the local point cloud data of the object is filtered out from the local point cloud data of the scene in the base coordinate system by setting a height threshold;
[0018] S104: transforming the local point cloud data of the object into the robot tool end coordinate system, and processing it using a parallel FPS algorithm and a spatial clustering algorithm to remove noise, thereby obtaining the noise-reduced local point cloud data of the object;
[0019] S105: The scene local point cloud data in step S102, the object local point cloud data in step S104, and the object local point cloud grasping center point detection result constraints are used as input information of the local object point cloud 6DoF grasping detection inference model LoG, and the grasping detection results relative to the terminal coordinate system preset in the model LoG are obtained: the grasping point information of the target object and the corresponding grasping quality score, and then the grasping point information relative to the robot tool terminal coordinate system is obtained according to the transformation relationship between the terminal coordinate system preset in the model LoG and the robot tool terminal coordinate system. The grasping point information includes the grasping center point coordinates and the grasping posture relative to the robot tool terminal coordinate system.
[0020] Furthermore, in step S1, the process of acquiring visual capture features is as follows:
[0021] S106: Using the grasping point information relative to the robot tool end coordinate system as a transformation matrix, the Gaussian points sampled at the origin of the tool end coordinate system are sequentially transformed to obtain corresponding transformation results;
[0022] S107: Input the transformation result and the Gaussian point into the grasping point feature compression coding network GraspGroupNet for processing to obtain high-dimensional feature information of the grasping point, namely the visual grasping feature.
[0023] Furthermore, in step S106, K three-dimensional points are sampled at the origin of the tool end coordinate system to characterize the spatial position of the clamping jaws. The K three-dimensional points are Gaussian points.
[0024] Furthermore, in step S107, the grasping point feature compression coding network GraspGroupNet includes a Grasp Group module and a Group module. In the Grasp Group module, the transformation result of the Gaussian point is processed as input, and then the output result is spliced with the grasping center point coordinates in the grasping point information as the input of the Group module. After processing by the Group module, a maximum pooling operation is performed on the output data channel to obtain high-dimensional feature information of the grasping point, that is, the visual grasping feature.
[0025] Furthermore, in step S1, the robot body related information includes the posture information of the tool end coordinate system at the current moment, the joint position of the gripper, the action output at the previous moment, and whether there is a grasping detection result under the current viewing angle; the robot body related information is processed through a linear layer with a bias to obtain the robot body features.
[0026] Furthermore, in step S2, in the multi-stage reward function, the proximity reward function is used to reduce the distance to the target object to be grasped, the grasp contact reward function is used to improve the firmness of the grasp and is triggered only when the gripper contacts the target object, and the target reward function is used to increase the probability of successfully grasping the object. The formula of the multi-stage reward function is as follows:
[0027] r=r dis +r contact +r lift
[0028] Among them, r dis Represents the proximity reward, which is represented by the distance between the current tool end coordinate system and the grasping point information; r contact Represents the grasping contact reward, which is expressed by the numerical value obtained when the gripper contacts the object and is set to a constant value; r lift represents the target reward, that is, the reward for reaching the target position, which is expressed by calculating the Euclidean distance between the object position and the target position;
[0029] In a deep reinforcement learning network, in addition to the multi-stage reward function, a reward of a first set value is also given when a grasping state or a stable state is detected, and when the grasping is completed, a reward of a second set value is obtained, where the second set value is greater than the first set value.
[0030] Furthermore, in step S2, the deep reinforcement learning network is implemented based on the improved SAC algorithm. Compared with the original SAC algorithm, the improved SAC algorithm has the following improvements: an additional grasping target prediction network and a grasping stage prediction network are introduced into both the action Actor network and the value function Critic network to accelerate the learning effect during the training process. Both the grasping target prediction network and the grasping stage prediction network are constructed using multiple fully connected layers, wherein the grasping target prediction network is used to predict the optimal grasping method, and the grasping stage prediction network is used to predict the grasping stage;
[0031] The loss function of the deep reinforcement learning network during training is as follows:
[0032]
[0033] Among them, L SACrepresents the loss function of the deep reinforcement learning network corresponding to the original SAC algorithm; L grasp_target and L grasp_stage is the loss function of the grasp target prediction network and the grasp phase prediction network; α and β are the corresponding weight coefficients, γ is the discount factor, and t represents the current round in the training process.
[0034] Furthermore, in step S3, the decision action obtained by the deep reinforcement learning network includes the gripper control signal w g and the tool end feed amount (Δx, Δy, Δz, Δα, Δβ, Δγ), where Δx, Δy, Δz, Δα, Δβ, and Δγ represent the translation and rotation angle along the xyz coordinate axis in the current tool end coordinate system, respectively.
[0035] Furthermore, in step S3, the gripper control signal w g The size of the jaw opening and closing is the output of the jaw movement. The output of the robot body movement is solved by the tool end feed (Δx, Δy, Δz, Δα, Δβ, Δγ). The solution process is as follows:
[0036] S31: Transform the tool end feed amount (Δx, Δy, Δz, Δα, Δβ, Δγ) to the base coordinate system as the reference system to obtain the Cartesian target pose information of the tool end. The formula is as follows:
[0037]
[0038] Among them, the tool end feed (Δx, Δy, Δz, Δα, Δβ, Δγ) is the target pose relative to the tool end coordinate system, which is expressed by the homogeneous transformation matrix Indicates that the transformation relationship from the current tool end coordinate system to the base coordinate system is used The transformation relationship of the target pose relative to the base coordinate system is obtained through matrix transformation
[0039] S32: Obtain the current joint angle through the joint status upload interface, and then solve it through the inverse kinematics solution interface to obtain the corresponding target joint angle;
[0040] S33: Finally, the target joint angle is passed through the underlying joint motion planning to obtain the motion output of the robot body.
[0041] Compared with the prior art, the present invention has the following advantages:
[0042] 1. Compared with the 6DoF grasping combined with path planning method, the method of the present invention can dynamically adjust its own motion and has a certain robustness to environmental changes.
[0043] 2. The present invention combines the grasping points as the feature information of the object. On the one hand, it fully draws on the existing mature unknown object grasping methods to improve the efficiency of deep reinforcement learning. On the other hand, the perception modal data is in the form of point cloud, which has smaller differences in the process of migration from simulation to reality, reducing the difficulty of migration.
[0044] 3. The algorithm structure of the present invention is clear, which is conducive to the development of a more intelligent grasping operation model. It has certain engineering advantages in terms of the demand and processing of observation data and the realization of action execution. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 3D visual reinforcement learning-based dynamic object grasping method according to an embodiment of the present invention;
[0046] Figure 2 is a schematic diagram of the structure of the GraspGroupNet network in an embodiment of the present invention;
[0047] Figure 3 is an example of a local point cloud of a scene obtained from the perspective of a camera in an embodiment of the present invention;
[0048] Figure 4 is an example of a local point cloud of a scene in a base coordinate system according to an embodiment of the present invention;
[0049] Figure 5 is an example of a local point cloud of an object in a base coordinate system according to an embodiment of the present invention;
[0050] Figure 6 is an example of a local point cloud of an object in the tool end coordinate system in an embodiment of the present invention;
[0051] Figure 7 This is an example of a local point cloud of an object after parallel FPS downsampling in an embodiment of the present invention;
[0052] Figure 8 2. It is a structural diagram of LoG feature extraction based on a local object point cloud 6DoF grasping and detection model in an embodiment of the present invention;
[0053] Figure 9 This is a structural diagram of the deep reinforcement learning algorithm SAC training based on the actor-critic framework in an embodiment of the present invention;
[0054] Figure 10 : This is a trend chart of specific indicators of the training process in an embodiment of the present invention, where (a) is the trend chart of the cumulative reward in each round, (b) is the change in success rate during training, (c) is the change in action strategy loss, and (d) is the change in evaluation value function loss. DETAILED DESCRIPTION
[0055] The following is a detailed description of an embodiment of the present invention. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process. However, the protection scope of the present invention is not limited to the following embodiment.
[0056] Example 1
[0057] like Figure 1 As shown, this embodiment provides a technical solution: a dynamic object grasping method based on 3D visual reinforcement learning, which mainly includes the following steps:
[0058] Step 1: Collection and processing of observation information
[0059] The depth camera collects visual information, converts the visual information into point cloud data, performs grasping point detection based on the point cloud data, obtains the grasping point information of the target object, and then extracts features from the grasping point information to obtain visual grasping features. At the same time, the robot body related information is obtained and processed to obtain the robot body features. The visual grasping features and the robot body features are observation information, which are spliced together as the state information of the deep reinforcement learning network.
[0060] In this step, to ensure the consistency and convenience of the sim2real algorithm migration, the observation information corresponding to the agent in the simulation and real scenarios includes the following two parts.
[0061] On the one hand, feature extraction of visual perception information involves converting the point cloud acquired from the first-person perspective of a camera (RGB-D camera, depth camera) to a base coordinate system. This converts the local point cloud information of the object into a base coordinate system, performs downsampling and clustering, and then performs 6DoF grasp detection based on the local point cloud to obtain grasp point information. In real-world environments, to meet the requirements for grasp point detection in dynamic environments, a more robust processing method is used during the deployment phase. Furthermore, a number of three-dimensional points (x, y, z) are sampled at the origin of the tool end coordinate system to represent the spatial position of the gripper. This sampling method uses a Gaussian random distribution with parameters: mean 0 and variance one-third of the maximum opening and closing size of the gripper. A total of K positions (Gaussian points) are collected, corresponding to a data dimension of (K, 3). The grasp point information obtained by detection (assuming the number is N) is converted to the robot tool end coordinate system because the end coordinate system setting in the grasp detection results obtained by the model LoG is inconsistent with the robot tool end coordinate system. This data is then saved as a homogeneous transformation matrix (4x4), corresponding to a data dimension of (N, 4, 4). The collected Gaussian points are then transformed in sequence, and the corresponding tensor information dimension is (N, K, 3). The Gaussian points (K, 3) are combined as the input (N+1, K, 3) of the subsequent grasping point hierarchical network (GraspGroupNet) to extract the high-dimensional feature information of the grasping points. The specific implementation method is shown in the GraspGroupNet network design. The final visual grasping feature dimension information is 256.
[0062] In this embodiment, the process of obtaining the grab point information is as follows:
[0063] 1. Point cloud acquisition and preprocessing
[0064] The visual information is obtained by the depth camera in combination with the internal and external parameters of the camera, and the 2D depth information is converted into point cloud data with three-dimensional spatial information, such as Figure 3 shown.
[0065] The transformation relationship between object pixel coordinates and world coordinates:
[0066]
[0067] Among them, the internal parameter matrix K is:
[0068]
[0069] Camera extrinsic matrix
[0070]
[0071] Through camera calibration and hand-eye calibration methods, confirm the camera intrinsic parameter matrix K and extrinsic parameter matrix According to the above transformation, the depth map parameters (u, v, z c ) can be used to derive the local point cloud data of the scene and convert it into point cloud data in the world coordinate system.
[0072] The present invention adopts the eye-in-hand setting method. The point cloud obtained by the camera is relative to the camera optical reference system, and is converted to the end reference system through the calibration relationship between the camera and the tool end. Then, the transformation relationship between the tool end and the base coordinate system is obtained by using the robot's forward kinematics, and the point cloud is transformed into the base coordinate system, such as Figure 4 shown.
[0073] Before starting to grasp, the relative relationship between the object and the robot remains unchanged. The local point cloud data of the scene is transformed into the base coordinate system. The local point cloud data of the object is filtered out by simple methods such as combining the height threshold, such as Figure 5 shown.
[0074] The acquired local point cloud of the object is transformed into the tool end reference system, such as Figure 6 As shown in , the downsampling operation is then performed, using the parallel farthest point sampling (FPS) algorithm, which is limited to 84 points; then based on density clustering, the downsampled point cloud (such as Figure 7 As shown in Figure 3, spatial clustering (Density-Based Spatial Clustering of Applications with Noise, DBSCAN) is performed to filter outlier noise points.
[0075] The present invention adopts a 6DoF grasping detection model LoG based on local object point cloud, and the corresponding point cloud coding structure of the model is as follows: Figure 8 As shown, similar to the PointMLP network structure, it extracts local high-dimensional feature information and then uses a multi-head module to directly predict grasp parameters. (This model has fast inference output and good grasp detection effect)
[0076] Specifically, the present invention uses the above-mentioned trained model for reasoning, and the corresponding reasoning model function interface input includes the scene local point cloud data with the robot end as the reference system and the above-mentioned filtered object local point cloud data, as well as the set object grasping center point detection result constraint (the default is 64), and directly predicts the grasping detection result relative to the end coordinate system preset in the model LoG: the grasping point information and the corresponding grasping quality score, and then transforms according to the transformation relationship between the end coordinate system preset in the model LoG and the robot tool end coordinate system to obtain the grasping point information relative to the robot tool end coordinate system. The grasping point information includes the grasping center point coordinates and the grasping posture relative to the robot tool end coordinate system.
[0077] The design and implementation of the GraspGroupNet grasp point feature compression encoding network takes into account the details between grasp points (such as grasping posture) and uses a hierarchical approach to further extract and analyze the characteristic information between grasping methods. During network training, data is processed in batches, loading multiple samples at a time. This corresponds to the first dimension of the data, denoted by B.
[0078] The structure of the network includes Grasp Group and Group modules. The designs of these two data encoding networks (Grasp Group and Group modules) both include two groups (one-dimensional convolution, layer normalization, activation function), followed by a maximum pooling operation. The data input of the Grasp Group module is the data after the fusion of the grasping point and the Gaussian point, that is, the transformation result of the Gaussian point, and the corresponding dimension size is (B, N+1, K, 3); then the output result is spliced with the N grasping center points again as the input of the Group module in the next stage. Finally, the maximum pooling operation is performed on the output data channel, and then the final data dimension of the network becomes (256 in this embodiment). The corresponding network diagram is as follows Figure 2 As shown in the figure, the fintra part is the above-mentioned GraspGroup module, and the finter part is the above-mentioned Group module.
[0079] Observational information, on the other hand, primarily involves information about the robot itself. This includes the current pose of the tool end coordinate system (rotations are expressed in Euler angles), the gripper's joint positions, the previous action output (including the target pose relative to the end coordinate system and the gripper's opening and closing position), and the presence or absence of a grasp detection result from the current viewing angle, totaling 20 dimensions. This information is then passed through a biased linear layer, resulting in a feature output with a dimension of 128.
[0080] The processed feature information of the two is spliced together, totaling 256+128 dimensions, which serves as the state information of the deep reinforcement learning algorithm.
[0081] Step 2: Deep reinforcement learning network design and processing
[0082] The state information from step 1 is input into a deep reinforcement learning network for processing to obtain a decision action. A multi-stage reward function is designed in the deep reinforcement learning algorithm, which includes a proximity reward function, a grasping contact reward function, and a target reward function. The deep reinforcement learning network is implemented based on an improved SAC algorithm.
[0083] In this step, the present invention selects the SAC reinforcement learning algorithm, which is less sensitive to hyperparameters than other model-free reinforcement learning algorithms, especially the training rate and reward discount factor. The full name of the SAC algorithm is Soft Actor-Critic, which is a deep reinforcement learning algorithm that combines maximum entropy reinforcement learning and the Actor-Critic framework. Using finite discrete Markov decision as the reinforcement learning analysis framework, compared to traditional reinforcement learning objectives, the objective function of SAC not only includes the cumulative reward of traditional reinforcement learning, but also introduces a weighted term of policy entropy. The corresponding optimal policy solution formula is as follows:
[0084]
[0085] Among them, π represents the strategy, T represents the maximum number of steps for the agent to interact with the environment, and ρ π represents the probability of random strategy distribution, function r represents the reward function, α represents the entropy weight coefficient, and H represents the strategy corresponding entropy information solution function.
[0086] Entropy information is used to describe the randomness of the strategy and is defined as follows:
[0087]
[0088] This method not only improves the reinforcement learning strategy toward the task goal, but also enhances the agent's exploration capabilities during training. Furthermore, by using a dual Q-network, the smaller of the two network outputs is selected when evaluating the action-policy network, reducing the impact of maximization bias during training.
[0089] The SAC algorithm uses the actor-critic training framework and is approximated by a neural network, which contains an action strategy actor, with π φ Symbolic representation; and four critic network designs: a state-value network V ψ , a target state value network and two action-value networks and
[0090] The corresponding network loss function solution formula is derived as follows:
[0091] For the state value network parameter update, the corresponding optimized loss function is defined as follows:
[0092]
[0093] Target state value network parameters Update: Where τ is the accounting coefficient.
[0094] Action-Value Network and The corresponding loss function is as follows:
[0095]
[0096] Among them, Q(s t ,a t )=r(s t ,a t )+γV ψ (s t+1 ).
[0097] For the actor policy network π φ , the corresponding loss function is as follows:
[0098]
[0099] In practical applications, in order to better balance exploration and utilization, the parameter α can also be adaptively updated, corresponding to the loss function:
[0100]
[0101] Where H0 is the initial value of entropy.
[0102] Training framework such as Figure 9 As shown in Figure 2, for each network parameter optimization, the gradient descent method is used to update the network parameters, where the cumulative reward is estimated using the minimum value of the two action value networks.
[0103] The implementation of the deep reinforcement learning algorithm SAC in the present invention is based on the Stable_baseline3 framework and combined with the target auxiliary network design, that is, additional network outputs are added to the Actor and Critic network designs. On the one hand, in the design of the action Actor network, in the last layer of the policy network (composed of two layers of perception [256,256]), a linear layer is used to predict the current stage and a linear layer is also constructed to predict the optimal grasping posture parameters (9-dimensional information: the first two columns of the rotation matrix and the corresponding xyz values). On the other hand, in the Critic network design, the input of the model includes feature extraction and action information of the state observation. Here, the design of the additional network only uses the feature information of the observation data, similar to the above-mentioned construction method, to predict the corresponding parameters. During the training process, the corresponding network updates are optimized together, and the weights of the corresponding loss function are analyzed as follows.
[0104] In the embodiment, a hierarchical reward function is designed in the deep reinforcement learning algorithm SAC, as follows:
[0105] In the grasping task, it can be regarded as a multi-stage task including approaching the object, grasping the object, and reaching the target position, so a multi-stage reward function design is adopted.
[0106] In simulation training, the position, posture, linear velocity, and angular velocity of an object can be directly obtained through the physical simulation interface. For multi-link robots, in addition to this information, it also contains the position, velocity, and acceleration information of each movable joint. In addition, the physical scene interface can also be used to obtain all contact point information. By determining whether the robot end gripper and the grasped object are included, the angle between the contact direction and the gripper's movement direction is calculated, which is used as a flag for grasping (is_obj_grasp). When the gripper between the two is greater than the set threshold max_angle, it is determined to be in the object grasping state.
[0107] To determine whether the grasping task is successful (is_success), in addition to the above, it is also necessary to check whether the robot and the object are in a stable state and whether the object is placed in the target position.
[0108] is_success=int(is_robot_static*is_obj_grasp*is_obj_static*is_obj_lift)
[0109] The robot's stability (is_robot_static) is determined by whether the angular velocity of the two joints of the robot's end gripper is greater than a set value. Similarly, the object's stability (is_obj_static) is determined by comparing its RMS velocity. Whether the object is lifted (is_obj_lift) is determined by the distance between the object's position and the target point.
[0110] In terms of visual information, if the agent, in its current state, determines from the first-person perspective that a suitable grasping method exists through 6DoF grasp detection based on a local point cloud, the is_info_exist flag is set to true, and a certain reward is also obtained to guide the agent toward the appropriate 6DoF grasping method. Furthermore, the present invention uses 6DoF grasp detection to guide autonomous grasping through reinforcement learning, thereby calculating the corresponding additional reward based on the detected grasping method. This is specifically achieved by filtering the current grasping posture from the loaded object-referenced grasping dataset based on its current perspective and its positional relationship with the object itself; and converting the resulting grasping posture to the tool end coordinate system.
[0111] The filtered grasping posture is transformed into the "target" of the robot's grasping motion through the positional relationship between the object and the robot's tool end, and the relative distance is solved in combination with the current tool end posture:
[0112]
[0113] in, Represents the grasping pose Grasp the pose, the corresponding normalized position and quaternion rotation parameters, similar to (p EE ,q EE ) represents the pose parameters of the robot tool end coordinate system.
[0114] Specifically, the approach reward (approach_reward) is designed based on a formula to guide the agent to reduce the distance to the grasp target; the grasp contact reward (grasp_reward) is set to improve the firmness of the grasp; and the goal reward (goal_reward) is used to increase the probability of successful grasping. Combined with the aforementioned rewards, certain rewards are also given when the grasp state and stable state are detected, and a larger reward is given when the grasp completion indicator is met.
[0115] Table 1 Multi-stage reward calculation and numerical calculation method
[0116]
[0117] The first column of the table represents different aspects of the reward function. Their meanings are described in the fourth column, while the specific numerical solutions are described in the third column. The total reward is calculated by adding these rewards together, keeping the weights consistent. The specific parameters are described in the second column. It's worth noting that some of these phased rewards are triggered by the criteria described above.
[0118] In this embodiment, in the SAC reinforcement learning algorithm, for the specific network design of the Actor and Critic, the present invention integrates the interface in the original framework, and both use two hidden layers [256,256] as the encoding of the state information. For the Actor, the subsequent action prediction is to sample the probability of the corresponding distribution in the Gaussian distribution (including the mean and ln-type variance parameters), and perform action sampling under this distribution. The connection between the encoding and the mean parameter prediction in the action distribution is determined by a fully connected layer, which directly maps the 256-dimensional information to the mean parameter output consistent with the size of the action space, and the logarithmic ln variance parameter is directly transformed through a matrix transformation from the 256-dimensional feature information to the parameter output consistent with the dimension of the action space.
[0119] The Critic incorporates two design concepts for value functions. For each value function, since its corresponding result is a constant, a multilayer perceptron network with two hidden layers is used. Its input is the sum of the state information dimension and the action dimension, and its output is a single dimension. This network, which describes the prediction of the value function, takes input that includes not only state information but also action information, consistent with the principle.
[0120] Furthermore, before the introduction of grasp point feature extraction based on visual information, reinforcement learning remains challenging for the current grasping task setting. Drawing on the auxiliary network design introduced by GA-DDPG, this paper introduces additional grasp target prediction networks and grasp phase prediction networks into both the action actor and value function critic network designs, accelerating the learning of the agent during training. The corresponding network designs are constructed using a similar approach using fully connected layers, with a 9-dimensional prediction information for the grasp target and a 4-dimensional prediction information for the grasp evaluation phase.
[0121] During the network training phase, in addition to calculating the error in the reinforcement learning network output using the aforementioned loss function calculation methods, the auxiliary network's corresponding error is calculated using additional observations of the actual state and detection results. The weights for the latter two error calculations (α and β set to 1) are related to the discount factor (γ set to 0.98) and the current training time.
[0122] Overall, the network's loss function is as follows:
[0123]
[0124] In summary, the hyperparameters related to SAC deep reinforcement learning training are shown in Table 2 below:
[0125] Table 2 Hyperparameters related to deep reinforcement learning training
[0126]
[0127] Step 3: Deploy reinforcement learning algorithms and issue decision actions
[0128] The real machine deployment of the corresponding algorithm of the present invention mainly includes the following aspects: a visual perception unit that uses an RGB-D camera to obtain point cloud information of the target object, a motion control unit composed of a collaborative robot and a two-finger gripper, and a decision-making unit composed of a deep reinforcement learning algorithm that issues motion commands.
[0129] The perception process of the visual perception unit is described in detail in step 1 "collection and processing of observation information".
[0130] In a real-world environment, the point cloud processing method for moving objects (target objects) is as follows. First, the currently acquired scene point cloud information is transformed to the base coordinate system, using the camera-to-tool end and tool end-to-base coordinate transformations. Based on the limited workspace range and height, extra point clouds, including those on the desktop, are removed. The point cloud is then downsampled using FPS point cloud sampling technology. Second, the grasping point information acquired in previous observations is transferred to the base coordinate system and similarly downsampled. It is then clustered with the current scene point cloud (ball-query). Further downsampling is performed to remove extra point clouds, and the processed point cloud is considered the point cloud region of the target object. Finally, the two point cloud data are concatenated and further downsampled using a spatial clustering algorithm. This is used as the local center sampling point of the moving target object, which is then transformed to the tool end coordinate system as input to the LoG inference model for 6DoF grasp pose detection of the target object.
[0131] The motion control unit is composed of a collaborative robot and a two-finger gripper.
[0132] For collaborative motion control, remote access is achieved through IP. The ctypes module is used in Python programming, combined with the SDK, to encapsulate and implement the following interfaces: joint status upload, current robot end position acquisition, inverse kinematics solution, and joint space control.
[0133] The two-finger gripper's communication method uses a 485 interface at the physical layer, UART serial communication at the data layer, and Modbus RTU protocol for data transmission at the application layer. The gripper's opening and closing control and status reading are further encapsulated in Python programming using the minimalmodbus module, which reads and writes data to corresponding registers, including control registers, force control registers, and current position registers.
[0134] The decision-making unit is used to use the above-mentioned deep reinforcement learning algorithm to obtain decision actions and issue them.
[0135] The action output of the reinforcement learning algorithm includes: gripper control and collaborative robot motion planning. Among them, the gripper control signal w g , which indicates the size of the gripper opening and closing; for collaborative robots, it is the tool end feed (Δx, Δy, Δz, Δα, Δβ, Δγ), which respectively indicate the translation and rotation angle along the xyz coordinate axis in the current tool end coordinate system.
[0136] In this embodiment, the motion output of the robot body is implemented by the following conversion method: first, the motion is transformed into the Cartesian target posture information of the tool end with reference to the base coordinate system; then, the current joint angle is obtained through the interface designed above, and then solved through the inverse kinematics solution interface to obtain the corresponding target joint angle information; finally, the target joint angle is planned through the underlying joint motion to realize the motion control of the robot body.
[0137] Example 2
[0138] In this embodiment, a six-degree-of-freedom collaborative robot of the CI-05 model produced by Jicui Intelligent Manufacturing Co., Ltd. is used, and a gripper of the LMG-90 model of Lebai Robotics is matched at the end as an experimental platform for simulation environment and real machine deployment.
[0139] The method of the present invention constructs a grasping environment under the SAPIEN physics engine and combines the RL algorithm to realize training and evaluation in a simulation environment.
[0140] In the simulation, the specific indicator trend diagram of the training process is as follows Figure 10 As shown, the simulation test results are shown in Table 3 below. Through the simulation environment construction and algorithm comparison results verification, it is proved that the method of the present invention has a certain effect, and has good generalization and consistency for different objects and different motion modes. Through experimental demonstration, it is found that compared with directly using the object's 3D point cloud as the observation information for reinforcement learning (the success rate is very poor), the present invention combines the grasping point as the object's feature information and the perception modal data (robot body features) as observation information, can dynamically adjust its own motion, has a certain robustness to environmental changes, and has a higher grasping success rate.
[0141] Table 3 Simulation test results on the CI-05 robot body
[0142]
[0143] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A dynamic object grasping method based on 3D visual reinforcement learning, characterized in that: The following steps are involved: S1: Collection and processing of observation information The depth camera collects visual information, converts the visual information into point cloud data, performs grasping point detection based on the point cloud data, obtains the grasping point information of the target object, and then extracts features from the grasping point information to obtain visual grasping features. At the same time, the robot body related information is obtained and processed to obtain the robot body features. The visual grasping features and the robot body features are observation information, which are spliced together as the state information of the deep reinforcement learning network. S2: Deep reinforcement learning network design and processing The state information in step S1 is input into the deep reinforcement learning network for processing to obtain the decision action; a multi-stage reward function is designed in the deep reinforcement learning network, which includes a proximity reward function, a grasping contact reward function, and a target reward function; S3: Robot body and gripper motion output Based on the decision-making actions obtained by the deep reinforcement learning network, the action output of the robot body and the action output of the gripper are obtained to complete the task of grasping the target object.
2. The dynamic object grasping method based on 3D visual reinforcement learning according to claim 1 is characterized in that: In step S1, the process of obtaining the grasping point information is as follows: S101: Acquire visual information through a depth camera, and convert the visual information into point cloud data with three-dimensional spatial information based on the transformation relationship between pixel coordinates and world coordinates, that is, obtain local point cloud data of the scene in the camera coordinate system; S102: transforming the scene local point cloud data in the camera coordinate system into the base coordinate system to obtain the scene local point cloud data in the base coordinate system, wherein the base coordinate system is the coordinate system where the robot base is located; S103: Before starting to grasp, the relative relationship between the target object and the robot remains unchanged, and the local point cloud data of the object is filtered out from the local point cloud data of the scene in the base coordinate system by setting a height threshold; S104: transforming the local point cloud data of the object into the robot tool end coordinate system, and processing it using a parallel FPS algorithm and a spatial clustering algorithm to remove noise, thereby obtaining the noise-reduced local point cloud data of the object; S105: The scene local point cloud data in step S102, the object local point cloud data in step S104, and the object local point cloud grasping center point detection result constraints are used as input information of the local object point cloud 6DoF grasping detection inference model LoG, and the grasping detection results relative to the terminal coordinate system preset in the model LoG are obtained: the grasping point information of the target object and the corresponding grasping quality score, and then the grasping point information relative to the robot tool terminal coordinate system is obtained according to the transformation relationship between the terminal coordinate system preset in the model LoG and the robot tool terminal coordinate system. The grasping point information includes the grasping center point coordinates and the grasping posture relative to the robot tool terminal coordinate system.
3. The dynamic object grasping method based on 3D visual reinforcement learning according to claim 2 is characterized in that: In step S1, the process of acquiring visual capture features is as follows: S106: Using the grasping point information relative to the robot tool end coordinate system as a transformation matrix, the Gaussian points sampled at the origin of the tool end coordinate system are sequentially transformed to obtain corresponding transformation results; S107: Input the transformation result and the Gaussian point into the grasping point feature compression coding network GraspGroupNet for processing to obtain high-dimensional feature information of the grasping point, namely the visual grasping feature.
4. The dynamic object grasping method based on 3D visual reinforcement learning according to claim 3 is characterized in that: In step S106, K three-dimensional points are sampled at the origin of the tool end coordinate system to represent the spatial position of the clamping jaws. The K three-dimensional points are Gaussian points.
5. The dynamic object grasping method based on 3D visual reinforcement learning according to claim 4 is characterized in that: In step S107, the grasping point feature compression coding network GraspGroupNet includes a Grasp Group module and a Group module. In the Grasp Group module, the transformation result of the Gaussian point is processed as input, and then the output result is spliced with the grasping center point coordinates in the grasping point information as the input of the Group module. After processing by the Group module, a maximum pooling operation is performed on the output data channel to obtain high-dimensional feature information of the grasping point, that is, the visual grasping feature.
6. The dynamic object grasping method based on 3D visual reinforcement learning according to claim 3, characterized in that: In step S1, the robot body related information includes the posture information of the tool end coordinate system at the current moment, the joint position of the gripper, the action output at the previous moment, and whether there is a grasping detection result under the current viewing angle; the robot body related information is processed through a linear layer with a bias to obtain the robot body features.
7. The dynamic object grasping method based on 3D visual reinforcement learning according to claim 6, characterized in that: In step S2, in the multi-stage reward function, the proximity reward function is used to reduce the distance to the target object, the grasping contact reward function is used to improve the firmness of the grasp and is triggered only when the gripper contacts the target object, and the target reward function is used to increase the probability of successfully grasping the object. The formula of the multi-stage reward function is as follows: r=r dis +r contact +r lift Among them, r dis Represents the proximity reward, which is represented by the distance between the current tool end coordinate system and the grasping point information; r contact Represents the grasping contact reward, which is expressed by the numerical value obtained when the gripper contacts the object and is set to a constant value; r lift represents the target reward, that is, the reward for reaching the target position, which is expressed by calculating the Euclidean distance between the object position and the target position; In a deep reinforcement learning network, in addition to the multi-stage reward function, a reward of a first set value is also given when a grasping state or a stable state is detected, and when the grasping is completed, a reward of a second set value is obtained, where the second set value is greater than the first set value.
8. The dynamic object grasping method based on 3D visual reinforcement learning according to claim 7, characterized in that: In step S2, the deep reinforcement learning network is implemented based on the improved SAC algorithm. Compared with the original SAC algorithm, the improved SAC algorithm has the following improvements: an additional grasping target prediction network and a grasping stage prediction network are introduced into both the action Actor network and the value function Critic network to accelerate the learning effect during the training process. Both the grasping target prediction network and the grasping stage prediction network are constructed using multiple fully connected layers, wherein the grasping target prediction network is used to predict the optimal grasping method, and the grasping stage prediction network is used to predict the grasping stage; The loss function of the deep reinforcement learning network during training is as follows: Among them, L SAC represents the loss function of the deep reinforcement learning network corresponding to the original SAC algorithm; L grasp_target and L grasp_stage is the loss function of the grasp target prediction network and the grasp phase prediction network; α and β are the corresponding weight coefficients, γ is the discount factor, and t represents the current round in the training process.
9. A dynamic object grasping method based on 3D visual reinforcement learning according to claim 1 or 8, characterized in that: In step S3, the decision action obtained by the deep reinforcement learning network includes the gripper control signal w g and the tool end feed amount (Δx, Δy, Δz, Δα, Δβ, Δγ), where Δx, Δy, Δz, Δα, Δβ, and Δγ represent the translation and rotation angle along the xyz coordinate axis in the current tool end coordinate system, respectively.
10. The dynamic object grasping method based on 3D visual reinforcement learning according to claim 8, characterized in that: In step S3, the gripper control signal w g The size of the jaw opening and closing is the output of the jaw movement. The output of the robot body movement is solved by the tool end feed (Δx, Δy, Δz, Δα, Δβ, Δγ). The solution process is as follows: S31: Transform the tool end feed amount (Δx, Δy, Δz, Δα, Δβ, Δγ) to the base coordinate system as the reference system to obtain the Cartesian target pose information of the tool end. The formula is as follows: Among them, the tool end feed (Δx, Δy, Δz, Δα, Δβ, Δγ) is the target pose relative to the tool end coordinate system, which is expressed by the homogeneous transformation matrix Indicates that the transformation relationship from the current tool end coordinate system to the base coordinate system is used The transformation relationship of the target pose relative to the base coordinate system is obtained through matrix transformation S32: Obtain the current joint angle through the joint status upload interface, and then solve it through the inverse kinematics solution interface to obtain the corresponding target joint angle; S33: Finally, the target joint angle is passed through the underlying joint motion planning to obtain the motion output of the robot body.