Robot grabbing method and system based on deep reinforcement learning
By constructing Markov decision model and designing reward functions, combining deep neural networks and dual replay cache mechanism, the problem of insufficient decision-making ability and low learning efficiency of robot crawling methods in complex environments is solved, and the efficient crawling task of robots in unstructured environments is realized.
Patent Information
- Application Number
- CN202510232858.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The mobile object robot grasping method based on reinforcement learning has insufficient independent decision-making ability in complex environments, making it difficult to cope with environmental uncertainty and randomness. The sample utilization rate during training is low, resulting in low learning efficiency.
Using a robot crawling method based on deep reinforcement learning, by constructing a Markov decision model, determining state space and action space, designing reward functions and dual replay cache mechanisms, and using deep neural networks to fit action value functions and strategy functions, to realize intelligent decision-making of robot mobile crawling.
Improve the autonomous decision-making ability of robot mobile grabbing in an unstructured environment, enhance environmental adaptability and learning efficiency, and realize efficient grabbing tasks for robots in complex environments.
Smart Images

Figure CN119973993A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robots and artificial intelligence technologies, and in particular to a robot grasping method and system based on deep reinforcement learning. Background Art
[0002] For the robot autonomous grasping technology of dynamic objects, the unstructured environment and the dynamic position of dynamic objects require more effective grasping decisions and a powerful visual system that interacts with dynamic objects in real time. Deep reinforcement learning (DRL) introduces deep neural networks (DNN) to solve the problem of robot dynamic object grasping. In the existing technology, some researchers divide dynamic grasping into two parts based on the dynamic grasping framework of reinforcement learning (RL): grasping strategy learning and trajectory prediction based on RL; some researchers learn the synergy between pushing and grasping strategies based on the CCA-MTFCN method of deep reinforcement learning, and propose a hard parameter sharing multi-task fully convolutional network (MTFCN) to model the action-value function to complete the task of removing all objects in a severely cluttered environment; some researchers extract state features based on the deep reinforcement learning model of the graph, and the graph reasoning module captures the internal relationship between features in different regions through the graph convolutional network to effectively explore invisible objects and improve the performance of collaborative grasping and pushing tasks; some researchers introduce prior knowledge information to optimize grasping actions and reduce the training time and interaction data required for assembly strategy learning algorithms; some researchers use domain randomization methods to periodically change the position and angle of objects in simulated scenes, thereby improving the diversity of training samples and the robustness of the algorithm; some researchers use the DDPG algorithm to automatically annotate moving objects and use Hindsight Experience Replay (HER) to improve the grasping efficiency value. However, there are some problems with the mobile object robot grasping method based on reinforcement learning, such as insufficient autonomous decision-making ability in complex environments and difficulty in coping with environmental uncertainty and randomness; low sample utilization during training, resulting in low learning efficiency, etc. Summary of the invention
[0003] The purpose of the present invention is to overcome the above-mentioned shortcomings and propose a robot grasping method and system based on deep reinforcement learning, which has strong autonomous decision-making ability, high environmental adaptability and high learning efficiency for robot mobile grasping in an unstructured environment.
[0004] The robot grasping method based on deep reinforcement learning of the present invention comprises the following steps:
[0005] Step 1: Install the robot and depth camera, initialize the environment model, and build a Markov decision model;
[0006] The constructed Markov decision model includes a state space, an action space, a reward function, a state transfer function, and a discount rate, which is represented by a five-tuple M=(S, A, p, r, γ), wherein s represents a set of possible states in the environment; A represents a set of possible actions in the environment; p: p(s'|s, a), p represents the probability of the following event: in the current state s, the robot performs action a, and the state of the environment changes to s'; r: r(s, a, s'), r is the reward function; γ is the discount rate;
[0007] Step 2: Set the control method of the robot, and determine the state space and action space in combination with the environment;
[0008] Determination of the state space: According to the position of the robot gripper (x t ,y t ,z t ) and the moving object coordinates (x t ',y t ',z t '), and the moving speed (v) of the moving object at time t is obtained by calculating the position at time t and time t-1. xt ',v yt ',v zt '), the robot gripper velocity (v xt ,v yt ,v zt ) is added to the state space, and the state space is defined as (x t ,y t ,z t ,v xt ,v yt ,v zt ,x t ',y t ',z t ',v xt ',v yt ',v zt ');
[0009] Determination of the action space: As the robot is a dual-arm robot, the action space is defined as a 12-dimensional space, that is, the joint motion speed of the dual arms. The action space is defined as (v l1 ,v l2 ,v l3 ,v l4 ,v l5 ,v l6 ,v r1 ,v r2 ,v r3 ,v r4 ,v r5 ,v r6 ), where v l1,v l2 ,v l3 ,v l4 ,v l5 ,v l6 The speed sequence of the robot's left arm, v r1 ,v r2 ,v r3 ,v r4 ,v r5 ,v r6 Execute the motion velocity sequence for the robot's right arm;
[0010] Step 3: Based on the deep reinforcement learning algorithm integrated with the dual replay cache mechanism, an improved deep reinforcement learning control strategy is provided for the robot grasping task. The specific steps include:
[0011] Step 3.1: Interaction between the robot and the environment: First, the interaction process between the robot and the environment is clarified, and the position coordinates of the moving object, the position of each joint of the robot, and the position of the gripper are obtained; then, the control strategy is trained based on the positive kinematics principle of the robot. In the initial stage of training, the robot's main purpose is to explore the environment space. After the robot runs stably, the main purpose is to stably grasp the object; after discovering the moving object, the control strategy controls the robot to gradually approach the moving object, and adjusts the robot's posture and position as needed to keep tracking; when the distance between the center of the gripper and the moving object is less than the preset threshold, the control strategy calculates the speed of each joint of the robot and sends a stop motion command to the robot, and then sends a grasping command to the gripper to grasp the moving object;
[0012] Step 3.2: Develop a training strategy: The goal of this strategy is to find the optimal strategy π * , so that the robot can obtain the maximum total reward after completing the task of grasping dynamic objects;
[0013] At the beginning of training, the action executed under the environment state is scored by the action value function, and its reward is used as an evaluation index, and the executed actions are stored; the action value function Q is:
[0014]
[0015] Among them, R(s,a,s') is the reward for executing action a in the current state s, resulting in the next state s', γ is the discount rate, α is the trade-off coefficient, and E calculates its expectation;
[0016] During the training process, set the target task and loss function, and continuously optimize the strategy π * The neural network parameters θ in are used to update the action value function, state value function, and double replay cache mechanism.
[0017] In the later stage of training, the actions stored in the cache are optimized according to the optimized parameters, and the action sequences that have obtained high rewards are used to conduct a new round of exploration of the environment to achieve adaptive dynamic grasping of the robot;
[0018] Step 3.3: Design reward function: Combine sparse reward with dense reward (dense reward: the reward obtained by the robot for each step of action; sparse reward: the reward obtained by the robot when it reaches the action or stage set according to the robot grasping task) to set the single-step reward r of reinforcement learning in the robot grasping process, expressed as:
[0019] r=r dense +r sparse
[0020] Among them, r dense is the intensive reward, r sparse For sparse rewards;
[0021] Step 3.4: Double replay cache mechanism: In reinforcement learning, through the double replay cache mechanism, the robot can learn from experience more efficiently and enhance the stability and efficiency of the learning process; the double replay cache mechanism is to set two replay cache areas A-memory and B-memory in the improved deep reinforcement learning control strategy, when the single-step reward is higher than the preset threshold ζ, the experience replay sample is stored in the A-memory memory cache area, otherwise it is placed in the B-memory memory cache area, and the two replay cache areas are sampled by dynamically adjusting the sampling probability to improve the efficiency of data use;
[0022] Step 4: Train the improved deep reinforcement learning control strategy and deploy it to the robot to perform the task of robot grasping moving objects.
[0023] The above-mentioned robot grasping method based on deep reinforcement learning, wherein: in step 3.2, the design target task is to assign tasks to the robot according to the information of the object to be grasped on the conveyor belt, and complete the robot's recognition, tracking, and grasping tasks of the moving object.
[0024] In the above-mentioned robot grasping method based on deep reinforcement learning, in step 3.2, the loss function is:
[0025]
[0026] Among them, in order to calculate the policy network π θ The loss is calculated using the state value function Q Φ1 、Action value function Q Φ1 The minimum value of *Under the maximum network parameter θ, the expectation is calculated. S is the action state storage sequence, D is the memory buffer space, N is the random sampling probability distribution, ξ is the random sampling probability. In order to emphasize that the action at the next moment is directly sampled from the environment by the policy network, the next moment action generated by a new round of sampling by the policy is defined as
[0027] In the above robot grasping method based on deep reinforcement learning, in step 3.2, the state value function is:
[0028]
[0029] Among them, U t is the sum of all single-step rewards from time t to the end of the round, E(A t ,S (t+1) …S n ,A n ) t Find the conditional expectation and eliminate the effect of the action on the state value.
[0030] The above robot grasping method based on deep reinforcement learning, wherein: in step 3.3, the dense reward r dense It is expressed as:
[0031] r dense =-w dis *l dis
[0032] Among them, l dis is the distance between the center point of the robot gripper and the grasping target point, w dis is the distance reward factor.
[0033] The above robot grasping method based on deep reinforcement learning, wherein: in step 3.3, the sparse reward r sparse , expressed as:
[0034] r sparse =r v +r g +r c +r t +r s
[0035] Among them, r v is the relative speed between the robot gripper and the object to be grasped; r g is the reward when the robot successfully grasps the object. In order to reduce the collision of the robot during the grasping process, r is set c is the reward when the robot collides with the environment, so as to optimize the robot's grasping strategy; set r tIn order to track the reward, the robot can be better guided to the correct grasping strategy and reduce the excessive time spent on exploring the action space in the early stage. During the training process, it is necessary to improve the grasping efficiency of the control algorithm to complete the robot grasping task and shorten the robot grasping time. For this purpose, a penalty term r based on the time step is added. s .
[0036] A robot grasping system based on deep reinforcement learning, comprising hardware facilities and software parts, wherein the hardware facilities include a robotic arm, a conveyor belt, and a depth camera, and the software part includes a visual detection module and a robot control algorithm module;
[0037] The visual detection module: firstly uses a depth camera to capture the RGB image and depth information of a randomly moving object on a workbench, identifies and determines the position of the moving object, and assigns the task to the left and right arms of the robot to perform the grasping task according to the type of the object; then, the position and speed data of the object are transmitted to the robot control algorithm module, which outputs the precise moving speed of each joint of the robot and guides the robot gripper to move toward the target object. When the gripper approaches the preset distance of the moving object, the gripper closes to complete the grasping action;
[0038] The robot control algorithm module controls the joint movement of the robot and the grasping operation of the gripper: first, the state of the moving object and the gripper is obtained, the speed of the gripper is calculated, and the robot is controlled to approach the target. The module controls the robot to grasp the moving object by outputting a 12-dimensional action space; in addition, the control module continuously interacts with the environment to obtain the action with the highest reward in the current state, so as to realize robot tracking and grasping of moving objects.
[0039] Compared with the prior art, the present invention has obvious beneficial effects. It can be seen from the above scheme that the present invention integrates the double replay buffer mechanism (DRB) and the reinforcement learning soft action-evaluation (SAC) algorithm to propose an improved deep reinforcement learning control strategy DRB-SAC. Through the construction of the Markov decision process model, the action space and state space are defined to provide a framework for the operation and observation environment, the objectives and constraints of the robot control task are determined, the training strategy and reward function are designed, and the deep neural network is used to fit the action value function and the strategy function to realize the intelligent decision-making of the robot mobile grasping and perform adaptive grasping. The dual playback cache mechanism of the present invention can improve the diversity of actual sampling trajectories during memory replay due to the rapid absorption of high-reward experience playback sampling. Two playback cache areas (A-memory, B-memory) are set. When the single-step reward (reward) is higher than the threshold ζ, the experience playback sampling is stored in the A-memory memory cache area, otherwise it is placed in the B-memory memory cache area, and the sampling probability is dynamically adjusted to achieve sampling of the two playback cache areas, thereby improving the efficiency of data use. In short, the present invention has the characteristics of strong autonomous decision-making ability of robot mobile grasping in unstructured environments, high environmental adaptability, and high learning efficiency.
[0040] The beneficial effects of the present invention are further illustrated below through specific implementation modes. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0042] The specific implementation methods, features and functions of the robot grasping method and system based on deep reinforcement learning proposed in the present invention are described in detail below in combination with the accompanying drawings and preferred embodiments.
[0043] See also Figure 1 , the robot grasping method based on deep reinforcement learning of the present invention, wherein: the method comprises the following steps:
[0044] Step 1: Install the robot and depth camera, initialize the environment model, and build a Markov decision model; this decision model is mainly used to solve the task coordination problem in the dual-arm robot collaborative environment. By observing the various task states and robot states of the environment, the optimal action that the decision model should perform in the current state and the degree of impact on subsequent tasks can be inferred based on the current state.
[0045] The constructed Markov decision model includes a state space, an action space, a reward function, a state transfer function, and a discount rate, which is represented by a five-tuple M=(S, A, p, r, γ), wherein s represents a set of possible states in the environment; A represents a set of possible actions in the environment; p: p(s'|s, a), p represents the probability of the following event: in the current state s, the robot performs action a, and the state of the environment changes to s'; r: r(s, a, s'), r is the reward function; γ is the discount rate;
[0046] Step 2: Set the control mode of the robot, combine it with the environment, observe the real-time state of the current environment (such as a depth camera), and determine the state space and action space;
[0047] Determination of the state space: According to the position of the robot gripper (x t ,y t ,z t ) and the moving object coordinates (x t ',y t ',z t '), and the moving speed (v) of the moving object at time t is obtained by calculating the position at time t and time t-1. xt ',v yt ',v zt '), the robot gripper velocity (v xt ,v yt ,v zt ) is added to the state space, and the state space is defined as (x t ,y t ,z t ,v xt ,v yt ,v zt ,x t ',y t ',z t ',v xt ',v yt ',v zt ');
[0048] Determination of the action space: As the robot is a dual-arm robot, the action space is defined as a 12-dimensional space, that is, the joint motion speed of the dual arms. The action space is defined as (v l1 ,v l2 ,v l3 ,v l4 ,v l5 ,v l6 ,v r1 ,v r2 ,v r3 ,v r4 ,v r5 ,vr6 ), where v l1 ,v l2 ,v l3 ,v l4 ,v l5 ,v l6 The speed sequence of the robot's left arm, v r1 ,v r2 ,v r3 ,v r4 ,v r5 ,v r6 Execute the motion velocity sequence for the robot's right arm;
[0049] Step 3: Based on the deep reinforcement learning algorithm integrated with the dual replay cache mechanism, an improved deep reinforcement learning control strategy is provided for the robot grasping task. The specific steps include:
[0050] Step 3.1: Interaction between the robot and the environment: First, the interaction process between the robot and the environment is clarified, and the position coordinates of the moving object, the position of each joint of the robot, and the position of the gripper are obtained; then, the control strategy is trained based on the positive kinematics principle of the robot. In the initial stage of training, the robot's main purpose is to explore the environment space. After the robot runs stably, the main purpose is to stably grasp the object; after discovering the moving object, the control strategy controls the robot to gradually approach the moving object, and adjusts the robot's posture and position as needed to keep tracking; when the distance between the center of the gripper and the moving object is less than the preset threshold, the control strategy calculates the speed of each joint of the robot and sends a stop motion command to the robot, and then sends a grasping command to the gripper to grasp the moving object;
[0051] Step 3.2: Develop a training strategy: The goal of this strategy is to find the optimal strategy π * , so that the robot can obtain the maximum total reward after completing the task of grasping dynamic objects;
[0052] At the beginning of training, the action executed under the environment state is scored by the action value function, and its reward is used as an evaluation index, and the executed actions are stored; the action value function Q is:
[0053]
[0054] Among them, R(s,a,s') is the reward for executing action a in the current state s, resulting in the next state s', γ is the discount rate, α is the trade-off coefficient, and E calculates its expectation;
[0055] During the training process, set the target task and loss function, and continuously optimize the strategy π *The neural network parameters θ in are used to update the action value function, state value function, and double replay cache mechanism.
[0056] The design target task is to assign tasks to the robot based on the information of the objects to be grasped on the conveyor belt, so as to complete the robot's tasks of identifying, tracking and grasping the moving objects.
[0057] The loss function is:
[0058]
[0059] Among them, in order to calculate the policy network π θ The loss is calculated using the state value function Q Φ1 、Action value function Q Φ1 The minimum value of * Under the maximum network parameter θ, the expectation is calculated. S is the action state storage sequence, D is the memory buffer space, N is the random sampling probability distribution, ξ is the random sampling probability. In order to emphasize that the action at the next moment is directly sampled from the environment by the policy network, the next moment action generated by a new round of sampling by the policy is defined as
[0060] The state value function is:
[0061]
[0062] Among them, U t is the sum of all single-step rewards from time t to the end of the round, E(A t ,S (t+1) …S n ,A n ) t Find the conditional expectation and eliminate the effect of the action on the state value.
[0063] In the later stage of training, the actions stored in the cache are optimized according to the optimized parameters, and the action sequences that have obtained high rewards are used to conduct a new round of exploration of the environment to achieve adaptive dynamic grasping of the robot;
[0064] Step 3.3: Design reward function: Combine sparse reward with dense reward (dense reward: the reward obtained by the robot for each step of action; sparse reward: the reward obtained by the robot when it reaches the action or stage set according to the robot grasping task) to set the single-step reward r of reinforcement learning in the robot grasping process, expressed as:
[0065] r=r dense +r sparse
[0066] Among them, rdense is the intensive reward, r sparse For sparse rewards;
[0067] The dense reward r dense It is expressed as:
[0068] r dense =-w dis *l dis
[0069] Among them, l dis is the distance between the center point of the robot gripper and the grasping target point, w dis is the distance reward factor.
[0070] The sparse reward r sparse , expressed as:
[0071] r sparse =r v +r g +r c +r t +r s
[0072] Among them, r v is the relative speed between the robot gripper and the object to be grasped; r g is the reward when the robot successfully grasps the object. In order to reduce the collision of the robot during the grasping process, r is set c is the reward when the robot collides with the environment, so as to optimize the robot's grasping strategy; set r t In order to track the reward, the robot can be better guided to the correct grasping strategy and reduce the excessive time spent on exploring the action space in the early stage. During the training process, it is necessary to improve the grasping efficiency of the control algorithm to complete the robot grasping task and shorten the robot grasping time. For this purpose, a penalty term r based on the time step is added. s .
[0073] Step 3.4: Double replay cache mechanism: In reinforcement learning, through the double replay cache mechanism, the robot can learn from experience more efficiently and enhance the stability and efficiency of the learning process; the double replay cache mechanism is to set two replay cache areas A-memory and B-memory in the improved deep reinforcement learning control strategy, when the single-step reward is higher than the preset threshold ζ, the experience replay sample is stored in the A-memory memory cache area, otherwise it is placed in the B-memory memory cache area, and the two replay cache areas are sampled by dynamically adjusting the sampling probability to improve the efficiency of data use;
[0074] The value of the threshold ζ is shown in Table 1.
[0075] Table 1 Dual playback cache mechanism sampling table
[0076]
[0077]
[0078] Step 4: Train the improved deep reinforcement learning control strategy and deploy it to the robot to perform the task of robot grasping moving objects.
[0079] A robot grasping system based on deep reinforcement learning, comprising hardware facilities and software parts, wherein the hardware facilities include a robotic arm, a conveyor belt, and a depth camera, and the software part includes a visual detection module and a robot control algorithm module;
[0080] The visual detection module: firstly uses a depth camera to capture the RGB image and depth information of a randomly moving object on a workbench, identifies and determines the position of the moving object, and assigns the task to the left and right arms of the robot to perform the grasping task according to the type of the object; then, the position and speed data of the object are transmitted to the robot control algorithm module, which outputs the precise moving speed of each joint of the robot and guides the robot gripper to move toward the target object. When the gripper approaches the preset distance of the moving object, the gripper closes to complete the grasping action;
[0081] The robot control algorithm module controls the joint movement of the robot and the grasping operation of the gripper: first, it obtains the state of the moving object and the gripper, calculates the speed of the gripper, and controls the robot to approach the target. This module controls the robot to grasp the moving object by outputting a 12-dimensional action space. In addition, the control module continuously interacts with the environment to obtain the action with the highest reward in the current state, so as to realize robot tracking and grasping of moving objects.
[0082] Performance Analysis:
[0083] 1. Physical Experiment Environment
[0084] In the physical experimental environment, a dual-arm UR3 robot equipped with a Robotiq 2F85 gripper and two Kinect V2.0 depth cameras mounted on a camera stand were used. The selected moving object was a regular cube paper box, and the moving object moved along the conveyor belt. The DRB-SAC algorithm model trained by the present invention was loaded into the virtual environment by running the control algorithm, and the robot arm joint coordinates at each time point were output. The algorithm model was completed and deployed on the real robot arm to perform related tasks.
[0085] 2 Experimental results and performance analysis
[0086] When the moving object follows the conveyor belt, the robot is controlled by the algorithm to grasp the moving object at each time coordinate point. When the object to be grasped moves, the speed command output by the control module controls the robot's gripper to continuously track the moving object. In this process, it continuously approaches the moving object and adjusts the gripper's gripping angle. When the distance between the two is less than or equal to the set gripping distance, the speed command becomes invalid, the speed of each robot joint is locked to 0, and the gripper closing command is executed.
[0087] Under the same environment and initialization conditions, 50 left and right arm grasping experiments were carried out using the TD3 and DRB-SAC control models respectively. The experimental results are shown in Table 2.
[0088] Table 2 Comparison of experimental results
[0089]
[0090] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still falls within the scope of the technical solution of the present invention.
Claims
1. A robot grasping method based on deep reinforcement learning, characterized in that: The method comprises the following steps: Step 1: Install the robot and depth camera, initialize the environment model, and build a Markov decision model; The constructed Markov decision model includes a state space, an action space, a reward function, a state transfer function, and a discount rate, which is represented by a five-tuple M=(S, A, p, r, γ), wherein s represents a set of possible states in the environment; A represents a set of possible actions in the environment; p: p(s'|s, a), p represents the probability of the following event: in the current state s, the robot performs action a, and the state of the environment changes to s'; r: r(s, a, s'), r is the reward function; γ is the discount rate; Step 2: Set the control mode of the robot, and determine the state space and action space in combination with the environment; Determination of the state space: According to the position of the robot gripper (x t ,y t ,z t ) and the moving object coordinates (x t ',y t ',z t '), and the moving speed (v) of the moving object at time t is obtained by calculating the position at time t and time t-1. xt ',v yt ',v zt '), the robot gripper velocity (v xt ,v yt ,v zt ) is added to the state space, and the state space is defined as (x t ,y t ,z t ,v xt ,v yt ,v zt ,x t ',y t ',z t ',v xt ',v yt ',v zt '); Determination of the action space: As the robot is a dual-arm robot, the action space is defined as a 12-dimensional space, that is, the joint motion speed of the dual arms. The action space is defined as (v l1 ,v l2 ,v l3 ,v l4 ,v l5 ,v l6 ,v r1 ,v r2 ,v r3 ,v r4 ,v r5 ,v r6 ), where v l1 ,v l2 ,v l3 ,v l4 ,v l5 ,v l6 The speed sequence of the robot's left arm, v r1 ,v r2 ,v r3 ,v r4 ,v r5 ,v r6 Execute the motion velocity sequence for the robot's right arm; Step 3: Based on the deep reinforcement learning algorithm integrated with the dual replay cache mechanism, an improved deep reinforcement learning control strategy is provided for the robot grasping task. The specific steps include: Step 3.1: Interaction between the robot and the environment: First, the interaction process between the robot and the environment is clarified, and the position coordinates of the moving object, the position of each joint of the robot, and the position of the gripper are obtained; then, the control strategy is trained based on the positive kinematics principle of the robot. In the initial stage of training, the robot's main purpose is to explore the environment space. After the robot runs stably, the main purpose is to stably grasp the object; after discovering the moving object, the control strategy controls the robot to gradually approach the moving object, and adjusts the robot's posture and position as needed to keep tracking; when the distance between the center of the gripper and the moving object is less than the preset threshold, the control strategy calculates the speed of each joint of the robot and sends a stop motion command to the robot, and then sends a grasping command to the gripper to grasp the moving object; Step 3.2: Develop a training strategy: The goal of this strategy is to find the optimal strategy π * , so that the robot can obtain the maximum total reward after completing the task of grasping dynamic objects; At the beginning of training, the action executed under the environment state is scored by the action value function, and its reward is used as an evaluation index, and the executed actions are stored; the action value function Q is: Among them, R(s,a,s') is the reward for executing action a in the current state s, resulting in the next state s', γ is the discount rate, α is the trade-off coefficient, and E calculates its expectation; During the training process, set the target task and loss function, and continuously optimize the strategy π * The neural network parameters θ in are used to update the action value function, state value function, and double replay cache mechanism. In the later stage of training, the actions stored in the cache are optimized according to the optimized parameters, and the action sequences that have obtained high rewards are used to conduct a new round of exploration of the environment to achieve adaptive dynamic grasping of the robot; Step 3.3: Design reward function: Combine sparse rewards with dense rewards to set the single-step reward r of reinforcement learning during the robot grasping process, expressed as: r=r dense +r sparse Among them, r dense is the intensive reward, r sparse For sparse rewards; Step 3.4: Double replay cache mechanism: In reinforcement learning, through the double replay cache mechanism, the robot can learn from experience more efficiently and enhance the stability and efficiency of the learning process; the double replay cache mechanism is to set two replay cache areas A-memory and B-memory in the improved deep reinforcement learning control strategy, when the single-step reward is higher than the preset threshold ζ, the experience replay sample is stored in the A-memory memory cache area, otherwise it is placed in the B-memory memory cache area, and the two replay cache areas are sampled by dynamically adjusting the sampling probability to improve the efficiency of data use; Step 4: Train the improved deep reinforcement learning control strategy and deploy it to the robot to perform the task of robot grasping moving objects.
2. The robot grasping method based on deep reinforcement learning as claimed in claim 1, characterized in that: In step 3.2, the design target task is to assign tasks to the robot based on the information of the object to be grasped on the conveyor belt, and complete the robot's tasks of identifying, tracking, and grasping the moving object.
3. The robot grasping method based on deep reinforcement learning as claimed in claim 1, characterized in that: In step 3.2, the loss function is: Among them, in order to calculate the policy network π θ The loss is calculated using the state value function Q Φ1 、Action value function Q Φ1 The minimum value of * Under the maximum network parameter θ, the expectation is calculated. S is the action state storage sequence, D is the memory buffer space, N is the random sampling probability distribution, ξ is the random sampling probability. In order to emphasize that the action at the next moment is directly sampled from the environment by the policy network, the next moment action generated by a new round of sampling by the policy is defined as 4. The robot grasping method based on deep reinforcement learning as claimed in claim 1, characterized in that: In step 3.2, the state value function is: Among them, U t is the sum of all single-step rewards from time t to the end of the round, E(A t ,S (t+1) …S n ,A n ) t Find the conditional expectation and eliminate the effect of the action on the state value.
5. The robot grasping method based on deep reinforcement learning as claimed in claim 1, characterized in that: In step 3.3, the dense reward r dense It is expressed as: r dense =-w dis *l dis Among them, l dis is the distance between the center point of the robot gripper and the grasping target point, w dis is the distance reward factor.
6. The robot grasping method based on deep reinforcement learning as claimed in claim 1, characterized in that: In step 3.3, the sparse reward r sparse , expressed as: r sparse =r v +r g +r c +r t +r s Among them, r v is the relative speed between the robot gripper and the object to be grasped; r g is the reward when the robot successfully grasps the object. In order to reduce the collision of the robot during the grasping process, r is set c is the reward when the robot collides with the environment, so as to optimize the robot's grasping strategy; set r t In order to track the reward, the robot can be better guided to the correct grasping strategy and reduce the excessive time spent on exploring the action space in the early stage. During the training process, it is necessary to improve the grasping efficiency of the control algorithm to complete the robot grasping task and shorten the robot grasping time. For this purpose, a penalty term r based on the time step is added. s .
7. A robot grasping system based on deep reinforcement learning using any one of the methods of claims 1 to 6, characterized in that: It includes hardware facilities and software parts. The hardware facilities include a robotic arm, a conveyor belt, and a depth camera. The software part includes a visual detection module and a robot control algorithm module. The visual inspection module: firstly uses a depth camera to capture the RGB image and depth information of a randomly moving object on the workbench, identifies and determines the position of the moving object, and assigns the task to the left and right arms of the robot for grasping tasks according to the type of the object; Subsequently, the position and speed data of the object are transmitted to the robot control algorithm module, which outputs the precise movement speed of each joint of the robot and guides the robot gripper to move toward the target object. When the gripper approaches the preset distance of the moving object, the gripper closes and the grasping action is completed. The robot control algorithm module controls the joint movement of the robot and the grasping operation of the gripper: first, the state of the moving object and the gripper is obtained, the speed of the gripper is calculated, and the robot is controlled to approach the target. The module controls the robot to grasp the moving object by outputting a 12-dimensional action space; in addition, the control module continuously interacts with the environment to obtain the action with the highest reward in the current state, so as to realize robot tracking and grasping of moving objects.
Citation Information
Patent Citations
Mechanical arm control method based on deep reinforcement learning
CN116533249A
Mechanical arm dynamic object grabbing method based on reinforcement learning
CN116945180A
Dynamic target rapid grabbing planning method and system based on deep reinforcement learning
CN117245666A
Sparse reward-oriented deep reinforcement learning mechanical arm grabbing method
CN118493388A
Cited By
3D vision automatic discharging system based on neural network adaptive learning
CN120588211A
Cooperative control method for double-arm robot
CN121157045A