A multi-robot collaborative control grasping method based on reinforcement learning technology
By adopting reinforcement learning technology in the multi-robot collaborative control system, a multi-agent deep Q network is built, which solves the problems of low collaboration efficiency and inflexible task allocation in multi-robot collaborative control, and achieves efficient and robust multi-robot collaborative grasping effect.
Patent Information
- Application Number
- CN202510398249.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing multi-robot collaborative control methods have problems such as low collaboration efficiency and inflexible task allocation, which affect the overall operating efficiency of the system.
The multi-robot collaborative control and grabbing method based on reinforcement learning technology is adopted. By establishing a multi-robot control mathematical model under a unified coordinate system, a greedy algorithm is used to build a task allocation model, and a multi-agent deep Q network is built for training, and a crawling control model is obtained to control the collaborative grasping action of multiple robots.
It realizes efficient collaborative crawling between multiple robots, improves the system's crawling efficiency and robustness, reduces manual dependence, reduces costs, and improves production efficiency and economic benefits.
Smart Images

Figure CN119910663B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and robotics technology, and in particular to a multi-robot collaborative control grasping method based on reinforcement learning technology. Background Art
[0002] With the continuous development of robotics technology, multi-robot systems have been widely used in industrial production, logistics distribution, environmental monitoring and other fields. However, there are many problems with the existing multi-robot collaborative control methods, such as: the collaboration efficiency between robots is low, and task conflicts are prone to occur; the task allocation is not flexible enough, and it cannot be dynamically adjusted according to the robot's own state, affecting the overall operation efficiency of the system. These problems seriously restrict the performance and application scope of multi-robot systems. Therefore, how to provide a multi-robot collaborative control grasping method based on reinforcement learning technology that can achieve efficient and accurate multi-robot collaborative grasping tasks is a technical problem that technicians in this field urgently need to solve. Summary of the invention
[0003] The present invention provides a multi-robot collaborative control grasping method based on reinforcement learning technology to solve the above technical problems.
[0004] In order to solve the above technical problems, the present invention provides a multi-robot collaborative control grasping method based on reinforcement learning technology, comprising the following steps:
[0005] Step 1: Establish a multi-robot control mathematical model in a unified coordinate system to determine the positions and postures of the end effectors of multiple robots and the positions and postures of the objects to be grasped;
[0006] Step 2: Use a greedy algorithm to build a multi-robot task allocation model, and assign a robot to perform a grasping task to each object to be grasped;
[0007] Step 3: Construct a reinforcement learning network, wherein the reinforcement learning network adopts a multi-agent deep Q network;
[0008] Step 4: training the reinforcement learning network to obtain a grasping control model;
[0009] Step 5: Use the grasping control model to control the collaborative grasping action of multiple robots.
[0010] Preferably, step 1 comprises:
[0011] Step 11: Confirm the scope of application of the algorithm;
[0012] Step 12: Convert the coordinate systems of multiple robots to the world coordinate system;
[0013] Step 13: Obtain the 6D postures of the end effectors of multiple robots and the 6D postures of the objects to be grasped in the world coordinate system.
[0014] Preferably, in step 11, the application scope of the algorithm includes: multiple robots perform gripping tasks at the same location, and the multiple robots are homogeneous, each including a six-axis robotic arm and an end effector.
[0015] Preferably, in step 13, the 6D posture of the end effector is expressed as in is the end effector coordinate in the world coordinate system, is the pitch angle, is the yaw angle, is the roll angle; the 6D posture of the object to be grasped is expressed as ,in, The center point coordinates of the object to be grasped in the world coordinate system, is the rotation posture of the object to be grasped.
[0016] Preferably, step 2 comprises:
[0017] Step 21: Define a mathematical model of a real-world scenario, wherein the real-world scenario includes A robotic arm and A crawling task;
[0018] Step 22: For each task and each robotic arm, calculate the corresponding grasping cost, and assign the task to the corresponding robotic arm based on the cost until all tasks are assigned or there are no free robotic arms.
[0019] Preferably, the reinforcement learning network is optimized using a centralized training and distributed execution method.
[0020] Preferably, step 3 comprises:
[0021] Step 31: construct a QMIX network, including a current network and a target network, wherein the current network and the target network have the same network structure;
[0022] Step 32: Establish a Q-value network for each robot, wherein the input of the Q-value network is the observation of each robot, and the output is the action space of the robot;
[0023] Step 33: Construct the multi-agent deep Q network based on the QMIX network and the Q-value network.
[0024] Preferably, the loss function of the reinforcement learning network is: ,in is the target value of the loss function, , The round reward generated by the reward function, is the discount factor, The Q value generated for the target network.
[0025] Preferably, the reward function includes: target proximity reward, inter-robot collision penalty, efficiency reward and grasping success reward.
[0026] Preferably, the parameters of the current network adopt a real-time update mode; and the parameters of the target network adopt a soft update mode.
[0027] Compared with the prior art, the multi-robot collaborative control grasping method based on reinforcement learning technology provided by the present invention has the following advantages:
[0028] 1. The present invention uses the continuous trial and error and learning of the reinforcement learning network to find the optimal action strategy for each robot, so that the robot can make decisions quickly and accurately during the grasping process, reducing unnecessary actions and time waste; at the same time, multiple robots can share experience through the reinforcement learning network, collaboratively optimize the overall grasping strategy, avoid mutual interference and conflict, and thus improve the grasping efficiency of the entire system;
[0029] 2. The present invention can dynamically adjust the robot's behavior strategy according to changes in the environment through the reinforcement learning network, so that the robot can quickly adapt to and complete the grasping task when facing different grasping objects, environmental conditions and task requirements; and when a robot in the system fails, other robots can dynamically adjust task allocation and behavior according to the reinforcement learning strategy to continue to complete the grasping task, ensuring the robustness and reliability of the system;
[0030] 3. The present invention can dynamically allocate tasks according to the robot's status and task requirements, so that each robot can efficiently participate in the grasping task and improve resource utilization; through the task allocation algorithm, the task allocation can be optimized to achieve the optimal state;
[0031] 4. The robot collaborative grasping system of the present invention has a high degree of automation, which can reduce dependence on manual labor, reduce labor costs, and improve production efficiency and economic benefits. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic diagram of the process framework of a multi-robot collaborative control grasping method based on reinforcement learning technology in a specific embodiment of the present invention;
[0033] Figure 2 It is a QMIX network structure diagram in a specific implementation mode of the present invention;
[0034] Figure 3A structural diagram of a current network (target network) in a specific implementation manner of the present invention;
[0035] Figure 4 A Q value network structure diagram in a specific implementation mode of the present invention;
[0036] Figure 5 This is a diagram of a hybrid network structure in a specific implementation manner of the present invention. DETAILED DESCRIPTION
[0037] In order to describe the technical solution of the above invention in more detail, specific embodiments are listed below to demonstrate the technical effects; it should be emphasized that these embodiments are used to illustrate the present invention but not to limit the scope of the present invention.
[0038] The multi-robot collaborative control grasping method based on reinforcement learning technology provided by the present invention is as follows: Figure 1 As shown, the following steps are included:
[0039] Step 1: Establish a multi-robot control mathematical model in a unified coordinate system to determine the positions of the end effectors of multiple robots and the positions of the objects to be grasped. This step can be performed by the multi-robot control module in the multi-robot collaborative control grasping system.
[0040] Step 2: Use a greedy algorithm to build a multi-robot task allocation model, and allocate a robot to perform a grasping task for each object to be grasped. This step can be performed by a task allocation module in a multi-robot collaborative control grasping system.
[0041] Step 3: Construct a reinforcement learning network, which uses a multi-agent deep Q-learning network (MADQN). This step can be performed by the reinforcement learning network module in the multi-robot collaborative control grasping system.
[0042] Step 4: Train the reinforcement learning network to obtain a grasping control model. This step can be performed by a reinforcement learning training module in the multi-robot collaborative control grasping system.
[0043] Step 5: Use the grasping control model to control the collaborative grasping action of multiple robots.
[0044] The present invention finds the optimal action strategy for each robot through continuous trial and error and learning of the reinforcement learning network, so that the robot can make decisions quickly and accurately during the grasping process, reducing unnecessary actions and time waste; at the same time, multiple robots can share experience through the reinforcement learning network, collaboratively optimize the overall grasping strategy, avoid mutual interference and conflict, and thus improve the grasping efficiency of the entire system.
[0045] In some embodiments, step 1 may include:
[0046] Step 11: Confirm the scope of application of the algorithm. In this embodiment, multiple robots perform gripping tasks at the same location, and they are all homogeneous, and are combined robots with a six-axis robotic arm and an end effector (such as a mechanical gripper). The robot serial number is set to For any robot, there is at least one robot that has the possibility of colliding with it within its working range with the base coordinate system as the origin and a radius of r. At this time, the set of these robots is suitable for the multi-robot collaborative control grasping algorithm.
[0047] Step 12: Convert the coordinate systems of multiple robots to the world coordinate system.
[0048] Take one of the robots as an example: Set the transformation matrix from the end effector coordinate system to the world coordinate system as ,but ,in, is the transformation matrix from the camera coordinate system to the world coordinate system, is the transformation matrix from the end effector coordinate system to the camera coordinate system.
[0049] Derived by Zhang Zhengyou calibration method , that is, the external parameter matrix of the camera, that is, the transformation matrix from the world coordinate system to the camera coordinate system, then .
[0050] Obtained by hand-eye calibration ,set up , Indicates the position In place The relative motion of the end effector is assumed to be , Indicates the position In place Since the eye is on the hand, the position of the camera relative to the end effector remains unchanged, so we can get , can be calculated by Tsai-Lenz two-step method ,Right now , the transformation matrix from the camera coordinate system to the end effector coordinate system, then .in, is the relative motion relationship between the robot end effector coordinate system and the base coordinate system, is the relative motion relationship between the calibration plate coordinate system and the camera coordinate system.
[0051] Through this method, all robot end effector coordinate systems can be unified into the same world coordinate system.
[0052] Step 13: The multi-robot system outputs parameters, that is, the 6D postures of the end effectors of the multiple robots and the 6D postures of the objects to be grasped in the world coordinate system are obtained.
[0053] First Take a robot as an example: Taking the robot base coordinate system as the reference coordinate system, the real-time coordinates of the end effector coordinate system can be calculated according to the kinematic formula, and then the world coordinate system can be calculated according to the transformation matrix. 6D posture of the end effector at the moment in is the end effector coordinate in the world coordinate system, is the pitch angle, is the yaw angle, is the roll angle; the 6D posture of the object to be grasped is expressed as ,in, The center point coordinates of the object to be grasped in the world coordinate system, is the rotation posture of the object to be grasped.
[0054] In some embodiments, step 2 may include:
[0055] Step 21: Define a mathematical model of the real-world scenario.
[0056] Set up the following scenario, there is The set of robotic arms is , there exists a Task set for a task , each task Need to be assigned to a robot arm , and its allocation cost is .
[0057]
[0058]
[0059]
[0060]
[0061] in, For robotic arm To the task The distance of the target object to the robot arm The location is , the position of the target object is ; It's a task The difficulty of grasping is determined by defining a function to score the target object based on its characteristics (shape, size, weight, etc.); It's a task Priority, based on the task Rating the urgency of is the weight coefficient.
[0062] Step 22: For each task And each robot , calculate the corresponding crawling cost and construct the cost matrix C={[{c}_{ij}]}_{m\times n} Repeatedly screen the tasks with the lowest cost, i.e. , the task Assign to robot arm , remove from the collection and , until all tasks are assigned or there are no free robotic arms.
[0063] In some embodiments, step 3 may include:
[0064] Step 31: Build the QMIX network, such as Figure 2 As shown, the QMIX network includes a current network and a target network, and the network structures of the current network and the target network are consistent, such as Figure 3 In some embodiments, the parameters of the current network are updated in real time based on the gradient descent of the loss function; the parameters of the target network are updated in soft mode, and the formula is: in, This updating method ensures the stability of the target network and avoids learning instability caused by drastic changes in the target network.
[0065] Step 32: Construct a single-agent deep Q network, that is, establish a Q-value network for each robot, where the input of the Q-value network is the observation of each robot, and the output is the action space of the robot.
[0066] The reinforcement learning network adopts centralized training and decentralized execution (CTDE). Each robot has its own Q-value network, whose input is the observation of each robot. The specific structure of the Q-value network is as follows: Figure 4 As shown, there are three layers of fully connected networks, and the activation function of the first two layers of fully connected networks is ReLU.
[0067] The formula for the ReLU activation function is: Q value network input layer parameters, where The jaws are open or closed: The action space output by the final Q-value network is closed: ,in is the three-axis speed of the gripper movement, is the three-axis angular velocity of the gripper movement. The linear velocity and angular velocity of each movement are measured in units of and , For the opening and closing action of the gripper, set to .
[0068] In order to reduce the difficulty of training and avoid gradient explosion, only one action's Q value is predicted each time. The specific algorithm is to set a greedy factor , generate a random number between 0 and 1. If the random number is less than the greedy factor, randomly select an action in the action space; if the random number is greater than the greedy factor, select the action with the largest Q value.
[0069] Step 33: Construct a hybrid network, and construct the multi-agent deep Q network based on the QMIX network and the Q-value network.
[0070] Hybrid networks are composed of Global state at the moment And the Q value sequence output by the Q value network of each robot As network input. Integrated into (batch_size, num_agents, 1), global state is (batch_size, num_state). Among them, batch_size is the number of batches, num_agents is the number of robots, and num_state is the sum of the global robot states: First, the global state Generated by Multi-Layer Perceptron (MLP) (batch_size, num_hidden, num_hidden), and then the global state Generated by MLP (batch_size, num_hidden). Then calculate . Use the same method to set the global state Generated by MLP (batch_size, num_hidden, 1) and (batch_size, 1). and Perform ReLU operation when generating Then perform ELU operation. Calculate .final, The shape is (batch_size, 1). Among them, num_hidden is the hidden layer parameter, and num_actions is the number of action spaces.
[0071] In the ELU activation function Take 1, the formula is: In each generation After that, Perform absolute value operations to ensure monotonicity constraints, and the parameters are always non-negative, that is, ,in is the collection of all actions.
[0072] In the reinforcement learning network, the reward function includes: target proximity reward, robot collision penalty, efficiency reward, and grasping success reward. The formula is:
[0073] in, is the weight coefficient.
[0074] Target proximity rewards: in, is the number of robots, For the i The position of the object to be grasped by the robot, For the i The position of the robot end effector.
[0075] Collision penalties between robots: in, is the distance between robots, A safe distance to avoid collision.
[0076] Efficiency Rewards: in, For the i The current moment of the robot, is the time decay coefficient.
[0077] Rewards for successful capture: In some embodiments, the loss function during the reinforcement learning network training process is: ,in is the target value of the loss function, , The round reward generated by the reward function, is the discount factor, The Q value generated for the target network.
[0078] In this embodiment, the specific network training hyperparameters are: learning rate is 0.005, training round epochs is 1000, discount factor is 0.95, batch_size is 32, greed factor is 0.05, and optimizer uses Adam. After training, the optimal model can be captured.
[0079] By adopting the above method, tasks can be dynamically allocated according to the status of the robots and task requirements, so that each robot can efficiently participate in the grasping task and improve resource utilization; through the task allocation algorithm, task allocation can be optimized to achieve the optimal state; the present invention has a high degree of automation, can reduce dependence on manual labor, reduce labor costs, and improve production efficiency and economic benefits.
[0080] In summary, the multi-robot collaborative control grasping method based on reinforcement learning technology provided by the present invention includes step 1: establishing a multi-robot control mathematical model in a unified coordinate system, determining the position and posture of the end effectors of multiple robots and the position and posture of the object to be grasped; step 2: using a greedy algorithm to construct a multi-robot task allocation model, and assigning a robot to perform a grasping task to each object to be grasped; step 3: constructing a reinforcement learning network, and the reinforcement learning network adopts a multi-agent deep Q network; step 4: training the reinforcement learning network to obtain a grasping control model; step 5: using the grasping control model to control the collaborative grasping action of multiple robots. The present invention finds the optimal action strategy for each robot through continuous trial and error and learning of the reinforcement learning network, so that the robot can make decisions quickly and accurately during the grasping process, reducing unnecessary actions and time waste; at the same time, multiple robots can share experience through the reinforcement learning network, collaboratively optimize the overall grasping strategy, avoid mutual interference and conflict, and thus improve the grasping efficiency of the entire system.
[0081] Obviously, those skilled in the art can make various changes and modifications to the invention without departing from the spirit and scope of the invention. Thus, if these modifications and variations of the invention fall within the scope of the claims of the invention and their equivalents, the invention is also intended to include these modifications and variations.
Claims
1. A multi-robot collaborative control grasping method based on reinforcement learning technology, characterized in that: The steps include: Step 1: Establish a multi-robot control mathematical model in a unified coordinate system to determine the positions and postures of the end effectors of multiple robots and the positions and postures of the objects to be grasped; Step 2: Use a greedy algorithm to build a multi-robot task allocation model, and assign a robot to perform a grasping task to each object to be grasped; The step 2 comprises: Step 21: Define a mathematical model of a real-world scenario, wherein the real-world scenario includes A robotic arm and A crawling task; Step 22: For each task and each robotic arm, calculate the corresponding grasping cost, and assign the task to the corresponding robotic arm based on the cost until all tasks are assigned or there are no free robotic arms; Step 3: Construct a reinforcement learning network, wherein the reinforcement learning network adopts a multi-agent deep Q network; The step 3 comprises: Step 31: construct a QMIX network, including a current network and a target network, wherein the current network and the target network have the same network structure; Step 32: Establish a Q-value network for each robot, wherein the input of the Q-value network is the observation of each robot, and the output is the action space of the robot; Step 33: construct the multi-agent deep Q network based on the QMIX network and the Q-value network; Step 4: training the reinforcement learning network to obtain a grasping control model; The loss function of the reinforcement learning network is: ,in is the target value of the loss function, , The round reward generated by the reward function, is the discount factor, The Q value generated for the target network Step 5: Use the grasping control model to control the collaborative grasping action of multiple robots.
2. The multi-robot collaborative control grasping method based on reinforcement learning technology as claimed in claim 1, characterized in that: Step 1 includes: Step 11: Confirm the scope of application of the algorithm; Step 12: Convert the coordinate systems of multiple robots to the world coordinate system; Step 13: Obtain the 6D postures of the end effectors of multiple robots and the 6D postures of the objects to be grasped in the world coordinate system.
3. The multi-robot collaborative control grasping method based on reinforcement learning technology as claimed in claim 2, characterized in that: In step 11, the application scope of the algorithm includes: multiple robots perform gripping tasks at the same location, and the multiple robots are homogeneous, each including a six-axis robotic arm and an end effector.
4. The multi-robot collaborative control grasping method based on reinforcement learning technology as claimed in claim 3 is characterized in that: In step 13, the 6D posture of the end effector is expressed as ,in is the end effector coordinate in the world coordinate system, is the pitch angle, is the yaw angle, is the roll angle; the 6D posture of the object to be grasped is expressed as ,in, The center point coordinates of the object to be grasped in the world coordinate system, is the rotation posture of the object to be grasped.
5. The multi-robot collaborative control grasping method based on reinforcement learning technology as claimed in claim 1, characterized in that: The reinforcement learning network is optimized by using centralized training and distributed execution methods.
6. The multi-robot collaborative control grasping method based on reinforcement learning technology as claimed in claim 1, characterized in that: The reward function includes: target proximity reward, inter-robot collision penalty, efficiency reward and grasping success reward.
7. The multi-robot collaborative control grasping method based on reinforcement learning technology as claimed in claim 6, characterized in that: The parameters of the current network adopt a real-time update mode; the parameters of the target network adopt a soft update mode.
Citation Information
Patent Citations
Mechanical arm dense object temperature priority grabbing method based on deep reinforcement learning
CN112405543A
Q-learning-based on-orbit maintenance task allocation method for multiple mechanical arms
CN117754571A