A Dexterous Two-Hand Cooperative Control Method Based on Multi-Agent Reinforcement Learning

By modeling each joint of a dexterous robot as an independent agent, and using multi-agent reinforcement learning method and sequential decision-making mechanism, the problems of conflict and non-stationarity between agents in the coordinated control of both hands are solved, achieving more efficient movement collaboration and learning stability.

CN119795175BActive Publication Date: 2025-07-01BEIJING UNION UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510113653.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-07-01
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle conflicts and non-stationarity between agents in the coordinated control of two-handed dexterous robots, resulting in low collaboration efficiency and learning stability.

Method used

By modeling each joint and finger of the dexterous robot as an independent agent, using a multi-agent reinforcement learning method, combining the sequential decision-making mechanism and a reward distribution strategy based on marginal contribution, a Q-value distribution network is designed to reduce the overestimation problem of traditional networks, and improving the exploration ability and convergence efficiency of the strategy through entropy regularization and importance sampling techniques.

Benefits of technology

It improves the efficiency of movement collaboration between multiple joints, dynamically adapts to complex action space and joint interaction needs, reduces action conflicts, improves operational accuracy and stability, and shows strong task universality and environmental adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119795175B_ABST
    Figure CN119795175B_ABST
Patent Text Reader

Abstract

The present invention relates to a dexterous two - hand collaborative control method based on multi - agent reinforcement learning, including: initializing the neural network and environment of the dexterous two - hands, and collecting data during environment interaction; based on the data during environment interaction, the joints make sequential decision - making actions according to the greedy policy; evaluating the decision - making actions through a Q - value distribution network, and simultaneously calculating the reward return; based on the reward return, introducing entropy regularization, and using the method of policy gradient optimization to update the joint policy; based on preset conditions, terminating the update of the joint policy, and outputting the optimal collaborative control strategy. The present invention is applicable to complex scenarios such as grasping, rotating, and assembling, and has high collaboration efficiency and environmental adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of reinforcement learning and control technology, and particularly to a dexterous two - hand collaborative control method based on multi - agent reinforcement learning. Background Art

[0002] Dexterous manipulators play an important role in robotics and are widely used in fields such as industrial manufacturing, medical assistance, and service robots. They achieve complex grasping, handling, and precision operation tasks through multi - degree - of - freedom joints and fingers. However, during the control process of dexterous manipulators, due to the complexity of high - degree - of - freedom movements and the dynamic coordination requirements between joints, high requirements are imposed on control accuracy and collaboration efficiency. Especially in two - hand collaborative operations, efficient motion planning and precise collaborative control need to be achieved.

[0003] Current control methods are mainly divided into model - based control methods and reinforcement - learning - based control methods. Model - based control methods rely on the manipulator dynamics model and achieve task goals through trajectory optimization and model predictive control (MPC). These methods perform well in structured environments, but in complex and unstructured scenarios, it is difficult for the model to accurately describe the dynamic interaction between the manipulator and the environment, resulting in insufficient adaptability and generalization ability.

[0004] Reinforcement - learning - based control methods can learn optimal policies by interacting with the environment, avoiding the dependence on an accurate model, and demonstrating strong task adaptability and decision - making flexibility. In the operation of dexterous manipulators, single - agent reinforcement learning has been widely applied to tasks such as single - hand grasping, rotation, and throwing. However, when applied to two - hand collaborative tasks, due to the high - dimensional action space and complex dynamic interaction between agents, existing methods are difficult to effectively handle conflicts and non - stationarity between agents, resulting in low collaboration efficiency and learning stability. Although multi - agent reinforcement learning (MARL) methods can model multi - agent systems, they still face problems such as unreasonable reward distribution and slow policy convergence in two - hand dexterous operations. Summary of the Invention

[0005] The object of the present invention is to provide a dexterous two - hand collaborative control method based on multi - agent reinforcement learning. By modeling each joint and finger of the dexterous manipulator as an independent agent, and based on the advantage function decomposition lemma, each agent makes decisions in sequence, assuming that subsequent agents take optimal actions, thereby improving the decision - making performance of the current agent. At the same time, the present invention designs a Q - value distribution network to reduce the over - estimation problem of traditional networks for agents and strengthen the collaboration effect. Finally, by sampling and evaluating the strategy of the Q - value distribution network, the policy instability in the high - dimensional continuous action space is reduced, further improving the operation accuracy and collaboration efficiency of the dexterous manipulator in complex task scenarios.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] A dexterous two - hand collaborative control method based on multi - agent reinforcement learning, comprising:

[0008] Initializing the neural network and environment of the dexterous two - hands, and collecting data during environment interaction;

[0009] Based on the data during environment interaction, the joints make sequential decision - making actions according to the greedy policy;

[0010] Evaluating the decision - making actions through the Q - value distribution network, and calculating the reward return at the same time;

[0011] Based on the reward return, introducing entropy regularization, and using the method of policy gradient optimization to update the joint policy;

[0012] Based on preset conditions, terminate the update of the joint policy, and output the optimal collaborative control policy.

[0013] Optionally, the neural network and environment of the dexterous two - hands include: the policy network, evaluation network and simulation environment of the dexterous two - hands;

[0014] The policy network includes: the policy network of each joint, with the input being the environment information and its own state, and the output being the joint action;

[0015] The evaluation network includes: a Q - value distribution network, with the input being the environment information, its own state and joint action, and the output being the evaluation of the joint action;

[0016] The simulation environment includes: starting a pre - set simulator and scenario.

[0017] Optionally, in the neural network and environment of the dexterous two - hands, there are also constraint conditions; the constraint conditions include:

[0018] Torque constraint: The torque applied by each joint should meet the limitation of the hardware capacity;

[0019] Action rate constraint: The amplitude of the joint action change within each time step shall not exceed a preset threshold to avoid system instability caused by high - frequency changes;

[0020] Joint action coordination constraint: The action combinations of all joints need to meet the coordination requirements of the dexterous two - hands and object dynamic interaction.

[0021] Optionally, collecting data during environment interaction includes:

[0022] At each time step, the system state and action of the dexterous two - hands collect data through interaction with the environment, including: the current state s t , the current action a t, immediate reward r t , next state s t+1 and task termination flag d t ;

[0023] Store the collected interaction data (s t , a t , r t , s t+1 , d t ) in the experience pool D; where, after exceeding the preset size of the experience pool, the first-in, first-out strategy is used to replace the old data.

[0024] Optionally, the data collected during environmental interaction further includes: performing importance sampling correction on the data stored in the experience pool; including:

[0025] For each sample (s, a) in the experience pool, calculate the importance weight where, π new (a|s) is the probability of the current policy, and π old (a|s) is the probability of the old policy during sampling;

[0026] Based on the importance weight, combined with the immediate reward r t and the temporal difference error, set the priority δ i = r i + γQ(s i+1 , a′) - Q(s i , a i ), and the sample priority is where, α is the priority adjustment parameter, used to balance the sampling probabilities of high-priority samples and low-priority samples;

[0027] Perform sampling according to the sample priority P i , and preferentially select the samples that contribute a preset amount to policy optimization for training.

[0028] Optionally, the joints make sequential decision-making actions according to the greedy strategy, including:

[0029] The joints set the decision-making order according to the task requirements: obtain the advantage value of each joint by using the advantage function decomposition theorem, and make sequential decisions from large to small according to the advantage value.;

[0030] The action a k of each joint k is determined by the local policy π k (a k |o k ); the input of the local policy is the local observation o k of the joint and the action a 1:k-1; among them, different joints of both hands are set as different agents, and each agent will have its own strategy, which is the local strategy;

[0031] The first joint a1 directly selects the action a1 = π1(o1) according to the local observation o1; the second joint a2 selects the action a2 = π2(o2, a1) based on the local observation o2 and the action a1 of the first joint; and so on until the last joint a 40 , and the joint action a = {a1, a2,..., a 40} is obtained; the joint action a is executed to drive the system from the current state s t to transfer to the next state s t+1 .

[0032] Optionally, to evaluate the decision-making action through the Q-value distribution network includes:

[0033] By constructing a Q-value distribution network, fitting the Q-value distribution according to the data generated by the interaction between the environment and the agent, and obtaining the reward of the joint, that is, the local reward, through sampling to reduce the overestimation problem of the traditional network for the agent and strengthen the cooperation effect; among them, the global reward is used to evaluate the performance of the dexterous hands as a whole during the task completion process;

[0034] Normalize and smooth the local reward.

[0035] Optionally, calculating the reward return includes:

[0036] Based on the local reward, calculate the local return function of each joint; where the local return function is used to characterize the contribution of the action of the current joint to the long-term return;

[0037] Accumulate the local return functions of all joints to obtain the global return function;

[0038] The local return function is:

[0039]

[0040] where r k (s, a) is the local reward of joint k, γ is the discount factor, represents the expectation of the next state s′ and action a′;

[0041] The global return function is:

[0042]

[0043] where N is the total number of joints and k is the kth joint.

[0044] Optionally, the joint policy π(a|s) is: the local policies π of all joints k (a k |o k ) combined The optimization objective of the joint policy is to maximize the long-term return of the system under the current policy Use the policy gradient method to optimize the policy parameters where is the gradient with respect to the policy parameter φ, J(π φ ) is the objective function, π φ (a|s) is the probability distribution of the policy, and Q(s,a) is the action-value function

[0045] Optionally, updating the joint policy includes:

[0046] Dynamically adjust the temperature parameter α, with the goal of maintaining the entropy value of the policy close to the target entropy value H target : Use gradient descent to update Combine policy gradient and entropy regularization, and update the policy parameter φ according to the following formula:

[0047] The beneficial effects of the present invention are:

[0048] By modeling each joint of the dexterous manipulator as an independent agent, adopting the multi-agent reinforcement learning method, and combining the sequential decision-making mechanism with the reward allocation strategy based on marginal contribution, the action cooperation between multiple joints is made more efficient. Compared with the traditional model-based control method and the single-agent reinforcement learning method, the present invention can dynamically adapt to the complex action space and joint interaction requirements during the task execution process, effectively reduce action conflicts, and improve the operation accuracy and stability of both hands in complex tasks such as grasping, rotating, and assembling

[0049] By means of the distributed return function estimation and importance sampling techniques, the problem of slow policy convergence in the high-dimensional action space is solved. At the same time, by introducing entropy regularization and smoothing mechanism, the exploration ability and convergence efficiency of the policy are improved. Dynamically adjusting the weight parameters of the reward function and the policy optimization objective enables the present invention to adapt to the requirements of different task stages, showing strong task versatility and environmental adaptability, especially excellent generalization ability in unstructured complex scenarios Description of the Drawings

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0051] Figure 1 It is a schematic diagram of the overall process of the embodiment of the present invention;

[0052] Figure 2 It is a schematic diagram of the overall architecture of the embodiment of the present invention;

[0053] Figure 3 It is a schematic diagram of the sequential decision-making architecture of the embodiment of the present invention. Detailed implementation manners

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0055] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0056] As Figure 1 - Figure 2 shown, a dexterous two-handed cooperative control method based on multi-agent reinforcement learning includes:

[0057] Initialize the neural network and environment of the dexterous two hands, and collect data during environment interaction;

[0058] Based on the data during environment interaction, the joints make sequential decision-making actions according to the greedy strategy;

[0059] Evaluate the decision-making actions through the Q-value distribution network, and calculate the reward return at the same time;

[0060] Based on the reward return, introduce entropy regularization, and use the method of policy gradient optimization to update the joint policy;

[0061] Based on preset conditions, terminate the update of the joint policy and output the optimal cooperative control policy.

[0062] Specifically, in this embodiment, the neural network and environment of the dexterous two hands include: the policy network, evaluation network, and simulation environment of the dexterous two hands;

[0063] The policy network includes: a policy network for each joint, with the input being environmental information and its own state, and the output being joint actions;

[0064] The evaluation network includes: a Q-value distribution network, with the input being environmental information, its own state, and joint actions, and the output being the evaluation of joint actions;

[0065] The simulation environment includes: starting a pre-set simulator and a scenario;

[0066] Initializing the neural network of the dexterous hands includes: initializing its weights by generating an orthogonal matrix, that is, assigning initial weights to each layer of the network through the method of the orthogonal matrix, so as to provide linearly independent parameter values for each layer of the neural network.

[0067] In the neural network of the dexterous hands and the environment, there are also set constraint conditions; the constraint conditions include:

[0068] Torque constraint: the torque applied to each joint should meet the limitations of the hardware capabilities;

[0069] Action rate constraint: the amplitude of the change in the joint action shall not exceed a preset threshold within each time step to avoid system instability caused by high-frequency changes;

[0070] Joint action coordination constraint: the action combinations of all joints need to meet the coordination requirements of the dexterous hands and the dynamic interaction with the object.

[0071] The agent observes the environment and obtains information: each joint acts as an independent agent, and obtains its own relevant information through the local observation space, including joint angle, angular velocity, relative position information with the object, adjacent joint states, and tactile feedback. Local observation only perceives its own relevant information, reduces global dependence, and constructs a decentralized multi-agent system.

[0072] Collect data and perform importance sampling: when interacting with the environment, the system collects the current state, action, immediate reward, next state, and task termination flag, and stores them in the experience pool. Through importance sampling, the sampling weights are calculated to give priority to high-priority samples for optimizing the policy.

[0073] As Figure 3 shown, make decisions to execute actions in the joint / finger order: the joints make decisions on actions step by step in the greedy policy order. Each joint selects an action based on local observation and the actions of the previous joints. The first joint makes a direct decision, and the subsequent joints complete the decision in combination with the previous actions. The actions of all joints are combined into a joint action to promote the transfer of the system state.

[0074] Distributed estimation of the return function:

[0075] By constructing a Q-value distribution network, fitting the Q-value distribution based on the data generated by the interaction between the environment and the agent, and obtaining the rewards of the joints through sampling, the overestimation problem of the traditional network for the agent is reduced, and the cooperation effect is strengthened. Finally, the rewards of each joint are accumulated to obtain the global return, providing a basis for policy optimization.

[0076] Update the policy and optimize the value function: The joint policy consists of the local policies of the joints and is optimized by the policy gradient method to maximize the long-term return of the policy. Entropy regularization is introduced to encourage the exploration of diverse actions and avoid local optima. The entropy coefficient and learning rate are dynamically adjusted to balance exploration and exploitation and ensure the stability of policy optimization.

[0077] Judge whether the task termination condition is satisfied: The task terminates when one of the following conditions is met: the position and orientation of the object reach the target range, the running time reaches the maximum number of steps, or the change in the policy and the reward value tend to be stable. If a hardware failure or abnormal action is detected, the task is immediately terminated to ensure safety, and the optimization result is output to end the loop.

[0078] Finally, output the result: After the task ends, the trained model and the training curve are output.

[0079] Furthermore, the state space includes: joint states, object states, tactile feedback information, and the action space of the joints;

[0080] The joint states include: the angular position and angular velocity of each joint;

[0081] The object states include: the position and orientation of the object in three-dimensional space;

[0082] The tactile feedback information is represented by the force and torque at the contact point;

[0083] The action space includes: the control variables of all joints; the control variables include the applied torque or the angular adjustment amount of each joint;

[0084] The initialization of the state space of the dexterous hands includes:

[0085] The angular position and angular velocity of all joints are set to the task initial values, the position and orientation of the object in three-dimensional space are initialized to the given values according to the task requirements, the tactile feedback information is initialized to zero, and constraint conditions are introduced in the action space.

[0086] Specifically, in this embodiment, the state space of the dexterous hands includes joint states, object states, and tactile feedback information, as follows:

[0087] In the joint states, the state of each joint is composed of the following parameters: joint angle θ k ∈[-π,π] (radians), representing the current rotational position of the joint; joint angular velocity Represents the change rate of joint angles. For the 20 controllable joints of two hands (20 for each hand), the total dimension of the joint state is 40×2 = 80.

[0088] In the object state, the state of the manipulated object includes the following parts: object position x object ∈[-10, 10] cm (in a three-dimensional Cartesian coordinate system); object attitude q object , represented by a quaternion (q w , q x , q y , q z ∈[-1, 1]); object velocity v object ∈[-5, 5] cm / s, representing the linear velocity of the object; object angular velocity ω object ∈[-5, 5] rad / s, representing the rotational velocity of the object. The total dimension of the object state is 3 (position) + 4 (quaternion attitude) + 3 (linear velocity) + 3 (angular velocity) = 13.

[0089] In the tactile feedback information, the tactile feedback information of each finger is described by the following parameters: contact force f tactile ∈[0, 5] represents the acting force at the contact point; contact torque τ tactile ∈[0, 5] represents the torque at the contact point. For the 10 fingers of two hands, the total dimension of the tactile feedback is 10×(3 + 1) = 40.

[0090] In the total state space, integrating the joint state, object state, and tactile feedback information, the total state space dimension of the dexterous hands is: total dimension = 80 (joint state) + 13 (object state) + 40 (tactile feedback information) = 133.

[0091] Since the dexterous hands adopt a multi-agent reinforcement learning framework, each joint can only obtain local state information related to itself. The local observation space includes:

[0092] The self-joint state includes the angle θ of the current joint k and the angular velocity with a dimension of 2.

[0093] The adjacent joint information includes the angles and angular velocities of adjacent joints, involving at most two joints before and after; the dimension is 2×2 = 4.

[0094] The relative displacement x of the relative position information between the object and the joint rel,k = x object - x k , where x k is the position at the end of the joint; the relative direction q rel,k = q object - q k in which qk is the quaternion of the joint direction; the dimension is 3 + 4 = 7.

[0095] The contact force f at the end of the joint of the haptic feedback information tactile,k and the torque τ tactile,k ; the dimension is 3 + 1 = 4.

[0096] In the total dimension of the local observation space, for each joint, the total dimension of the local observation space is: the dimension of the local observation space = 2 (own joint state) + 4 (adjacent joint information) + 7 (relative position information) + 4 (haptic feedback information) = 17.

[0097] In the single-joint action space, the action of each joint is a continuous value, representing the applied torque τ k ∈[-2, 2]. In the joint action space, for the 40 joints of two hands, its joint action space is:

[0098]

[0099] The torque τ applied by each joint k must be within the range allowed by the hardware to ensure that it does not exceed the load-bearing capacity of the joint; the change rate Δτ of the joint action k shall not exceed 0.5 Nm at each time step to avoid system instability caused by high-frequency changes; the combined actions of all joints need to meet the dynamic interaction coordination between the dexterous hands and the object to prevent conflicting actions.

[0100] At the beginning of the task, the neural network and the environment need to be initialized. The angles θ k and angular velocities of all joints are set to the initial values of the task, usually zero or default settings. The position x object and attitude q object of the object are initialized to given values according to the task requirements. The haptic feedback information is initialized to zero, indicating that the system has not yet come into contact with the object. The relative position information and direction information are calculated by the differences between the end of the joint and the position and direction of the object.

[0101] The action space is composed of the control variables of all joints, specifically including the applied torque or angle adjustment amount of each joint. The action of a single joint is represented as a continuous variable τ k , and its value range is [τ min , τ max , where τ min and τ max are respectively the minimum and maximum values of the torque allowed to be applied to this joint, and the value range is determined by the hardware performance and task requirements. For all joints of the dexterous hands, its joint action space is defined as:

[0102] A = A1 × A2 × … × A N

[0103] where A k = [τ min , τ max represents the action space of k joints, and N is the total number of joints of the dexterous hands.

[0104] To ensure the controllability of joint actions, the present invention introduces constraint conditions in the action space, including the following:

[0105] Torque constraint: The torque τ k applied to each joint should meet the limitations of the hardware capabilities to ensure that the actions do not exceed the load-bearing capacity of the mechanical structure;

[0106] Action rate constraint: The amplitude of the action change of the joint |Δτ k | should not exceed the preset threshold within each time step to avoid system instability caused by high-frequency changes;

[0107] Joint action coordination constraint: The action combinations of all joints need to meet the coordination requirements of the dexterous hands and the dynamic interaction with the object to prevent conflicting actions.

[0108] During the action execution, each joint selects an action a m according to its local observation space O m and the current policy π m (a m |O m ), where a m represents the specific control input of the m-th joint. The joint actions act on the system through the joint action a = {a1, a2,..., a N} to drive the system state to transfer from the current state s to the next state s'. The transfer of the system state is determined by the following physical laws:

[0109] s′ = T(s, a)

[0110] where T(s, a) is the state transition function, which describes the evolution process of the system under the joint action a, and s and s' are the current state and the next state respectively. In the present invention, the state transition function T(s, a) includes two parts: joint state update and object state update. The joint state update is based on the following formula:

[0111]

[0112] where Δτ k is the joint action input, and f(τ k ) is the function of the torque on the joint angular velocity. The state update of the manipulated object is described by the following formula:

[0113] x object (t + 1) = x object (t) + v object ·Δt

[0114] q object (t + 1) = q object (t) + ω object ·Δt

[0115] Wherein, v object and ω object respectively represent the linear velocity and angular velocity of the object, which are jointly determined by the joint action a and the tactile feedback.

[0116] To balance exploration and exploitation and accelerate the algorithm convergence, the present invention introduces random noise into the action space by adding random noise when executing the policy such that the actual action of each joint is:

[0117] a′ m = a m + ∈ m

[0118] Wherein, σ is the standard deviation of the noise, controlling the exploration degree. The purpose of introducing randomness is to prevent the policy from converging to the local optimal solution prematurely and improve the exploration ability of the dexterous hands in the high-dimensional action space.

[0119] The definition and constraints of the space provide a theoretical basis for the motion planning of the dexterous hands, ensuring that each joint can meet the physical and task requirements of the system when selecting actions, and at the same time providing direct input for the subsequent design of the reward function and policy optimization.

[0120] The global reward function R(s, a) is used to evaluate the overall task completion of the dexterous hands, specifically including the following three items:

[0121] Reward the dexterous hands for moving the object to the target position, defined as r pos = -α||x object - x goal || 2 , wherein, α = 10, x goal is the target position. Reward the dexterous hands for rotating the object to the target attitude, defined as r rot = -β||q object - q goal || 2 , wherein, β = 5, q goal is the target attitude. Penalize the joints for applying excessive torque to encourage energy-saving operations, defined as wherein, γ = 0.1. The global reward function is the weighted sum of the above three parts:

[0122] R(s,a) = r pos + r rot + r energy

[0123] Allocate the global reward to each joint, and the marginal contribution φ of each joint k represents its increment to the global reward, defined as φ k = Q(s,a 1:k ) - Q(s,a 1:k-1 ); Allocate the global reward to each joint according to the marginal contribution, and the local reward r k The calculation formula is

[0124] Furthermore, the data collected during environmental interaction includes:

[0125] At each time step, the system state and actions of the dexterous hands collect data through interaction with the environment, including: the current state s t , the current action a t , the immediate reward r t , the next state s t+1 and the task termination flag d t ;

[0126] Store the collected interaction data (s t ,a t ,r t ,s t+1 ,d t ) in the experience pool D; Among them, after exceeding the preset size of the experience pool, the first-in-first-out strategy is used to replace the old data.

[0127] Specifically, in this embodiment, at each time step, the system state and actions of the dexterous hands collect data through interaction with the environment, specifically including the following content: the current state s t includes joint state, object state, tactile feedback information; the current action a t is the action generated by the joint policy of the joint; the immediate reward r t is calculated according to the global reward function R(s t ,a t ); the next state s t+1 is the state of the system after executing the action a t ; the flag d t indicating whether the termination condition is reached represents whether the task is completed or failed. Store the collected interaction data (s t ,a t ,r t ,s t+1 ,d tStored in the experience pool D; the size of the experience pool is set to 100,000 pieces of data, and when this limit is exceeded, the first-in-first-out strategy is adopted to replace the old data.

[0128] Since the policy is continuously updated during training, there may be biases in the quality of different samples and learning efficiency. Therefore, importance sampling correction needs to be performed on the data to ensure that high-quality sample data is repeatedly learned.

[0129] For each sample (s, a) in the experience pool, calculate its importance weight where π new (a|s) is the probability of the current policy, and π old (a|s) is the probability of the old policy at the time of sampling. Combining the immediate reward r t and the temporal difference error, set the priority δ for the sample i = r i + γQ(s i+1 , a′) - Q(s i , a i ), and the sample priority is where α is the priority adjustment parameter, which is used to balance the sampling probabilities of high-priority and low-priority samples. Sample according to the sample priority P i , and preferentially select samples that contribute more to policy optimization for training. The size of each batch sampling in the experience pool is 64.

[0130] Furthermore, the joints sequentially select decision actions in order, including:

[0131] The joints of the dexterous hands set the decision order of the greedy policy according to the task requirements: obtain the advantage value of each joint by using the advantage function decomposition theorem, and make sequential decisions from large to small according to the advantage value.

[0132] The action a k of each joint k is determined by its local policy π k (a k |o k ); the input of the local policy is the local observation o k of the joint and the action a 1:k-1 of the previous joint. The joint policy is the combination of the local policies of all joints, which is

[0133] The first joint a1 directly selects the action a1 = π1(o1) according to its local observation o1; the second joint a2 selects the action a2 = π2(o2, a1) based on the local observation o2 and the action a1 of the first joint; execute in this logic sequentially until the last joint a 40 , and obtain the joint action a = {a1, a2,..., a 40}. Execute the combined action a to drive the system from the current state s t to transfer to the next state s t+1 .

[0134] Furthermore, for the decision-making action, the obtained rewards from evaluation include:

[0135] By constructing a Q-value distribution network, according to the data generated by the interaction between the environment and the agent, fitting the Q-value distribution, obtaining the rewards of the joints through sampling, and normalizing and smoothing the rewards.

[0136] Furthermore, in this embodiment, in order to improve the cooperation ability of each joint, a Q-value network is constructed, and the specific steps are as follows:

[0137] Calculate the reward at the current time step according to the global reward function R(s t ,a t ).

[0138] In order to learn the return distribution, it is necessary to extend the Bellman equation to the distribution form. The traditional Bellman equation describes the recursive relationship of Q(s,a):

[0139]

[0140] Among them, R(s,a) is the immediate reward, s' is the next state after executing the action a, and a' is the action selected under the state s'. In distributed Q-learning, the Bellman equation of the return distribution is expressed as:

[0141] Z(s,a) = R(s,a) + γZ(s′,a′)

[0142] Where denotes distribution equality. In other words, the return distribution of the current state-action pair is equal to the sum of the immediate reward and the discounted future return distribution. This form reflects the recursive property of the return distribution.

[0143] To concretize the above formula, this embodiment defines the distributed Bellman operator J π :

[0144]

[0145] Among them, J π combines the immediate reward and the future return distribution for updating the current return distribution.

[0146] In distributed Q-learning, in this embodiment, by minimizing the predicted distribution Z θ (s,a) and the target distribution Z targetTrain the neural network using the distance between (s,a). The target distribution Z target (s,a) is defined as:

[0147]

[0148] where y is the target value, calculated based on the distributed Bellman equation: y = r + γQ θ (s′,a′), where Q θ (s′,a′) is the expected value of the next state-action pair; σ target is the standard deviation of the target distribution, usually calculated from the uncertainty of Z(s′,a′). The target distribution represents an approximation of the true return distribution of the current state-action pair.

[0149] To optimize the neural network, this embodiment needs to define a loss function to measure the difference between the predicted distribution Z θ (s,a) and the target distribution Z target (s,a). Usually, the KL divergence is used as the metric:

[0150]

[0151] Assume that both Z θ (s,a) and Z target (s,a) are Gaussian distributions, and their probability density functions are respectively: where Q θ (s,a) is the mean of the predicted distribution; σ ψ (s,a) is the standard deviation of the predicted distribution; y is the mean of the target distribution; σ target is the standard deviation of the target distribution.

[0152] Estimate the reward of the joint according to the Q-value network. The local reward r k The calculation formula is To prevent the reward value from being too large or too small, perform normalization processing to limit the reward value within the range of [-1,1]: where ∈ is a small value (such as 10 -6 ) to prevent the denominator from being zero. To avoid drastic fluctuations in the reward distribution process, a smoothing mechanism is adopted: where α ∈ [0,1] is the smoothing coefficient, is the reward value at the previous time step.

[0153] Furthermore, calculating the reward return includes:

[0154] Based on the local reward, calculate the local return function for each joint; where the local return function is used to characterize the contribution of the action of the current joint to the long-term return;

[0155] Accumulate the local reward functions for all joints to obtain the global reward function.

[0156] Specifically, in this embodiment, in a multi-agent system, the local reward function z k (s,a) of each joint is used to describe the contribution of the current joint's action to the long-term reward, and is specifically defined as follows:

[0157] The local reward function z k (s,a) of each joint k consists of an immediate reward and a future reward, and is recursively calculated as:

[0158]

[0159] where r k (s,a) is the local reward of joint k; γ is the discount factor (set to γ = 0.99); represents the expectation for the next state s′ and action a′.

[0160] The global reward function Z(s,a) is generated by accumulating the local reward functions of all joints:

[0161]

[0162] where N is the total number of joints (N = 40).

[0163] Furthermore, the joint policy π(a|s) is: the combination of the local policies π k (a k |o k ) of all joints The optimization objective of the joint policy is to maximize the long-term reward of the system under the current policy Use the policy gradient method to optimize the policy parameters

[0164] To balance exploration and exploitation, an entropy term H(π) is added to the optimization objective to encourage the randomness of the policy: where α is the temperature parameter, which controls the trade-off between exploration and exploitation (initially set to α = 0.2). The entropy of the policy is defined as H(π φ (·|s)) = -∫π φ (a|s)logπ φ (a|s)da; a larger entropy value αH(π) encourages wider exploration, while a smaller entropy value tends to focus on high-reward actions.

[0165] The temperature parameter α is a parameter that can be optimized to control the weight of the entropy term. At the beginning of optimization, more states need to be explored, so generally α is relatively high. In the later stage of optimization, rewards are generally the main focus, and the proportion of the entropy term will decrease. Similar to the annealing algorithm, it is called the temperature parameter.

[0166] To improve the stability of training, the learning rate η of the policy network π gradually decreases with the training time where is the initial learning rate, β is the learning rate decay parameter, and t is the number of training steps.

[0167] Furthermore, updating the joint policy includes:[[]]

[0168] Dynamically adjusting the temperature parameter α with the goal of maintaining the entropy value of the policy close to the target entropy value H target :[[]] Updating using gradient descent Combining policy gradient and entropy regularization, update the policy parameter φ according to the following formula:[[]]

[0169] Output the finally optimized joint policy model and value function model, which are used to describe the motion planning and control strategy of dexterous hands in complex task scenarios. The output models include the following:[[]]

[0170] The joint policy is composed of local policies of each joint, in the form of Output the global value function V(s) and the local return function z of each joint i The change curves of (s,a). Output the change curve of the reward function during the training process.

[0171] This embodiment proposes a modeling and calculation method for the distributed return function, decomposes the global return into local returns of each joint, estimates them using a recursive method, and combines the upper and lower bounds of the dual-objective return estimation mechanism to reduce the policy instability in the high-dimensional continuous action space, thereby significantly improving the task execution efficiency and the stability of policy optimization.

[0172] By preferentially sampling high-quality data according to importance, preferentially selecting data that contributes more to policy optimization for training, significantly improving the sample utilization efficiency, accelerating policy convergence, and solving the problem of slow policy optimization in the high-dimensional action space.

[0173] Introduce entropy regularization in the policy optimization process, balance exploration and exploitation, avoid the policy falling into local optima, and adapt to the different stage requirements of the task by dynamically adjusting the reward weight parameter and entropy coefficient. For example, pay attention to the weight of position error during the grasping stage, and pay more attention to the weight of attitude error during the rotation stage to achieve more accurate task control.

[0174] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A dexterous two-handed collaborative control method based on multi-agent reinforcement learning, characterized in that: include: Initialize the neural network and environment of the dexterous hands and collect data when interacting with the environment; Based on the data from the interaction with the environment, the joints make sequential decisions according to the greedy strategy; The joints make sequential decision actions according to the greedy strategy, including: The joints are set up in a decision order according to the task requirements: the advantage value of each joint is obtained by using the advantage function decomposition theorem, and sequential decisions are made from large to small according to the advantage value; Each joint Action By local strategy Decision; the input of the local strategy is the local observation of the joint and the action of the preceding joint ; Different joints of the two hands are set as different agents, and each agent has its own strategy, which is the local strategy; The first joint Directly based on local observation Select Action ; Second joint Based on local observation and the first joint movement Select Action ; Follow this logic in sequence until the last joint , get the joint action ; Perform joint actions , pushing the system from its current state Transfer to next state ; Use the Q-value distribution network to evaluate the decision action and calculate the reward return; Based on the reward return, entropy regularization is introduced, and the joint strategy is updated using the policy gradient optimization method; Based on the preset conditions, the update of the joint strategy is terminated and the optimal collaborative control strategy is output.

2. The dexterous two-handed collaborative control method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The neural network and environment of dexterous hands include: dexterous hands strategy network, evaluation network and simulation environment; The strategy network includes: a strategy network for each joint, the input of which is environment information and its own state, and the output is joint action; The evaluation network includes: a Q-value distribution network, the input of which is environmental information, its own state and joint action, and the output of which is the evaluation of the joint action; The simulation environment includes: starting a preset simulator and a scenario; Initializing the neural network of the dexterous hands includes: initializing its weights by generating an orthogonal matrix, that is, assigning initial weights to each layer of the network by the orthogonal matrix method, thereby providing linearly independent parameter values ​​for each layer of the neural network.

3. The dexterous two-handed collaborative control method based on multi-agent reinforcement learning according to claim 2 is characterized in that: The neural network and environment of the dexterous hands also include constraints; the constraints include: Torque constraint: The torque applied by each joint should meet the limitations of hardware capabilities; Motion rate constraint: The movement amplitude of the joint must not exceed the preset threshold in each time step to avoid system instability caused by high-frequency changes; Coordination constraints of joint actions: The combination of actions of all joints must meet the coordination requirements of dynamic interaction between dexterous hands and objects.

4. The dexterous two-handed collaborative control method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The data collected during the environment interaction includes: In each time step, the system state and action of the dexterous hands collect data by interacting with the environment, including: current state , Current Action , Instant Rewards , next state and task termination flag ; The collected interaction data Stored in experience pool In which, when the preset size of the experience pool is exceeded, the first-in-first-out strategy is used to replace the old data.

5. The dexterous two-handed collaborative control method based on multi-agent reinforcement learning according to claim 4 is characterized in that: The data collected during the environment interaction also includes: performing importance sampling correction on the data stored in the experience pool; including: For each sample in the experience pool , calculate the importance weight ,in, is the probability of the current strategy, is the probability of the old strategy at the time of sampling; Based on the importance weights, combined with immediate rewards and time difference error, set the priority for the sample, the sample priority is ,in, The priority adjustment parameter is used to balance the sampling probability of high-priority samples and low-priority samples. , is the discount factor; Based on sample priority Sampling is performed, and samples that contribute to the strategy optimization preset are given priority for training.

6. The dexterous two-handed collaborative control method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The evaluating the decision action by the Q-value distribution network includes: By constructing a Q-value distribution network, fitting the Q-value distribution according to the data generated by the interaction between the environment and the agent, and obtaining the rewards of the joints, i.e., local rewards, through sampling methods, this can reduce the overestimation problem of the traditional network for the agent and enhance the collaborative effect; among them, the global reward is used to evaluate the overall performance of the dexterous hands in the process of completing the task; Normalize and smooth the local rewards.

7. The dexterous two-handed collaborative control method based on multi-agent reinforcement learning according to claim 6 is characterized in that: Calculating the reward return includes: Based on the local reward, a local reward function of each joint is calculated; wherein the local reward function is used to characterize the contribution of the action of the current joint to the long-term reward; Accumulate the local reward functions of all joints to obtain the global reward function; The local reward function is: in, For joints Partial rewards, is the discount factor, Indicates the next state and actions expectations; The global reward function is: in, is the total number of joints, For the joints.

8. The dexterous two-handed collaborative control method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The joint strategy is: local strategy for all joints Combination of ; The optimization goal of the joint strategy is to maximize the long-term return of the system under the current strategy ; Optimize policy parameters using policy gradient method ;in, For the strategy parameters Find the gradient, is the objective function, For hope, is the probability distribution of the strategy, is the logarithmic probability of the strategy, is the action value function.

9. The dexterous two-handed collaborative control method based on multi-agent reinforcement learning according to claim 8 is characterized in that: Updating the joint strategy includes: Dynamically adjust temperature parameters , the goal is to maintain the entropy value of the strategy close to the target entropy value : ; Update using gradient descent ; Combine policy gradient and entropy regularization to update policy parameters according to the following formula : ; Among them, the temperature parameter is used to explore and balance the control strategy.

Citation Information

Patent Citations

  • Robot agent reinforcement learning training method and system in complex scene

    CN119129642A