Intelligent grabbing control method and system based on multi-agent reinforcement learning

By treating each finger of a dexterous hand as an independent agent, using the method of multi-agent reinforcement learning, it adaptively adjusts the grab strategy in an uncertain environment, solving the problem of insufficient crawling accuracy and accuracy in the existing technology, achieving higher crawling success rate and stability.

CN120056123AActive Publication Date: 2025-05-30SHANDONG UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510419585.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-30
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing grasp control of human-like dexterous hands grasping objects in uncertain environments, there is a problem of insufficient grasping accuracy and accuracy, making it difficult to adapt to the physical properties and dynamic changes of objects.

Method used

The intelligent grasping control method based on multi-agent reinforcement learning is adopted, and each finger is regarded as an independent agent. Through the collaboration and information sharing of the multi-agent system, each finger can adaptively adjust the grasping strategy according to real-time feedback and handle the grasping tasks of different physical characteristics and dynamic changes.

Benefits of technology

It significantly improves the accuracy and stability of the crawling process, reduces the risk of crawling failure, is suitable for item grabbing in uncertain environments, and has the ability to learn independently and adapt dynamically.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120056123A_ABST
    Figure CN120056123A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent grabbing control method and system based on multi-agent reinforcement learning, and the method comprises the steps: taking each finger of a humanoid dexterous hand as an independent agent, generating a control action based on a respective strategy network, and forming a multi-agent system; empirical data generated in the interaction process of each agent and the environment in the multi-agent system is stored by utilizing a playback pool mechanism, the sampling priority of a playback pool of each agent is dynamically adjusted according to the task target, the learning progress and the task importance, and the playback pools share the empirical data; a staged reward mechanism is adopted, and each agent in the multi-agent system is jointly guided to optimize the position, the joint angle and the contact force in different stages of a grabbing task through individual rewards and global rewards; a multi-agent depth deterministic strategy gradient algorithm is used for training, and a trained multi-agent system is used for intelligent grabbing control. The precision and stability of the grabbing process are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent grasping control, and particularly relates to an intelligent grasping control method and system based on multi-agent reinforcement learning. Background Art

[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] At present, the technology of robots is developing rapidly, and its application fields have shifted from traditional industrial production to home life. As a robot hand that can simulate human hand movements and has multi-degree-of-freedom and high-precision control capabilities, compared with traditional industrial robot grasping tools, the anthropomorphic dexterous hand has higher flexibility and adaptability, and can handle objects with complex shapes and uneven weights. Therefore, it has important potential in a variety of application scenarios. The anthropomorphic dexterous hand has a five-finger structure similar to that of a human hand. Its grasping control can not only improve the operation ability of the robot in a complex environment, but also significantly improve the adaptability and efficiency of the robot in practical applications. It has a wide range of application prospects in the fields of industrial production, medical treatment, service robots, and logistics distribution.

[0004] However, in practical applications, the design of the grasping controller for the anthropomorphic dexterous hand is still challenging. Especially in an uncertain environment, when the shape, size, weight, etc. of the object to be grasped are unknown, its grasping accuracy (such as grasping success rate, grasping pose stability, etc.) and precision (grasping force accuracy, running trajectory accuracy, etc.) are significantly insufficient, seriously affecting the grasping efficiency and popularization and application of the multi-fingered anthropomorphic dexterous hand.

[0005] Traditional anthropomorphic dexterous hand grasping control methods are mostly based on rules and experience, relying on manual parameter adjustment, and usually cannot effectively cope with complex unstructured environments. These methods often have difficulty adapting to the physical properties of the object (such as material, shape, surface smoothness, etc.), initial state (such as the placement posture of the object, initial velocity, etc.), motion state (such as velocity, acceleration, angular acceleration, etc.), and task objectives (such as stable grasping, rapid lifting or moving of the object, etc.) when grasping an object. Therefore, the flexibility, stability, and adaptability of these control methods are often poor, especially when facing unknown or dynamic environments, it is easy to cause grasping failures.

[0006] In recent years, with the rapid development of artificial intelligence technologies such as deep learning and reinforcement learning, data-driven humanoid dexterous hand grasping control methods have gradually attracted attention. These methods can be trained with a large amount of data, enabling the humanoid dexterous hand to adaptively adjust its strategy during the grasping process, thereby improving the accuracy and precision of grasping. Especially in the application of reinforcement learning, the humanoid dexterous hand can gradually optimize the grasping strategy through interaction with the environment to handle complex grasping tasks. However, these methods still have some problems, especially the need for a large amount of data, accurate annotation, high training time, as well as poor learning generalization ability, online control performance, and control accuracy, stability, and adaptability during the training process. These drawbacks limit the universality and efficiency of reinforcement learning in practical applications. On the other hand, traditional methods usually face the problems of requiring a large amount of data and high training time, and lack sufficient learning generalization ability and online control performance, which is particularly obvious in grasping tasks in complex environments. Summary of the Invention

[0007] To solve the above problems, the present invention proposes an intelligent grasping control method and system based on multi-agent reinforcement learning. By treating each finger of the humanoid dexterous hand as an independent agent and utilizing the cooperation and information sharing of the multi-agent system, each finger can adaptively adjust the grasping strategy according to real-time feedback, handle different physical characteristics of the object, and cope with dynamically changing grasping tasks, significantly improving the accuracy and stability of the grasping process, avoiding the sliding or instability of the object during object control and grasping, reducing the risk of grasping failure, and being particularly suitable for grasping items in uncertain environments.

[0008] According to some embodiments, the present invention adopts the following technical solutions:

[0009] An intelligent grasping control method based on multi-agent reinforcement learning, comprising the following steps:

[0010] Treat each finger of the humanoid dexterous hand as an independent agent, obtain its own local state information from the global state, and generate control actions based on their respective policy networks to form a multi-agent system;

[0011] Use the replay pool mechanism to store the experience data generated during the interaction between each agent in the multi-agent system and the environment. The sampling priority of each agent's replay pool is dynamically adjusted according to its task objective, learning progress, and task importance, and the experience data is shared among the replay pools;

[0012] Adopt a phased reward mechanism, and jointly guide each agent in the multi-agent system to optimize the position, joint angle, and contact force respectively at different stages of the grasping task through individual rewards and global rewards;

[0013] Train using the multi-agent deep deterministic policy gradient algorithm, and use the trained multi-agent system for intelligent grasping control.

[0014] As an alternative implementation, the local state information includes joint angles, finger positions, contact forces, and the characteristics of the target object, and is used to guide the grasping actions of the corresponding fingers.

[0015] As an alternative implementation, use the replay pool mechanism to store the experience data generated during the interaction between each agent in the multi-agent system and the environment. The process of dynamically adjusting the sampling priority of each agent's replay pool according to its task objective, learning progress, and task importance includes: each agent's replay pool preferentially selects relevant experience data according to its task objective, and the task objective priority is determined according to the ratio of the experience score related to the current task objective to the experience scores related to all task objectives;

[0016] The replay pool preferentially selects experience data with a temporal difference error greater than the set value in the initial stage, and preferentially samples experience data with diversity in the later stage;

[0017] In a multi-task learning environment, the replay pool preferentially samples experience data of more important tasks.

[0018] As a further step, the process of dynamic adjustment includes: in the initial stage of training, each agent relies more on experience with a temporal difference error greater than the set value; in the later stage of training, gradually increase the proportion of diversity sampling, and the dynamic sampling weight is:

[0019]

[0020] where Q quality (e t ) is the quality score of experience e t , which is calculated from the temporal difference error and experience diversity, S task (e t ) represents the relevance of experience e t to the current task objective, and N is the total number of all samples in the experience pool.

[0021] As an alternative implementation, the process of sharing experience data between replay pools includes: when some agents' replay pools lack certain experiences, obtain relevant experience data from other agents' replay pools through the experience sharing mechanism. The experience sharing mechanism allows each agent to accelerate learning by sharing data in the experience pool without direct interaction.

[0022] As an alternative implementation, a phased reward mechanism is adopted. The process of jointly guiding each agent in the multi-agent system to optimize the position, joint angle, and contact force at different stages of the grasping task through individual rewards and global rewards includes: The phased reward mechanism consists of two stages. The first stage is the joint position reward and finger position reward stage, and the second stage is the contact force reward stage. During the grasping process, the agent automatically switches the reward calculation method according to the task objectives of the current stage, and each stage is jointly guided by individual rewards and global rewards.

[0023] As a further defined implementation, the individual reward is determined by comprehensively considering the joint angle, finger position, and contact force to ensure that each finger can exhibit optimal behavior in the grasping task. Specifically:

[0024]

[0025] Among them, p i (t) is the current position of the i-th finger, p target is the target position, θ i (t) is the joint angle of the i-th finger, θ initial is the initial angle of the joint, f i (t) is the contact force of the i-th finger, f target is the desired contact force, ω position , ω joint , ω force are the corresponding weight coefficients used to control the influence of different factors.

[0026] As a further defined implementation, the global reward is the weighted sum of the individual rewards of all agents, representing the task completion situation of the entire multi-agent system. The global reward is used to reflect the overall effect of all fingers collaborating to complete the task. The global reward is:

[0027]

[0028] Among them, r i (t) is the individual reward of the i-th finger, and N represents the number of fingers in the system.

[0029] As a further defined implementation, in the first stage, when the finger is not in contact with the object, the reward is determined by the difference between the finger position and the joint angle. The closer the finger position is to the target position, the greater the position reward; the larger the joint angle, the greater the joint reward. In the second stage, when the finger is in contact with the object, the reward is determined only by the contact force. The closer the contact force is to the desired contact force, the greater the reward.

[0030] As a further defined implementation, the reward stage is determined based on whether the contact force of the finger exceeds a set threshold. If the contact force of the finger is less than or equal to the set threshold, it is in the first stage; otherwise, it is in the second stage.

[0031] As a further defined implementation, the overall reward considers the task completion of each finger and the grasping effect of the entire system, and adjusts the priorities of different task objectives through weight coefficients. Specifically:

[0032]

[0033] Among them, r i (t) is the individual reward of the i-th finger at time step t, ω individual , ω global are weight coefficients that control the contribution ratios of individual rewards and global rewards in the overall reward. At the initial stage of training, the individual reward occupies a larger proportion, and as training progresses, the weight of the global reward is gradually increased.

[0034] As an alternative implementation, the process of training using the multi-agent deep deterministic policy gradient algorithm includes:

[0035] Set the number of agents, initialize the environmental state, reward function, and experience replay pool;

[0036] In each time step, each agent selects an action according to the local state through the policy network. The actions of all agents are combined into a joint action, the joint action is executed, a new state and reward are obtained, and the current experience is stored in the experience replay pool of each agent;

[0037] Each agent samples based on the importance and diversity of the experience, preferentially samples the experience that is most helpful for the current task, optimizes the policy network through the policy gradient algorithm, updates the value network through the temporal difference error, and evaluates the value of the current action;

[0038] Judge the current stage. If the finger touches the object, switch to the second stage, and the reward is calculated only based on the contact force; otherwise, it is in the first stage, and the reward is calculated based on the finger position and joint angle;

[0039] Judge whether to switch stages according to the contact force. If the contact force exceeds the threshold, enter the second stage;

[0040] At the end of each round of training, output the total reward, judge whether the convergence condition is reached, and if so, stop training.

[0041] As a further defined implementation, the process of optimizing the policy network through the policy gradient algorithm and updating the value network through the temporal difference error includes:

[0042] The process of updating the policy network includes:

[0043]

[0044] Among them, θ i is the policy network parameter of the i-th agent, and π i (a i |s i ) is the policy of the i-th agent, which is the probability of selecting action a i when the state s i is given. Q i (s i ,a i ) is the value network function, which represents the expected return of the agent taking action a i under the state s i . J represents the maximized expected return of the i-th agent, denotes the expectation under the policy π i ;

[0045] The value network is updated by the temporal difference error:

[0046] L(θ i ) = E[(Q i (s i ,a i ) - (r i + γQ i (s i+1 ,a i+1 ))) 2 ;

[0047] Among them, γ is the discount factor, which represents the importance of future rewards, and r i represents the immediate reward obtained by the agent from the environment at the i-th time step.

[0048] An intelligent grasping control system based on multi-agent reinforcement learning includes:

[0049] A multi-agent system construction module, configured to regard each finger of the humanoid dexterous hand as an independent agent, obtain its own local state information from the global state, and generate control actions based on their respective policy networks to form a multi-agent system;

[0050] A replay pool establishment module, configured to use the replay pool mechanism to store the experience data generated during the interaction between each agent in the multi-agent system and the environment. The sampling priority of each agent's replay pool is dynamically adjusted according to its task objectives, learning progress, and task importance, and the experience data is shared among the replay pools;

[0051] The reward mechanism optimization module is configured to adopt a phased reward mechanism, and jointly guide each agent in the multi-agent system to optimize the position, joint angle, and contact force at different stages of the grasping task through individual rewards and global rewards;

[0052] The intelligent grasping control module is configured to be trained using the multi-agent deep deterministic policy gradient algorithm, and use the trained multi-agent system for intelligent grasping control.

[0053] A humanoid dexterous hand applies the above method or includes the above system.

[0054] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0055] In the present invention, by treating each finger as an independent agent and utilizing the cooperation and information sharing of the multi-agent system, each finger can adaptively adjust the grasping strategy according to real-time feedback, handle different physical characteristics of the object, and cope with dynamically changing grasping tasks. It can achieve grasping behaviors with anthropomorphic characteristics, and at the same time has the abilities of autonomous learning and dynamic adaptation, significantly improving the accuracy and stability of the grasping process, avoiding the sliding or instability of the object during the object control and grasping process, and reducing the risk of grasping failure.

[0056] Through the collaborative cooperation of the multi-agent system, the present invention optimizes the task allocation and enhances the robustness and fault tolerance of the system. When a single finger fails, other agents can timely adjust their actions to ensure the completion of the grasping task. This mechanism significantly improves the adaptability and fault tolerance of the system and can still maintain a high grasping success rate in a dynamically changing environment.

[0057] The present invention adopts a parallel learning method, which greatly accelerates the optimization of the grasping strategy, significantly improves the learning efficiency and system adaptability. By adapting to the dynamic state of the object in real time, the present invention breaks through the limitations of traditional methods in grasping control under uncertain environments, provides a new idea for the grasping control of a five-finger humanoid dexterous hand, and has strong practical application potential and technical value.

[0058] By introducing the adaptive experience replay pool and personalized replay pool mechanisms, the present invention can effectively manage the experience data of each agent, dynamically adjust the experience sampling strategy according to the task objective, learning progress, and task importance, thereby accelerating the learning process and improving the stability and efficiency of training; through joint action generation and agent cooperation, multiple fingers (agents) can cooperate to complete the grasping task, and at the same time, the personalized experience replay pool is adopted to promote the experience sharing among agents, improving the cooperation efficiency among multi-agents.

[0059] By introducing a replay pool mechanism with diverse sampling, the agent can avoid overfitting, enhance the generalization ability of the system, and enable the agent to maintain good adaptability and stability in more complex and dynamic environments.

[0060] The present invention adopts a phased reward mechanism to optimize the finger position, joint angle, and contact force respectively at different stages of the grasping task, enabling the agent to adaptively adjust the strategy under different task objectives, thereby improving the grasping accuracy and stability, and ultimately achieving a higher grasping success rate.

[0061] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given in conjunction with the accompanying drawings and described in detail as follows. Brief Description of the Drawings

[0062] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0063] Figure 1 It is a flowchart of the method implementation for an embodiment;

[0064] Figure 2 It is a schematic diagram of the replay pool mechanism for agent individuation in an embodiment;

[0065] Figure 3 It is a schematic diagram of the hardware system integration in an embodiment, where 1 - 5 respectively represent the thumb, index finger, middle finger, ring finger, and little finger (i.e., five agents), 6 represents the robotic arm, 7 represents the anthropomorphic dexterous hand body, 8 represents the upper-level planning controller of the robot, and 9 represents 23 kinds of irregular objects to be grasped.

[0066] Figure 4 It is a flowchart of the application steps of the anthropomorphic dexterous hand in an embodiment. Detailed Description of the Specific Embodiments

[0067] The present invention will be further described below in conjunction with the drawings and embodiments.

[0068] It should be noted that the following detailed descriptions are all illustrative and are intended to provide a further description of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0069] Note that the terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly dictates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0070] In the case of no conflict, the embodiments in this application and the features in the embodiments may be combined with each other.

[0071] Embodiment 1

[0072] An intelligent grasping control method based on multi-agent reinforcement learning. In this embodiment, a five-finger anthropomorphic dexterous hand is taken as an example for illustration, which includes the following key points:

[0073] Regarding the five fingers as independent agents, they obtain their own local state information from the global state and generate control actions based on their respective policy networks; these actions are integrated through a joint cooperation module to form an overall grasping operation on the target object.

[0074] The individualized replay pool mechanism for agents is used to store the state, action, reward, and next state information generated during the interaction between the agent and the environment. This mechanism can dynamically adjust the storage and sampling strategies of the replay pool according to the learning progress, task objectives, and importance of the samples of each agent, thereby improving the utilization efficiency of experience, accelerating the training process, and enhancing the stability of the model.

[0075] In a multi-process parallel training environment, based on the strategy of centralized training and distributed execution, and multi-agent reinforcement learning optimizes the collaborative strategy of the five fingers through a global reward function, enabling it to flexibly adapt to the shape, position, and dynamic changes of the object in a complex environment.

[0076] The following is a detailed introduction:

[0077] Step 1: Construct a multi-agent system for a five-finger anthropomorphic dexterous hand

[0078] Regard each finger of the five-finger anthropomorphic dexterous hand as an independent agent (thumb, index finger, middle finger, ring finger, and little finger). Each agent obtains its own local state information according to the global state in the environment. These local state information includes joint angles, finger positions, contact forces, and characteristics of the target object, etc., which are used to guide the grasping actions of the fingers.

[0079] Step 2: The individualized replay pool mechanism for agents

[0080] In reinforcement learning, the experience replay pool is an important component. Storing experiences reasonably can accelerate the learning process and improve the stability of the model. In a humanoid dexterous hand multi-agent system, a large amount of experience data is generated when multiple fingers interact with the environment. These experience data will be stored in the experience replay pool and provide feedback during the training process. Each experience record is defined as:

[0081] e t =(S t ,A t ,R t ,S t+1 ) (1)

[0082] where e t represents the experience stored at time step t, S t represents the global state of the agent at time step t, A t represents the action selected and executed in state S t , R t represents the reward obtained after executing action A t , and S t+1 represents the next global state to which the environment transfers after the agent executes A t . These data are stored in the replay pool. As training progresses, new experiences will gradually replace old ones. The key to the adaptive experience replay pool lies in evaluating the importance and diversity of each experience and adjusting the storage and sampling strategies based on these evaluation results to ensure the efficiency and stability of training.

[0083] In a humanoid dexterous hand system or a multi-task learning environment, the task objectives, learning progress, and task priorities of each finger may be different. Therefore, each finger requires a personalized replay pool that can dynamically adjust the sampling strategy according to its current task requirements and learning status. To achieve this, as Figure 2 shown, this embodiment proposes a replay pool mechanism for multi-agent individuation, whose core goal is: in an adaptive manner, ensure that each agent preferentially selects the experiences that are most helpful for its learning from the replay pool.

[0084] ① Personalized replay pool design

[0085] The replay pool of each agent is dynamically adjusted according to factors such as its task objective, learning progress, and task importance. Specifically, the sampling strategy of the replay pool will preferentially select the experience data that is most relevant to the current task objective and learning stage of the finger.

[0086] a. Association of task objectives: The replay pool for each finger will preferentially select relevant experiences according to its task objective. For example, if a finger's task is grasping stability, the replay pool will preferentially store experience data related to the stability of grasping an object. In this way, the experiences in the replay pool can always be highly relevant to the task objective.

[0087] Task objective priority adjustment formula:

[0088]

[0089] Where Q task (e t ) is the experience score related to the current task objective, and P task (e t ) is the priority sampling probability of experience e t .

[0090] b. Influence of finger learning progress: The learning progress of a finger will directly affect the update strategy of its experience pool. In the initial stage of learning, a finger usually needs more importance-weighted experiences to quickly correct its strategy, so the replay pool will preferentially select those experience data with larger (Temporal Difference Error, TD) errors. In the later stage of training, the finger's strategy has tended to be stable. At this time, the replay pool samples more diverse experiences to enhance the finger's adaptability in different tasks and environments. The learning progress formula is defined as follows:

[0091] P progress (e t ) = α·|δ t | + β·S task (e t ) (3)

[0092] Where α and β are weight coefficients that respectively control the influence of TD error and task relevance in the learning process, δ t is the TD error, and S task (e t ) is the score of the correlation between the current experience and the task objective.

[0093] c. Task importance and replay pool: In a multi-task learning environment, the sampling priority of the replay pool will be dynamically adjusted according to the importance and difficulty of the tasks. For example, for more important tasks, the replay pool will preferentially sample the experiences of that task, while for less important tasks, the sampling frequency of their experiences will be reduced. This ensures that in a multi-task learning environment, the finger can preferentially learn the experiences of important tasks and improve the overall performance. It is defined as follows:

[0094] S task (e t) = f(TaskPriority, TaskDifficulty, TaskLearningProgress) (4)

[0095] Among them, TaskPriority is the priority of the task, TaskDifficulty is the difficulty of the task, and TaskLearningProgress is the learning progress of the finger on this task.

[0096] ② Experience sharing among tasks

[0097] In a multi-agent system, multiple agents usually cooperate to complete common tasks, so their experience pools can be shared with each other. In the replay pool of finger individuation, when some key experiences are lacking in the replay pool of certain fingers, relevant experience data can be obtained from the replay pools of other fingers through the experience sharing mechanism. By sharing experiences, fingers can more quickly supplement and improve their own knowledge bases, promoting the collaboration and learning efficiency of the entire system. The experience sharing mechanism allows fingers to accelerate learning by sharing data in the experience pool without direct interaction. Experience sharing among fingers not only helps to reduce unnecessary repeated learning but also improves the overall collaborative effect of the system. The experience sharing formula is defined as follows:

[0098]

[0099] Among them, P shared (e t ) represents whether agent t shares experience e t , and N(t) represents the set of other fingers collaborating with agent t.

[0100] ③ Dynamic priority sampling of the replay pool

[0101] As the finger learning progresses, the sampling priority in the replay pool will be dynamically adjusted according to the finger's task objectives, learning progress, and task importance. In the initial stage of training, the finger will rely more on experiences with larger TD errors; while in the later stage of training, the proportion of diversity sampling will be gradually increased to avoid overfitting and improve the generalization ability of the agent. The dynamic sampling weight formula is defined as follows:

[0102]

[0103] Among them, Q quality (e t ) is the quality score of experience e t , calculated from the TD error and experience diversity, S task (e t ) represents the relevance of experience e t to the current task objective, and N is the total number of all samples in the experience pool.

[0104] ④Independent management of the task pool

[0105] In a multi-task learning or multi-finger environment, the replay pool for each task can be independently managed. Each task replay pool adjusts the sampling strategy according to the progress, difficulty, and importance of the task, ensuring that the agent can focus on the experience of the current task while not neglecting the learning needs of other tasks. The formula for the independent task pool is defined as follows:

[0106] P task (e t ) = f(TaskPriority, TaskProgress, TaskDifficulty) (7)

[0107] Where TaskPriority is the priority of the task, TaskProgress is the learning progress of the finger on this task, and TaskDifficulty is the difficulty of the task.

[0108] Step 3: Stage-based reward mechanism

[0109] The reward mechanism proposed in the present invention aims to optimize the multi-agent grasping task, guiding the agent (finger) to gradually adjust its behavior during the grasping process through individual rewards and global rewards. The reward mechanism is designed in two stages: the first stage (joint position reward and finger position reward) and the second stage (contact force reward). During the grasping process, the agent automatically switches the reward calculation method according to the task objectives of the current stage.

[0110] ①Individual reward

[0111] The individual reward is designed by comprehensively considering the joint angle, finger position, and contact force, ensuring that each finger can exhibit optimal behavior in the grasping task. The simplified formula for the individual reward is as follows:

[0112]

[0113] Where p i (t) is the current position of the i-th finger, p target is the target position, θ i (t) is the joint angle of the i-th finger, θ initial is the initial angle of the joint, f i (t) is the contact force of the i-th finger, f target is the desired contact force, ω position , ω joint , ω force are the corresponding weight coefficients, controlling the influence of different factors.

[0114] It should be noted that in the first stage (when not yet in contact with the object), the reward is determined by the differences in finger position and joint angle. The closer the finger position is to the target position, the greater the position reward. The larger the joint angle (before contacting the object), the greater the joint reward. In the second stage (when the finger is in contact with the object), the reward is determined only by the contact force. The closer the contact force is to the desired contact force, the greater the reward.

[0115] ② Global Reward

[0116] The global reward is the weighted sum of all individual finger rewards and represents the task completion of the entire multi-agent system. The global reward is used to reflect the overall effect of all fingers collaborating to complete the task. The formula for calculating the global reward is as follows:

[0117]

[0118] where r i (t) is the individual reward of the i-th finger. N represents the number of fingers in the system.

[0119] ③ Phase Switching Mechanism

[0120] The reward mechanism dynamically switches the calculation method of the reward based on whether the finger is in contact with the object. If the contact force f i (t) of the finger exceeds the set threshold f threshold , it means the finger has come into contact with the object, and at this time, it switches to the contact force reward phase.

[0121] is_contacted(t) = Ι(f i (t) > f threshold ) (10)

[0122] where Ι(·) is the exponential function, indicating whether to enter the second phase. If the contact force exceeds the threshold, it returns 1, indicating entering the second phase; otherwise, it returns 0, still in the first phase.

[0123] ④ Reward Optimization and Balance

[0124] The individual reward and the global reward are combined into a comprehensive overall reward R t This reward is a quantitative feedback on the overall performance of the multi-agent system (fingers). The overall reward takes into account the task completion of each finger and the grasping effect of the entire system, and adjusts the priorities of different task objectives through weight coefficients.

[0125]

[0126] where r i (t) is the individual reward of the i-th finger at time step t, ω individual , ω globalis the weight coefficient that controls the contribution ratio of the individual reward and the global reward in the overall reward. These two variables will be gradually adjusted as the training progresses. In the initial stage of training, the agent needs to pay more attention to the joint positions and finger positions, so the individual reward occupies a larger proportion. As the training progresses, the system gradually turns to optimize the contact force and grasping stability, and at this time, the weight of the global reward increases.

[0127] Step 4: Training of the MADDPG algorithm

[0128] This step describes the training process based on the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm, focusing on the policy network update and the value network update. Through these updates, the agents can optimize their behavioral strategies according to the rewards they obtain, and achieve cooperation and the completion of task goals in the multi-agent grasping task.

[0129] ① Policy network

[0130] In MADDPG, each agent selects actions through a deterministic policy. The policy network generates corresponding actions based on the current state, with the goal of maximizing the long-term return of each agent. The policy update of each agent is carried out through the following formula:

[0131]

[0132] where θ i is the policy network parameter of the i-th agent, π i (a i |s i ) is the policy of the i-th agent, which is the probability of selecting action a i given state s i , Q i (s i , a i ) is the value network function, which represents the expected return of the agent taking action a i in state s i , J represents the maximized expected return of the i-th agent, represents the expectation under policy π i .

[0133] ② Value network

[0134] The value network of each agent is used to evaluate the value of the current state and action, guiding the agent on how to select the best action. The value network is updated through the TD error:

[0135] L(θ i ) = E[(Q i (si , a i ) - (r i +γQ i (s i+1 , a i+1 ))) 2 (13)

[0136] Among them, γ is the discount factor, representing the importance of future rewards. Through the TD error, the value network is continuously updated to ensure that it can accurately evaluate the value of each state-action pair. The optimization of the value network helps the agent better estimate future rewards, thereby improving the performance of the policy network.

[0137] In summary, for the intelligent grasping controller based on multi-agent reinforcement learning, its training process is as follows:

[0138] 1. Set the number of agents, and initialize the environmental state, reward function, and experience pool.

[0139] 2. At each time step, perform the following operations:

[0140] a. Each agent selects an action according to the local state through the policy network.

[0141] b. Joint action generation: The actions of all agents are combined into a joint action.

[0142] c. Execute the joint action to obtain a new state and reward.

[0143] d. Store the current experience in the experience replay pool of each agent.

[0144] 3. Experience sampling and replay pool update:

[0145] a. Each agent samples based on the importance and diversity of the experience, and preferentially samples the experience that is most helpful for the current task.

[0146] b. Update the policy network: Optimize the policy network through the policy gradient algorithm to increase the long-term reward.

[0147] c. Update the value network: Update the value network through the TD error to evaluate the value of the current action.

[0148] 4. Stage-based reward mechanism:

[0149] Judge the current stage (joint position stage or contact force stage): If the finger touches the object, switch to the contact force stage, and the reward is calculated only based on the contact force. Otherwise, it is in the joint position stage, and the reward is calculated based on the finger position and joint angle.

[0150] 5. Stage switching:

[0151] Determine whether to switch phases based on the contact force. If the contact force exceeds the threshold, enter the second phase (contact force reward).

[0152] 6. End condition:

[0153] At the end of each round of training, output the total reward. Determine whether the convergence condition (reward is stable or the maximum number of steps is reached) is met. If so, stop training.

[0154] As a typical embodiment, as Figure 3 shown, it is applied to a five-finger humanoid dexterous hand. The five-finger humanoid dexterous hand body 7 is connected to the robotic arm 6. The upper-level planning controller 8 of the robot is used to control the actions of the five-finger humanoid dexterous hand body 7, and can control the thumb 1, index finger 2, middle finger 3, ring finger 4, and little finger 5 to cooperate in grasping various objects.

[0155] The upper-level planning controller 8 of the robot is equipped with a storage medium, and the storage medium stores the method provided in the first embodiment or the trained multi-agent system (model).

[0156] During a specific grasping task, as Figure 4 shown, it includes the following steps:

[0157] Set the number of agents and set the reward function;

[0158] Initialize the experience replay pool;

[0159] Obtain the global state, which includes the actions, position states, and immediate rewards of all agents, and store the experience data in the agent experience replay pool;

[0160] Select the actions of each finger / agent;

[0161] Combine the actions of each agent to obtain a joint action and execute it;

[0162] Judge the current reward phase and determine whether to switch. If so, switch to the second phase; otherwise, maintain the first phase;

[0163] Update the policy network and the value network;

[0164] Judge whether the grasping is successful. If so, end the grasping task; otherwise, return to the step of selecting the actions of each finger / agent.

[0165] Embodiment 2

[0166] An intelligent grasping control system based on multi-agent reinforcement learning, including:

[0167] The multi-agent system construction module is configured to regard each finger of the humanoid dexterous hand as an independent agent, obtain its own local state information from the global state, and generate control actions based on their respective policy networks to form a multi-agent system;

[0168] The return visit pool establishment module is configured to use the replay pool mechanism to store the experience data generated during the interaction between each agent in the multi-agent system and the environment. The sampling priority of each agent's replay pool is dynamically adjusted according to its task objective, learning progress, and task importance, and the experience data is shared among the replay pools;

[0169] The reward mechanism optimization module is configured to adopt a phased reward mechanism, and jointly guide each agent in the multi-agent system to optimize the position, joint angle, and contact force respectively at different stages of the grasping task through individual rewards and global rewards;

[0170] The intelligent grasping control module is configured to be trained using the multi-agent deep deterministic policy gradient algorithm, and use the trained multi-agent system for intelligent grasping control.

[0171] Embodiment III

[0172] A humanoid dexterous hand applying the method of Embodiment I or including the system of Embodiment II.

[0173] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0174] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0175] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0177] The foregoing are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made by those skilled in the art without creative efforts within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intelligent grasping control method based on multi-agent reinforcement learning, characterized in that: The following steps are involved: Each finger of the humanoid dexterous hand is regarded as an independent intelligent agent, which obtains its own local state information from the global state and generates control actions based on its own strategy network to form a multi-agent system. The replay pool mechanism is used to store the experience data generated by each agent in the multi-agent system during the interaction with the environment. The sampling priority of each agent replay pool is dynamically adjusted according to its task goal, learning progress and task importance, and the experience data is shared between replay pools. A phased reward mechanism is adopted to guide each agent in the multi-agent system to optimize the position, joint angle and contact force at different stages of the grasping task through individual rewards and global rewards. The multi-agent deep deterministic policy gradient algorithm is used for training, and the trained multi-agent system is used for intelligent grasping control.

2. The intelligent grasping control method based on multi-agent reinforcement learning as claimed in claim 1, characterized in that: The replay pool mechanism is used to store the experience data generated by each agent in the multi-agent system during the interaction with the environment. The sampling priority of each agent's replay pool is dynamically adjusted according to its task goal, learning progress and task importance. The process includes: the replay pool of each agent prioritizes the relevant experience data according to its task goal, and the task goal priority is determined by the ratio of the experience score related to the current task goal to the experience score related to all task goals; The playback pool gives priority to sampling empirical data with a time series difference error greater than a set value in the early stage, and gives priority to sampling empirical data with diversity in the later stage; In a multi-task learning environment, the replay pool prioritizes sampling of experience data for more important tasks; Or further, the dynamic adjustment process includes: in the early stage of training, each agent relies more on the experience with a temporal difference error greater than the set value; in the later stage of training, the proportion of diversity sampling is gradually increased, and the dynamic sampling weight is: Among them, Q quality (e t ) is the experience e t The quality score of is calculated by the temporal difference error and empirical diversity, S task (e t ) represents experience e t The relevance to the current task goal, N is the total number of samples in the experience pool.

3. The intelligent grasping control method based on multi-agent reinforcement learning as claimed in claim 1, characterized in that: The process of sharing experience data between replay pools includes: when some agents lack certain experience in their replay pools, they obtain relevant experience data from the replay pools of other agents through the experience sharing mechanism. The experience sharing mechanism allows each agent to accelerate learning by sharing data in the experience pool without direct interaction.

4. The intelligent grasping control method based on multi-agent reinforcement learning as claimed in claim 1, characterized in that: The process of adopting a staged reward mechanism to jointly guide each agent in the multi-agent system to optimize the position, joint angle and contact force at different stages of the grasping task through individual rewards and global rewards includes: the staged reward mechanism includes two stages, the first stage is the joint position reward and finger position reward stage, and the second stage is the contact force reward stage. During the grasping process, the agent automatically switches the reward calculation method according to the task goal of the current stage, and each stage is jointly guided by individual rewards and global rewards; Or further, the individual reward is determined by comprehensively considering the joint angle, finger position and contact force to ensure that each finger can perform optimal behavior in the grasping task, specifically: Among them, p i (t) is the current position of the ith finger, p target is the target position, θ i (t) is the joint angle of the ith finger, θ initial is the initial angle of the joint, f i (t) is the contact force of the ith finger, f target is the desired contact force, ω position ,ω joint ,ω force is the corresponding weight coefficient, which is used to control the influence of different factors; The global reward is the weighted sum of the individual rewards of all agents, representing the task completion of the entire multi-agent system. The global reward is used to reflect the overall effect of all fingers working together to complete the task. The global reward is: Among them, r i (t) is the individual reward of the ith finger, and N represents the number of fingers in the system.

5. The intelligent grasping control method based on multi-agent reinforcement learning as claimed in claim 4, characterized in that In the first stage, when the finger is not in contact with the object, the reward is determined by the difference between the finger position and the joint angle. The closer the finger position is to the target position, the greater the position reward. The greater the joint angle, the greater the joint reward; In the second stage, when the finger contacts the object, the reward is determined only by the contact force. The closer the contact force is to the expected contact force, the greater the reward. The reward stage is determined according to whether the contact force of the finger exceeds the set threshold. If the contact force of the finger is less than or equal to the set threshold, it is in the first stage, otherwise it is in the second stage.

6. The intelligent grasping control method based on multi-agent reinforcement learning as claimed in claim 1, characterized in that: The overall reward takes into account the task completion of each finger and the grasping effect of the entire system, and adjusts the priority of different task objectives through weight coefficients, specifically: Among them, r i (t) is the individual reward of the ith finger at time step t, ω individual ,ω global It is the weight coefficient, which controls the contribution ratio of individual reward and global reward in the total reward. In the early stage of training, individual reward accounts for a larger proportion. As the training progresses, the weight of global reward gradually increases.

7. The intelligent grasping control method based on multi-agent reinforcement learning as claimed in claim 1, characterized in that: The process of training using a multi-agent deep deterministic policy gradient algorithm includes: Set the number of agents, initialize the environment state, reward function and experience replay pool; At each time step, each agent selects an action based on the local state through the policy network. The actions of all agents are combined into a joint action, which is executed to obtain a new state and reward, and the current experience is stored in the experience replay pool of each agent. Each agent samples based on the importance and diversity of experience, giving priority to the experience that is most helpful for the current task, optimizing the policy network through the policy gradient algorithm, updating the value network through the temporal difference error, and evaluating the value of the current action; Determine the current stage. If the finger touches the object, switch to the second stage and the reward is calculated based on the contact force only. Otherwise, it is in the first stage and the reward is calculated based on the finger position and joint angle. The contact force determines whether to switch the phase. If the contact force exceeds the threshold, the second phase is entered. At the end of each round of training, the total reward is output to determine whether the convergence condition has been reached. If so, training is stopped.

8. The intelligent grasping control method based on multi-agent reinforcement learning as claimed in claim 7, characterized in that: The process of optimizing the policy network through the policy gradient algorithm and updating the value network through the temporal difference error includes: The process of updating the policy network includes: Among them, θ i is the policy network parameter of the ith agent, π i (a i |s i ) is the strategy of the ith agent, given the state s i When selecting action a i The probability of Q i (s i ,a i ) is the value network function, which indicates that the agent is in state s i Take action a i The expected return of the ith agent, J represents the maximum expected return of the ith agent, E πi In the strategy π i expectations; The value network is updated through the temporal difference error update: Among them, γ is the discount factor, which indicates the importance of future rewards, and ri represents the immediate reward obtained by the agent from the environment at the i-th time step.

9. An intelligent grasping control system based on multi-agent reinforcement learning, characterized in that it includes: A multi-agent system building module is configured to treat each finger of the humanoid dexterous hand as an independent agent, obtain its own local state information from the global state, and generate control actions based on their respective policy networks to form a multi-agent system; A replay pool establishment module is configured to use the replay pool mechanism to store the experience data generated by each agent in the multi-agent system during the interaction with the environment. The sampling priority of each agent replay pool is dynamically adjusted according to its task goal, learning progress and task importance, and the experience data is shared between replay pools. The reward mechanism optimization module is configured to adopt a phased reward mechanism, which guides each agent in the multi-agent system to optimize the position, joint angle and contact force at different stages of the grasping task through individual rewards and global rewards; The intelligent grasping control module is configured to use the multi-agent deep deterministic policy gradient algorithm for training and use the trained multi-agent system for intelligent grasping control.

10. A humanoid dexterous hand, characterized in that: A method according to any one of claims 1 to 8 or a system according to claim 9.

Citation Information

Patent Citations

  • Multi-agent deep reinforcement learning strategy optimization method based on attention mechanism

    CN113392935A

  • Five-finger dexterous robot arm control method based on multi-agent deep reinforcement learning

    CN116330290A

  • Intelligent agent control method and device, equipment and storage medium

    CN117518907A

  • Sparse reward-oriented deep reinforcement learning mechanical arm grabbing method

    CN118493388A

  • Multi-unmanned aerial vehicle communication resource allocation method based on multi-agent reinforcement learning

    CN118632356A