Dexterous hand adaptive control method and device for multi-agent collaborative reinforcement learning
By constructing a simulation environment and linkage rules for a multi-degree-of-freedom bionic dexterous hand, configuring independent intelligent agents, and adopting a multi-agent proximal strategy optimization algorithm, the problem of multi-finger collaborative grasping control of a high-degree-of-freedom dexterous hand was solved, achieving improvements in high precision, stability, and adaptability.
Patent Information
- Application Number
- CN202510956975.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-14
AI Technical Summary
Existing technologies have insufficient control accuracy and stability in the multi-finger collaborative grasping control of high-degree-of-freedom dexterous hands, making it difficult to meet real-time requirements. In addition, there is a lack of a reinforcement learning training framework for multi-agent collaborative operations, resulting in a low grasping success rate.
By constructing a simulation environment for a multi-degree-of-freedom bionic dexterous hand, linkage rules are generated to constrain finger movements, and each finger is configured as an independent intelligent agent. A multi-agent proximal strategy optimization algorithm is used to train a collaborative grasping strategy model, and control instructions are generated in combination with real-time sensor data to drive the bionic dexterous hand to perform adaptive grasping operations.
It improves the grasping accuracy, stability and adaptability to dynamic environments in high-degree-of-freedom scenarios, reduces interference between finger movements, and improves the grasping success rate.
Smart Images

Figure CN120773073A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robotics technology, and in particular to a method and device for adaptive control of a dexterous hand using multi-agent collaborative reinforcement learning. Background Art
[0002] Adaptive grasping control methods for bionic dexterous hands have been a research hotspot in the field of humanoid robotics. Existing dexterous hand control methods primarily rely on precise kinematic and dynamic models to design control strategies and optimize control parameters. However, when dexterous hands have a high degree of freedom (e.g., at least three degrees of freedom per finger, for a total of more than 15 degrees of freedom), the dynamic interactions between fingers become extremely complex. Traditional control models struggle to accurately describe the dynamic behavior of multi-finger coordination, resulting in insufficient control accuracy and stability.
[0003] Currently, common dexterous hand control methods include centralized control and joint-level distributed control. Due to the high computational complexity of centralized control, the response delay often exceeds 100ms, making it difficult to meet real-time requirements and having poor adaptability to dynamic environments. While joint-level distributed control can improve local response speed, it lacks a finger-level collaboration mechanism and cannot optimize force distribution and motion coordination among multiple fingers. In high-degree-of-freedom scenarios, traditional methods are prone to motion interference between fingers due to the lack of effective collaborative constraints, causing collisions or grasping failures, and the grasping success rate is less than 70% in unstructured scenarios. In addition, the application of reinforcement learning technology in robot control is gradually increasing, but the standard PPO algorithm is only for single agents and lacks a reinforcement learning training framework for multi-agent collaborative operations, making it difficult to meet the needs of multi-finger collaborative grasping.
[0004] Therefore, the existing technology has obvious limitations in the multi-finger collaborative grasping control of high-degree-of-freedom dexterous hands. There is an urgent need for an adaptive grasping control method that can achieve multi-finger collaboration, adapt to complex environments, and have high efficiency and stability.
[0005] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0006] The embodiments of the present application provide a method and device for adaptive control of a dexterous hand based on multi-agent collaborative reinforcement learning, which can realize multi-finger collaborative adaptive grasping of a bionic dexterous hand, improve the accuracy, stability and adaptability of grasping to dynamic environments in high-degree-of-freedom scenarios, and reduce motion interference between fingers.
[0007] The present invention provides a method for adaptive control of a dexterous hand using multi-agent collaborative reinforcement learning, including:
[0008] A simulation environment including a multi-degree-of-freedom bionic dexterous hand and a target object is constructed using a physics engine;
[0009] generate linkage rules for constraining the actions of each finger based on the multi-finger collaborative kinematics characteristics of the bionic dexterous hand; the linkage rules include at least one of joint angle constraints, end distance constraints, collaborative trajectory constraints, and dynamic coupling constraints;
[0010] configure each finger as an independent agent, and use a multi-agent near-end policy optimization algorithm to constrain the actions of each finger in combination with the linkage rules, to train a multi-finger collaborative grasping strategy model;
[0011] deploy the trained collaborative grasping strategy model to an entity bionic dexterous hand control system, generate control instructions by real-time acquisition of sensor data, and drive the bionic dexterous hand to perform adaptive grasping operations according to the control instructions.
[0012] Preferably, in some embodiments of the present application, the simulation environment containing the multi-degree-of-freedom bionic dexterous hand and the target object is constructed by a physics engine, comprising:
[0013] The degrees of freedom of the bionic dexterous hand are configured, wherein the thumb is provided with 3 degrees of freedom, and the remaining fingers are each provided with 4 degrees of freedom;
[0014] Load the dexterous hand model and the target object model in the physics engine, and initialize the joint parameters, mass distribution and friction damping coefficient of the dexterous hand model;
[0015] Set the state space and the action space; wherein the state space includes joint angles, angular velocities, finger end positions, object states and contact information, and the total dimension is 60; the action space is defined as the normalized torque output of each joint, and the total dimension is 19.
[0016] Preferably, in some embodiments of the present application, the linkage rules for constraining the actions of each finger are generated based on the multi-finger collaborative kinematics characteristics of the bionic dexterous hand, comprising:
[0017] Constrain the action output by a joint angle limiting function to ensure that the joint angle is within the mechanical limiting range;
[0018] Based on the spatial relationship of the end positions, when the distance between the ends of any two fingers is less than a safety threshold, adjust the action in reverse by gradient to increase the distance;
[0019] Constrain the symmetrical relationship between the position of the thumb and the average position of the other four fingers by a collaborative trajectory penalty term;
[0020] Calculate the end velocity based on the Jacobian matrix, and scale the action amplitude when the relative velocity exceeds a threshold.
[0021] Preferably, in some embodiments of the present application, each finger is configured as an independent agent, and a multi-agent proximal strategy optimization algorithm is used in combination with the linkage rules to constrain the movements of each finger, and a multi-finger collaborative grasping strategy model is trained, including:
[0022] Assigning a corresponding grasping role to each finger, wherein the grasping role includes a dominant grasping role, a primary auxiliary role, and a secondary auxiliary role;
[0023] Dynamically adjust the role weight corresponding to the capture role according to the capture stage;
[0024] Guide each finger to perform corresponding collaborative behavior through the role-specific reward function;
[0025] A dynamic shared information flow mechanism is adopted to dynamically adjust the agent's observation space according to the grasping task stage.
[0026] Preferably, in some embodiments of the present application, the dynamic shared information flow mechanism is adopted to dynamically adjust the agent's observation space according to the grasping task stage, including:
[0027] In the approach phase, global object information is received through the thumb and index finger, and only the object position is received through the remaining fingers;
[0028] During the gripping phase, the position information of the adjacent fingers is added through the index and middle fingers, and the position of the thumb is received through the ring and pinky fingers;
[0029] In the stable stage, the ring finger and little finger receive the contact force information of the whole hand, and the remaining fingers receive the position and contact force of the adjacent fingers.
[0030] Preferably, in some embodiments of the present application, the multi-agent proximal strategy optimization algorithm includes:
[0031] Adopting a shared value function, the global state is input to evaluate the synergy effect;
[0032] The dynamic policy gradient clipping range is set based on the difference in the character's degree of freedom, where the thumb clipping range is smaller than that of the other fingers;
[0033] Correlated action noise is generated through a collaborative exploration mechanism that shares noise seeds.
[0034] Preferably, in some embodiments of the present application, the trained collaborative grasping strategy model is deployed to the physical bionic dexterous hand control system, control instructions are generated by real-time acquisition of sensor data, and the bionic dexterous hand is driven to perform adaptive grasping operations according to the control instructions, including:
[0035] Loading the trained collaborative grasping strategy model and deploying it to the physical bionic dexterous hand control system, establishing a mapping interface between sensor data and the collaborative grasping strategy model input;
[0036] Real-time collection of sensor data, including joint angle and angular velocity data collected by joint angle sensors, finger end contact force data collected by tactile sensors, and target object position and posture data collected by visual sensors;
[0037] Normalizing the sensor data and inputting it into the collaborative grasping strategy model to output torque control instructions for each joint;
[0038] The output torque control instruction is constrained in real time through linkage rules to generate an anti-interference control signal.
[0039] Preferably, in some embodiments of the present application, the real-time constraint on the output torque control command by the linkage rule includes:
[0040] Detect whether the joint angle exceeds the limit. If it is detected that the limit is exceeded, the torque is cut according to the mechanical limit;
[0041] Calculate the real-time distance between each finger end. If the real-time distance between the finger ends is less than the preset safety threshold, the gradient reverse adjustment is triggered.
[0042] Dynamically scale the torque output amplitude based on the object's motion state.
[0043] Preferably, in some embodiments of the present application, the method of deploying the trained collaborative grasping strategy model to a physical bionic dexterous hand control system, generating control instructions by real-time acquisition of sensor data, and driving the bionic dexterous hand to perform adaptive grasping operations according to the control instructions further includes:
[0044] When the tactile sensor detects that the continuous collision force exceeds the threshold, the current control instruction is interrupted;
[0045] Initiate the emergency recovery strategy and control all fingers to reposition along the preset safety trajectory.
[0046] Accordingly, an embodiment of the present application provides a multi-agent collaborative reinforcement learning adaptive control device for a dexterous hand, comprising:
[0047] An environment construction module, used to construct a simulation environment containing a multi-degree-of-freedom bionic dexterous hand and a target object through a physics engine;
[0048] a rule generation module for generating linkage rules for constraining the movements of each finger based on the multi-finger collaborative kinematic characteristics of the bionic dexterous hand; the linkage rules comprising at least one of a joint angle constraint, an end distance constraint, a collaborative trajectory constraint, and a dynamic coupling constraint;
[0049] A strategy training module is used to configure each finger as an independent agent, use a multi-agent proximal strategy optimization algorithm combined with the linkage rules to constrain the movements of each finger, and train a multi-finger collaborative grasping strategy model;
[0050] The control module is used to deploy the trained collaborative grasping strategy model to the physical bionic dexterous hand control system, generate control instructions by real-time acquisition of sensor data, and drive the bionic dexterous hand to perform adaptive grasping operations according to the control instructions.
[0051] The present application provides a method and device for adaptive control of a dexterous hand based on multi-agent collaborative reinforcement learning. First, a simulation environment is constructed through a physical engine to provide an interactive platform for multi-agent reinforcement learning, which can accurately simulate the interaction state between the dexterous hand and the target object. Then, the linkage rules generated based on the collaborative kinematic characteristics of multiple fingers constrain the finger movements from multiple aspects such as joint angle and end distance, effectively avoiding motion interference. Then, each finger is treated as an independent agent, and the collaborative grasping strategy model is trained in combination with the multi-agent proximal strategy optimization algorithm and the linkage rules, so that each finger can achieve efficient collaboration while maintaining independence. Finally, after the trained collaborative grasping strategy model is deployed to the physical system, control instructions are generated and driven to execute by real-time acquisition of sensor data to achieve adaptive grasping. Therefore, the adaptive control scheme for the dexterous hand based on multi-agent collaborative reinforcement learning provided by the present application can improve the grasping accuracy, stability and adaptability of the high-degree-of-freedom bionic dexterous hand in complex scenarios, reduce interference such as collisions between fingers, and thus improve the grasping success rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0053] Figure 1 This is a diagram of the application environment of the adaptive control method for dexterous hands using multi-agent collaborative reinforcement learning provided by an embodiment of the present application;
[0054] Figure 2 1 is a flow chart of a method for adaptive control of a dexterous hand using multi-agent collaborative reinforcement learning provided by an embodiment of the present application;
[0055] Figure 3 Schematic diagram of the structure of the bionic dexterous hand provided in an embodiment of the present application;
[0056] Figure 4This is another flow chart of the adaptive control method for a dexterous hand using multi-agent collaborative reinforcement learning provided by an embodiment of the present application;
[0057] Figure 5 It is a structural diagram of the adaptive control device for dexterous hands with multi-agent collaborative reinforcement learning provided in an embodiment of the present application. DETAILED DESCRIPTION
[0058] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0059] Figure 1 FIG1 is an application environment diagram of a multi-agent collaborative reinforcement learning method for a dexterous hand adaptive control method in an embodiment. Figure 1 The method for adaptive control of a dexterous hand using multi-agent collaborative reinforcement learning is applied to a system for adaptive control of a dexterous hand using multi-agent collaborative reinforcement learning. The system includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal. The mobile terminal can be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 can be implemented as an independent server or a server cluster consisting of multiple servers. The server 120 is configured to execute a multi-agent collaborative reinforcement learning adaptive control method for a dexterous hand, including: constructing a simulation environment including a multi-degree-of-freedom bionic dexterous hand and a target object through a physical engine; generating linkage rules for constraining the movements of each finger based on the multi-finger collaborative kinematic characteristics of the bionic dexterous hand; the linkage rules include at least one of joint angle constraints, end distance constraints, collaborative trajectory constraints and dynamic coupling constraints; configuring each finger as an independent agent, using a multi-agent proximal strategy optimization algorithm combined with linkage rules to constrain the movements of each finger, and training a multi-finger collaborative grasping strategy; deploying the trained collaborative grasping strategy to the physical bionic dexterous hand control system, generating control instructions by real-time acquisition of sensor data, and driving the bionic dexterous hand to perform adaptive grasping operations according to the control instructions.
[0060] In order to solve the above-mentioned technical problems and overcome the defects of the existing technology, the embodiments of the present application provide a method and device for adaptive control of a dexterous hand based on multi-agent collaborative reinforcement learning, which can realize multi-finger collaborative adaptive grasping of a bionic dexterous hand, improve the grasping accuracy, stability and adaptability to dynamic environments in high-degree-of-freedom scenarios, and reduce motion interference between fingers.
[0061] See also Figure 2 , Figure 2 This is a flow chart of a method for adaptive control of a dexterous hand using multi-agent collaborative reinforcement learning provided by an embodiment of the present application. This method can be applied to both a terminal and a server. This embodiment mainly uses the method for adaptive control of a dexterous hand using multi-agent collaborative reinforcement learning as an example to illustrate the application of the method to a terminal. The method for adaptive control of a dexterous hand using multi-agent collaborative reinforcement learning provided by an embodiment of the present application can specifically include the following steps:
[0062] S1. Use a physics engine to build a simulation environment containing a multi-degree-of-freedom bionic dexterous hand and a target object.
[0063] Specifically, for step S1, constructing a simulation environment is the basis for achieving adaptive control of the dexterous hand. The physics engine can accurately simulate the physical interaction between the dexterous hand and the target object, including collision detection, dynamic response, etc. This process requires accurate modeling of the geometric parameters, mass distribution, joint angle range, etc. of the dexterous hand. A highly simulated training environment is created through the physics engine, allowing the dexterous hand to learn and optimize under conditions close to reality. For example, the 3D model of the dexterous hand is loaded using the MuJoCo physics engine, and the joint parameters, mass distribution, friction damping coefficient, etc. are initialized to ensure the authenticity of the simulation environment.
[0064] S2. Generate linkage rules for constraining the movements of each finger based on the multi-finger collaborative kinematic characteristics of the bionic dexterous hand; the linkage rules include at least one of joint angle constraints, end distance constraints, collaborative trajectory constraints, and dynamic coupling constraints;
[0065] Specifically, for step S2, the multi-finger collaborative kinematic characteristics are the key to achieving efficient grasping of the dexterous hand. By analyzing the kinematic model of the dexterous hand, the design of linkage rules can effectively avoid motion interference between fingers, such as joint overruns and finger collisions. The linkage rules include joint angle constraints, end distance constraints, collaborative trajectory constraints, and dynamic coupling constraints. These linkage rules are integrated into the training process through action clipping and reward and penalty mechanisms to ensure the coordination and safety of the dexterous hand in high-degree-of-freedom scenarios. For example, the joint angle limit function is used to ensure that the joint angle is within the mechanical limit range, and the end distance constraint is used to avoid collisions between fingers.
[0066] S3. Configure each finger as an independent agent, use a multi-agent proximal strategy optimization algorithm combined with linkage rules to constrain the movements of each finger, and train a multi-finger collaborative grasping strategy model;
[0067] Specifically, for step S3, each finger is configured as an independent agent, enabling distributed control among fingers and improving the response speed and flexibility of the system. The Multi-Agent Proximal Policy Optimization (MAPPO) algorithm is an effective reinforcement learning method that can stabilize the training of collaborative strategy models in a multi-agent environment. By assigning an independent agent to each finger, the force distribution and motion coordination among multiple fingers can be better optimized. Combined with linkage rules to constrain the actions of each finger, the safety and effectiveness of the training process can be ensured. For example, through role assignment and dynamic weight adjustment, the finger agents are guided to learn collaborative behaviors, improving the efficiency and stability of the grasping.
[0068] S4. Deploy the trained collaborative grasping strategy model to the physical bionic dexterous hand control system, generate control instructions through real-time acquisition of sensor data, and drive the bionic dexterous hand to perform adaptive grasping operations according to the control instructions;
[0069] Specifically, for step S4, deploying the trained collaborative grasping strategy model to the physical dexterous hand control system is a key step to realize practical application. By real-time acquisition of sensor data such as joint angles, angular velocities, and finger tip contact forces, accurate control instructions can be generated. These instructions are real-time constrained by linkage rules to generate anti-interference control signals, driving the dexterous hand to perform adaptive grasping operations. For example, when the tactile sensor detects a continuous collision force exceeding the threshold, the current control instruction is interrupted and an emergency recovery strategy is initiated to ensure the safety of the operation.
[0070] This embodiment simulates the environment by constructing a highly simulated physical engine, ensuring that the dexterous hand is trained under near-real conditions; designs linkage rules based on kinematic characteristics to effectively avoid motion interference and ensure training safety; uses multi-agent reinforcement learning algorithms to optimize multi-finger collaborative grasping strategies and improve grasping efficiency and stability; and combines real-time sensor data feedback to achieve precise control and adaptive grasping, thereby improving the grasping success rate of high-degree-of-freedom dexterous hands in complex scenarios, providing an efficient and stable grasping control solution for industrial assembly, medical assistance, and other fields.
[0071] Preferably, in some embodiments, step S1 "constructing a simulated environment containing a multi-degree-of-freedom bionic dexterous hand and a target object through a physical engine" can specifically include:
[0072] S11. Configure the degree-of-freedom distribution rules of the bionic dexterous hand, wherein the thumb is set to have 3 degrees of freedom and the remaining fingers are each set to have 4 degrees of freedom;
[0073] Specifically, for step S11, when constructing the simulation environment, it is first necessary to configure the degree of freedom distribution rules of the bionic dexterous hand. The bionic dexterous hand in this embodiment has 5 fingers, with a total of 19 degrees of freedom. Among them, the thumb is set with 3 degrees of freedom, and each of the remaining fingers is set with 4 degrees of freedom. The above-mentioned degree of freedom distribution fully imitates the biomechanical characteristics of the human hand, thereby enhancing the flexibility and grasping diversity of the dexterous hand. The thumb is equipped with 3 degrees of freedom, corresponding to the flexion and extension and adduction / abduction of the metacarpophalangeal joint, and the flexion and extension of the interphalangeal joint. The other fingers are equipped with 4 degrees of freedom for each finger, corresponding to the flexion and extension and adduction / abduction of the metacarpophalangeal joint, the flexion and extension of the proximal interphalangeal joint, and the flexion and extension of the distal interphalangeal joint.
[0074] S12. Loading the dexterous hand model and the target object model into the physics engine, and initializing the joint parameters, mass distribution, and friction damping coefficient of the dexterous hand model;
[0075] Specifically, for step S12, a physics engine (such as MuJoCo) is used to load the 3D models of the dexterous hand and the target object. These models need to be accurately calibrated to ensure the authenticity of the simulation environment. Initialize the joint parameters, mass distribution, and friction damping coefficient of the dexterous hand model. The precise setting of these parameters is crucial for simulating real physical behavior. For example, the geometric parameters of the connecting rod are optimized according to the average size of human fingers to ensure that the connecting rod length and structure of the dexterous hand conform to biomechanical properties. The mass distribution reasonably distributes the mass of each connecting rod to simulate the real inertia of hand movement. Friction and damping set the joint friction coefficient and damping coefficient to ensure that the movement is smooth and conforms to the real physical properties.
[0076] S13. Set the state space and action space; the state space includes joint angles, angular velocities, finger end positions, object states, and contact information, with a total dimension of 60; the action space is defined as the normalized torque output of each joint, with a total dimension of 19;
[0077] Specifically, for step S13, in order to achieve effective reinforcement learning training, it is necessary to clearly define the state space and action space, which can be specifically as follows: the state space includes joint angles, angular velocities, finger end positions, object states, and contact information, with a total dimension of 60. Among them, the joint state includes 19 joint angles and 19 angular velocities, a total of 38 dimensions; the end position includes 5 finger end positions, a total of 15 dimensions; the object state includes the object center position, quaternion posture and linear velocity, a total of 10 dimensions; the contact information includes the contact force of 5 fingertips, a total of 5 dimensions. All state values are mapped to the range of [-1,1] through linear normalization to improve the stability and convergence speed of training. The action space is defined as the normalized torque output of each joint, with a total dimension of 19. The torque is mapped to the joint angle change through the physics engine to ensure control accuracy and response speed.
[0078] This embodiment provides a highly simulated training platform for multi-agent reinforcement learning by accurately loading and initializing the dexterous hand model and related parameters in the physics engine. By defining the state space and action space in detail, it comprehensively describes the dexterous hand and its interaction state with the target object, and provides the necessary input and output formats for the reinforcement learning algorithm, ensuring that the dexterous hand has sufficient flexibility and grasping diversity.
[0079] In a specific embodiment, for step S1, as Figure 3 As shown, Figure 3 The schematic diagram of the structure of the bionic dexterous hand provided in this embodiment, the structural design of the homemade bionic dexterous hand fully imitates the biomechanical properties of the human hand, with 5 fingers and a total of 19 degrees of freedom. The ring finger, middle finger, index finger and little finger are each equipped with 4 degrees of freedom, corresponding to the metacarpophalangeal joint (MCP, 2 degrees of freedom: flexion and extension and adduction / abduction), the proximal interphalangeal joint (PIP, 1 degree of freedom: flexion and extension), and the distal interphalangeal joint (DIP, 1 degree of freedom: flexion and extension). The thumb is equipped with 3 degrees of freedom, corresponding to the metacarpophalangeal joint (MCP, 2 degrees of freedom: flexion and extension and adduction / abduction) and the interphalangeal joint (IP, 1 degree of freedom: flexion and extension). This distribution of degrees of freedom enhances the flexibility and grasping diversity of the dexterous hand, and is particularly suitable for complex in-hand operation tasks. The joint angles are denoted as θij, where i = 1, 2, …, 5 represents the finger number (1 for the thumb, 2-5 for the index, middle, ring, and pinky fingers), and j = 1, 2, 3, 4 represents the joint number (only j = 1, 2, 3 for the thumb). A 3D model of the dexterous hand was loaded using the MuJoCo physics engine, and the model parameters were precisely calibrated to ensure simulation realism.
[0080] Among them, the geometric parameters of the dexterous hand connecting rod are as follows:
[0081] Thumb: proximal connecting rod length l 11 =0.04m, middle connecting rod l 12 =0.04m, distal connecting rod length l 13 =0.03m. Other fingers: proximal connecting rod l i1 =0.02m, middle connecting rod l i2 =0.04m, distal connecting rod l i3 =0.04m, end connecting rod l i4 = 0.03m. These lengths are optimized based on average human finger size, with the thumb slightly longer to enhance opposition and the end link shorter to improve grip accuracy.
[0082] Mass distribution:
[0083] Thumb: connecting rod mass m 11 =0.025kg,m 12 =0.015kg,m13 =0.015kg. Other fingers: m i1 =0.02kg,m i2 =0.015kg,m i3 =0.01kg,m i4 =0.005kg, the mass of the end link is small to reduce inertia.
[0084] Friction and damping: The joint friction coefficient μ = 0.5, and the damping coefficient d = 0.001 simulate the motion resistance of real joints to ensure smooth movements and conformity to biomechanical characteristics.
[0085] Joint angle range:
[0086] Thumb MCP flexion and extension:
[0087] Thumb MCP adduction / abduction:
[0088] Thumb IP flexion and extension:
[0089] Other fingers MCP flexion and extension:
[0090] Other fingers MCP adduction / abduction:
[0091] Other finger PIP and DIP flexion and extension:
[0092] The angle range is based on the mechanical limit design of the homemade dexterous hand to ensure natural and safe movements and prevent hyperextension or hyperflexion.
[0093] To calculate the finger end position p i =[x i ,y i ,z i ] T , DH parameters are used to describe the kinematic chain of the dexterous hand. The transformation matrix T of each joint ij for:
[0094]
[0095] Among them, DH parameter example (index finger MCP flexion and extension joint): connecting rod length l 21 =0.04m, connecting rod deflection angle α 21 =0, connecting rod offset d 21 =0, joint angle θ 21 The finger tip position is calculated by multiplying the transformation matrix:
[0096]
[0097] For the thumb, the transformation matrix involves only three joints; for the other fingers, it involves four joints. The above formula ensures accurate calculation of the end position, providing reliable spatial positioning data for subsequent linkage rules and grasping strategies.
[0098] The state space is designed to provide comprehensive environmental information to each finger agent while maintaining computational efficiency. The total dimension is approximately 60, specifically including:
[0099] Joint state: 19 joint angles θ ij (3 for the thumb and 4 for each of the other fingers), 19 angular velocities There are 38 dimensions in total, reflecting the movement state of the fingers.
[0100] End position: 5 finger end positions p i =[x i ,y i ,z i ] T (3 dimensions, a total of 15 dimensions), providing spatial positioning information through forward kinematics calculation.
[0101] Object state: object center position p obj =[x obj ,y obj ,z obj ](3D), quaternion posture (4D), linear velocity (3D), a total of 10 dimensions, providing dynamic information of the grasping target.
[0102] Contact information: The dexterous hand is equipped with tactile sensors that record the contact force of the five fingertips (1 dimension for each finger, a total of 5 dimensions).
[0103] To improve training stability, all state values are mapped to [-1, 1] through linear normalization:
[0104] Angle normalization:
[0105] Position normalization:
[0106] The state update frequency is 100Hz, synchronized with the action execution.
[0107] The action space is defined as the joint torque of each finger, denoted as a ij∈[-1,1], with a total dimension of 19. Torques are mapped to joint angle changes using the physics engine, with a time step of Δt = 0.01s. Action normalization is performed: the actual torque range [-2,2] Nm is mapped to [-1,1] to match the output of the reinforcement learning algorithm. The action space is designed based on the drive method of the custom-made dexterous hand (driven by a brushless DC motor), with a maximum motor torque of 2 Nm and a resolution of 0.01 Nm, to ensure control accuracy and responsiveness. After the action is executed, MuJoCo calculates the joint dynamics and updates the finger states and object interactions.
[0108] The simulation environment is implemented through the OpenAI Gym interface, and the MuJoCo physics engine is responsible for calculating kinematics, dynamics, and collision detection. The environment supports multi-process parallel training and uses Isaac Gym for efficient data collection, with each process simulating an independent scene (for example, 1,000 parallel environments generating 100,000 samples per second). The initial state randomization strategy includes:
[0109] Initial finger posture: joint angle θ ij Evenly distributed within their respective ranges, such as thumb MCP flexion and extension Other fingers MCP adduction / abduction Simulates a naturally unfolded or slightly bent state.
[0110] Initial state of the object: position is [-0.1, 0.1] 3 Randomization within m and random rotation of postures increase task diversity and challenge. The environment interaction frequency is set to 100Hz (update every 0.01s) to ensure real-time performance and simulation accuracy. The environment supports real-time rendering, generating 3D visualizations of fingers and objects using OpenGL, facilitating debugging and analysis of grasping behavior during training. Furthermore, the environment integrates logging functionality to record the status, actions, and rewards of each interaction for subsequent analysis and optimization.
[0111] Preferably, in some embodiments, step S2 of "generating linkage rules for constraining the movements of each finger based on the multi-finger collaborative kinematic characteristics of the bionic dexterous hand" may specifically include:
[0112] S21. Constrain the motion output through the joint angle limit function to ensure that the joint angle is within the mechanical limit range;
[0113] Specifically, for step S21, during the reinforcement learning training process, the action output needs to be strictly constrained within the physical limitations of the joints to prevent unrealistic actions in the simulation environment and protect the physical dexterous hand from damage. The joint angle limit function clips the action output value to ensure that it conforms to the pre-defined joint angle range. For example, the flexion and extension range of the MCP joint of the thumb is [-30°, 30°], the adduction / abduction range is [-15°, 15°], and the flexion and extension range of the IP joint is [0°, 90°]. The flexion and extension range of the MCP joints of other fingers is [-30°, 30°], the adduction / abduction range is [-10°, 10°], and the flexion and extension range of the PIP and DIP joints is [0°, 120°]. These limit values are determined based on the mechanical design and biomechanical characteristics of the dexterous hand. The clipping function adjusts the action values that are out of range to the nearest legal value to ensure that the joint angles are always within a safe range.
[0114] S22. Based on the spatial relationship of the end positions, when the distance between any two finger ends is less than a safety threshold, adjust the action by gradient reversal to increase the distance;
[0115] Specifically, for step S22, in order to avoid collisions between fingers, it is necessary to monitor the distance between the ends of each finger in real time. When the distance between the ends of any two fingers is less than a preset safety threshold (for example, 5 mm), the action is adjusted by reverse gradient to move the finger ends away from each other. Specifically, by calculating the gradient information of the positions of the two finger ends, the action output is adjusted to increase the distance between the two. This is achieved by introducing a penalty term in the reward function. When it is detected that the end distance is too small, the penalty mechanism is triggered, and the action strategy is updated by reverse gradient to avoid collisions between fingers.
[0116] S23. Constrain the symmetry between the thumb position and the average position of the other four fingers through the collaborative trajectory penalty term;
[0117] Specifically, for step S23, in order to achieve a stable grasp, it is necessary to ensure that the thumb and the other four fingers form a symmetrical clamping relationship, similar to the thumb's opposing effect during human grasping. The collaborative trajectory penalty term constrains the action output by calculating the deviation between the thumb position and the average position of the other four fingers. Specifically, a symmetry index is defined. When the deviation between the thumb position and the average position of the other four fingers exceeds a certain threshold, a penalty term is applied to guide the thumb to move to a symmetrical position. This helps to form a stable grasping force distribution and enhance the stability and reliability of the grasp.
[0118] S24. Calculate the terminal velocity based on the Jacobian matrix, and scale the motion amplitude when the relative velocity exceeds a threshold;
[0119] Specifically, for step S24, in a high-degree-of-freedom dexterous hand, the rapid movement of the fingers may cause dynamic interference or collision. To avoid this, the relative velocity of each finger end is calculated based on the Jacobian matrix. The Jacobian matrix maps the movement in the joint space to the spatial movement of the end effector. By calculating the magnitude and direction of the end velocity, the relative movement between the fingers can be monitored in real time. When the relative velocity exceeds a preset threshold (for example, 0.5 m / s), the speed is reduced by scaling the movement amplitude, thereby reducing the risk of dynamic interference.
[0120] This embodiment ensures that the movements of the dexterous hand are always within a safe range through joint angle limit functions, avoiding joint overload or unnatural bending; based on spatial relationship monitoring of the end position, the movement is adjusted by gradient reversal to effectively prevent collisions between fingers; the collaborative trajectory penalty term enhances the symmetry and stability of the grasp, simulating the thumb opposition effect of human grasping; and the speed monitoring and scaling mechanism based on the Jacobian matrix reduces the risk of dynamic interference.
[0121] In a specific embodiment, for step S2, this embodiment designs and implements a set of linkage rules based on the multi-finger collaborative kinematic characteristics of the 19-DOF bionic dexterous hand. These rules constrain the movements of each finger agent, prevent motion interference during training (such as finger collisions, excessive joint angles, or grip conflicts), and provide safety for multi-finger collaboration. The linkage rules comprehensively consider spatial constraints, speed constraints, and coordination requirements, and are integrated into the simulation environment through action clipping and reward and penalty mechanisms to ensure efficient coordination of the 19-DOF system.
[0122] The following four linkage rules are designed to comprehensively consider space, speed, and coordination constraints to provide a secure foundation for multi-finger collaboration:
[0123] Rule 1: Joint Angle Constraints
[0124] To prevent joints from exceeding limits, the updated angles of the motion must fit within their respective ranges:
[0125] θ′ ij =clip(θ ij ,θ min,ij ,θ max,ij )
[0126] Among them, θ min,ij ,θ max,ij Definition by finger and joint type (e.g. thumb MCP adduction / abduction ), the clip function is:
[0127] clip(x,x min ,x max )=max(x min ,min(x,x max ))
[0128] This rule ensures that the movements of all 19 joints comply with mechanical limits, avoiding unnatural bending or hyperextension and protecting the dexterous hand hardware.
[0129] Rule 2: End distance constraint
[0130] To prevent collisions between fingers, the distance between any two finger tips must be greater than the minimum safety distance:
[0131]
[0132] If the distance violates the constraint, adjust the action to increase the distance:
[0133]
[0134] Here, the gradient term is calculated using the chain rule:
[0135]
[0136] The adjustment step size η = 0.1 was optimized through experiments to ensure that the motion correction was smooth and not excessive.
[0137] Rule 3: Cooperative trajectory constraints
[0138] To achieve a stable grasp, the thumb (finger 1) is required to form a symmetrical grip with the average position of the other four fingers, simulating the thumb's opposition in human grasping:
[0139]
[0140] By penalizing deviations:
[0141]
[0142] This rule enhances the balance of grasping force and is particularly suitable for the high dexterity requirements of 19-DOF dexterous hands.
[0143] Rule 4: Dynamically coupled constraints
[0144] To prevent collisions caused by rapid movement, limit the relative speed of the finger ends:
[0145]
[0146] The terminal velocity is calculated from the Jacobian matrix, assuming
[0147]
[0148] If the speed exceeds the limit, scale the action:
[0149]
[0150] This rule ensures smooth motion of the 19-DOF system in high-dimensional motion space and reduces the risk of dynamic interference.
[0151] To strengthen the execution of linkage rules, an interference penalty term is added to the reward function to optimize the high complexity of 19 degrees of freedom:
[0152]
[0153] The penalty weights k1 = k2 = 10 and k3 = 1 were determined experimentally to balance interference avoidance and grasping performance. The interference penalty is combined with the collaboration reward to ensure that multi-finger movements are both safe and coordinated.
[0154] Preferably, in some embodiments, step S3 of "configuring each finger as an independent agent, using a multi-agent proximal strategy optimization algorithm combined with linkage rules to constrain the movements of each finger, and training a multi-finger collaborative grasping strategy model" may specifically include:
[0155] S31. Assigning a corresponding grasping role to each finger, the grasping role includes a leading grasping role, a main auxiliary role and a secondary auxiliary role;
[0156] Specifically, for step S31, in the multi-agent reinforcement learning framework, assigning a specific grasping role to each finger is the key to achieving efficient collaboration. The specific roles can be assigned as follows:
[0157] Thumb (Agent 1): The dominant gripper, responsible for providing the primary gripping force. The thumb's three degrees of freedom (MCP flexion / extension, adduction / abduction, and IP flexion / extension) enable it to flexibly adjust its posture and adapt to different object shapes.
[0158] The index and middle fingers (agents 2 and 3) play a primary supporting role, providing support and forming a stable grip triangle with the thumb. Each finger's four degrees of freedom (MCP flexion and extension, adduction / abduction, PIP flexion and extension, and DIP flexion and extension) enable precise contact point adjustment.
[0159] The ring and pinky fingers (agents 4 and 5) are secondary auxiliary actors responsible for stabilizing the object and preventing it from rotating or sliding. The four degrees of freedom of each finger allow it to fine-tune its posture and adapt to the surface of the object.
[0160] S32. Dynamically adjust the role weight corresponding to the capture role according to the capture stage;
[0161] Specifically, for step S31, the grasping process can be divided into three main stages: approach stage, clamping stage, and stabilization stage. In different stages, the role weight of each finger needs to be dynamically adjusted to adapt to the task requirements. In the approach stage, the weight of the thumb and index finger is higher, and the key contact points of the object are approached preferentially to quickly form a preliminary clamping. In the clamping stage, the weight of the index finger and middle finger is increased to form a stable clamping triangle with the thumb, ensuring uniform distribution of grasping force. In the stabilization stage, the weight of the ring finger and little finger is increased to fine-tune the posture and enhance the stability of grasping to prevent the object from slipping.
[0162] S33. Guiding each finger to perform corresponding cooperative behavior through role-specific reward function;
[0163] Specifically, for step S31, a role-specific reward function is designed for each role to guide the finger agent to learn the corresponding cooperative behavior, for example, the thumb reward encourages the thumb to approach the key contact points of the object to form the main clamping force; the index finger and middle finger reward encourages the formation of a stable clamping triangle to ensure uniform distribution of grasping force; the ring finger and little finger reward encourages light contact to stabilize the object.
[0164] S34. Adopting a dynamic shared information flow mechanism to dynamically adjust the observation space of the agent according to the grasping task stage;
[0165] Specifically, for step S31, the dynamic shared information flow mechanism dynamically adjusts the observation space of the agent according to the different stages of the grasping task to optimize the information sharing efficiency. In the approach stage, the thumb and index finger receive global object information (position, posture, speed), and the other fingers only receive object position. This reduces the information processing burden of the ring finger and little finger, and preferentially supports the rapid approach of the thumb and index finger. In the clamping stage, the index finger and middle finger increase the position information of the adjacent finger tip, and the ring finger and little finger receive the position of the thumb. This ensures that the index finger and middle finger cooperate with the thumb to form a clamping triangle, while reducing the information redundancy of other fingers. In the stabilization stage, the ring finger and little finger receive the contact force information of the whole hand, and the rest of the fingers receive the position and contact force of the adjacent finger. This helps the ring finger and little finger fine-tune the action to keep the object stable.
[0166] The embodiment assigns clear grasping roles to each finger, combines dynamic role weight adjustment and role-specific reward function, and effectively guides the learning of multi-agent cooperative grasping strategy. The dynamic shared information flow mechanism adjusts the observation space of the agent according to the grasping stage, optimizes the information sharing efficiency, and reduces the computational burden.
[0167] Preferably, in some embodiments, the reward function is designed in the following way:
[0168] A composite reward function is constructed including grasping success reward, role cooperation reward, interference penalty, and action efficiency penalty;
[0169] The reward weight is dynamically adjusted according to the training stage, focusing on interference avoidance in the early stage and collaboration efficiency in the later stage.
[0170] Specifically, in multi-agent reinforcement learning, designing a well-designed reward function is key to guiding agents to learn effective strategies. The composite reward function comprehensively considers four aspects: grasp success, role collaboration, interference avoidance, and action efficiency.
[0171] Reward for successful grasping: When at least three fingers are in contact with the object and the contact force reaches a certain threshold (for example, 0.5N), and the object speed is lower than a certain value (for example, 0.01m / s), it is considered a successful grasp and a higher reward is given.
[0172] Role collaboration rewards: Based on the roles assigned to the fingers, each finger is encouraged to perform corresponding collaborative behaviors. For example, the thumb leads the grip, the index and middle fingers provide support, and the ring and pinky fingers are responsible for stability.
[0173] Interference penalty: When the distance between fingers is less than a safety threshold (e.g. 5 mm) or the joint angle exceeds the limit, a penalty is imposed.
[0174] Movement efficiency penalty: Limit unnecessary movement range and energy consumption. For example, penalties can be imposed based on movement range.
[0175] In order to balance interference avoidance in the early stages of training and collaboration efficiency in the later stages, the reward weight needs to be dynamically adjusted according to the training stage. In the early stages (0-3000 rounds), the focus is on interference avoidance to ensure the safety of the system; in the middle stages (3000-9000 rounds), emphasis is placed on grasping success and the support of the main auxiliary roles (index finger and middle finger); in the later stages (9000-15000 rounds), the stabilizing effect of the secondary auxiliary roles (ring finger and little finger) and the overall collaboration efficiency are enhanced. In addition, the interference penalty weight can be dynamically adjusted according to the collision rate. For example, when the collision rate is below a certain threshold (for example, 10%), the interference penalty weight is reduced to focus on grasping optimization; in the early stages, a high penalty is maintained to ensure safety.
[0176] This example constructs a composite reward function that comprehensively considers grasp success, character collaboration, interference avoidance, and action efficiency, effectively guiding a multi-agent reinforcement learning algorithm to optimize grasping strategies. Dynamically adjusting reward weights gradually balances interference avoidance and collaboration efficiency based on the training phase, ensuring a secure training process and a highly efficient final strategy.
[0177] Preferably, in some embodiments, a dynamic shared information flow mechanism is used to dynamically adjust the agent's observation space according to the grasping task stage, including:
[0178] In the approach phase, global object information is received through the thumb and index finger, and only the object position is received through the remaining fingers;
[0179] Specifically, during the approach phase, the dexterous hand needs to quickly locate and approach the target object. At this point, the thumb and index finger, as the primary grasping agents, need to receive global object information, including the object's position, posture, and velocity, in order to quickly adjust their posture and approach the object.
[0180] Specifically, the observation space of the thumb and index finger includes: their own state (joint angle, angular velocity, end position, contact force), global object information (object center position, quaternion attitude, linear velocity). The other fingers (middle finger, ring finger and pinky finger) only receive the object's position information at this stage to reduce the information processing burden and give priority to supporting the rapid approach of the thumb and index finger. The observation space of the middle finger, ring finger and pinky finger includes their own state (joint angle, angular velocity, end position, contact force) and object position (object center position). Through this information distribution mechanism, it can be ensured that the thumb and index finger can respond quickly during the approach phase, while the other fingers remain relatively still to avoid unnecessary motion interference.
[0181] During the gripping phase, the position information of the adjacent fingers is added through the index and middle fingers, and the position of the thumb is received through the ring and pinky fingers;
[0182] Specifically, during the gripping phase, the index and middle fingers need to work together with the thumb to form a stable gripping triangle. Therefore, the observation space of the index and middle fingers is augmented with the position information of the tips of the adjacent fingers (thumb and other fingers) so that they can adjust their own positions and form a uniform gripping force distribution.
[0183] The observation space of the index and middle fingers includes their own states (joint angles, angular velocities, end positions, contact forces), global object information (object center position, quaternion attitude, linear velocity), and the positions of neighboring fingers (the end positions of the thumb and adjacent fingers). The ring and pinky fingers begin to play a role at this stage, receiving information about the thumb's end position in order to adjust their own posture and provide additional support to prevent the object from rotating or slipping. The observation space of the ring and pinky fingers includes their own states (joint angles, angular velocities, end positions, contact forces), object position (object center position), and thumb position (thumb end position). This information distribution mechanism enhances the synergy between the index and middle fingers and the thumb, while enabling the ring and pinky fingers to adjust their posture based on the thumb's position, improving the stability and reliability of grasping.
[0184] In the stable stage, the ring finger and little finger receive the contact force information of the whole hand, and the remaining fingers receive the position and contact force of the adjacent fingers;
[0185] Specifically, during the stabilization phase, the grasping action has already been established, and further adjustments are needed to the force distribution of each finger to ensure a stable grasp of the object. At this point, the observation space of the ring finger and pinky finger is supplemented with contact force information from the entire hand to fine-tune their posture and provide stable support.
[0186] The observation space of the ring and pinky fingers includes their own states (joint angles, angular velocities, end position, and contact forces), the contact forces of the entire hand (contact force information for all fingers), and the object's position and attitude (object center position and quaternion attitude). The remaining fingers (thumb, index, and middle fingers) continue to receive information on the positions and contact forces of their neighbors during this phase to further optimize force distribution and ensure uniform grip. The observation space of the thumb, index, and middle fingers includes their own states (joint angles, angular velocities, end position, and contact forces), the end positions of their neighbors, contact force information, and the object's center position and quaternion attitude.
[0187] In order to further optimize the efficiency of information sharing, a dynamic neighbor selection algorithm is used to dynamically select the most relevant neighbor fingers based on the end distance. Specifically, the neighbor finger set of each finger is determined by the k-nearest neighbor algorithm (k-NN), thereby reducing the information dimension. For example, neighbor selection: for each finger, the k fingers closest to it are selected as the neighbor finger set. Information sharing: Only the finger position and contact force information within the neighbor finger set is shared to reduce information redundancy and improve computational efficiency. The dynamic neighbor selection algorithm can flexibly adjust the scope of information sharing according to the real-time needs of the grasping task, ensuring that the intelligent agent receives the most relevant information at different stages, thereby improving the efficiency of multi-finger collaboration.
[0188] This embodiment optimizes the efficiency of information sharing by dynamically adjusting the observation space of the intelligent agent according to the different stages of the grasping task. In the approach stage, the thumb and index finger receive global object information, quickly locate and approach the target object; the other fingers only receive the object position, reducing the information processing burden. In the clamping stage, the index finger and middle finger add the adjacent finger position information to form a stable clamping triangle with the thumb; the ring finger and little finger receive the thumb position to provide additional support. In the stabilization stage, the ring finger and little finger receive the contact force information of the whole hand and fine-tune the posture to enhance the grasping stability; the remaining fingers receive the adjacent finger position and contact force information to optimize the force distribution. In addition, the dynamic adjacent finger selection algorithm further reduces information redundancy and improves computational efficiency. These technical means work together to significantly improve the efficiency and stability of the dexterous hand in multi-finger collaborative grasping tasks, ensuring a smooth transition and high success rate of the grasping process.
[0189] Preferably, in some embodiments, the multi-agent proximal strategy optimization algorithm includes:
[0190] Adopting a shared value function, the global state is input to evaluate the synergy effect;
[0191] The dynamic policy gradient clipping range is set based on the difference in the character's degree of freedom, where the thumb clipping range is smaller than that of the other fingers;
[0192] Correlated action noise is generated through a collaborative exploration mechanism that shares noise seeds.
[0193] Specifically, for the multi-agent proximal policy optimization algorithm, a shared value function is a key component in multi-agent reinforcement learning. It integrates the states and actions of all agents and evaluates the synergistic effect of the current policy. Specifically, the input to the shared value function is the global state, including the joint angles, angular velocities, end position, object state, and contact information of all fingers, with a total dimension of 60. In this way, the value function captures the overall effect of multi-finger collaborative grasping, rather than just the performance of individual fingers. Due to the different roles and degrees of freedom of different fingers, their policy gradient updates require different processing. The thumb, as the dominant grasper, has a greater impact on grasping performance, requiring more refined policy updates. The policy gradient clipping range for the thumb is set to a smaller value (e.g., 0.1), while the clipping ranges for other fingers can be set to larger values (e.g., 0.2). To improve the efficiency of multi-agent exploration, a collaborative exploration mechanism using a shared noise seed is adopted. All agents share a noise seed, generating related but not identical motion noise. This ensures that the finger movements maintain a certain degree of coordination during exploration and avoids disordered random exploration.
[0194] This embodiment comprehensively evaluates the synergistic effect of multiple agents through a shared value function to ensure the global optimality of policy updates; a dynamic policy gradient clipping mechanism based on differences in role degrees of freedom optimizes the policy update process of different fingers and improves training stability and convergence speed; a collaborative exploration mechanism with shared noise seeds enhances the efficiency and synergy of multi-agent exploration and avoids the waste of training resources caused by disordered exploration.
[0195] In a specific embodiment, for step S3, this embodiment configures each finger as an independent agent and uses the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm to train the collaborative grasping strategy of the 19-DOF bionic dexterous hand. The innovation of this step is reflected in the following aspects:
[0196] Multi-finger collaboration rules: Design multi-finger collaboration rules based on kinematics and tasks, clearly defining the thumb as the dominant grasping role (providing the main clamping force), the index finger and middle finger as the main auxiliary roles (forming a support triangle), and the ring finger and little finger as secondary auxiliary roles (stabilizing the object). Optimize grasping efficiency and force distribution through dynamic priority adjustment and role collaboration rewards.
[0197] Dynamic shared information flow: A mechanism is proposed to dynamically adjust shared information according to the grasping task stage (approach, clamping, stabilization). The intelligent agent selectively shares object status, neighboring finger positions, and contact forces based on its role and task requirements, enhancing multi-finger collaboration, reducing information redundancy, and improving computational efficiency.
[0198] MAPPO collaborative optimization: Through shared value functions, dynamic policy gradient clipping and collaborative exploration mechanisms, it stably trains 19-degree-of-freedom high-dimensional collaborative strategies and overcomes the optimization challenges of high-dimensional action spaces.
[0199] Dynamic Adaptive Rewards: Design a reward function that includes grasping success, role collaboration, interference penalty, and action efficiency, and balance collaboration and safety through adaptive weight scheduling.
[0200] Linkage rule integration: Combined with the linkage rules in step 2, this constrains the agent's actions, prevents motion interference, and ensures a safe and efficient training process.
[0201] (1) Agent configuration and grasping role allocation
[0202] Each finger acts as an independent agent, for a total of five agents (one for the thumb, one each for the index finger, middle finger, ring finger, and pinky finger). This optimization optimizes multi-finger collaboration efficiency by clearly assigning grasping roles. This role assignment is based on the biomechanical properties of the human hand and incorporates the structural characteristics of the 19-DOF dexterous hand:
[0203] Thumb (Agent 1): The thumb plays a dominant role in the grasping process, providing the primary gripping force and prioritizing access to key contact points (e.g., opposite the center of the object), simulating the thumb's opposing action in human grasping. The thumb's three degrees of freedom (MCP flexion / extension, adduction / abduction, and IP flexion / extension) enable it to flexibly adjust its posture and adapt to varying object shapes.
[0204] The index and middle fingers (agents 2 and 3) play a primary supporting role, providing support and forming a stable grip triangle with the thumb to ensure even distribution of grip force. Four degrees of freedom (MCP flexion and extension, adduction / abduction, PIP flexion and extension, and DIP flexion and extension) support precise contact point adjustment.
[0205] The ring and pinky fingers (agents 4 and 5) are secondary auxiliary actors, responsible for stabilizing the object and preventing rotation or slipping. They have a smaller range of motion and prioritize maintaining contact to enhance grip stability. Their four degrees of freedom allow for fine-tuning of their posture to adapt to the surface of the object.
[0206] The agent configuration is as follows:
[0207] Observation space (approximately 25 dimensions per agent, dynamically adjusted):
[0208] Self-state: joint angle θ ij (3 or 4 dimensions), angular velocity (3 or 4 dimensions), end position p i (3D), contact force f i (1 dimension).
[0209] Global shared information: target object position p obj (3D), quaternion attitude (4D), linear velocity (3D).
[0210] Local shared information (dynamic): the end positions of adjacent fingers (3D / finger, neighboring finger set Adjust according to the mission phase) or contact force f j (1 dimension / finger).
[0211] Action space: thumb 3D torque a 1j , the 4-dimensional torque a of other fingers ij , a total of 19 dimensions, ranging from [-1,1].
[0212] Role constraints: Role assignment is achieved through action weights and reward functions:
[0213] Action weighting: Thumb ω1 = 1.5, encouraging greater range of motion to dominate gripping; index and middle fingers ω2 = ω3 = 1.0, supporting precise support; ring and pinky fingers ω4 = ω5 = 0.7, limiting movement to focus on stability. Action scaling:
[0214] a i,t =ω i ·a i,t
[0215] Role Rewards: Specific rewards are designed for each role to ensure that the thumb prioritizes gripping, the index and middle fingers provide support, and the ring and pinky fingers stabilize the object. Role assignments guide the agent through dynamic weights and rewards, such as the thumb and index finger leading the movement during the approach phase, while the other fingers contribute during the stabilization phase, mimicking the collaborative grasping pattern of the human hand.
[0216] (2) Multi-finger collaboration rule design
[0217] To achieve efficient multi-finger coordination, the following coordination rules were designed and incorporated into the MAPPO training process to ensure coordinated movements among the fingers and optimize grip force distribution and stability. These rules are based on the kinematic characteristics of the 19-DOF dexterous hand and the collaborative model of human grasping, highlighting multi-finger coordination:
[0218] Rule 1: Thumb-dominated grip The thumb preferentially approaches the opposite side of the object to form the main gripping force, simulating the key role of the thumb in human grasping. The reward function is used to encourage the thumb to approach the key contact point of the object:
[0219] R thumb=-k thumb ·p 1,t -(p obj,t -r obj ·n obj ), k thumb =1.0;
[0220] Among them, p 1,t is the end position of the thumb, p obj,t is the center position of the object, r obj is the radius of the object, n obj is the surface normal of the object (obtained through MuJoCo collision detection). This bonus ensures that the thumb is positioned on the opposite side of the object, forming a strong grip.
[0221] Rule 2: The index and middle fingers support the triangle. The index and middle fingers play the main supporting role, close to the side of the object, and form a stable grip triangle with the thumb to ensure that the grip force is evenly distributed. The support behavior is encouraged through the reward function:
[0222]
[0223] Here, I(·) is the indicator function, which ensures that the index finger and middle finger maintain a sufficient distance from the thumb (>0.05m) to form a triangular clamping structure and enhance the grasping stability.
[0224] Rule 3: The ring finger and pinky finger stabilize the object. The ring finger and pinky finger play a secondary auxiliary role, responsible for stabilizing the object and preventing rotation or slipping. The movement range is small. The reward function is used to encourage stabilization behavior:
[0225]
[0226] Among them, f i,t For contact force, the bonus only takes effect when the contact force is > 0.2N, ensuring that the ring finger and pinky finger maintain light contact, stabilizing the object without interfering with the main grip.
[0227] Rule 4: Dynamic Collaboration Prioritization
[0228] Dynamically adjust finger priorities according to the grasping task stage (approach, clamping, stabilization) to optimize collaboration efficiency.
[0229] Approaching stage The thumb and index finger take priority, action weights ω1 = 1.5, ω2 = 1.2, and other ω i =0.8, which encourages fast approach to objects.
[0230] Clamping stage (at least 2 fingers in contact, contact force > 0.5N): index finger and middle finger are preferred, ω2=ω3=1.2, ω1=1.0, other ω i =0.7, strengthen triangular clamping.
[0231] Stable stage (all fingers in contact, object speed < 0.01m / s): ring finger and little finger take priority, ω4=ω5=1.0, other ω i =0.8, enhancing the stabilizing effect.
[0232] Prioritization is achieved through dynamic weight adjustment:
[0233] ω i (t) = ω i,base +Δω i ·f stage (t),f stage (t)∈{0,1},ω i,base
[0234] =[1.5,1.0,1.0,0.7,0.7],Δω i =[0.3,0.2,0.2,0.3,0.3]
[0235] Among them, f stage (t) is determined based on the task phase (approach, grip, stabilization) and is judged based on object distance and contact force. These rules guide the agent to learn collaborative behaviors through reward functions and action weights, simulating the collaborative grasping mode of the human hand and significantly improving multi-finger coordination.
[0236] (3) Dynamic and adaptive shared information flow mechanism
[0237] This embodiment proposes a dynamic and adaptive shared information flow mechanism that dynamically adjusts the shared information content based on the grasping task stage and finger roles, optimizing the efficiency of multi-finger collaboration and reducing information redundancy. The design of the shared information flow takes into account the high-dimensional state space of the 19-DOF dexterous hand and the complexity of multi-finger collaboration, ensuring that the agent receives the most relevant information at different stages:
[0238] Approaching stage The thumb and index finger lead the movement, sharing global object information to coordinate approach:
[0239]
[0240] Among them, q obj,t is the quaternion posture of the object. The other fingers only receive the object position, and the observation dimension is reduced:
[0241]
[0242] This mechanism reduces the information processing burden on the ring finger and little finger, and prioritizes the rapid approach of the thumb and index finger.
[0243] Clamping phase (at least two fingers in contact, contact force > 0.5N): index and middle fingers take priority, and information about the position of the adjacent finger ends is added to coordinate the clamping:
[0244]
[0245] Neighboring finger set Determined by the k-nearest neighbor algorithm (k=2):
[0246]
[0247] The thumb continues to receive global information, while the ring finger and pinky finger receive the object position and thumb position:
[0248]
[0249] This mechanism ensures that the index and middle fingers work together with the thumb to form a gripping triangle, while reducing information redundancy of other fingers.
[0250] Stable phase (all fingers in contact, object speed < 0.01m / s): The ring finger and pinky finger are prioritized, and all finger contact forces are shared to optimize stability:
[0251]
[0252] The thumb, index finger, and middle finger receive the adjacent finger positions and contact forces:
[0253]
[0254] Contact force sharing helps the ring and pinky fingers fine-tune their movements and keep the object stable.
[0255] The dynamic neighbor selection algorithm combines the task phase and optimizes information sharing through k-nearest neighbors (k=2):
[0256]
[0257] This algorithm dynamically selects the most relevant neighboring fingers based on the end-to-end distance, reducing information dimensionality (from 15 to 9 per agent) and improving computational efficiency by 10%. Furthermore, the shared information flow is differentiated based on role assignment, with the thumb and index finger receiving more global information while assisting the fingers in focusing on local information, enhancing targeted collaboration.
[0258] (4) Collaborative optimization implementation of the MAPPO algorithm
[0259] MAPPO uses a shared value function, dynamic policy gradient pruning, and collaborative exploration mechanism to stably train a 19-degree-of-freedom high-dimensional collaborative strategy. The algorithm design is optimized for multi-finger collaboration requirements:
[0260] Actor Network: Each agent has an independent actor network (multi-layer perceptron, MLP, 256×256 neurons, ReLU activation), with input observation o i,t , output action distribution is a Gaussian distribution (the mean is output by the network and the variance is adaptive).
[0261] Shared value function: All agents share a centralized value network (MLP, 256×256 neurons) with the global state s as input t (60 dimensions), output value estimate V φ (s t ). The shared value function captures the multi-finger synergy effect through global information and reduces the non-stationarity of the multi-agent environment.
[0262] Objective function: Each agent i optimizes the clipping strategy objective:
[0263]
[0264] Among them, the probability ratio Advantage Estimate It can be calculated by the generalized advantage estimate (GAE):
[0265]
[0266] Among them, γ = 0.99 is the discount factor, λ = 0.95 is the GAE parameter, r t For shared rewards (defined later). The dynamic clipping range is adjusted according to the role:
[0267]
[0268] The thumb clipping range ∈1=0.15, other fingers ∈ i = 0.2, ensuring more stable strategy updates for high-DOF fingers.
[0269] Value Function Loss: Optimizing Shared Value Network:
[0270]
[0271] Entropy regularization: To encourage exploration, add entropy loss:
[0272]
[0273] The total loss is:
[0274]
[0275] The dynamic entropy coefficient c2(t) is adaptively adjusted according to the grasping success rate. When the success rate is high, exploration is reduced and retrieval is accelerated. In the early stage, high exploration is maintained to avoid local optimality.
[0276] Cooperative exploration mechanism: a multi-agent cooperative exploration strategy is proposed, which coordinates the actions of agents by sharing noise seeds:
[0277] n i (t) = Gaussian(mu = 0, sigma t , seed = global_seed)
[0278] where sigma t = 0.1 * exp(-0.001 * t) decays with the number of training steps, and the shared seed ensures that the finger actions remain coordinated during exploration, reducing chaotic interference. Cooperative exploration generates correlated noise through a global seed, simulating the cooperative exploration behavior of human hands when grasping, and improving the efficiency of multi-finger collaboration.
[0279] (5) Dynamic adaptive reward function design
[0280] The reward function considers grasping targets, role collaboration, interference avoidance, and action efficiency, and is optimized for high-dimensional action spaces of 19 degrees of freedom and multi-finger collaboration requirements:
[0281]
[0282] where:
[0283] indicates a successful grasp (at least 3 fingers in contact, contact force > 0.5N, object velocity < 0.01m / s), otherwise 0.
[0284] R thumb ,R support ,R stabilize defined by the collaboration rules, respectively encouraging the thumb to pinch, the index and middle fingers to support, and the ring and little fingers to stabilize.
[0285] R interference defined in step 2, penalizing interference behavior.
[0286] Action penalty term∑ i ω i ||a i,t || 2 adjusted according to the role weight omega i , limiting energy consumption, with higher action costs for the thumb and lower action costs for auxiliary fingers.
[0287] The weight is dynamically adjusted to balance collaboration and safety:
[0288] Initial (0-3000 rounds): w5 = 1.0, w1 = 1.0, w2 = 0.8, prioritize interference avoidance to ensure safety of the 19-degree-of-freedom system.
[0289] Mid-term (3000-9000 episodes): w5=0.5, w1=1.5, w3=0.3, w2=1.0, emphasizing grip success and support from primary auxiliary digits (index, middle).
[0290] Late-term (9000-15000 episodes): w5=0.2, w4=0.4, w3=0.3, w1=1.2, enhancing stabilization from secondary auxiliary digits (ring, pinky) and overall coordination efficiency.
[0291] Adaptive weight scheduling: dynamically adjusting interference penalty weights based on collision rate:
[0292] w5(t) = w 5,init max(0.2, exp(-0.5·collision_rate t )), w 5,init =1.0
[0293] When collision rate decreases, reduce interference penalty, focus on grip optimization; initial high penalty to ensure safety.
[0294] Role coordination rewards (R thumb , R support , R stabilize ) combined with dynamic weight scheduling, clearly distinguish finger roles, optimize grip force distribution, improve clamping force uniformity by 10% (evaluated by contact force variance). Reward function through phased adjustment, guide agents from initial interference avoidance to efficient cooperation, simulate the natural gripping process of human hands.
[0295] (6) Training process
[0296] The training process is designed to efficiently and stably optimize multi-finger coordination strategies:
[0297] Initialization: randomly initialize actor and shared value network (Xavier initialization), set experience buffer (capacity 1e6), target network soft update parameter τ=0.01.
[0298] Data collection: 5 agents interact in parallel, 200 steps per episode, collect trajectories (state, observation, action, reward, next state). Actions are clipped by step 2 linkage rules to ensure no interference. Observations are generated by dynamic shared information flow, role weights ω i Adjust the action.
[0299] Policy update: sample trajectories at the end of each episode, perform 10 PPO updates (batch size 256), use Adam optimizer (learning rate 1e-4). The update process includes:
[0300] Calculate advantage estimate Update actor network to optimize
[0301] Update the shared value network to minimize L critic .
[0302] Adding entropy loss Dynamically adjust c2(t).
[0303] Collaborative exploration: Actions add Gaussian noise (standard deviation σ t =0.1·exp(-0.001·t)), and collaboration is ensured by sharing seeds:
[0304]
[0305] Preferably, in some embodiments, step S4 of "deploying the trained collaborative grasping strategy model to the physical bionic dexterous hand control system, generating control instructions by real-time acquisition of sensor data, and driving the bionic dexterous hand to perform adaptive grasping operations according to the control instructions" may specifically include:
[0306] S41. Load the trained collaborative grasping strategy model and deploy it to the physical bionic dexterous hand control system, and establish a mapping interface between sensor data and collaborative grasping strategy model input;
[0307] Specifically, in step S41, when deploying the trained collaborative grasping strategy model to the physical bionic dexterous hand control system, it is first necessary to establish a mapping interface between sensor data and model input. Specifically, the physical dexterous hand is equipped with a variety of sensors, including joint angle sensors, tactile sensors, and visual sensors. The data collected by these sensors needs to be preprocessed and normalized to match the input format of the training model.
[0308] The mapping interface can be implemented through the following steps: real-time sensor data acquisition; filtering and noise reduction processing of the acquired data to ensure data accuracy and stability; mapping the data to the model input range, typically [-1, 1] or [0, 1]; and converting the processed data into the format required by the model input, such as converting joint angle and angular velocity data into a vector, tactile sensor data into a vector, and visual sensor data into a vector. This mapping interface ensures that the trained model can directly use sensor data collected by the physical dexterous hand for reasoning and control.
[0309] S42 real-time collection of sensor data, sensor data including joint angle and angular velocity data collected by the joint angle sensor, finger end contact force data collected by the tactile sensor, and target object position and posture data collected by the visual sensor;
[0310] Specifically, for step S42, in the physical dexterous hand control system, real-time data collection is the key to achieving adaptive grasping. Specifically, the system needs to collect the following types of sensor data:
[0311] Joint angle and angular velocity data are collected through encoders or angle sensors installed at each joint to reflect the movement state of the finger; finger end contact force data are collected through tactile sensors to reflect the contact between the finger and the object; target object position and posture data are collected through visual sensors (such as cameras or depth sensors) to reflect the position and posture of the target object.
[0312] The steps to implement real-time data acquisition are as follows: when the system starts, initialize all sensors to ensure that they are working properly; set the frequency of data acquisition, usually 100Hz, to ensure the real-time and accuracy of the data; ensure that the data of different types of sensors are synchronized in time to avoid data inconsistency problems; and transmit the collected data to the control system through a high-speed communication interface (such as USB or Ethernet).
[0313] S43. The sensor data is normalized and input into the collaborative grasping strategy model to output torque control instructions for each joint;
[0314] Specifically, in step S43, before inputting sensor data into the collaborative grasping strategy model, the data needs to be normalized to ensure consistency and stability. For example, the joint angle sensor collects angle and angular velocity data for each joint, normalized to the range [-1, 1]; the tactile sensor collects contact force data from the fingertips, normalized to the range [0, 1]; and the visual sensor collects position and posture data of the target object, normalized to the range [-1, 1]. This normalized data is input into the collaborative grasping strategy model, which calculates the torque control commands for each joint through forward propagation.
[0315] S44. Real-time constraints are applied to the output torque control command through linkage rules to generate an anti-interference control signal;
[0316] Specifically, in step S44, after generating the torque control command, the command needs to be constrained in real time through linkage rules to ensure that the dexterous hand's movements are safe and effective. Specific constraints include:
[0317] Joint Angle Constraints: Ensure joint angles are within mechanical limits. For example, if the calculated torque would cause the joint angle to exceed the limit, the torque value is clipped to ensure the joint angle is within a safe range.
[0318] End distance constraint: Ensure the distance between the tips of the fingers is greater than a safety threshold. If the inter-finger distance is detected to be less than the safety threshold, adjust the torque commands to move the finger tips away from each other.
[0319] Dynamic coupling constraint: Ensure the relative velocity of the finger tips is within a safe range. If the relative velocity is detected to exceed a threshold, scale the torque commands to reduce the velocity.
[0320] These constraints are implemented by monitoring and adjusting the torque commands in real-time, ensuring that the dexterous hand does not experience motion interference or collisions while performing grasping operations.
[0321] Preferably, in some embodiments, the step S4 "deploy the trained collaborative grasping strategy model to the physical bionic dexterous hand control system, generate control commands by real-time acquisition of sensor data, and drive the bionic dexterous hand to perform adaptive grasping operations according to the control commands", further includes:
[0322] S45. When the tactile sensor detects a sustained collision force exceeding a threshold, interrupt the current control command;
[0323] S46. Start an emergency recovery strategy to control all fingers to unfold and reset according to a pre-set safety trajectory;
[0324] Specifically, in order to ensure the safety and reliability of the system, when the tactile sensor detects a sustained collision force exceeding a threshold, the system needs to immediately interrupt the current control command and start an emergency recovery strategy. The specific implementation steps are as follows: real-time monitoring of tactile sensor data, calculating the contact force of each finger. If the contact force exceeds the threshold (e.g. 10N) for a certain period of time (e.g. 0.1 seconds), collision detection is triggered; immediately interrupt the current control command and stop executing the current grasping action; control all fingers to unfold and reset according to a pre-set safety trajectory. For example, the joint angles of the fingers can be gradually adjusted to the initial position to ensure that the dexterous hand returns to normal state; after the dexterous hand completes the reset, the system can re-evaluate the environmental state and decide whether to restart the grasping task.
[0325] This embodiment ensures the effective application of the strategy by deploying the trained collaborative grasping strategy model to the physical bionic dexterous hand control system and establishing a mapping interface between sensor data and model input. Real-time acquisition and normalization of sensor data provide accurate input for the model, enabling the dexterous hand to adjust the grasping strategy in real-time according to environmental changes. Real-time constraints on control commands through linkage rules generate anti-interference control signals, ensuring the safety and effectiveness of the action. In addition, the emergency recovery strategy can quickly respond when detecting abnormal collisions, protecting the dexterous hand from damage.
[0326] Preferably, in some embodiments, step S44 "real-time constraint on output torque control instruction by linkage rule" includes:
[0327] S441. Detect whether the joint angle is out of limit, and if so, clip the torque according to mechanical limit;
[0328] Specifically, for step S441, when the dexterous hand performs a grasping operation, it is necessary to ensure that the angle of each joint is within its mechanical limit range to avoid hardware damage or unnatural motion. The specific implementation steps are as follows: real-time acquisition of angle data of each joint is realized through the encoder or angle sensor installed at the joint. Compare the acquired joint angle with the preset mechanical limit range. For example, the MCP joint of the thumb has a flexion range of [-30°, 30°], a adduction / abduction range of [-15°, 15°], and an IP joint flexion range of [0°, 90°]; the MCP joint of other fingers has a flexion range of [-30°, 30°], an adduction / abduction range of [-10°, 10°], and a PIP and DIP joint flexion range of [0°, 120°]. If it is detected that the joint angle is out of limit, the torque is clipped according to the mechanical limit.
[0329] S442. Calculate the real-time distance between the fingertips of each finger, and if it is detected that the real-time distance between the fingertips is less than the preset safety threshold, trigger gradient reverse adjustment;
[0330] Specifically, for step S442, in order to prevent collision between fingers during grasping, it is necessary to monitor the distance between the fingertips of each finger in real time. The specific implementation steps are as follows: calculate the real-time position of each finger tip through forward kinematics. Calculate the Euclidean distance between the fingertips of each two fingers. Compare the calculated distance with the preset safety threshold (e.g. 5mm). If the distance is less than the safety threshold, trigger gradient reverse adjustment. By calculating the gradient information, adjust the action output to increase the distance between the fingertips.
[0331] S443. Dynamically scale the torque output amplitude based on the motion state of the object;
[0332] Specifically, for step S443, in order to adapt to different grasping tasks and object characteristics, it is necessary to dynamically adjust the torque output amplitude according to the motion state of the object. The specific implementation steps are as follows: Real-time acquisition of the object's position, speed, acceleration and other motion state data. These data can be obtained through visual sensors (such as cameras or depth sensors) or force sensors. According to the object's motion state, judge the difficulty of the current grasping task. For example, if the object speed is large or the acceleration is high, it means that the object is moving more violently and requires a larger torque for stable grasping. Dynamically scale the torque output amplitude according to the object's motion state. For example, when the object speed exceeds a certain threshold, increase the scaling factor to increase the torque output; when the object speed is low, reduce the scaling factor to reduce the torque output to avoid damage to the object due to excessive force.
[0333] This embodiment ensures the safe operation of the dexterous hand joints by real-time monitoring of joint angles and clipping of excess torque; effectively avoids collisions between fingers by calculating the real-time distance between the finger ends and triggering gradient reverse adjustment; and dynamically scales the torque output amplitude based on the object's motion state to adapt to the needs of different grasping tasks.
[0334] In a specific embodiment, Figure 4 As shown, Figure 4 Another flow chart of the adaptive control method for a dexterous hand based on multi-agent collaborative reinforcement learning provided in this embodiment specifically includes the following steps: first, the physical interaction between the dexterous hand and the target object is simulated in a simulation environment to generate state data in real time; then, these data are extracted to form a state space, which is input into the dynamic information flow module after normalization processing. The module dynamically adjusts the information sharing content of each finger agent according to the grasping stage; each agent generates a preliminary action strategy in the corresponding Actor network based on the received information; at the same time, the reward calculation module constructs a composite reward function to evaluate the quality of the current action strategy; the Critic network evaluates the overall strategy value based on the global state; then the preliminary action strategy is tailored and adjusted in combination with the linkage rules in the motion constraint module to generate a final control signal and feed it back to the simulation environment to drive the dexterous hand movement; the entire process is iterated, and the grasping strategy is continuously optimized until the expected grasping effect is achieved.
[0335] To sum up, the adaptive control method of dexterous hands based on multi-agent collaborative reinforcement learning provided in this embodiment can significantly improve the stability of the adaptive grasping of bionic dexterous hands through the multi-agent collaborative PPO reinforcement learning algorithm and the multi-finger motion interference constraint. By utilizing the multi-finger collaborative motion control strategy, precise control of the grasping motion of the dexterous hands can be achieved. Especially in high-degree-of-freedom scenarios, motion interference constraints can effectively avoid collisions between fingers and effectively improve the grasping success rate. It is suitable for complex scenarios such as industrial precision assembly, medical surgical assistance, and service robot object handling.
[0336] It should be understood that although Figure 2 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 2 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0337] To facilitate the implementation of the multi-agent collaborative reinforcement learning adaptive control method for a dexterous hand in the embodiments of the present application, the embodiments of the present application also provide a multi-agent collaborative reinforcement learning adaptive control device for a dexterous hand based on the multi-agent collaborative reinforcement learning adaptive control method. The meanings of the terms herein are the same as those in the multi-agent collaborative reinforcement learning adaptive control method for a dexterous hand, and the specific implementation details can be found in the description of the method embodiment.
[0338] See also Figure 5 , Figure 5 This is a schematic diagram of the structure of a dexterous hand adaptive control device for multi-agent collaborative reinforcement learning provided in an embodiment of the present application. The dexterous hand adaptive control device for multi-agent collaborative reinforcement learning may specifically include an environment construction module 201, a rule generation module 202, a strategy training module 203, and a control module 204. Specifically, it may be as follows:
[0339] An environment construction module 201 is used to construct a simulation environment including a multi-degree-of-freedom bionic dexterous hand and a target object through a physics engine;
[0340] A rule generation module 202 is configured to generate linkage rules for constraining the movements of each finger based on the multi-finger collaborative kinematic characteristics of the bionic dexterous hand; the linkage rules include at least one of a joint angle constraint, an end distance constraint, a collaborative trajectory constraint, and a dynamic coupling constraint;
[0341] Strategy training module 203, configured to configure each finger as an independent agent, use multi-agent proximal strategy optimization algorithm combined with linkage rules to constrain the movements of each finger, and train a multi-finger collaborative grasping strategy model;
[0342] The control module 204 is used to deploy the trained collaborative grasping strategy model to the physical bionic dexterous hand control system, generate control instructions by collecting sensor data in real time, and drive the bionic dexterous hand to perform adaptive grasping operations according to the control instructions.
[0343] Regarding the specific definition of the adaptive control device for dexterous hands with multi-agent collaborative reinforcement learning, please refer to the definition of the adaptive control method for dexterous hands with multi-agent collaborative reinforcement learning above, which will not be repeated here. Each module in the above-mentioned adaptive control device for dexterous hands with multi-agent collaborative reinforcement learning can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0344] The multi-agent collaborative reinforcement learning adaptive control device for the dexterous hand provided in this embodiment can improve the grasping accuracy, stability and adaptability of the high-degree-of-freedom bionic dexterous hand in complex scenarios, reduce interference such as collisions between fingers, and thereby improve the grasping success rate.
[0345] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A multi-agent collaborative reinforcement learning adaptive control method for a dexterous hand, characterized by: The steps include: A simulation environment including a multi-degree-of-freedom bionic dexterous hand and a target object is constructed using a physics engine; Based on the multi-finger collaborative kinematic characteristics of the bionic dexterous hand, generating linkage rules for constraining the movements of each finger; the linkage rules include at least one of joint angle constraints, end distance constraints, collaborative trajectory constraints, and dynamic coupling constraints; Each finger is configured as an independent agent, and the multi-agent proximal strategy optimization algorithm is combined with the linkage rules to constrain the movements of each finger and train a multi-finger collaborative grasping strategy model; The trained collaborative grasping strategy model is deployed to the physical bionic dexterous hand control system, and control instructions are generated by real-time acquisition of sensor data. The bionic dexterous hand is driven to perform adaptive grasping operations according to the control instructions.
2. The adaptive control method for dexterous hands based on multi-agent collaborative reinforcement learning according to claim 1 is characterized in that: The simulation environment including the multi-degree-of-freedom bionic dexterous hand and the target object is constructed by using a physics engine, including: Configuring the degree of freedom distribution rule of the bionic dexterous hand, wherein the thumb is set with 3 degrees of freedom and the remaining fingers are set with 4 degrees of freedom each; Loading the dexterous hand model and the target object model into the physics engine, and initializing the joint parameters, mass distribution, and friction damping coefficient of the dexterous hand model; Set up a state space and an action space; wherein, the state space includes joint angles, angular velocities, finger end positions, object states, and contact information, with a total dimension of 60 dimensions; the action space is defined as the normalized torque output of each joint, with a total dimension of 19 dimensions.
3. The adaptive control method for dexterous hands based on multi-agent collaborative reinforcement learning according to claim 1 is characterized in that: The generation of linkage rules for constraining the movements of each finger based on the multi-finger collaborative kinematic characteristics of the bionic dexterous hand includes: The joint angle limit function is used to constrain the motion output to ensure that the joint angle is within the mechanical limit range; Based on the spatial relationship of the end positions, when the distance between any two finger ends is less than the safety threshold, the action is adjusted in the reverse gradient to increase the distance; The symmetry between the thumb position and the average position of the other four fingers is constrained by the collaborative trajectory penalty term; The terminal velocity is calculated based on the Jacobian matrix, and the action amplitude is scaled when the relative velocity exceeds a threshold.
4. The adaptive control method for dexterous hands based on multi-agent collaborative reinforcement learning according to claim 1 is characterized in that: The method configures each finger as an independent agent, uses a multi-agent proximal strategy optimization algorithm combined with the linkage rules to constrain the movements of each finger, and trains a multi-finger collaborative grasping strategy model, including: Assigning a corresponding grasping role to each finger, wherein the grasping role includes a dominant grasping role, a primary auxiliary role, and a secondary auxiliary role; Dynamically adjust the role weight corresponding to the capture role according to the capture stage; Guide each finger to perform corresponding collaborative behavior through the role-specific reward function; A dynamic shared information flow mechanism is adopted to dynamically adjust the agent's observation space according to the grasping task stage.
5. The adaptive control method for dexterous hands based on multi-agent collaborative reinforcement learning according to claim 4 is characterized in that: The dynamic shared information flow mechanism is used to dynamically adjust the agent's observation space according to the grasping task stage, including: In the approach phase, global object information is received through the thumb and index finger, and only the object position is received through the remaining fingers; During the gripping phase, the position information of the adjacent fingers is added through the index and middle fingers, and the position of the thumb is received through the ring and pinky fingers; In the stable stage, the ring finger and little finger receive the contact force information of the whole hand, and the remaining fingers receive the position and contact force of the adjacent fingers.
6. The adaptive control method for dexterous hands based on multi-agent collaborative reinforcement learning according to claim 1, characterized in that: The multi-agent proximal strategy optimization algorithm includes: Adopting a shared value function, the global state is input to evaluate the synergy effect; The dynamic policy gradient clipping range is set based on the difference in the character's degree of freedom, where the thumb clipping range is smaller than that of the other fingers; Correlated action noise is generated through a collaborative exploration mechanism that shares noise seeds.
7. The adaptive control method for dexterous hands based on multi-agent collaborative reinforcement learning according to claim 1 is characterized in that: The trained collaborative grasping strategy model is deployed to the physical bionic dexterous hand control system, control instructions are generated by real-time acquisition of sensor data, and the bionic dexterous hand is driven to perform adaptive grasping operations according to the control instructions, including: Loading the trained collaborative grasping strategy model and deploying it to the physical bionic dexterous hand control system, establishing a mapping interface between sensor data and the collaborative grasping strategy model input; Real-time collection of sensor data, including joint angle and angular velocity data collected by joint angle sensors, finger end contact force data collected by tactile sensors, and target object position and posture data collected by visual sensors; Normalizing the sensor data and inputting it into the collaborative grasping strategy model to output torque control instructions for each joint; The output torque control instruction is constrained in real time through linkage rules to generate an anti-interference control signal.
8. The adaptive control method for dexterous hands based on multi-agent collaborative reinforcement learning according to claim 7 is characterized in that: The real-time constraint of the output torque control command by the linkage rule includes: Detect whether the joint angle exceeds the limit. If it is detected that the limit is exceeded, the torque is cut according to the mechanical limit; Calculate the real-time distance between each finger end. If the real-time distance between the finger ends is less than the preset safety threshold, the gradient reverse adjustment is triggered. Dynamically scale the torque output amplitude based on the object's motion state.
9. The adaptive control method for dexterous hands based on multi-agent collaborative reinforcement learning according to claim 7, characterized in that: The method further includes deploying the trained collaborative grasping strategy model to the physical bionic dexterous hand control system, generating control instructions by real-time acquisition of sensor data, and driving the bionic dexterous hand to perform adaptive grasping operations according to the control instructions. When the tactile sensor detects that the continuous collision force exceeds the threshold, the current control instruction is interrupted; Initiate the emergency recovery strategy and control all fingers to reposition along the preset safety trajectory.
10. A multi-agent collaborative reinforcement learning adaptive control device for a dexterous hand, characterized by: include: An environment construction module, used to construct a simulation environment containing a multi-degree-of-freedom bionic dexterous hand and a target object through a physics engine; a rule generation module for generating linkage rules for constraining the movements of each finger based on the multi-finger collaborative kinematic characteristics of the bionic dexterous hand; the linkage rules comprising at least one of a joint angle constraint, an end distance constraint, a collaborative trajectory constraint, and a dynamic coupling constraint; A strategy training module is used to configure each finger as an independent agent, use a multi-agent proximal strategy optimization algorithm combined with the linkage rules to constrain the movements of each finger, and train a multi-finger collaborative grasping strategy model; The control module is used to deploy the trained collaborative grasping strategy model to the physical bionic dexterous hand control system, generate control instructions by real-time acquisition of sensor data, and drive the bionic dexterous hand to perform adaptive grasping operations according to the control instructions.
Citation Information
Cited By
Hybrid model based on reinforcement learning and multiple experts
CN121649987A
Reinforcement learning method for tendon rope driven dexterous hand operation
CN122231923A