Musculoskeletal virtual dexterous hand shape control method and system based on layering strategy

By decoupling the control task of the musculoskeletal virtual dexterous hand into high-level motion planning and low-level muscle actuation through a hierarchical strategy, and combining deep reinforcement learning and supervised learning, the problems of low learning efficiency and insufficient control precision in the control of the musculoskeletal virtual dexterous hand are solved, and efficient and stable control results are achieved.

CN120901956AActive Publication Date: 2025-11-07SOUTHEAST UNIV

Patent Information

Application Number
CN202511156590.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-07
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

Existing technologies for controlling musculoskeletal virtual dexterous hands suffer from low learning efficiency, unstable strategies, and insufficient control precision, especially in the face of control challenges under high-dimensional redundancy and nonlinear dynamic characteristics.

Method used

A hierarchical strategy is adopted to decouple the control task into high-level motion planning and low-level muscle drive. By combining deep reinforcement learning and supervised learning, control modules are designed and trained in joint space and muscle space respectively, and dynamic calculations are performed using the Lagrange dynamics equations.

Benefits of technology

It significantly improves learning efficiency and strategy stability, enhances control precision, and achieves efficient and reliable control of high-dimensional redundant bionic robot systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120901956A_ABST
    Figure CN120901956A_ABST
Patent Text Reader

Abstract

The invention discloses a musculoskeletal virtual dexterous hand shape control method and a musculoskeletal virtual dexterous hand shape control system based on a layering strategy. The core of the method is to decouple a complex muscle control task into two layers of high-layer kinematics planning and bottom-layer dynamics mapping. Performing motion planning in a low-dimensional joint space by adopting a deep reinforcement learning algorithm in a high-level hand shape simulation layer to generate a target joint angle sequence; in the muscle control layer of the bottom layer, through a neural network subjected to supervised learning training, joint instructions planned on the upper layer are efficiently and accurately mapped into activation signals for driving multiple paths of redundant muscles. Through the hierarchical decoupling strategy, the high-level controller avoids the problem of direct exploration in the native high-dimensional muscle space, and the bottom-level controller is specially responsible for solving the problem of complex nonlinear mapping from a moving target to redundant muscle driving. According to the method, the learning efficiency and stability of the control strategy are remarkably improved, and the final precision of hand shape simulation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence and robot technology, in particular to a muscle-skeleton virtual dexterous hand hand type control method and system based on hierarchical strategy. BACKGROUND

[0002] Under the background of rapid development of artificial intelligence and robot technology, dexterous hands have been widely used in humanoid robots, virtual reality and medical rehabilitation due to their advantage of simulating complex human hand movements. The development of dexterous hand control technology has evolved from simple two-fingered grippers to five-fingered dexterous hand systems that can highly simulate human hand functions.

[0003] Currently, there are two main technical paths to achieve dexterous hand control. The first is to develop mechanical structure dexterous hands such as Shadow Hand. This type of mechanical hand usually relies on motors, cables and other actuators to achieve multi-degree-of-freedom control. Although significant progress has been made, mechanical hands still face challenges in flexibility, compliance and bionics due to their inherent rigid or semi-rigid materials and traditional driving transmission methods, making it difficult to completely simulate the compliant characteristics of biological tissues and efficient transmission.

[0004] To overcome the limitations of the above mechanical hands, the second research path, i.e. constructing a muscle-skeleton virtual dexterous hand model based on biomechanical properties, has become an important direction to solve the compliance and flexibility of dexterous control. This method constructs a highly bionic human hand system containing bones, joints and muscle tendon units through computational modeling. However, the control of this muscle-skeleton virtual dexterous hand also faces great technical challenges. Its motion control presents high dimensionality, significant redundancy and complex nonlinear dynamics. Specifically, the number of muscles far exceeds the degrees of freedom of joints, resulting in a high degree of redundancy in the system; at the same time, muscles can only produce one-way driving characteristics of tension, and their force output is closely related to the current length and contraction speed, which further increases the difficulty of designing control algorithms.

[0005] To address the above complex control problems, the industry has begun to use data-driven methods such as deep reinforcement learning. One direct control idea is to use an end-to-end reinforcement learning model that tries to directly learn a mapping strategy from the state information of the dexterous hand to all muscle activation signals. However, practice shows that this end-to-end learning method, although direct, has a huge search space for strategy optimization due to the need to directly explore in the muscle activation space containing dozens of degrees of freedom, which in turn leads to a series of problems such as low learning efficiency, unstable strategy and insufficient control accuracy. During the training process, the learning curve often fluctuates sharply, the convergence speed is slow, and the control accuracy and response speed ultimately reached are also difficult to meet the requirements of fine operation.

[0006] Therefore, how to design a new control strategy to effectively deal with the high-dimensional redundancy and nonlinear dynamics of the musculoskeletal system, while ensuring control accuracy, significantly improving learning efficiency and stability, is a technical problem that needs to be solved in the current field. SUMMARY

[0007] In view of the low learning efficiency, unstable strategy and insufficient control accuracy of the existing technology in using end-to-end reinforcement learning to control the musculoskeletal virtual dexterous hand, the application proposes a musculoskeletal virtual dexterous hand shape control method and system based on hierarchical strategy. The application decouples the complex control task into high-level motion planning and low-level muscle driving, significantly reducing the difficulty of exploration of reinforcement learning. Compared with the prior art, the application has significant advantages in learning efficiency, strategy stability and hand shape imitation accuracy, and provides an efficient and reliable control scheme for high-dimensional redundant bionic robot systems.

[0008] The application adopts the following technical scheme:

[0009] A musculoskeletal virtual dexterous hand shape control method based on hierarchical strategy, according to the current state of the dexterous hand and a target hand shape posture, a first control module for motion planning in joint space is used to generate target joint motion instructions; according to the target joint motion instructions, a second control module for mapping joint motion to muscle activation is used to generate muscle activation signals for driving multiple muscles of the dexterous hand; the muscle activation signals are input to a preset musculoskeletal dynamics model, and the force generated by the muscle and the actual joint motion are calculated according to the Lagrange dynamics equation, so as to update the state of the dexterous hand to imitate the target hand shape posture. Specifically, the following steps are included:

[0010] Step S1, target hand shape posture acquisition and preprocessing. The three-dimensional key point data of the reference hand shape are acquired, and the coordinate system mapping and inverse kinematics method are used to convert them into target joint angle vectors suitable for the musculoskeletal virtual dexterous hand model.

[0011] Step S2, high-level kinematics planning. In the hand shape imitation layer, according to the current joint state of the dexterous hand and the target joint angle vector obtained in step S1, a first control module based on deep reinforcement learning is used for motion planning to output an intermediate target joint angle sequence.

[0012] Step S3, bottom-level dynamics mapping. In the muscle control layer, according to the intermediate target joint angle sequence output in step S2, a second control module based on supervised learning is used to accurately map the motion planning instructions to the muscle activation signals for driving the multiple muscles of the dexterous hand.

[0013] Step S4, musculoskeletal dynamics execution. The muscle activation signal generated in step S3 is input into a preset musculoskeletal dynamics model. The model first calculates the muscle force generated by each muscle unit according to the Hill muscle model, and then substitutes the muscle force into the Lagrangian dynamics equation of the system:

[0014]

[0015] In this way, the actual joint acceleration driven by the muscle is solved, and the joint angular velocity and joint angle of the dexterous hand are updated in turn, thereby completing the physical process of hand shape imitation.

[0016] Step S5, training the first control module (hand shape imitation layer) and the second control module (muscle control layer). Through training, the first control module can generate an optimal intermediate target joint angle sequence, and the second control module can generate an optimal muscle activation signal according to the sequence, and finally drive the dexterous hand to accurately and stably complete the hand shape imitation task.

[0017] Further, the specific steps of step S1 are as follows:

[0018] Step S11, data source acquisition: acquiring original data containing multiple sets of three-dimensional hand key point coordinates from a public hand shape posture data set. In an embodiment of the present application, the data set is the FreiHand data set.

[0019] Step S12, target point mapping: in order to unify the coordinate system, an optimal rigid body transformation matrix is solved by using the Kabsch algorithm, and a standardized pose key point set in the muscle-skeletal virtual dexterous hand model coordinate system is mapped from the key point coordinate set in the original data.

[0020] Step S13, joint angle conversion: for the standardized pose key points obtained in step S12, an iterative optimization algorithm based on the Jacobian matrix is used to solve the inverse kinematics problem. The problem is modeled as a nonlinear least squares optimization problem, and the optimization objective is represented by the following formula:

[0021]

[0022] Where θ is the joint angle vector to be solved, k is the number of end effectors (finger tips), F i (θ) is the forward dynamics function, which represents the three-dimensional space coordinates of the i-th end effector when the joint angle is θ, P i,target is the target position of the i-th end effector. By iteratively solving the optimization problem, the standardized pose key points are converted into the target joint angle vector of the dexterous hand model.

[0023] Further, the specific steps of step S2 are as follows:

[0024] Step S21, constructing a state space: a state space is constructed for the first control module, and the state vector of the state space is spliced with the body perception information and task-related information of the dexterous hand, specifically including: the current angles of all joints, the current angular velocities of all joints, the target joint angle vector, and the error vector of the current posture and the target posture.

[0025] Step S22, generating a target joint angle: input the state vector in step S21 into the first control module, and the policy network (Actor network) of the module outputs an action vector, which is the intermediate target joint angle to be tracked in the next control period.

[0026] Further, the specific steps of step S3 are as follows:

[0027] Step S31, calculating a target joint torque: according to the error between the intermediate target joint angle output by step S2 and the current joint angle of the dexterous hand, a proportional-differential (PD) controller is used to calculate the required joint torque for driving each joint to track the intermediate target:

[0028]

[0029] wherein K p is the proportional gain, K d is the differential gain, θ target is the intermediate target joint angle, θ current and are the current joint angle and angular velocity, respectively. Then, the torque is converted into a target joint angular acceleration according to the rigid body dynamics equation.

[0030] Step S32, mapping muscle activation signals: input the target joint angular acceleration calculated in step S31 and the current muscle state of the dexterous hand as inputs, and output a multi-dimensional vector through the second control module (a pre-trained supervised learning neural network), wherein each dimension of the vector corresponds to the activation level of a muscle block, ranging from 0 to 1.

[0031] Further, the specific steps of step S4 are as follows:

[0032] Step S41, calculating muscle force: input the muscle activation signal generated in step S32 into a mechanics analysis module based on the Hill three-element muscle force model. According to the current activation level, muscle fiber length, and contraction speed of the muscle, the active force f CE and the passive force f PE generated by each muscle tendon unit are calculated, so as to obtain the final muscle force f m transferred to the skeleton.

[0033] Step S42, Update the kinetic state: Update the muscle forces f calculated in step S41. m Substituting the pre-defined Lagrange equations for the musculoskeletal system:

[0034]

[0035] Where q is a generalized coordinate. It's speed. It is acceleration; M(q) is the mass matrix. It is the matrix of Coriolis force and centrifugal force, f m and f c These are muscle force and constraint force, respectively, α m J indicates the degree of muscle activation. m and J c It is the Jacobian matrix that maps muscle force and constraint force to generalized coordinates, τ. ext It is an external force.

[0036] By solving this equation, the actual joint acceleration of the system driven by muscle force can be obtained. Finally, by measuring the actual joint acceleration Perform numerical integration to update the joint angular velocities of the dexterous hand sequentially. The joint angle q is used to complete the state update of one control cycle, thereby realizing the physical simulation of hand movement.

[0037] Furthermore, the specific steps of step S5 are as follows:

[0038] Step S51: Training of the first control module: The first control module is trained using the Proximal Policy Optimization (PPO) algorithm based on the Actor-Critic framework. A composite reward function is designed to guide the training. This reward function is a weighted sum of four components: attitude error reward, success reward, control cost penalty, and velocity penalty. It aims to maximize imitation accuracy while minimizing motion energy consumption and jitter. The formula is as follows:

[0039] R t =w pose R pose,t +R bonus,t +w ctrl R ctrl,t +w vel R vel,t

[0040] Where w pose R is the attitude error reward coefficient. pos,t The reward for attitude error is used to drive the agent to reduce the error between the current attitude and the target attitude; R bonus,tReward for success, a one-time positive incentive given when the pose error is less than a preset threshold; w ctrl Control penalty coefficient, R ctrl,t Control cost penalty, used to suppress unnecessary muscle over-activation; w vel Velocity penalty coefficient, R vel,t Velocity penalty, used to suppress oscillation near the target, and promote model stability.

[0041] Step S52, training of the second control module: the mapping problem from joint torque to muscle activation is reformulated as a constrained quadratic programming (QP) problem, whose optimization objective is to minimize the difference between the joint acceleration generated by muscle activation and the target joint acceleration, and to minimize the energy consumption of the overall muscle activation. The objective function is expressed as:

[0042]

[0043] where α is the muscle activation vector, M(q) is the mass matrix, Coriolis force and centrifugal force terms, w reg Regularization term weight, used to penalize excessively high muscle activation. Further, the QP problem is converted into a regression problem under the supervised learning framework, and a training set is constructed by sampling data from the training experience pool of the first control module, where the input features are the target joint angular acceleration, and the labels are the optimal muscle activation levels obtained by solving the QP problem. The neural network of the second control module is trained by minimizing the prediction error.

[0044] A muscle-skeleton virtual dexterous hand handshape control system based on a hierarchical strategy, comprising: a first control module configured to perform motion planning in joint space according to a current state of the dexterous hand and a target handshape pose, to generate target joint motion instructions; a second control module configured to map joint motion to muscle activation according to the target joint motion instructions, to generate muscle activation signals for driving multiple muscles of the dexterous hand; and a muscle-skeleton dynamics module configured to receive the muscle activation signals generated by the second control module, and calculate actual joint motion according to a preset Lagrangian dynamics equation, to update the state of the dexterous hand.

[0045] Further, the first control module is a deep reinforcement learning network, and the second control module is a supervised learning network, and the system further comprises a proportional-differential controller; the working process is as follows:

[0046] (1) the first control module (deep reinforcement learning network) is configured to receive the current state of the dexterous hand and the target handshape pose, and output an intermediate target joint angle;

[0047] (2), the system further comprises a proportional-differential (PD) controller configured to calculate a target joint torque and a target joint acceleration according to an error between the intermediate target joint angle and the current joint angle of the dexterous hand;

[0048] (3), the second control module (supervised learning network) is configured to receive the target joint torque calculated by the PD controller and map it to a muscle activation signal used to drive the plurality of muscles of the dexterous hand;

[0049] (4), the muscle activation signal is input into a preset musculoskeletal dynamics model, which first calculates the muscle force generated by each muscle unit according to the Hill muscle model, then substitutes the muscle force into the Lagrangian dynamics equation of the system to solve the actual joint acceleration driven by the muscle, and sequentially updates the joint angular velocity and joint angle of the dexterous hand, so that it interacts with the physical rules in the simulation environment to complete the hand shape imitation task.

[0050] The beneficial effects of the present application are that by decoupling the complex muscle control task into high-level kinematic planning and low-level dynamics mapping, the high-level strategy can be learned in a low-dimensional, more task-oriented joint space, while the low-level can accurately handle the high-dimensional, redundant muscle dynamics problem through supervised learning. This hierarchical strategy significantly reduces the exploration difficulty of reinforcement learning, thereby greatly improving the learning efficiency, strategy stability and final motion control accuracy of the model. BRIEF DESCRIPTION OF DRAWINGS

[0051] Figure 1 A framework diagram for the muscle-skeletal virtual dexterous hand shape generation of the present application is provided.

[0052] Figure 2 A structure diagram of the Actor-Critic neural network in the hand shape imitation layer is provided.

[0053] Figure 3 A neural network structure diagram of the muscle control layer is provided. DETAILED DESCRIPTION

[0054] The present application will be further illustrated below in conjunction with the drawings and specific embodiments, and it should be understood that the specific embodiments are only used to illustrate the present application and not to limit the scope of the present application.

[0055] REFERENCE Figure 1A hand shape control framework of the musculoskeletal virtual dexterous hand is shown in the present application. In one embodiment of the present application, the musculoskeletal virtual dexterous hand model is constructed according to human hand anatomy to achieve highly biomimetic dynamic simulation. The model contains 20 joints, 26 degrees of freedom and 39 muscle-tendon units. This complex and redundant structure is the direct source of the control problem to be solved in the present application. The technical solution of the present application can include the following steps:

[0056] Step one, target hand shape posture acquisition and preprocessing. Obtain the three-dimensional key point data of the reference hand shape from the hand shape posture data set, and convert it into the target joint angle vector suitable for the musculoskeletal virtual dexterous hand model of the present application through coordinate system mapping and inverse kinematics method.

[0057] Step two, high-level kinematics planning. Referring to Figure 2 In the hand shape imitation layer, according to the current joint state of the dexterous hand and the target joint angle vector obtained in step one, a first control module (Actor-Critic) based on deep reinforcement learning is used for motion planning to output an intermediate target joint angle sequence.

[0058] Step three, control module training and dynamics execution. The first and second control modules described in steps two and three are trained respectively. After training, the generated muscle activation signal is input into the musculoskeletal dynamics system. The system calculates the muscle force according to the Hill muscle model and substitutes it into the Lagrange dynamics equation to solve the actual joint acceleration to update the joint state, thereby driving the dexterous hand to accurately and stably complete the hand shape imitation task.

[0059] Step four, control module training and execution. The first and second control modules described in steps two and three are trained offline or online, and finally the trained model is used to drive the dexterous hand to accurately and stably complete the hand shape imitation task.

[0060] The specific implementation of step one is as follows:

[0061] In order to obtain a usable target hand shape posture, first, the original data containing multiple sets of three-dimensional hand key point coordinates are obtained from the public hand shape posture data set (FreiHand data set). Since the original data coordinate system is different from the musculoskeletal virtual dexterous hand model coordinate system used in the present application, alignment is required. This embodiment uses Kabsch algorithm to solve an optimal rigid transformation matrix to map the key point coordinate set in the original data to a standardized pose key point set in the dexterous hand model coordinate system.

[0062] After obtaining the aligned key points, inverse kinematics is used to convert them into a target joint angle vector. This process is modeled as a nonlinear least squares optimization problem, whose optimization objective is to find a set of joint angles θ that minimizes the position error between the model key points and the target key points:

[0063]

[0064] where k is the number of end effectors (fingertips), F i (θ) is the forward dynamics function, which represents the 3D spatial coordinates of the ith end effector when the joint angles are θ, P i,target is the target position that the ith end effector is expected to reach.

[0065] The neural network model of the hand imitation layer in step two is as shown in Figure 2 , which is specifically implemented as:

[0066] The hand imitation layer is responsible for high-level motion planning. Its state space is designed to include a comprehensive vector of dexterous hand physical state and task goal information, specifically including: current joint angles, current joint angular velocities, target joint angles, and error vectors between current poses and target poses.

[0067] Both the Actor network and the Critic network use a multi-layer perceptron (MLP) structure. In a specific embodiment, both networks contain two hidden layers with 256 neurons each, and use LeakyReLU as the activation function. After the state vector is input into the Actor network, the network outputs an action, which is the intermediate target joint angle that the PD controller needs to track in the next control period. The Critic network evaluates the value of the state vector, and its output is used to guide the update of the Actor network during the training process.

[0068] In step three, the neural network model of the PD controller and the muscle control layer is as shown in Figure 3 , which is specifically implemented as:

[0069] This step aims to convert the motion intention planned by the upper layer into muscle activation signals at the bottom layer. First, a PD controller calculates the joint torque τ required to drive each joint according to the error between the intermediate target joint angle θ target output by the hand imitation layer and the current joint angle θ current :

[0070]

[0071] where K p is the proportional gain, K d is the derivative gain, and θ targetθ current and are the current joint angle and angular velocity, respectively. Refer to Figure 3 , the target joint acceleration calculated from the joint torque τ is input to the muscle activation mapping network of the muscle control layer.

[0072] This network is responsible for solving the muscle redundancy problem. In one specific embodiment, this network is a fully connected neural network containing three hidden layers (with 512, 512, 256 neurons, respectively). The network outputs a 39-dimensional muscle activation vector α, each dimension corresponding to the activation level of a muscle. The output layer uses a Tanh function followed by a ReLU function to ensure that the activation values are constrained within the [0, 1] interval. The output muscle activation vector is used in the musculoskeletal dynamics system. Specifically, according to the Hill muscle model, the muscle force f m generated by each muscle is calculated using the activation level. Subsequently, these muscle forces are substituted into the Lagrangian dynamics equation of the system:

[0073]

[0074] where q is the generalized coordinate, is the velocity, is the acceleration. M(q) is the mass matrix, is the Coriolis force and centrifugal force matrix, f m and f c are the muscle force and constraint force, respectively, α m represents the muscle activation level, J m and J c are the Jacobian matrices that map the muscle force and constraint force to the generalized coordinates, τ ext is the external force.

[0075] The actual joint acceleration is obtained by solving this equation and updating the joint angular velocity and joint angle in turn, thereby realizing the accurate calculation and simulation of the physical movement of the virtual dexterous hand.

[0076] In step four, the training process of the two control modules is specifically implemented as:

[0077] 1. The training of the hand shape imitation layer uses the Proximal Policy Optimization (PPO) algorithm to train the Actor-Critic network shown in Figure 2 . In order to guide the agent learning, a composite reward function is designed:

[0078] R t = w pose R pose,t + R bonus,t + w ctrlR ctrl,t +w vel R vel,t

[0079] where w pose is the pose error reward coefficient, R pos,t is the pose error reward, used to drive the agent to reduce the error between the current pose and the target pose; R bonus,t is the success reward, a one-time positive incentive given when the pose error is less than a preset threshold; w ctrl is the control penalty coefficient, R ctrl,t is the control cost penalty, used to suppress unnecessary muscle overactivation; w vel is the speed penalty coefficient, R vel,t is the speed penalty, used to suppress oscillation near the target and promote model stability.

[0080] 2. Training of the muscle control layer: the mapping problem of joint torque to muscle activation is converted into a constrained quadratic programming (QP) regression problem, whose loss function aims to minimize the difference between the actual generated acceleration and the target acceleration, and to regularize the muscle activation energy penalty:

[0081]

[0082] where α is the muscle activation vector, M(q) is the mass matrix, Coriolis force and centrifugal force terms, w reg is the weight of the regularization term, used to penalize excessive muscle activation. The dataset generated by solving this QP problem is used to supervise the learning training of the muscle activation mapping network shown in FIG. 2. Figure 3

[0083] It should be noted that the above content only illustrates the technical idea of the present application, and cannot limit the protection scope of the present application. For ordinary skilled persons in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which fall within the protection scope of the claims of the present application.​

Claims

1. A hierarchical strategy based musculoskeletal virtual dexterous hand handshape control method, characterized in that, According to the current state of the dexterous hand and a target hand posture, a first control module for motion planning in joint space is used to generate target joint motion instructions; According to the target joint motion instructions, a second control module for mapping joint motion to muscle activation is used to generate muscle activation signals for driving multiple muscles of the dexterous hand; The muscle activation signals are input into a preset musculoskeletal dynamics model, and the force generated by the muscles and the actual joint motion are calculated according to the Lagrange dynamics equation, so as to update the state of the dexterous hand to imitate the target hand posture.

2. The method of claim 1, wherein, Specifically, the following steps are included: Step S1, target hand posture acquisition and preprocessing; three-dimensional key point data of a reference hand posture are acquired, and are converted into a target joint angle vector suitable for a musculoskeletal virtual dexterous hand model through coordinate system mapping and inverse kinematics method; Step S2, high-level kinematics planning; in the hand posture imitation layer, a first control module based on deep reinforcement learning is used for motion planning according to the current joint state of the dexterous hand and the target joint angle vector obtained in step S1, and an intermediate target joint angle sequence is output; Step S3, bottom-level dynamics mapping; in the muscle control layer, a second control module based on supervised learning is used to accurately map the motion planning instructions into muscle activation signals for driving multiple muscles of the dexterous hand according to the intermediate target joint angle sequence output in step S2; Step S4, musculoskeletal dynamics execution; The muscle activation signals generated in step S3 are input into a preset musculoskeletal dynamics model; The model first calculates the muscle force generated by each muscle unit according to the Hill muscle model, then substitutes the muscle force into the musculoskeletal dynamics model to solve the actual joint acceleration driven by the muscles, and sequentially updates the joint angular velocity and joint angle of the dexterous hand, thereby completing the physical process of hand posture imitation; Step S5, training the first control module and the second control module respectively; Through training, the first control module can generate an optimal intermediate target joint angle sequence, and the second control module can generate an optimal muscle activation signal according to the sequence, so as to finally drive the dexterous hand to accurately and stably complete the hand posture imitation task.

3. The method of claim 2, wherein, The specific steps of step S1 are as follows: Step S11, data source acquisition: acquiring original data containing multiple sets of three-dimensional hand key point coordinates from a public hand posture dataset; Step S12, target point mapping: using Kabsch algorithm to solve an optimal rigid transformation matrix, and mapping the key point coordinate set in the original data to a standardized pose key point set in the musculoskeletal virtual dexterous hand model coordinate system; Step S13, joint angle conversion: for the standardized pose key points obtained in step S12, an iterative optimization algorithm based on Jacobian matrix is used to solve the inverse kinematics problem; the problem is modeled as a nonlinear least squares optimization problem, and the optimization objective is represented by the following formula: Where θ is the joint angle vector to be solved, k is the number of end effectors (i.e., fingertips), and F... i (θ) is the positive dynamic function, representing the three-dimensional spatial coordinates of the i-th end effector at joint angle θ. i,target Let be the target position that the i-th end effector expects to reach.

4. The method of claim 3, wherein, The specific steps of step S2 are as follows: Step S21, constructing a state space: a state space is constructed for the first control module, and a state vector of the state space is spliced with body perception information and task-related information of the dexterous hand, specifically including: current angles of all joints, current angular velocities of all joints, a target joint angle vector, and an error vector between a current pose and a target pose; Step S22, generating a target joint angle: inputting the state vector in step S21 into the first control module, and outputting an action vector from a policy network of the first control module, the action vector being an intermediate target joint angle to be tracked in a next control period.

5. The method of claim 4, wherein, The specific steps of the step S3 are as follows: Step S31, calculating a target joint torque: according to an error between the intermediate target joint angle output in step S2 and a current joint angle of the dexterous hand, a proportional-derivative controller is used to calculate a required joint torque for driving each joint to track the intermediate target: where K p is a proportional gain, K d is a derivative gain, θ target is an intermediate target joint angle, θ current and are the current joint angle and angular velocity, respectively; subsequently, the torque is converted to a target joint angular acceleration according to the rigid body dynamics equation; Step S32, mapping a muscle activation signal: taking the target joint angular acceleration calculated in step S31 and a current muscle state of the dexterous hand as inputs, and outputting a multi-dimensional vector from the second control module, each dimension of the vector corresponding to an activation level of a muscle, ranging from 0 to 1.

6. The method of claim 5, wherein, The specific steps of the step S4 are as follows: Step S41, calculating a muscle force: inputting the muscle activation signal generated in step S32 into a mechanics analysis module based on a Hill three-element muscle force model; According to the current activation degree of the muscle, muscle fiber length and contraction speed, the active force f generated by each muscle tendon unit is calculated CE and the passive force f PE , so as to obtain the muscle force f finally transmitted to the bone m ; Step S42, updating the dynamics state: updating the dynamics state of the musculoskeletal system according to the muscle forces f calculated in step S41 m Substitute into the preset musculoskeletal dynamics model, i.e. the musculoskeletal system Lagrange dynamics equation: where q is the generalized coordinate, is the velocity, is the acceleration; M(q) is the mass matrix, is the Coriolis and centrifugal force matrix, f m and f c are muscle force and constraint force, respectively, a m denotes the muscle activation level, J m and J c are the Jacobian matrices that map muscle force and constraint force to the generalized coordinates, τ ext is the external force; By solving this equation, the actual joint acceleration of the system driven by muscle force is obtained Finally, by numerically integrating the actual joint acceleration , the joint angular velocity and joint angle q of the dexterous hand are updated in turn, thus completing the state update of a control cycle and realizing the physical simulation of hand movement.

7. The method of claim 6, wherein, The specific steps of the step S5 are as follows: Step S51, training of the first control module: using a proximal policy optimization (PPO) algorithm based on an Actor-Critic framework to train the first control module; A composite reward function is designed to guide the training, and the reward function is a weighted sum of four components of a pose error reward, a success reward, a control cost penalty, and a speed penalty, aiming to maximize the imitation accuracy while minimizing the motion energy consumption and jitter, and the formula is as follows: R t = w pose R pose,t + R bonus,t + w ctrl R ctrl,t + w vel R vel,t where w pose is the pose error reward coefficient, R pos,t is the pose error reward, used to drive the agent to reduce the error between the current pose and the target pose; w bonus,t is the success reward, a one-time positive incentive given when the pose error is less than a pre-set threshold; w ctrl is the control penalty coefficient, R ctrl,t is the control cost penalty, used to suppress unnecessary muscle over-activation; w vel is the velocity penalty coefficient, R vel,t is the velocity penalty, used to suppress oscillation near the target and promote model stability; Step S52, training of the second control module: the mapping problem from joint torque to muscle activation is re-expressed as a constrained quadratic programming (QP) problem, and the optimization objective is to minimize the difference between the joint acceleration generated by the muscle activation and the target joint acceleration, and to minimize the energy consumption of the overall muscle activation; the objective function is expressed as: where α is the muscle activation vector, M(q) is the mass matrix, Coriolis and centrifugal force terms, w reg is the regularization term weight, used to penalize excessively high muscle activations; further, the QP problem is converted into a regression problem under the supervised learning framework, a training set is constructed by sampling data from the training experience pool of the first control module, wherein the input features are the target joint angular accelerations, and the labels are the optimal muscle activation levels obtained by solving the QP problem, and the neural network of the second control module is trained by minimizing the prediction error.

8. A hierarchical policy based musculoskeletal virtual dexterous hand handshape control system, comprising: It comprises: The first control module is configured to perform motion planning in a joint space according to a current state of the dexterous hand and a target hand posture, to generate a target joint motion instruction; The second control module is configured to map joint motion to muscle activation according to the target joint motion instruction, to generate muscle activation signals for driving multiple muscles of the dexterous hand; The musculoskeletal dynamics module is configured to receive the muscle activation signals generated by the second control module, and calculate actual joint motion according to a preset Lagrangian dynamics equation, to update the state of the dexterous hand.

9. The system of claim 8, wherein, The first control module is a deep reinforcement learning network, the second control module is a supervised learning network, and the system further comprises a proportional-derivative controller; and a working process thereof is as follows: (1) The first control module is configured to receive the current state of the dexterous hand and the target hand posture, and output an intermediate target joint angle; (2) The system further comprises a proportional-differential controller configured to calculate a target joint torque and obtain a target joint acceleration according to the error between the intermediate target joint angle and the current joint angle of the dexterous hand; (3) The second control module is configured to receive the target joint torque calculated by the PD controller and map it to a muscle activation signal used to drive the multiple muscles of the dexterous hand; (4) The muscle activation signal is input into a preset muscle-skeleton dynamics model, which first calculates the muscle force generated by each muscle unit according to the Hill muscle model, then substitutes the muscle force into the Lagrange dynamics equation of the system to solve the actual joint acceleration driven by the muscle, and in turn updates the joint angular velocity and joint angle of the dexterous hand, so that it interacts with the physical rules in the simulation environment to complete the hand posture imitation task.

Citation Information

Patent Citations

  • Human body electromyographic signal direct-drive joint torque mapping method

    CN114227673A

  • Dexterous hand finger joint angle control method, device and system and medium

    CN119489449A

  • Thumb independent optimization five-finger manipulator control method based on hierarchical strategy control

    CN119567270A

  • Dexterous hand control system and method

    CN120363174A

  • A bionic hand and method thereof

    IN202011000596A

Cited By

  • Musculoskeletal robot control method for table tennis playing scene

    CN121340214A

  • A control method for musculoskeletal robots in table tennis scenarios

    CN121340214B

  • Dexterous hand motion planning and control method based on visual language motion model

    CN121696972A