Dexterous palm inner operation control method based on reinforcement learning and real-time pose feedback
By using reinforcement learning and real-time pose feedback methods in the simulation environment, a smart hand control strategy is constructed, which solves the problem of insufficient control flexibility of smart hand in complex environments, and realizes efficient adaptation and fine operation of smart hand in unstructured environments.
Patent Information
- Application Number
- CN202510452131.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Currently, robots have insufficient control flexibility in complex environments, difficult to adapt to unstructured environments, and the results of reinforcement learning and training are difficult to transfer from simulation to reality.
A clever in-palm operation control method based on reinforcement learning and real-time pose feedback is adopted to optimize the control strategy by building experimental scenarios in the simulation environment, building a reinforcement learning strategy network, and using the pose information output from FoundationPose.
It improves the adaptability and intelligence level of dexterity hands, can independently learn and perform fine operations in different task scenarios, reduces human intervention, and improves the stability and generalization ability of task execution.
Smart Images

Figure CN120056125A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a dexterous hand operation control method, in particular to a dexterous hand in-palm operation control method based on reinforcement learning and real-time pose feedback. Background Art
[0002] In the face of the growing demand for automation and the trend of intelligent development, the application of robot dexterous hands has gradually expanded to a wider range of fields. As an important actuator for fine operations, robot dexterous hands have flexible multi-finger coordination capabilities and can complete complex tasks such as grasping and in-palm operations, which have important application values in fields such as aerospace, intelligent manufacturing, and medical rehabilitation.
[0003] However, the current dexterous hand control methods still mainly rely on hard-coded control. Hard-coded control depends on preset logical rules and is difficult to adapt to the dynamically changing task requirements. This control method has great limitations when facing complex and changeable unstructured environments.
[0004] In view of this, there is an urgent need for a robot dexterous hand control method based on reinforcement learning and real-time pose feedback, which can be efficiently trained in a simulation environment and ensure its stability and adaptability in the real environment through Sim-to-Real transfer technology, so as to improve the intelligence level of robot dexterous hands and expand their application scope in industrial production and life services. Summary of the Invention
[0005] Object of the Invention: The present invention aims to solve the problems of insufficient control flexibility of current robot underactuated dexterous hands in complex environments, difficulty in adapting to unstructured environments, and difficulty in migrating the training results of reinforcement learning from simulation to reality, and provides a dexterous hand in-palm operation control method based on reinforcement learning and real-time pose feedback.
[0006] Technical Solution: A dexterous hand in-palm operation control method based on reinforcement learning and real-time pose feedback according to the present invention specifically includes the following steps:
[0007] (1) Build an experimental scenario for underactuated dexterous hand in-palm operation, deploy an underactuated dexterous hand model, an in-palm object model, and an interaction environment in a simulation environment; and collect training data;
[0008] (2) For the underactuated dexterous hand in-palm operation task, construct an in-palm operation reinforcement learning policy network based on the PPO or DDPG algorithm, and design a refined state space, action space, and reward function according to the task requirements and the structural characteristics of the underactuated dexterous hand;
[0009] (3) Use the training data collected in the simulation environment in step (1) and the reinforcement learning policy network established in step (2) for simulation training to learn the optimal policy; write a reinforcement learning configuration file based on the manager and a task environment configuration file to achieve efficient interaction between the reinforcement learning algorithm and the physical simulation platform;
[0010] (4) Use the reinforcement learning policy network trained in step (3), take the in-palm object pose information output by FoundationPose and the dexterous hand joint information directly obtained from the dexterous hand real machine as state inputs, calculate and output the optimal control policy, and perform the motion control of the dexterous hand.
[0011] Furthermore, the training data in step (1) includes the state information of the underactuated dexterous hand and the state information of the in-palm object; the state information of the underactuated dexterous hand includes the joint angles, joint velocities, joint accelerations, and joint positions; the state information of the in-palm object includes the position information and the attitude information.
[0012] Furthermore, the specific steps of step (1) are as follows:
[0013] (11) Use the NVIDIA Isaac Sim physical simulation platform to build an experimental scenario for in-palm operation of the underactuated dexterous hand, and deploy the underactuated dexterous hand model and interaction objects in the scenario to provide a high-fidelity simulation environment;
[0014] (12) For the joint drive characteristics of the underactuated dexterous hand, modify its URDF model file, adjust the constraint conditions and dynamic parameters of each joint, etc., to ensure that the dexterous hand in the simulation environment can truly reflect its motion characteristics in actual operation;
[0015] (13) Based on the modified URDF model file, construct the USD model of the dexterous hand to ensure that the dexterous hand can be correctly loaded in the USD scenario and can perform normal physical interactions with other objects;
[0016] (14) Import the constructed USD model into the NVIDIA Isaac Sim physical simulation environment, verify the interactivity of the dexterous hand in the simulation environment, ensure the physical consistency and stability of the simulation results, and provide a reliable simulation environment for subsequent reinforcement learning training; and collect training data in the simulation environment.
[0017] Furthermore, the reinforcement learning policy network in step (2) includes a state space, an action space, and a reward function; the state space includes the joint angles, joint velocities, joint accelerations, joint position information of the dexterous hand, and the 6D pose information of the object in the palm; the action space includes all control strategies for each joint of the dexterous hand, and a continuous action control method is adopted to ensure that the dexterous hand can perform operations smoothly; the reward function configures reward items based on the reward mechanism of task completion to help the agent evaluate the advantages and disadvantages of the current state and action, and continuously optimize the strategy accordingly.
[0018] Furthermore, the reward function is:
[0019] R = λ 1 R success + λ 2 R pos + λ 3 R ori
[0020] where λ 1 controls the weight of the successful control reward, and λ 2 , λ 3 control the weights of the position tracking reward and the orientation tracking reward respectively. In the combined reward function, the successful control reward provides the ultimate goal, while the position and orientation rewards guide exploration;
[0021] The successful control reward R success is:
[0022] dθ = ||quat error magnitude(s o , s g )||
[0023]
[0024] where s g is the target state of the object in the palm, s o is the current state of the object in the palm. The state information is represented by a homogeneous matrix of quaternion plus position coordinates. quat_error_magnitude calculates the error between the two states, and τ is the error threshold between the current orientation and the target orientation;
[0025] The position tracking reward R pos is:
[0026] R pos = -||p o - p g || 2
[0027] where p g is the target position of the object in the palm, p ois the current position of the object in the palm, using L 2 The norm calculates the Euclidean length between the current position and the target position of the object, and taking the negative value facilitates maximizing the reward in subsequent reinforcement learning training;
[0028] The direction tracking reward R ori is:
[0029] β = ||q o - q g || 2
[0030]
[0031] where q g is the target direction of the object in the palm, q o is the current direction of the object in the palm, ∈ is a small positive number to prevent division by zero errors, and the L2 norm is used to calculate the error between the current position and the target position of the object, and taking the reciprocal facilitates maximizing the reward in subsequent reinforcement learning training.
[0032] Furthermore, the manager-based reinforcement learning configuration file in step (3) includes a command manager, an action manager, an observation manager, an event manager, a termination manager, and a reward manager; the command manager defines high-level task commands and places the target object in the dexterous hand in each training episode; the action manager defines the basic operations that the dexterous hand can perform in the experimental scenario; the observation manager defines the observation items of the dexterous hand during the task process, including joint state information and target object pose information, providing perceptual data for reinforcement learning decision-making; the event manager defines and manages events in different modes to assist in managing state transitions during the training process; the termination manager defines the task termination conditions for determining whether the current training episode should terminate, including object dropping, operation failure, or reaching the target state; the reward manager designs a reward function, configures reward items, and sets weight coefficients according to task requirements to encourage the agent to achieve the task goal and punish movements that have an adverse impact on the task, improving the stability and efficiency of the policy.
[0033] Furthermore, the task environment configuration file in step (3) uses pure state input to drive decision-making and conducts reinforcement learning training in multiple task scenarios, enabling the reinforcement learning policy to adapt to different object types, initial poses, and environmental constraints, achieving an improvement in multi-scenario generalization ability and enhancing the autonomous adaptability and generalization control ability of the underactuated dexterous hand.
[0034] Furthermore, the specific steps of step (4) are as follows:
[0035] (41) Calibrate the RGBD camera using Matlab: Fix the position of the depth camera and move the camera calibration board. Use the Matlab camera calibration tool to calibrate the collected calibration board images to obtain the camera internal parameter matrix;
[0036] (42) Construct the OBJ model file of the object to be estimated as a reference for pose estimation; if the CAD model file of the object cannot be obtained, utilize the generalization ability of large-scale synthetic training of FoundationPose to complete the pose estimation of the object;
[0037] (43) Use the depth camera to obtain the RGB image and depth image in the experimental environment. Select the first frame of the RGB image and perform mask annotation using SAM to optimize the feature extraction and pose estimation of the target object;
[0038] (44) Adopt the Mask R-CNN detector to obtain the 2D bounding box of the object. Initialize the translation through the median value of the 3D points within the bounding box to determine the initial position of the object. Centered on the target object, uniformly sample multiple rotation angles to generate the rotation matrix for global pose initialization as the initial pose; input the image data collected by the RGBD camera, and FoundationPose outputs the pose update information of the object within the palm;
[0039] (45) Adopt the trained reinforcement learning policy network, use the current state information as input, calculate and output the optimal control policy to achieve precise motion control of the dexterous hand;
[0040] (46) The 6D pose information of the object within the palm output by the FoundationPose algorithm, that is, the position information and orientation information of the object within the palm relative to the camera; the real-time joint information of the dexterous hand directly collected from the dexterous hand real machine, including joint angles, joint speeds, and end positions;
[0041] (47) Adopt the reinforcement learning policy network, calculate the optimal action policy according to the current state information, convert the optimal control policy output by the reinforcement learning policy network into low-level control signals, and send them to the dexterous hand actuator; control the dexterous hand to perform fine motions to complete the expected in-palm operation tasks;
[0042] (48) During the execution process, the real-time state of the dexterous hand and the pose of the object within the palm change in real time. The policy network updates the state observation information at each time step, dynamically adjusts, and continuously outputs the control policy.
[0043] Beneficial effects: Compared with the prior art, the beneficial effects of the present invention are:
[0044] 1. The present invention uses reinforcement learning to autonomously explore and optimize control strategies through a trial-and-error learning and reward-driven mechanism during the continuous interaction of the dexterous hand, thereby improving its adaptability and intelligence level; through the proximal policy optimization reinforcement learning method, the dexterous hand can autonomously learn appropriate operation strategies in different task scenarios, reduce human intervention, and improve the stability and generalization ability of task execution.
[0045] 2. Since the movement of the underactuated dexterous hand involves continuous control of multiple degrees of freedom and multiple joints sharing a single drive source, and the action space is complex, traditional reinforcement learning algorithms are difficult to efficiently search for optimal strategies. Therefore, a probability distribution is used to represent the strategy, and actions are selected through sampling rather than searching for the best action in a finite set of actions, which can naturally handle the continuous action space. Thus, the present invention uses a reinforcement learning policy network based on the PPO or DDPG algorithm for policy output, and optimizes the policy distribution through the policy gradient method, so that the new policy does not deviate too far from the old policy, which helps to maintain the stability of learning, reduce the policy search complexity, and improve the training efficiency.
[0046] 3. Since reinforcement learning training requires a large amount of trial-and-error data, if deployed on a real machine for training, the time and economic costs are too high. Therefore, the present invention uses a method of building an experimental scenario in a physical simulation platform for parallel reinforcement learning training, which not only greatly improves the training efficiency but also saves economic costs. After the RGBD camera obtains the environmental RGB image and depth image information, and the deep learning network outputs the 6D pose information of the object in the palm, only by inputting the pose information into the reinforcement learning policy network trained in the simulation task scenario can the optimal control strategy of the dexterous hand be obtained. The optimal control strategy output by the reinforcement learning policy network is converted into low-level control signals to drive the dexterous hand to autonomously complete tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flowchart of the present invention;
[0048] Figure 2 is a schematic diagram of the reinforcement learning training process of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0049] The present invention will be further described in detail below with reference to the accompanying drawings.
[0050] As Figure 1 shown, the present invention proposes a dexterous hand in-palm operation control method based on reinforcement learning and real-time pose feedback, which includes the following steps:
[0051] Step 1: Use the NVIDIA Isaac Sim physical simulation platform to build an experimental scenario for underactuated dexterous hand in-hand operation. Deploy the underactuated dexterous hand model, in-hand object model, and interaction environment in the simulation environment. Based on the physics engine, set parameters such as collision detection and dynamic constraints to ensure the high fidelity of the simulation environment. Conduct data collection in the simulation environment. The collected training data includes the state information of the dexterous hand in real time, such as joint angles, joint velocities, and end effector positions, as well as the position information and pose information (6D pose information) of the target object.
[0052] Use the NVIDIA Isaac Sim physical simulation platform to build an experimental scenario for underactuated dexterous hand in-hand operation, and deploy the underactuated dexterous hand model and interaction objects in it to construct a high-fidelity simulation environment. For the joint drive characteristics of the underactuated dexterous hand, adjust its URDF model file, optimize the joint constraint conditions and dynamic parameters to ensure that the motion characteristics of the dexterous hand in the simulation environment can truly reflect the actual operation situation. Based on the modified URDF model file, construct the corresponding USD model so that the dexterous hand can be correctly loaded in the USD scene and perform normal physical interactions with other objects. Subsequently, import the completed USD model into the NVIDIA Isaac Sim physical simulation environment to verify its interaction performance, including key indicators such as collision detection and dynamic response, and ensure that the simulation results meet high standards in terms of physical consistency and stability, thus providing reliable data support for subsequent reinforcement learning training.
[0053] For the structural characteristics and operation requirements of the underactuated dexterous hand, based on the high-precision physics engine built into the simulation platform, configure the scenario, including collision detection, friction coefficient setting, rigid body properties, etc., to ensure the authenticity of the simulation environment, obtain a simulation effect highly similar to the real scenario, and provide reliable environmental support for reinforcement learning training. Conduct systematic data collection in the constructed simulation environment. The collected training data is divided into two categories. The first category is the state information s of the dexterous hand itself a , including the real-time angles, angular velocities, and angular accelerations of each joint, as well as the dynamic information such as the position and pose of the end effector of the dexterous hand. The second category is the state information s of the in-hand object b , including three-dimensional position coordinate information and rotational pose information, which are used to reflect the dynamic changes and feedback information of the in-hand object during the operation.
[0054] Step 2: For the characteristics of the underactuated dexterous hand in-hand operation task, such as high state dimension, strong dynamic constraints, and complex interactions, propose Figure 2The method for constructing an efficient operation decision-making network based on the reinforcement learning algorithm. The policy network outputs actions that act on the simulation environment, obtains state information and reward information from the simulation environment for experience storage. The evaluation network outputs the value function after the current action is applied, and the advantage function determines whether the current action network is superior to the old action network, calculates the gradient of the loss with respect to the weights, and updates the policy network parameters to improve the operation autonomy and adaptability of the dexterous hand in the simulation environment. Through reinforcement learning, dynamic decision-making and action planning of the dexterous hand in complex in-hand operation tasks are realized, overcoming the problem that traditional hard-coding methods are difficult to adapt to the random changes of complex environments.
[0055] Considering the special structure of the underactuated dexterous hand, where multiple joints share a single drive source, with a continuous action space and a high degree of freedom, traditional reinforcement learning methods have problems such as high action search complexity and slow convergence speed when facing a continuous action space, making it difficult to achieve efficient decision-making. Therefore, a reinforcement learning algorithm framework based on the Proximal Policy Optimization (PPO) or Deep Deterministic Policy Gradient (DDPG) algorithm is built to reduce the search complexity of the action space and improve the learning efficiency and decision-making quality.
[0056] The reinforcement learning policy network for the underactuated dexterous hand's in-hand operation task includes the design of a refined state space, action space, and reward function. State space S: S ∈ R n including information such as the joint angles, joint velocities, joint accelerations, and joint positions of the dexterous hand, as well as the 6D pose information of the object in the hand; Action space A: A ∈ R n All control strategies for each joint of the dexterous hand, using a continuous action control method to ensure that the dexterous hand can perform operations smoothly. Reward function R: Design a reward mechanism based on task completion. Three reward items are configured in the in-hand operation task to help the intelligent agent evaluate the quality of the current state and action, and continuously optimize the strategy accordingly. The reward function is configured as follows:
[0057] Success reward R success :
[0058] dθ = ||quaterror magnitude(s o , s g )||
[0059]
[0060] where s g is the target state of the object in the hand, s ois the current state of the object in the palm. The state information is represented by a homogeneous matrix of quaternion plus position coordinates. quat_error_magnitude calculates the error between two states, and τ is the error threshold between the current direction and the target direction. The success reward has a clear goal and does not rely on complex error calculations, making the training more efficient. If the agent successfully reaches the goal, it can obtain a fixed high reward, which helps to stabilize the learning.
[0061] Position tracking reward R pos :
[0062] R pos = -||p o - p g || 2
[0063] where p g is the target position of the object in the palm, and p o is the current position of the object in the palm. The L 2 norm is used to calculate the Euclidean length between the current position and the target position of the object. Taking the negative value is convenient for maximizing the reward in subsequent reinforcement learning training. The position tracking reward enables the agent to obtain continuous feedback and can get rewards without completely reaching the goal, improving the stability of training.
[0064] Orientation tracking reward R ori :
[0065] β = ||q o - q g || 2
[0066]
[0067] where q g is the target orientation of the object in the palm, and q o is the current orientation of the object in the palm. ∈ is a small positive number to prevent division by zero errors. The L2 norm is used to calculate the error between the current position and the target position of the object. Taking the reciprocal is convenient for maximizing the reward in subsequent reinforcement learning training. The orientation tracking reward encourages the agent to minimize the rotation error as much as possible and quickly adjust to the correct direction, which is especially suitable for tasks of dexterous in-palm operation.
[0068] Final reward function:
[0069] R = λ 1 R success + λ 2 R pos + λ 3 R ori
[0070] where λ 1Control the weight of the success reward (usually large); λ 2 , λ 3 Control the weights of the position tracking reward and the orientation tracking reward respectively. In the combined reward function, the success reward provides the ultimate goal, while the position and orientation rewards guide exploration, taking into account both stability and convergence speed.
[0071] Step 3: Achieve efficient interaction between the reinforcement learning algorithm and the physical simulation platform.
[0072] Write a reinforcement learning configuration file based on the manager architecture. This file integrates multiple functional modules, including a command manager, an action manager, an observation manager, etc., to simulate different task conditions and improve the flexibility and adaptability of training. In this environment, use the training data collected in the simulation environment in Step 1 and the reinforcement learning policy network established in Step 2 for simulation training to learn the optimal policy. Among them, the command manager is used to define high-level task instructions, initialize the task scenario in each training episode and place the target object in the dexterous hand; the action manager defines the basic operations of the dexterous hand in the experimental scenario to ensure that it can execute reasonable control strategies; the observation manager is responsible for collecting the perception data of the dexterous hand, including joint state information, pose information of the target object, etc., to provide accurate environmental feedback for reinforcement learning decision-making; the event manager is used to define and manage events in different modes to ensure reasonable switching of states during the training process; the termination manager sets the task termination conditions, such as the object falling, the operation failing or the task goal being achieved, to determine whether the current training episode should end; the reward manager designs the reward function according to the task requirements and configures different reward items and weight coefficients to encourage the agent to perform operations beneficial to the task goal, while punishing invalid or harmful behaviors, thereby improving the stability and efficiency of the strategy. Maximize the reward during training to guide the output of the optimal policy.
[0073] In addition, by setting different task scenarios and writing corresponding task environment configuration files, perform reinforcement learning training under various environmental conditions, and use pure state input to drive decision-making, so that the reinforcement learning policy can adapt to different object types, initial poses and environmental constraints, realize the improvement of multi-scenario generalization ability, and improve the autonomous adaptability and generalization control ability of the underactuated dexterous hand.
[0074] Step 4: Use the 6D pose estimation algorithm FoundationPose to perform in-palm operation transfer based on pose estimation.
[0075] Use Matlab to calibrate the RGBD camera. Fix the position of the depth camera and move the calibration board. Use the Matlab camera calibration tool to process the captured calibration board images to obtain the camera internal parameter matrix. Construct the OBJ model file of the in-hand object as a reference to support subsequent pose estimation. If the CAD model of the object cannot be obtained, the large-scale synthetic training ability of FoundationPose can be used to achieve pose estimation with a small number of reference images.
[0076] In the experimental environment, use the depth camera to capture RGB images and depth images. Select the first frame of the RGB image and use the Segment Anything Model (SAM) for mask annotation to optimize the feature extraction and pose estimation of the target object. Use the MaskR-CNN detector to obtain the 2D bounding box of the object, and use the median of the 3D points inside the box for translation initialization to determine the initial position of the object. At the same time, centered on the target object, uniformly sample multiple rotation angles to generate the rotation matrix for global pose initialization as the initial pose. Input the image data captured by the RGBD camera, and the FoundationPose outputs the pose update information of the in-hand object.
[0077] Use the reinforcement learning policy network trained in step 3, with the current state information as the input, including the 6D pose information of the in-hand object output by FoundationPose and the dexterous hand joint information directly captured from the real machine. The reinforcement learning policy network trained in step 3 calculates and outputs the optimal action, and then converts it into a low-level control signal and sends it to the dexterous hand actuator to achieve fine movements such as fingertip operation, grasping, rotation, and pushing to complete the in-hand operation task.
[0078] During the execution process, the state of the dexterous hand and the pose of the in-hand object change in real time. The policy network updates the state observation information at each time step and dynamically adjusts the control strategy to ensure that the dexterous hand adapts to the changing operation environment.
[0079] The present invention adopts a reinforcement learning policy network to enable the underactuated dexterous hand to get rid of the limitations of traditional hard-coded control, be able to autonomously learn and adapt to different in-hand operation tasks, and improve the intelligent level of decision-making; use the simulation environment to train the reinforcement learning policy network and then migrate and deploy it to the real machine, reducing the cost and risk of real machine training and improving the learning efficiency. It can be used in fine operation scenarios such as industrial assembly, medical rehabilitation, service robots, etc., enhancing the application value of robots in complex unstructured environments and providing technical support for human-computer interaction and intelligent manufacturing.
[0080] The present invention has been described in detail in conjunction with specific embodiments, but these descriptions should not be construed as limiting the present invention. Those skilled in the art understand that, without departing from the spirit and scope of the present invention, various equivalent substitutions, modifications or improvements can be made to the technical solutions and their embodiments of the present invention, and all of these fall within the scope of the present invention. The protection scope of the present invention shall be subject to the appended claims.
Claims
1. A dexterous palm operation control method based on reinforcement learning and real-time posture feedback, characterized in that: The following steps are involved: (1) Build an experimental scenario for underactuated dexterous hand manipulation in the palm, deploy the underactuated dexterous hand model, palm object model, and interactive environment in a simulation environment, and collect training data; (2) For the palm manipulation task of underactuated dexterous hands, a palm manipulation reinforcement learning strategy network based on the PPO or DDPG algorithm is constructed, and a refined state space, action space, and reward function are designed according to the task requirements and the structural characteristics of the underactuated dexterous hands; (3) using the training data collected in the simulation environment of step (1) and the reinforcement learning strategy network established in step (2) to perform simulation training and learn the optimal strategy; writing a manager-based reinforcement learning configuration file and a task environment configuration file to achieve efficient interaction between the reinforcement learning algorithm and the physical simulation platform; (4) Using the reinforcement learning strategy network trained in step (3), the palm object pose information output by FoundationPose and the dexterous hand joint information directly obtained from the dexterous hand machine are used as state inputs to calculate and output the optimal control strategy to perform motion control of the dexterous hand.
2. The dexterous palm operation control method based on reinforcement learning and real-time posture feedback according to claim 1 is characterized in that: The training data in step (1) includes state information of the under-actuated dexterous hand and state information of the object in the palm; the state information of the under-actuated dexterous hand includes angles of each joint, velocities of each joint, accelerations of each joint, and positions of each joint; the state information of the object in the palm includes position information and posture information.
3. The dexterous palm operation control method based on reinforcement learning and real-time posture feedback according to claim 1 is characterized in that: Step (1) The specific steps are as follows: (11) Using the NVIDIA Isaac Sim physical simulation platform, we built an experimental scenario for underactuated dexterous hand manipulation, and deployed an underactuated dexterous hand model and interactive objects in the scenario to provide a high-fidelity simulation environment. (12) According to the joint driving characteristics of the underactuated dexterous hand, its URDF model file is modified, and the constraint conditions and dynamic parameters of each joint are adjusted to ensure that the dexterous hand in the simulation environment can truly reflect its motion characteristics in actual operation; (13) Based on the modified URDF model file, build the USD model of the dexterous hand to ensure that the dexterous hand can be correctly loaded in the USD scene and can interact normally with other objects; (14) Import the constructed USD model into the NVIDIA Isaac Sim physical simulation environment to verify the interactivity of the dexterous hand in the simulation environment, ensure the physical consistency and stability of the simulation results, and provide a reliable simulation environment for subsequent reinforcement learning training; And collect training data in a simulation environment.
4. The dexterous palm operation control method based on reinforcement learning and real-time posture feedback according to claim 1 is characterized in that: The reinforcement learning strategy network in step (2) includes a state space, an action space and a reward function; the state space includes the joint angles, joint velocities, joint accelerations, joint position information of the dexterous hand and the 6D posture information of the object in the palm; the action space includes all control strategies for the joints of the dexterous hand, and adopts a continuous action control method to ensure that the dexterous hand can perform operations smoothly; the reward function is based on a reward mechanism for task completion, and configures reward items to help the intelligent agent evaluate the pros and cons of the current state and action, and continuously optimize the strategy accordingly.
5. The dexterous palm operation control method based on reinforcement learning and real-time posture feedback according to claim 4 is characterized in that: The reward function is: R=λ1R success +λ2R pos +λ3R ori Among them, λ1 controls the weight of the success reward, λ2 and λ3 control the weights of the position tracking reward and direction tracking reward respectively. The success reward in the combined reward function provides the final goal, while the position and direction rewards guide exploration; Success Reward R success for: dθ=||quaterror magnitude(s o ,s g )|| Among them, s g is the target state of the object in the palm, s o is the current state of the object in the palm. The state information is represented by a homogeneous matrix of quaternion plus position coordinates. quat_error_magnitude calculates the error between two states. τ is the error threshold between the current direction and the target direction. Location Tracking Reward R pos for: R pos =-||p o -p g || 2 Among them, p g is the target position of the object in the palm, p o is the current position of the object in the palm. The L2 norm is used to calculate the Euclidean length between the current position of the object and the target position. Negating it facilitates maximizing the reward in subsequent reinforcement learning training. Direction Tracking Reward R ori for: β=||q o -q g || 2 Among them, q g is the target direction of the object in the palm, q o is the current direction of the object in the palm, ∈ is a small positive number to prevent division by zero errors, and the L2 norm is used to calculate the error between the current position of the object and the target position. The reciprocal is taken to maximize the reward in subsequent reinforcement learning training.
6. The dexterous palm operation control method based on reinforcement learning and real-time posture feedback according to claim 1 is characterized in that: The manager-based reinforcement learning configuration file in step (3) includes a command manager, an action manager, an observation manager, an event manager, a termination manager, and a reward manager; the command manager defines high-level task commands to place the target object into the dexterous hand in each training round; the action manager defines the basic operations that the dexterous hand can perform in the experimental scenario; the observation manager defines the observation items of the dexterous hand during the task, including joint state information and target object posture information, to provide perception data for reinforcement learning decision-making; the event manager defines and manages events in different modes to help manage state switching during training; The termination manager defines the task termination conditions, which are used to determine whether the current training round should be terminated, including object falling, operation failure, or reaching the target state; The reward manager designs a reward function, configures reward items and sets weight coefficients according to task requirements, encourages the agent to achieve task goals and punishes movements that have adverse effects on the task, thereby improving the stability and efficiency of the strategy.
7. The dexterous palm operation control method based on reinforcement learning and real-time posture feedback according to claim 1 is characterized in that: The non-task environment configuration file in step (3) adopts pure state input to drive decision making, and performs reinforcement learning training in multiple task scenarios, so that the reinforcement learning strategy can adapt to different object types, initial postures and environmental constraints, achieve the improvement of multi-scenario generalization ability, and improve the autonomous adaptability and generalization control ability of the under-actuated dexterous hand.
8. The dexterous palm operation control method based on reinforcement learning and real-time posture feedback according to claim 1 is characterized in that: The specific steps of step (4) are as follows: (41) Use Matlab to calibrate the RGBD camera: fix the depth camera position, move the camera calibration plate, and use the Matlab camera calibration tool to calibrate the collected calibration plate image to obtain the camera intrinsic parameter matrix; (42) Construct an OBJ model file of the estimated object as a reference to perform pose estimation; If the CAD model file of the object is not available, the generalization capability of FoundationPose large-scale synthetic training is used to complete the pose estimation of the object; (43) Use a depth camera to obtain RGB images and depth images in the experimental environment, select the first frame of RGB image, and use SAM for mask annotation to optimize the feature extraction and pose estimation of the target object; (44) The Mask R-CNN detector is used to obtain the 2D bounding box of the object, and the translation initialization is performed through the median of the 3D points in the bounding box to determine the initial position of the object. Taking the target object as the center, multiple rotation angles are uniformly sampled to generate a rotation matrix for global pose initialization as the initial pose. The image data collected by the RGBD camera is input, and FoundationPose outputs the pose update information of the object in the palm. (45) The trained reinforcement learning strategy network is used to calculate and output the optimal control strategy using the current state information as input to achieve precise motion control of the dexterous hand; (46) 6D pose information of the object in the palm output by the FoundationPose algorithm, that is, the position and orientation information of the object in the palm relative to the camera; real-time joint information of the dexterous hand collected directly from the dexterous hand machine, including joint angles, joint velocities, and end positions; (47) A reinforcement learning strategy network is used to calculate the optimal action strategy based on the current state information, and the optimal control strategy output by the reinforcement learning strategy network is converted into a low-level control signal and sent to the dexterous hand actuator; the dexterous hand is controlled to perform fine movements to complete the expected palm operation task; (48) During the execution process, the real-time state of the dexterous hand and the position of the object in the palm change in real time. The strategy network updates the state observation information at each time step, dynamically adjusts and continuously outputs the control strategy.
Citation Information
Patent Citations
Mechanical arm intelligent control rapid training method based on deep reinforcement learning
CN112338921A
Five-finger dexterous robot arm control method based on multi-agent deep reinforcement learning
CN116330290A
Dexterous hand and mechanical arm reinforcement learning cooperative control method based on heuristic trajectory
CN117733850A
Dynamic interactive representation-based dexterous manipulator grabbing method
CN117798919A
Virtual-real fusion-based edge end simulation method and system
CN118092215A
Cited By
Method and system for constructing intelligent simulation data set of electric power scene
CN120822432A
Multi-fingered dexterous hand operation reinforcement learning method based on human action prediction model
CN121492070A
Reinforcement learning method for multi-fingered hand operation based on human motion prediction model
CN121492070B
Indoor 3D space synthetic data platform for intelligent training with body
CN121725138A
Indoor 3D space synthetic data platform for embodied intelligence training
CN121725138B