Dexterous mechanical arm control method based on uncertainty perception fused with long-term and short-term reward strategy gradient
By introducing a long-short reward strategy gradient method with uncertainty awareness into the control of a dexterous robotic arm, the problems of slow convergence speed and poor control smoothness of traditional algorithms in dynamic environments are solved, achieving high-precision and stable robotic arm control that meets the real-time and lifespan requirements of industrial applications.
Patent Information
- Application Number
- CN202511451318.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2025-12-09
AI Technical Summary
Existing technologies for controlling dexterous robotic arms suffer from high degrees of freedom, amplifying the structural defects of traditional algorithms. This results in slow convergence speed, poor control smoothness, and difficulty in balancing real-time requirements with equipment lifespan assurance. In particular, path tracking errors are high and hardware wear is severe in dynamic environments.
We adopt a gradient method for long-short reward strategies based on uncertainty perception. By collecting the state information of the robotic arm in real time, we use a forward model network to predict the mean and uncertainty, dynamically generate reward fusion weights, and use a reliability judgment module to block short-term rewards when the error exceeds the threshold, so as to ensure that the strategy network depends on long-term reward updates and achieve smoothness and stability of action commands.
It improves the control accuracy and convergence speed of the robotic arm in complex dynamic environments, enhances motion smoothness, reduces hardware losses, and meets the industrial application requirements of millimeter-level positioning accuracy and millisecond-level response latency.
Smart Images

Figure CN121083646A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of reinforcement learning and dexterous manipulator control, and particularly relates to a dexterous manipulator control method based on uncertainty perception and fusion of long-term and short-term reward policy gradient. BACKGROUND
[0002] Dexterous manipulators have great potential in industrial operations, medical surgeries and human-computer interaction due to their high flexibility and environmental adaptability. Model-based control algorithms can effectively control the manipulator to track the desired trajectory, and the control accuracy is also good.
[0003] Although the DDPG, TD3 and other methods based on policy gradient can handle continuous action space, they require millions of training steps to converge due to high sample complexity, and the low exploration efficiency leads to high path tracking error in dynamic environment. More importantly, the optimization mechanism that relies solely on long-term cumulative reward lacks real-time perception of environmental dynamics, inducing high-frequency jitter of control commands. Although the existing improvement scheme such as model predictive reinforcement learning (MBPO) introduces an environmental model, it causes policy divergence in complex tasks due to the prediction deviation of dynamics, and the reward function designed by hand has weak generalization ability, which cannot meet the real-time requirements and device life protection.
[0004] The root cause is that the high degree of freedom of dexterous manipulators amplifies the structural defects of traditional algorithms: on the one hand, fixed exploration strategies cannot adapt to state space mutations in nonlinear deformation processes, resulting in insufficient sample utilization; on the other hand, the non-smoothness of control commands easily induces micro-crack propagation in material stress concentration areas. Under the dual constraints of millimeter-level positioning accuracy and millisecond-level response delay required by industrial applications, the existing technology has a "impossible triangle" among convergence speed, control smoothness and hardware loss.
[0005] In view of the above problems in the prior art, it is urgent to propose a dexterous manipulator control method based on uncertainty perception and fusion of long-term and short-term reward policy gradient. SUMMARY
[0006] To solve the above technical problems, the application proposes a dexterous manipulator control method based on uncertainty perception and fusion of long-term and short-term reward policy gradient to solve the problems existing in the prior art.
[0007] To achieve the above purpose, the application provides a dexterous manipulator control method based on uncertainty perception and fusion of long-term and short-term reward policy gradient, comprising the following steps:
[0008] Real-time acquisition of manipulator state information, pre-processing to obtain a unified state vector;
[0009] According to the state vector, a predicted mean and a predicted uncertainty of a next state are obtained through a forward model network;
[0010] According to the predicted mean and the predicted uncertainty, a long-term and short-term reward fusion weight is dynamically generated through an attention network;
[0011] According to the fusion weight, a long-term reward policy gradient and a short-term reward policy gradient are fused, and a policy network parameter is updated;
[0012] In the process of updating the policy network, according to the forward model prediction error and the real state error, a reliability judgment module is used to shield the short-term reward when the error exceeds a threshold, so that the policy network continues to update only by relying on the long-term reward;
[0013] According to the action instruction output by the updated policy network, a robot arm is driven to perform a movement.
[0014] Optionally, the process of collecting the robot arm state information in real time and obtaining a unified state vector after preprocessing includes:
[0015] The end position, attitude, joint angular velocity and contact force of the robot arm collected by the sensor are filtered, normalized and fused to form a unified state vector.
[0016] Optionally, the process of obtaining a predicted mean and a predicted uncertainty of a next state according to the state vector through a forward model network includes:
[0017] The current state vector and the action vector are input into the forward model network, a random inactivation layer is set after each hidden layer, a number of forward propagations are performed, and a predicted mean and a predicted variance of a next state are output, wherein the predicted variance is used as the predicted uncertainty.
[0018] Optionally, the process of dynamically generating a long-term and short-term reward fusion weight according to the predicted mean and the predicted uncertainty through an attention network includes:
[0019] The state vector and the predicted variance are spliced and input into the attention network, and a fusion weight between zero and one is output through a Sigmoid activation function, which is used for nonlinearly adjusting the proportion of long-term reward and short-term reward in the policy gradient.
[0020] Optionally, the process of fusing a long-term reward policy gradient and a short-term reward policy gradient according to the fusion weight, and updating a policy network parameter includes:
[0021] A long-term reward policy gradient is calculated through a value network, a short-term reward policy gradient is calculated through a reward prediction network, a weighted sum of the two is obtained according to the fusion weight, a mixed policy gradient is obtained, and a policy network parameter is updated.
[0022] Optionally, in the process of policy network updating, according to the forward model prediction error and the real state error, the reliability judgment module is used to shield the short-term reward when the error exceeds the threshold, so that the policy network continues to update only by relying on the long-term reward, and the process comprises the following steps:
[0023] The error between the predicted state output by the forward model and the real next state is compared, and when the error is greater than a set threshold, the fusion weight output by the attention network is forced to be zero, so that the policy network update only relies on the long-term reward policy gradient.
[0024] Optionally, the parameter updating process of the attention network comprises the following steps:
[0025] The attention network parameters are updated by the gradient descent method to maximize the evaluation value of the value network to the action output by the policy network, so that the fusion weight guides the policy network to output high-value actions.
[0026] The application also provides a computer device comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to realize the steps of the method.
[0027] The application also provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to realize the steps of the method.
[0028] The application also provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to realize the steps of the method.
[0029] Compared with the prior art, the application has the following advantages and technical effects:
[0030] The application introduces a prediction uncertainty quantification means in the forward model network, sets an attention network to dynamically adjust the long-term and short-term reward fusion weight according to the uncertainty, and further uses a reliability judgment module to shield the short-term reward when the model prediction error exceeds the threshold, so that the policy network updating process has both short-term environmental sensitivity and long-term stability, thereby improving the control precision, convergence speed and motion smoothness of the robot arm in a complex dynamic environment. BRIEF DESCRIPTION OF DRAWINGS
[0031] The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The embodiments of this application and their explanations are used to explain this application, and do not constitute an improper limitation on this application. In the drawings:
[0032] Figure 1 The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The embodiments of this application and their explanations are used to explain this application, and do not constitute an improper limitation on this application. In the drawings: DETAILED DESCRIPTION
[0033] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0034] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0035] Example 1
[0036] like Figure 1 As shown, this embodiment provides a dexterous robotic arm control method based on uncertainty perception and fusion of long- and short-term reward policy gradients, including the following steps:
[0037] Real-time acquisition of robotic arm status information, followed by preprocessing to obtain a unified state vector;
[0038] Based on the state vector, the predicted mean and prediction uncertainty of the next state are obtained through the feedforward model network;
[0039] Based on the predicted mean and prediction uncertainty, a long-term and short-term reward fusion weight is dynamically generated through an attention network.
[0040] Based on the fusion weights, the long-term reward policy gradient and the short-term reward policy gradient are fused to update the policy network parameters;
[0041] During the policy network update process, based on the prediction error of the forward model and the actual state error, the reliability judgment module blocks short-term rewards when the error exceeds the threshold, so that the policy network continues to update only by relying on long-term rewards.
[0042] The robotic arm is driven to perform movements based on the action instructions output by the updated policy network.
[0043] As an optional implementation, this embodiment provides a dexterous robotic arm control system, including:
[0044] Robotic arm body: A dexterous robotic arm structure built based on the theory of multibody system dynamics, used to simulate the bending and torsional behavior of a dexterous robotic arm;
[0045] Sensor module: Real-time acquisition of end effector pose, joint angles, joint speeds, and contact force information of the robotic arm, including inertial measurement unit, vision sensor, force / torque sensor and high-precision joint encoder;
[0046] State preprocessing module: integrated on the FPGA chip, used for filtering, coordinate conversion, normalization processing and state vector fusion calculation of the multi-source, heterogeneous raw data collected by the sensor module;
[0047] Control module: deploying trained policy network, value network, forward model network, reward prediction network and attention network, used for receiving the current state vector of the preprocessing module and outputting control instructions for each joint driving unit; the module also contains a dynamic weight adjuster, which is realized by the attention network;
[0048] Uncertainty perception module: a Dropout layer is set after each hidden layer in the forward model network, used for T times of random forward propagation during inference, outputting the predicted mean value and the predicted variance of the next state, and the variance is the quantitative indicator of model uncertainty.
[0049] Reliability judgment module: real-time comparison of the prediction state of the forward model network and the actual next state error. When the error is greater than the preset threshold, it is determined that the current environment dynamic is intense or the model prediction is unreliable, and an instruction is sent to the dynamic weight adjuster to force the weight factor λ to zero at the current time step, ensuring that when the model prediction is unreliable, the system automatically degrades to rely only on the robust long-term reward policy, avoiding control instability or mechanical arm shaking caused by model distortion;
[0050] Attention network: a parameterized neural network taking the current state and the predicted variance as input, and outputting a dynamic weight ; the dynamic weight is used to nonlinearly fuse the long-term and short-term reward policy gradients; its training goal is to maximize the long-term cumulative value, so that the system can intelligently adjust the degree of dependence on short-term rewards according to the state and uncertainty;
[0051] Actuator module: including servo driver and motor, used for receiving control instructions and driving the joints of the dexterous manipulator to move.
[0052] As an optional implementation, the embodiment provides a control algorithm for the above system, which comprises:
[0053] S1, system initialization; constructing a manipulator dynamics model, initializing the policy network (Actor network) and the value network (Critic network) parameters , , target Actor network and target Critic network Parameters: , , randomly initialize forward model network , reward network , attention network parameters , , , empty experience replay set ;
[0054] It should be noted that each network specifically refers to:
[0055] Actor network : policy network, input is the current state , output deterministic action , influenced by long-term cumulative reward and short-term next state reward to achieve policy optimization, adjust policy parameters ;
[0056] Critic network : using time difference algorithm, updating from the current action value to the target value in iterations, here referred to as long-term reward, action value function is the expected cumulative discounted return following the policy :
[0057] ;
[0058] ;
[0059] Where is the discount factor, is the target Critic network Q value;
[0060] Forward model network : constructed by a feedforward neural network with a parameter , input is state and action , add a Dropout layer after each hidden layer, perform T times forward propagation on the same input , get T predicted states, finally output predicted mean and variance ;
[0061] Reward prediction function : constructed by a feedforward neural network with a parameter feedforward neural network construction, learning network parameters with Adam optimizer, input is forward network predicted state, reward is output;
[0062] The output of the forward model network is used as the input of the reward prediction network, which predicts the reward of the next state under the given current state and current action , referred to as short-term reward.
[0063] S2, state collection and action execution; at each time step t, the current state vector is obtained through the sensor , based on the current Actor network , select action , execute action , get reward from the environment and the next state , while recording the prediction variance of the forward model network for this state transition , and the state transition process and the forward model prediction variance are stored in the experience replay set D
[0064] S3, data sampling and network forward; randomly select N state transitions from D respectively into the Actor target network, Critic network, Critic target network, forward network, reward prediction network and attention network;
[0065] S4, network parameter update;
[0066] Specifically, in this embodiment, the network update functions are as follows:
[0067] The Critic network parameters are updated using the following formula:
[0068] ;
[0069] where is the target Critic network Q value, N is the number of samples collected from the experience replay set D, and the optimal parameter ω is obtained by minimizing ;
[0070] The forward model network parameters are updated according to the following formula:
[0071] ;
[0072] where N is the number of samples collected from the experience replay set D, and the optimal parameter is obtained by minimizing ;
[0073] The reward prediction network parameters are updated according to the following formula:
[0074] ;
[0075] Where N is the number of samples collected from the experience replay set D, and the optimal parameters are obtained by minimizing ;
[0076] The Actor network parameters are updated based on the attention mechanism:
[0077] For each sample in the batch, first forward propagation is performed:
[0078] ;
[0079] Then the mixed gradient is calculated and the Actor is updated:
[0080] ;
[0081] Where is the Critic network reward (long-term reward), is the reward prediction network output (short-term reward), and the optimal parameters are obtained by maximizing ;
[0082] The attention network parameters are updated ;
[0083] The value loss is calculated:
[0084] ;
[0085] The attention network is updated by gradient descent:
[0086] ;
[0087] S5, soft update the target network parameters of the Actor and Critic:
[0088] ;
[0089] ;
[0090] Where is the soft update parameter;
[0091] Repeat the above S2-S4, deploy the trained Actor network and attention network to the control system to realize high-precision dynamic control of the robot arm.
[0092] In the implementation process, the mechanical arm state includes end position, attitude, joint angular velocity, contact force, etc., which is collected in real time by a sensor and normalized and then input into the Actor network, and the output is the control torque or angle increment of each joint; the attention network structure is set as a feedforward network with one hidden layer, using ReLU activation function, and the output layer uses Sigmoid function.
[0093] The training stage is carried out in a simulation environment (Elastica), and the Cosserat beam model is used to simulate the dynamic behavior of the mechanical arm, and each neural network is trained through a large amount of interaction data. After training, the Actor network is deployed to the actual mechanical arm control system to realize real-time closed-loop control.
[0094] Simulation experiment:
[0095] The first simulation experiment in this embodiment is that the end of the mechanical arm needs to touch a static target in three-dimensional space; the second experiment is that the tip of the mechanical arm needs to track a target moving along a fixed trajectory in three-dimensional space; the third experiment is that the tip of the mechanical arm needs to track a target with a random motion trajectory in three-dimensional space.
[0096] Here, the mechanical arm is modeled as a cosserat beam with a length of 1m, a radius of 5cm, and a Young's modulus of 10Mpa. The bottom of the mechanical arm is fixed, vertically standing in three-dimensional space, and freely moving under driving. The target object is a ball with a radius of 5cm, which is the same as the radius of the mechanical arm.
[0097] The three parts of the neural network constitute the core of the controller, among which the Actor is composed of a fully connected neural network with a size of 44*400*300*12, and the structure of the Critic is 56*400*300*1. The structure of the forward model network is 44*400*300*12, and a Dropout layer is added after each fully connected layer, while the reward network only has one hidden layer with 400 nodes, and the structure is 44*400*1. The observed state is normalized to a 44-dimensional input to the Actor, and the 12-dimensional output action is used to control the movement of the soft arm. The network parameters are trained using the Adam optimizer, and the learning rate of the Actor network is , the learning rates of the Critic network, the environment model and the reward network are the same, and are set to . Other hyperparameter tuning is as follows: discount factor , batch size , experience replay set capacity is 2000000, and the target network soft update parameter is .
[0098] It is worth noting that this paper is not to explore the optimal weight factor value, but to explore the trend of the algorithm performance when the short-term reward technique is introduced. Therefore, only four representative weights are selected here to control the weight factor = 0.2, 0.5, 0.8, 1, and the last group is the control based on the DDPG algorithm.
[0099] In the first experiment, all methods were trained for 5 independent runs of 1 million steps. All algorithms can find the control strategy to touch the static target, except for the algorithm = 1, because its policy gradient only depends on the short-term reward. = 0.2 reward is similar to DDPG, but the convergence speed is slower. = 0.5 and DDPG have no significant difference in learning curve, but the performance is slightly better than DDPG. = 0.8, the reward return is 13000, while DDPG is trained for about 80 million steps to reach a reward value of 12000.
[0100] From the experimental results, it can be seen that the greater the weight of the short-term reward, the faster the algorithm converges. But when the weight is too low, such as , the short-term reward part has no obvious effect on the dynamic perception of the environment. In addition, when the weight factor is 1, the analysis of the weight factor changes from quantitative change to qualitative change, so that the policy learning only depends on the short-term reward, lacks the guidance of the long-term reward, and the algorithm cannot converge.
[0101] In the second experiment, the movement of the target ball is more complex than in the first experiment. The end of the robot arm needs to track the moving ball in three-dimensional space. The target ball moves back and forth at a constant speed of 0.2 meters per second. The reward function and action range are the same as in the first experiment, the neural network structure is the same as in the first experiment, except that the state dimension changes to 37, and the hyperparameter tuning method is different. The parameters of the Critic network, the forward model network, and the reward network are trained with the Adam optimizer at the same learning rate of 1e-4. Here the learning rate of the Actor network is set to 2e-4. In addition, the discount factor is , the batch size , and the target network soft update parameter . The size of the experience replay set D is set to 1000000.
[0102] Each algorithm is trained for 5 independent runs of 1 million steps. From the experimental results, it can be seen that all algorithms have similar final reward values, except for = 1. =0.8 has obvious advantages, it obtains about 13000 reward value at around 300,000 steps, while DDPG converges to similar reward value at around 600,000 steps. That is, =0.2 converges to similar reward value with less training time. =0.2 and =0.5 algorithm performance is similar to DDPG. In addition, the controller based on =1 algorithm performs the worst in tracking dynamic ball task, and cannot successfully learn the strategy to track the moving target.
[0103] Both touch and tracking tasks show that the greater the weight of short-term reward, the shorter the convergence time required by the algorithm. This shows that model-based reinforcement learning (short-term reward technique) is more effective than performance-based reinforcement learning (traditional reinforcement learning algorithm is guided by a single performance indicator). Therefore, the algorithm guided by long-term reward and short-term reward has a significant effect.
[0104] In addition, in the touch and tracking tasks of the robotic arm, for the random movement of the target ball, adjusting the weight factor such as =0.2, 0.5 and =0.8 has the same training effect, so =0.8 as a representative is compared with DDPG, and the algorithm has similar convergence speed to the traditional DDPG algorithm.
[0105] The target ball is random and irregular in motion, and the environment is more complex and dynamic. The application of short-term reward technique depends on the modeling of the environment, and the forward model network cannot always learn an accurate environmental dynamic model in a complex dynamic environment. Research shows that forward model-based RL is mainly successful in these restrictive fields (only need to learn a simple environmental dynamic model), and the forward model network tends to use the previously visited area during training, which leads to model bias when facing complex dynamic environments. Due to model bias and numerical instability, RL cannot obtain a good strategy for more dynamic and complex tasks.
[0106] In summary, the strategy gradient algorithm (PGLS) affected by long-term reward and short-term reward at the same time has good application prospects in continuous control problems, especially in the control of high-dimensional space dexterous robots.
[0107] Embodiment Two
[0108] The embodiment also discloses a computer device comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to realize the steps of the method in Embodiment One.
[0109] Embodiment Three
[0110] The embodiment also discloses a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the steps of the method in the embodiment one.
[0111] Embodiment four
[0112] The embodiment also discloses a computer program product, which comprises a computer program, and the computer program is executed by a processor to realize the steps of the method in the embodiment one.
[0113] The above is only the preferred specific implementation of the present application, but the protection scope of the present application is not limited to this, any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A dexterous robotic arm control method based on uncertainty perception and fusion of long- and short-term reward policy gradients, characterized in that, Includes the following steps: Real-time acquisition of robotic arm status information, followed by preprocessing to obtain a unified state vector; Based on the state vector, the predicted mean and prediction uncertainty of the next state are obtained through the feedforward model network; Based on the predicted mean and prediction uncertainty, a long-term and short-term reward fusion weight is dynamically generated through an attention network. Based on the fusion weights, the long-term reward policy gradient and the short-term reward policy gradient are fused to update the policy network parameters; During the policy network update process, based on the prediction error of the forward model and the actual state error, the reliability judgment module blocks short-term rewards when the error exceeds the threshold, so that the policy network continues to update only by relying on long-term rewards. The robotic arm is driven to perform movements based on the action instructions output by the updated policy network.
2. The method according to claim 1, characterized in that, The process of acquiring real-time robotic arm status information and obtaining a unified state vector through preprocessing includes: The end-effector position, posture, joint angular velocity, and contact force collected by the sensors are filtered, normalized, and fused to form a unified state vector.
3. The method according to claim 1, characterized in that, Based on the state vector, the process of obtaining the predicted mean and prediction uncertainty of the next state through the feedforward model network includes: The current state vector and action vector are input into the feedforward model network. A random deactivation layer is set after each hidden layer. Several forward propagations are performed, and the predicted mean and prediction variance of the next state are output. The prediction variance is used as the prediction uncertainty.
4. The method according to claim 3, characterized in that, Based on the predicted mean and prediction uncertainty, the process of dynamically generating long- and short-term reward fusion weights through an attention network includes: The state vector and prediction variance are concatenated and then fed into the attention network. The Sigmoid activation function outputs a fusion weight between zero and one, which is used to nonlinearly adjust the proportion of long-term and short-term rewards in the policy gradient.
5. The method according to claim 1, characterized in that, The process of fusing the long-term reward policy gradient with the short-term reward policy gradient and updating the policy network parameters according to the fusion weights includes: The long-term reward policy gradient is calculated using a value network, and the short-term reward policy gradient is calculated using a reward prediction network. The two are then weighted and summed according to their fusion weights to obtain the hybrid policy gradient and update the policy network parameters.
6. The method according to claim 1, characterized in that, During the policy network update process, based on the prediction error of the forward model and the error of the actual state, the reliability judgment module blocks short-term rewards when the error exceeds a threshold, so that the policy network continues to be updated only by relying on long-term rewards. This process includes: The error between the predicted state output by the forward model and the actual next state is compared. When the error exceeds a set threshold, the fusion weights output by the attention network are forced to zero, so that the policy network update depends only on the gradient of the long-term reward policy.
7. The method according to claim 1, characterized in that, The parameter update process of the attention network includes: With the goal of maximizing the evaluation value of the policy network's output actions by the value network, the parameters of the attention network are updated using gradient descent, so that the fused weights guide the policy network to output high-value actions.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1-7.
Citation Information
Cited By
Robot motion control method and system based on cerebellum reinforcement learning
CN121756368A