Arm-hand robot grabbing method based on deep reinforcement learning

By improving the strategy network and experience pool design of the deep reinforcement learning algorithm, the problems of slow convergence speed and poor stability in the arm-mobile robot grabbing task are solved, and more efficient task success rate and environmental adaptability are achieved.

CN120533737APending Publication Date: 2025-08-26CHONGQING UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510950337.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-06-05
Filing Date
2025-07-10
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In high-degree-of-freedom grab tasks, the mobile phone robot faces problems such as low sample efficiency, slow convergence speed, unstable algorithms and limitations in network structure, resulting in poor task learning effect.

Method used

Using a method based on deep reinforcement learning, the policy network is improved through the sparse causal time self-attention mechanism and LSTM series structure, combined with post-hoc experience replay and value network update of adaptive conservative Q learning, the policy network and experience pool design are optimized, and the convergence speed and stability of the algorithm are improved.

Benefits of technology

It effectively improves the success rate and stability of the arm-mobile robot crawling task, improves the convergence speed of the algorithm and the stability of the policy network, and enhances the adaptability and generalization ability to complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120533737A_ABST
    Figure CN120533737A_ABST
Patent Text Reader

Abstract

The invention discloses an arm-hand robot grabbing method based on deep reinforcement learning. The arm-hand robot grabbing method comprises the following core steps: 1) providing a strategy network structure based on a sparse causal time self-attention mechanism; (2) a post experience recombination method is provided, so that the utilization efficiency of successful samples by the algorithm is improved; 3) designing a self-adaptive conservative Q learning value network updating method, and dynamically adjusting the intensity of a regularization item through a self-adaptive adjustment mechanism based on a time sequence difference error and an average reward; the method can effectively improve the convergence speed of the algorithm, balance exploration and stability requirements in the training process, and improve the stability of the algorithm. In addition, according to the strategy network structure of the method, through local window sparse connection and an LSTM series structure, the modeling capacity of single-step, local and overall action characteristics is effectively enhanced, and finally the success rate of the grabbing task can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is applied to an arm-hand robot system with a high degree of freedom, and specifically relates to an arm-hand robot grasping method based on deep reinforcement learning. Background Art

[0002] Traditional robotic grasping methods rely on precise modeling and optimization. Although they are suitable for structured environments, they have poor adaptability in complex dynamic scenes. Deep learning-based grasping methods have improved the intelligence level of visual perception and grasping posture prediction, but they are highly dependent on large-scale labeled data. Multimodal grasping methods integrate visual and tactile information to improve perception accuracy, but their real-time and generalization capabilities are still insufficient. Deep reinforcement learning-based grasping methods combine the advantages of perception and decision-making and have the ability to learn autonomously and optimize strategies. This method can optimize grasping strategies through unsupervised interaction and demonstrate strong adaptability and generalization capabilities in unstructured environments. Although it faces challenges such as low sample efficiency, high training costs, and difficulties in virtual-real migration, its advantages in intelligence and flexibility make it one of the most promising directions in grasping method research.

[0003] Both deep reinforcement learning and robotic grasping have received significant attention and produced numerous achievements. However, due to the high degrees of freedom of arm-hand robots, deep reinforcement learning in these applications faces challenges such as low sample efficiency, slow convergence, susceptibility to local optima, algorithmic instability, and limitations in network architecture. These issues result in slow convergence and low success rates for task learning, and the learned trajectories may be unsafe or suboptimal.

[0004] Currently, mainstream deep reinforcement learning algorithms include DDPG, TD3, PPO, and SAC, all of which are continuously optimized in terms of sample efficiency, policy stability, and generalization. DDPG is suitable for continuous action spaces but is sensitive to hyperparameters; TD3 improves stability through delayed updates and policy perturbations; and PPO has a good policy update mechanism but is limited in high-dimensional action spaces. In contrast, SAC, based on a maximum entropy policy optimization framework, improves sample efficiency while maintaining policy stability, making it particularly suitable for tasks in high-dimensional, complex action spaces, such as arm-hand robotic grasping. Summary of the Invention

[0005] Currently, due to the high degrees of freedom of arm-hand robots, reinforcement learning in such applications faces problems such as low sample efficiency, slow convergence, algorithmic instability, and network structure limitations. These issues lead to slow convergence of task learning and poor task performance. Therefore, this invention addresses these issues in the grasping scenario of arm-hand robots, providing an arm-hand robot grasping method based on deep reinforcement learning.

[0006] A grasping method for an arm-hand robot based on deep reinforcement learning, the method comprising the following steps:

[0007] 1) Define the state space S of the arm-hand robot (a 50-dimensional vector with 18 joint angles s qpos With angular velocity s qvel , the 3D vector position s of the grasping center of the five-fingered hand grip-pos and a 4-dimensional vector towards s grip-quat , the 3D vector position s of the grasped object object-pos and a 4-dimensional vector towards s object-quat ), action space A (a 12-dimensional vector, including the 6-dimensional vector robot end tool center posture a robot The incremental form of control, the 6-dimensional vector position a of the dexterous hand hand ) and the reward function R(s t ,a t ), where s t is the state at time t, a t is the action to be performed by the agent at time t (where s t ∈S,a t ∈A);

[0008] 2) Initialize the value network Q θ , policy network and related parameters (such as the regularization coefficient range λ min and λ max , reward threshold R threshold wait);

[0009] 3) In each Episode, through the main strategy network π(s t ) Generate action a t , obtain state transition data (s t ,a t ,r t ,s t+1 );

[0010] 4) Use the post-experience replay technology to assign a new target g to the experience data and modify the corresponding reward r. The experience becomes (s t ||g,a t ,rt ′,s t+1 ||g) and stored in the hierarchical experience pool (buffer_1, buffer_2, buffer_3). The reorganization module reorganizes the data in the experience pool to generate new experience samples for training;

[0011] 5) Calculate the average reward of the last 4 episodes. Enter the update phase, randomly sample data from the experience pool, calculate the target TD value and TD error, normalize the TD error, and dynamically adjust the regularization weight η according to the average reward value and TD error. θ Optimize and update the policy network by maximizing the value network value Q(s,π(s))

[0012] Compared with the prior art, the present invention has the following technical effects:

[0013] Compared with the native experience pool and post-experience replay method of the SAC algorithm, this method can effectively improve the convergence speed of the algorithm; compared with conservative Q-learning, this method is based on the adaptive adjustment mechanism of temporal difference error and average reward, dynamically adjusts the strength of the regularization term, balances exploration and stability requirements during training, and improves the stability of the algorithm; in addition, the policy network structure of this method effectively enhances the modeling ability of single-step, local and overall action features through local window sparse connection and LSTM series structure, which can ultimately effectively improve the success rate of grasping tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 This is the overall block diagram of the algorithm;

[0015] Figure 2 This is a schematic diagram of the improved strategy network structure;

[0016] Figure 3 Comparison diagram of the expanded receptive fields of full connection and sparse connection;

[0017] Figure 4 Reorganize the schematic diagram for reinforcement learning trajectory;

[0018] Figure 5 An example diagram of the agent grasping something and then becoming unstable, falling, and then grasping it again;

[0019] Figure 6 This is an example diagram of the nut assembly task;

[0020] Figure 7 This is a comparison chart of the algorithm convergence speed reward curve;

[0021] Figure 8 This is a comparison chart of the algorithm stability reward curve; DETAILED DESCRIPTION

[0022] The present invention will be described in further detail below with reference to the accompanying drawings and specific embodiments.

[0023] A grasping method for an arm-hand robot based on deep reinforcement learning, the method comprising the following steps:

[0024] 1) Define the state space S of the arm-hand robot (a 50-dimensional vector with 18 joint angles s qpos With angular velocity s qvel , the 3D vector position s of the grasping center of the five-fingered hand grip-pos and a 4-dimensional vector towards s grip-quat , the 3D vector position s of the grasped object object-pos and a 4-dimensional vector towards s object-quat ), action space A (a 12-dimensional vector, including the 6-dimensional vector robot end tool center posture a robot The incremental form of control, the 6-dimensional vector position a of the dexterous hand hand ) and the reward function R(s t ,a t ), where s t is the state at time t, a t is the action to be performed by the agent at time t (where s t ∈S,a t ∈A);

[0025] 2) Initialize the value network Q θ , policy network and related parameters (such as the regularization coefficient range λ min and λ max , reward threshold R threshold wait);

[0026] 3) In each Episode, through the main strategy network π(s t ) Generate action a t , obtain state transition data (s t ,a t ,r t ,s t+1 );

[0027] 4) Use the post-experience replay technology to assign a new target g to the experience data and modify the corresponding reward r. The experience becomes (s t ||g,a t ,r t ′,s t+1||g) and stored in the hierarchical experience pool (buffer_1, buffer_2, buffer_3). The reorganization module reorganizes the data in the experience pool to generate new experience samples for training;

[0028] 5) Calculate the average reward of the last 4 Episodes. Enter the update phase, randomly sample data from the experience pool, calculate the target TD value and TD error, normalize the TD error, and dynamically adjust the regularization weight η according to the average reward value and TD error; minimize the loss function of the regularization term to adjust Q θ Optimize and update the policy network by maximizing the value network value Q(s,π(s))

[0029] The overall block diagram of the algorithm of the present invention is as follows Figure 1 As shown, combining steps 1), 2), 3), 4), and 5) can give the complete process of the improved algorithm:

[0030]

[0031]

[0032] The following mainly describes steps 3), 4), and 5) in detail:

[0033] In step 3, the agent uses the main policy network π(s) in each episode. t ) Generate action a t , obtain state transition data (s t ,a t ,r t ,s t+1 ). To address the problem that traditional fully connected policy network structures are difficult to capture complex temporal dependencies, a policy network structure based on sparse causal temporal self-attention mechanism is proposed;

[0034] The attention computation complexity is reduced by sparse connection of local windows, and connected in series with LSTM to construct a new strategy structure. The improved new strategy network structure is as follows Figure 2 The core of the entire policy network structure lies in the Sparse Causal Temporal Attention mechanism (Sparse CT-MSA), which consists of multiple submodules and captures local and global temporal dependencies by expanding the receptive field layer by layer.

[0035] The specific steps include:

[0036] 1) Local Window Mechanism: The core of the sparse causal temporal attention mechanism, it is primarily used to reduce the complexity of attention computation. This mechanism divides the input time series into several non-overlapping windows of size W and performs attention computation within each window, limiting the attention computation to the window range. This design significantly reduces computational complexity while focusing on capturing strong dependencies between local time steps. The formula for calculating attention within a window is:

[0037]

[0038] Among them, the query Q (Query), key K (Key) and value V (Value) are generated by the linear projection of the input feature X:

[0039] Q h =XW q ,K h =XW k ,V h =XW v (2)

[0040] in, is a learnable linear projection matrix, d k is the dimension of the key vector.

[0041] 2) Causal masking mechanism: This is another key component in the design of the sparse causal temporal attention mechanism, used to ensure the causal constraints of the model. Through the mask matrix, the model can block information from future time steps, ensuring that the calculation of the current time step only depends on the current and previous time steps. The mask matrix is ​​defined as:

[0042]

[0043] In the attention calculation, the mask matrix is ​​added to the dot product result of the key and the query to shield the influence of future time steps. The updated attention matrix calculation formula is:

[0044]

[0045] This design ensures causality in time series tasks and provides a guarantee for causal modeling of policy networks and value networks in reinforcement learning.

[0046] 3) Layer Normalization and Skip Connections: The Sparse Causal Temporal Attention module incorporates layer normalization (LayerNorm) and residual connection mechanisms. Layer normalization helps stabilize the training process, while skip connections effectively alleviate the vanishing gradient problem in deep networks and preserve input information.

[0047] 4) Multilayer Perceptron Submodule: Following the Mask Calculation Module are two layers of multilayer perceptrons, which utilize nonlinear activation functions to further enhance feature representation. This design enhances the model's nonlinear modeling capabilities, making it more flexible in time series modeling.

[0048] 5) Layer-by-layer window size expansion: In the multi-layer design of the sparse causal temporal attention mechanism, the window size is expanded layer by layer. Lower layers primarily capture short-term dependencies, while higher layers expand the receptive field to capture long-term global dependencies. This design optimizes computational efficiency while ensuring the accuracy of dependency modeling.

[0049] Based on the existing causal temporal self-attention mechanism, this paper proposes a sparsification operation to further optimize the connection pattern: Unlike the traditional fully connected design, the sparsification operation reduces the computational complexity by limiting the connections within each window to only the dependencies between key time steps. Specifically, the connections within the window only retain the most critical time step dependencies, rather than fully connecting all time steps. The specific sparsification operation is as follows: Figure 3 shown.

[0050] In this sparse design, Block 1, with a window size of 2, primarily processes short-term dependencies, Block 2, with a window size of 4, handles medium-term dependencies, and Block 3, with a window size of 8, captures dependencies over longer time spans. This design achieves a time complexity of O(ρTWC), where ρ represents sparsity (typically ρ<1); T is the sequence length, W is the window size, and C is the feature dimension.

[0051] Before the policy network was improved, the input required for each update of the policy network was usually "single frame" data, that is, batch tuples (s t ,a t ,r t ,s t+1 After the policy network is improved, because the input needs to match the sequence input format of LSTM and attention mechanism, the input of the improved policy network needs "multi-frame" data, that is, the input needs to be converted into batch tuples (s t ,a t ,r t ,s t+1 )…(s t+8 ,a t+8 ,r t+8 ,s t+9 ).

[0052] The action trajectory of the entire reinforcement learning can be divided into single-step action features, local action features, and global action features. Combining the LSTM and the sparse causal temporal attention mechanism, the policy network can better extract the feature information in such trajectories. In this policy network, LSTM is responsible for extracting the temporal dependence features in single-step, local, and global actions from the entire reinforcement learning trajectory, while the sparse causal temporal attention mechanism efficiently models the dependence relationship between local action features through sparse connections within the local window.

[0053] Step 4) By improving the storage strategy of the original experience pool of the SAC algorithm, a method of post hoc experience recombination is proposed, a phased experience pool structure is designed, and the "recombination" mechanism is combined to improve the utilization efficiency of the algorithm for successful samples. The specific steps are as follows:

[0054] When designing the experience pool, the Future strategy is selected. The Future strategy randomly selects k states s t that satisfy t < i < T from the trajectory sequence where the current state s i is located as the target set G. In the grasping operation, the stage targets are first selected, which are the target g reach-stage in the approaching stage, the target g grasp-stage in the grasping stage, and the target g lift-stage in the lifting stage; therefore, the target set G of the Future strategy can be defined as:

[0055] G = [g reach-stage , g grasp-stage , g lift-stage (5)

[0056] Then, the reward is calculated for each stage target using Equation (6).

[0057] r t = [s t+1 = g] (6)

[0058] Through the above design, when the agent has not reached the current stage target after the exploration within a limited number of time steps, the reward value of the stage target after the current time step t can be used as the reward value of the current time step t, thus completing the modification of the experience pool data. The experience data processed by the HER algorithm is as shown in Equation (7):

[0059] (s t ||g, a t , r t , s t+1 ||g) (7)

[0060] If the agent's exploration in a particular episode concludes with experience that is considered either near-stage or successfully captured, then according to the HER algorithm, this experience is converted into experience at the next higher level. However, the Hindsight Experience Replay (HER) experience pool does not modify the top-level successfully retrieved experience. It represents task success, but it accounts for the smallest portion of the entire experience pool. This leads to the algorithm underutilizing this "successful" experience during training. This is where the "recombination" mechanism comes in, taking some of the successfully retrieved and captured experiences and "recombining" them.

[0061] like Figure 4 As shown in the figure, we select the position k of the two trajectories (experiences) as the split point (the moment k is the moment when the agent just grasps the target object), and divide the two trajectories into the first half and the second half. The first half of the successful grasping trajectory (s1, a1, r1, s2), ..., (s k ,a k ,r k ,s k+1 ), intercept a section of the trajectory that is successfully lifted with the corresponding starting point s k The second half of the trajectory is spliced ​​to the first half of the successful trajectory to generate a "recombined" trajectory.

[0062] In summary, combining hindsight experience replay (HER) and adding the recombination method can obtain a new experience pool.

[0063] In step 5), when updating the value network, we designed an adaptive conservative Q-learning value network update method to address the problem of policy oscillation in the later stages of algorithm training. This method dynamically adjusts the strength of the regularization term through an adaptive adjustment mechanism based on temporal difference error and average reward. The specific steps are as follows:

[0064] The optimization objective function of the original conservative Q learning is as shown in formula (8), where the coefficient η of the regularization term is the weight of the regularization term, which is a hyperparameter;

[0065]

[0066] The regularization term in formula (8) consists of two formulas, one is Latent term, the latent term is the Q-value expectation on the latent distribution. Generally, the latent action distribution can be generated by adding random noise to the actions generated by the behavior policy distribution. This allows for efficient exploration of actions outside the behavior policy distribution in high-dimensional space. The commonly used formula for generating latent actions is:

[0067]

[0068] Where a is the action generated by the behavior policy; ∈ is Gaussian noise; The mean is 0 and the variance is σ 2 Gaussian distribution of a′; the action after adding noise, often used for strategy smoothing or target strategy generation;

[0069] The other regularization term in formula (8) is Actual term, the actual term is the Q-value expectation on the behavior strategy distribution;

[0070] The TD error is defined as the difference between the current Q network output and the TD target value:

[0071] TD target =r+γ·Q target (s′,π target (s′)) (10)

[0072] TD error =Q(s,a)-TD target (11)

[0073] Among them TD target is the temporal difference (TD) target value; r is the reward of the current step; γ is the discount factor that controls the impact of future rewards; Q target is the output of the target Q network, representing the state-action value function; s′ is the next state; π target (s′) is the action selected by the target policy in the next state;

[0074] In order to facilitate the dynamic adjustment of regularization weights, TD error Perform normalization:

[0075]

[0076] Where N is the size of the sampling batch, the clamp(x,0,1) function can be used to error The mean of is limited to the range of [0,1];

[0077] The regularization term weight η is based on normalized_TD error Dynamic adjustment, the formula is:

[0078]

[0079] Where: R avg It is the average reward of multiple episodes of the agent. In actual application, it is the average reward of 4 episodes. threshold is the reward threshold, λ min and λ maxare the minimum and maximum values ​​of the regularization weight, set to 0 and 0.2 respectively;

[0080] About R threshold In practical applications, the reward function for a complex task is designed in stages. After the agent interacts with the environment, the reward at a single time step is normalized to facilitate algorithm comparison. Each stage is usually assigned a scale, which is the reward after the stage is completed. The scale of each stage satisfies the following conditions:

[0081]

[0082] n represents the last stage before the task is successfully completed, so in order to allow the agent to explore more stably when the task is close to completion, R threshold The selection of can be based on formula (15):

[0083]

[0084] Episode steps Represents the number of time steps in each episode, that is, the total number of steps that the agent needs to interact with the environment in an episode;

[0085] When R avg ≤R threshold At the beginning of training, the reward value is low. In order to enhance the exploration ability, the regularization weight η = λ min When R avg >R threshold When it is the late stage of training, the reward value is high, the agent gradually converges, and the regularization weight η is based on normalized_TD error Dynamic adjustment;

[0086] Combined with the dynamically adjusted η, the optimization objective function of adaptive conservative Q learning is:

[0087]

[0088] Example 1

[0089] The task scenarios provided by robosuite can be used to simulate the grasping operation of this method. Before simulation, the relevant reward function needs to be designed. Taking the Lift scenario as an example, the reward function design process is as follows:

[0090] For the grasping task of an arm-hand robot, a reward function combining sparse rewards and dense rewards is designed. The specific logic is as follows: when the task is completed (the sparse reward condition is met), a sparse reward is directly given; if the task is not completed and dense rewards are enabled, the approach reward, heading reward, and grasping reward are used to guide the agent to gradually complete the task. The overall definition of the reward function is as follows:

[0091]

[0092] Among them, R sparse represents sparse rewards; R shaping Denotes dense reward. The following is a detailed description of the components of the reward function.

[0093] R sparse Sparse rewards are used to provide clear feedback after the task is completed, indicating that the task goal has been achieved. For the grasping operation, when the arm-hand robot lifts the target object, a fixed reward value of 2.25 is provided. It is defined as follows:

[0094]

[0095] R shaping Intensive rewards guide the agent to complete the task step by step by refining the intermediate process of the grasping task. Intensive rewards are composed of the following three parts:

[0096] 1) Proximity reward R reaching

[0097] To encourage the agent to gradually approach the target, a distance-based proximity reward is designed. The proximity reward is designed based on the Euclidean distance d between the grasping center and the center of the target, and is specifically defined as:

[0098] R reaching =1-tanh(10.0·d) (19)

[0099] Where d is the Euclidean distance from the grasping center of the dexterous hand to the target object. As the grasping center of the dexterous hand gradually approaches the target object, the reward value gradually increases, thereby guiding the arm-hand robot to gradually learn the approach operation.

[0100] 2) Towards reward R orientation

[0101] To ensure that the end effector can adjust to the desired posture with the target object, a heading reward based on the rotation angle deviation is designed. The heading reward is calculated based on the relative rotation angle θ between the grasping center and the target object in the x-axis direction. x To design, specifically defined as:

[0102] R orientation =(1-tanh(0.1·|θx -θ desired |))·0.25 (20)

[0103] Among them, θ desired =120° is the desired target angle, when θ desired =120° makes the z-axis of the grasp center perpendicular to the tabletop. This design reduces the agent's exploration in the orientation dimension. By rewarding orientations with smaller deviations, the end effector of the robot arm is encouraged to adjust to the ideal posture.

[0104] 3) Grab reward R grasp

[0105] To further guide the agent to complete the grasping action, a fixed reward of 0.25 is provided when the thumb tip of the dexterous hand and the tip of any of the other four fingers touch the block at the same time. It is defined as:

[0106]

[0107] Based on the above intensive reward part, intensive reward R shaping It can be defined as:

[0108] R shaping =R reaching +R orientation +R grasp (twenty two)

[0109] The final reward value R is normalized and scaled as needed, and the formula is:

[0110]

[0111] The function of reward_scale is to normalize and scale reward values, ensuring that the reward range is suitable for the optimization requirements of the reinforcement learning algorithm. In practical applications, rewards generally need to be normalized, so reward_scale is set to 1 here to ensure that the reward value of a single time step does not exceed 1. This setting ensures consistency in reward ranges across tasks, which not only helps the model learn stably but also ensures fairness in comparative experiments.

[0112] By completing the above reward function design and training the algorithm process of the present invention in the task scenario in the simulation environment, the relevant grasping operations of the arm-hand robot can be realized, such as Figure 5 The block grabbing task shown and Figure 6 Nut assembly task shown.

[0113] also Figure 7 The following is a comparison of the convergence speed of the algorithms. Figure 8The figure shows a comparison of algorithm stability. The results of the two figures show that the present invention has a faster convergence speed and better stability.

[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A grasping method for an arm-hand robot based on deep reinforcement learning, characterized in that: The method comprises the steps of: 1) Define the arm-hand robot state space S, action space A and reward function R(s) t ,a t ), where s t is the state at time t, a t is the action to be performed by the agent at time t, where s t ∈S,a t ∈A; The state space S is: a 50-dimensional vector, the angles s of 18 joints qpos With angular velocity s qvel , the 3D vector position s of the grasping center of the five-fingered hand grip-pos and a 4-dimensional vector towards s grip-quat , the 3D vector position s of the grasped object object-pos and a 4-dimensional vector towards s object-quat ; The action space A is: a 12-dimensional vector, including the 6-dimensional vector robot end tool center posture a robot The incremental form of control, the 6-dimensional vector position a of the dexterous hand hand ; 2) Initialize the value network Q θ , policy network and related parameters (such as the regularization coefficient range λ min and λ max , reward threshold R threshold wait); 3) In each Episode, through the main strategy network π(s t ) Generate action a t , obtain state transition data (s t ,a t ,r t ,s t+1 ); 4) Use the post-experience replay technology to assign a new target g to the experience data and modify the corresponding reward r. The experience becomes (s t ||g,a t ,r′ t ,s t+1 ||g) and stored in the hierarchical experience pool (buffer_1, buffer_2, buffer_3). The reorganization module reorganizes the data in the experience pool to generate new experience samples for training; 5) Calculate the average reward of the last four episodes and enter the update phase. Randomly sample data from the experience pool, calculate the target TD value and TD error, normalize the TD error, and dynamically adjust the regularization weight η according to the average reward value and TD error. The loss function of the regularization term is minimized to adjust Q. θ Optimize and update the policy network by maximizing the value network value Q(s,π(s)) 2. The arm-hand robot grasping method based on deep reinforcement learning according to claim 1 is characterized in that: In step 3, the agent uses the main policy network π(s) in each episode. t ) Generate action a t , obtain state transition data (s t ,a t ,r t ,s t+1 ), here we propose a policy network structure based on sparse causal temporal self-attention mechanism to solve the problem that traditional fully connected policy network structure is difficult to capture complex temporal dependencies; The computational complexity of attention is reduced through sparse connections in local windows, and a new policy structure is constructed in series with LSTM. The core of the entire policy network structure lies in the sparse causal temporal attention mechanism, which consists of multiple sub-modules and captures local and global temporal dependencies by expanding the receptive field layer by layer. The specific steps include: 1) Local Window Mechanism: The core part of the sparse causal temporal attention mechanism, mainly used to reduce the complexity of attention calculation. This mechanism divides the input time series into several non-overlapping windows of size W, and performs attention calculation within each window, limiting the attention calculation to the window range. This design significantly reduces the computational complexity while focusing on capturing the strong dependencies between local time steps. The calculation formula for attention within the window is: Among them, the query Q (Query), key K (Key) and value V (Value) are generated by the linear projection of the input feature X: Q h =XW q ,K h =XW k ,V h =XW v (2) in, is a learnable linear projection matrix, d k is the dimension of the key vector; 2) Causal masking mechanism: This is another key component in the design of the sparse causal temporal attention mechanism, used to ensure causality constraints in the model. Through the mask matrix, the model can mask information from future time steps, ensuring that the calculation of the current time step depends only on the current and previous time steps. The mask matrix is ​​defined as: In the attention calculation, the mask matrix is ​​added to the dot product result of the key and the query to shield the influence of future time steps. The updated attention matrix calculation formula is: This design ensures causality in time series tasks and provides a guarantee for causal modeling of policy networks and value networks in reinforcement learning; 3) Layer Normalization and Skip Connections: Layer Normalization and Residual Connections are introduced into the Sparse Causal Temporal Attention module. Layer Normalization helps stabilize the training process, while Skip Connections effectively alleviate the vanishing gradient problem in deep networks and preserve input information. 4) Multilayer Perceptron Submodule: Following the Mask Calculation Module are two layers of multilayer perceptrons, which utilize nonlinear activation functions to further enhance feature expression. This design enhances the model's nonlinear modeling capabilities, making it more flexible in time series modeling. 5) Layer-by-layer window size expansion: In the multi-layer design of the sparse causal temporal attention mechanism, the window size is expanded layer by layer. The lower layers mainly capture short-term dependencies, while the higher layers capture long-term global dependencies by expanding the receptive field. This design optimizes computational efficiency while ensuring the accuracy of dependency modeling. In the sparse design, Block 1 with a window size of 2 mainly processes short-term dependencies, Block 2 with a window size of 4 processes medium-term dependencies, and Block 3 with a window size of 8 is responsible for capturing dependencies with longer time spans. With this design, the module's time complexity is O(ρTWC), where ρ represents sparsity, typically ρ<1; T is the sequence length, W is the window size, and C is the feature dimension. Before the policy network was improved, the input required for each update of the policy network was usually "single frame" data, that is, batch tuples (s t ,a t ,r t ,s t+1 ); After the policy network is improved, because the input needs to match the sequence input format of LSTM and attention mechanism, the input of the improved policy network requires "multi-frame" data, that is, the input needs to be converted into batch tuples (s t ,a t ,r t ,s t+1 )…(s t+8 ,a t+8 ,r t+8 ,s t+9 ); The entire reinforcement learning action trajectory can be divided into single-step action features, local action features, and global action features. The combination of LSTM and sparse causal temporal attention mechanism policy network can better extract feature information from such trajectories; in this policy network, LSTM is responsible for extracting the time-dependent features of single-step, local, and global actions from the entire reinforcement learning trajectory, while the sparse causal temporal attention mechanism efficiently models the dependency relationship between local action features through sparse connections within the local window.

3. The arm-hand robot grasping method based on deep reinforcement learning according to claim 1 is characterized in that: In step 4), by improving the storage strategy of the SAC algorithm's native experience pool, a post-experience reorganization method is proposed. A phased experience pool structure is designed, and combined with a "reorganization" mechanism, the algorithm's efficiency in utilizing successful samples is improved. The specific steps are as follows: Select the Future strategy when designing the experience pool. The Future strategy randomly selects k states s that satisfy t < i < T from the trajectory sequence where the current state s is located. t As the target set G. In the grasping operation, the stage targets are first selected, which are the target g in the approaching stage, i the target g in the grasping stage, reach-stage and the target g in the lifting stage. grasp-stage Therefore, the target set G of the Future strategy can be defined as: lift-stage ; G=[g reach-stage ,g grasp-stage ,g lift-stage ] (5) Then use formula (6) to calculate the reward for each stage goal; r t =[s t+1 =g] (6) Through the above design, when the agent has not reached the goal of the current stage after the exploration within a limited time step, the reward value of the stage goal after the current time step t can be used as the reward value of the current time step t. In this way, the modification of the experience pool data is completed. The experience data processed by the HER algorithm is as follows: (s t ||g,a t ,r t ,s t+1 ||g) (7) If the agent's exploration in a certain episode ends with experience that is close to the stage or successfully grasped, then according to the HER algorithm process, these experiences will be converted into experience at the next level. However, the HER experience pool does not modify the top-level successfully grasped experience. It is considered task success experience, but this part accounts for the smallest proportion of the entire experience pool, which also leads to the algorithm under-utilizing this "successful" experience during training. At this time, "recombination" is introduced, which takes out some of the successfully grasped and successfully grasped experiences and "recombines" them. The position k of the two trajectories is selected as the split point. The k moment is the moment when the agent just grasps the target object. The two trajectories are divided into the first half and the second half. The first half of the trajectory (s1, a1, r1, s2), ..., (s k ,a k ,r k ,s k+1 ), intercept a section of the trajectory that is successfully lifted with the corresponding starting point s k The second half of the trajectory is spliced ​​to the first half of the successful trajectory to generate a "recombined" trajectory; Combining hindsight experience replay (HER) and adding recombination methods can obtain a new experience pool.

4. The arm-hand robot grasping method based on deep reinforcement learning according to claim 1 is characterized in that: In step 5), when updating the value network, we designed an adaptive conservative Q-learning value network update method to address the problem of policy oscillation in the later stages of algorithm training. This method dynamically adjusts the strength of the regularization term through an adaptive adjustment mechanism based on temporal difference error and average reward. The specific steps are as follows: The optimization objective function of the original conservative Q learning is as shown in formula (8), where the coefficient η of the regularization term is the weight of the regularization term, which is a hyperparameter; The regularization term in formula (8) consists of two formulas, one is Latent term, the latent term is the Q-value expectation on the latent distribution. Generally, the latent action distribution can be generated by adding random noise to the actions generated by the behavior policy distribution. This allows for efficient exploration of actions outside the behavior policy distribution in high-dimensional space. The commonly used formula for generating latent actions is: Where a is the action generated by the behavior policy; ∈ is Gaussian noise; The mean is 0 and the variance is σ 2 Gaussian distribution of a′; the action after adding noise, often used for strategy smoothing or target strategy generation; The other regularization term in formula (8) is Actual term, the actual term is the Q-value expectation on the behavior strategy distribution; The TD error is defined as the difference between the current Q network output and the TD target value: TD target =r+γ·Q target (s′,π target (s′)) (10) TD error =Q(s,a)-TD target (11) Among them TD target is the temporal difference (TD) target value; r is the reward of the current step; γ is the discount factor that controls the impact of future rewards; Q target is the output of the target Q network, representing the state-action value function; s′ is the next state; π target (s′) is the action selected by the target policy in the next state; In order to facilitate the dynamic adjustment of regularization weights, TD error Perform normalization: Where N is the size of the sampling batch, the clamp(x,0,1) function can be used to error The mean of is limited to the range of [0,1]; The regularization term weight η is based on normalized_TD error Dynamic adjustment, the formula is: Where: R avg It is the average reward of multiple episodes of the agent. In actual application, it is the average reward of 4 episodes. threshold is the reward threshold, λ min and λ max are the minimum and maximum values ​​of the regularization weight, set to 0 and 0.2 respectively; About R threshold The method selected in practical applications is: the reward function design of a complex task is designed in stages. After the agent interacts with the environment, the reward of a single time step should be normalized to facilitate algorithm comparison. Usually, each stage will have a proportional coefficient scale, which is usually the reward after the stage is completed. The coefficients of each stage meet the following requirements: n represents the last stage before the task is successfully completed, so in order to allow the agent to explore more stably when the task is close to completion, R threshold The selection can be based on formula (15): Episode steps Represents the number of time steps in each episode, that is, the total number of steps that the agent needs to interact with the environment in an episode; When R avg ≤R threshold At the beginning of training, the reward value is low. In order to enhance the exploration ability, the regularization weight η = λ min ; When R avg >R threshold When it is the late stage of training, the reward value is high, the agent gradually converges, and the regularization weight η is based on normalized_TD error Dynamic adjustment; Combined with the dynamically adjusted η, the optimization objective function of adaptive conservative Q learning is:

Citation Information

Cited By

  • Humanoid robot lower limb walking control method and device

    CN121069743A