A Robot Tracking Control Learning Method Based on Post-Screening Experience Replay

Through the robot tracking control learning method based on post-search screening experience replay, the problem of single reward method in robotic arm trajectory tracking is solved, and a more efficient and stable trajectory tracking effect is achieved.

CN119427356BActive Publication Date: 2025-06-24DONGGUAN UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411641640.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-06-24
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

In the track tracking of robotic arm, the reward method is single during experience replay and cannot be adjusted according to the actual state of the robotic arm, resulting in poor tracking effect and stability of the trained model.

Method used

The robot tracking control learning method based on post-search screening experience playback is adopted. By initializing the target network parameters and experience pool, the action complex state is determined based on the trajectory smoothness and node covariance, the appropriate experience pool optimization method and experience selection method are selected, and the soft update weight of the network is adjusted according to the reward evaluation value.

Benefits of technology

The tracking effect and tracking stability of the trained model are improved, the agent's exploration of meaningless actions is avoided, the logic of reward feedback and output actions is enhanced, and the accuracy and stability of the model's tracking of the robotic arm trajectory is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119427356B_ABST
    Figure CN119427356B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of robotic arm trajectory tracking, and particularly to a robot tracking control learning method based on post hoc screening experience replay, including: initializing the target network parameters and the experience pool, and storing the state transition tuples in the experience pool; when the number of state transition tuples in the experience pool is greater than the preset number of state transition tuples, determining the action complex state according to the trajectory smoothness and node covariation degree of the robotic arm of the robot; determining the experience pool optimization method according to the action complex state; selecting the experiences with the predicted position deviation greater than or equal to the standard predicted position deviation as the screened experiences, and determining the experience selection method according to the number of the screened experiences; updating the critic network and the actor network, and respectively performing soft updates on the target networks of the critic network and the actor network; adjusting the soft update weights of the target network of the critic network and the target network of the actor network according to the reward evaluation value, and the present invention improves the tracking effect and tracking stability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robotic arm trajectory tracking, and in particular to a robot tracking control learning method based on post hoc screening experience replay. Background Art

[0002] In the trajectory tracking of a robotic arm, when training an agent through a deep reinforcement learning control algorithm, due to defects such as long training time and low sample utilization rate, the learning efficiency of the agent is low. Therefore, experience replay is often used to improve sample utilization rate. However, an inappropriate reward method will cause the agent to learn with bias, resulting in poor tracking effect and tracking stability of the trained model. Therefore, how to improve the tracking effect and tracking stability of the trained model is a technical problem that needs to be solved urgently by those skilled in the art.

[0003] Patent Publication No. CN118493388A discloses a deep reinforcement learning robotic arm grasping method for sparse rewards, including: First, analyze the characteristics of the robotic arm grasping task, model it as a Markov decision problem, design a binary sparse reward to reduce the complexity of the reward function design and lower the design cost; Second, use the DDPG algorithm as the main deep reinforcement learning training algorithm framework to build an Actor-Critic structure network to process the continuous state-action space; Then, design a post hoc experience replay mechanism, use the G-HGG algorithm to generate auxiliary targets, use a pre-trained action network to screen actions and add exploration noise and energy functions to process the cumulative experience pool to enhance experience utilization rate and improve training efficiency and grasping success rate; Finally, build a robotic arm model and scene information, and use interactive data for optimized training to achieve robotic arm target grasping. It can be seen that the above technical solution has the following problems: When performing experience replay, the reward method is single, and it is impossible to adaptively adjust the reward method according to the actual state of the robotic arm. Moreover, only the next state is set as the tracking target for reward, resulting in that the reward function cannot accurately describe the action when the complexity of the tracking task is high, resulting in poor tracking effect and tracking stability of the trained model. Summary of the Invention

[0004] Therefore, the present invention provides a robot tracking control learning method based on post hoc screening experience replay to overcome the problems in the prior art that when performing experience replay, the reward method is single, it is impossible to adaptively adjust the reward method according to the actual state of the robotic arm, and only the next state is set as the tracking target for reward, resulting in that the reward function cannot accurately describe the action when the complexity of the tracking task is high, resulting in poor tracking effect and tracking stability of the trained model.

[0005] To achieve the above object, the present invention provides a robot tracking control learning method based on post hoc screening experience replay, including:

[0006] Initialize the target network parameters and the experience pool, and store the state transition tuples in the experience pool;

[0007] When the number of state transition tuples in the experience pool is greater than the preset number of state transition tuples, determine the action complexity state according to the trajectory smoothness of the robot manipulator and the node covariation degree;

[0008] Determine the experience pool optimization method according to the action complexity state, that is, determine the reward value of the pre-updated experience in the experience pool according to the position error and the estimated position deviation, or determine the reward value of the pre-updated experience in the experience pool according to the estimated position deviation;

[0009] Select the experiences with the estimated position deviation greater than or equal to the standard estimated position deviation as the screened experiences, and determine the experience selection method as screening experience selection for the experience combination or the reference sequence according to the number of screened experiences. Record the screened experiences selected by the experience selection method as the pre-updated experiences;

[0010] Update the critic network and the actor network, and perform soft updates on the target networks of the critic network and the actor network respectively;

[0011] Adjust the soft update weights of the target networks of the critic network and the actor network according to the reward evaluation value.

[0012] Furthermore, the action complexity state includes:

[0013] The first action complexity state where the trajectory smoothness is in the first preset trajectory smoothness range or the node covariation degree is in the first preset node covariation degree range;

[0014] The second action complexity state where the trajectory smoothness is in the second preset trajectory smoothness range and the node covariation degree is in the second preset node covariation degree range;

[0015] Among them, each value in the first preset trajectory smoothness range is greater than each value in the second preset trajectory smoothness range, and each value in the first preset node covariation degree range is greater than each value in the second preset node covariation degree range.

[0016] Furthermore, the experience pool optimization method is determined according to the action complexity state, where

[0017] In the first action complexity state, the experience pool optimization method is to determine the reward value of the pre-updated experience in the experience pool according to the position error and the estimated position deviation;

[0018] In the second action complexity state, the experience pool optimization method is to determine the reward value of the pre-updated experience in the experience pool according to the estimated position deviation.

[0019] Further, select the state transition tuples with the predicted position deviation greater than or equal to the standard predicted position deviation as the screened experiences;

[0020] The predicted position deviation at time t is determined according to the difference between the position error at time t and the position error at time t + 1.

[0021] Further, the experience selection method is determined according to the number of screened experiences, where

[0022] if the number of screened experiences is greater than or equal to the preset number of screened experiences, perform screened experience selection for the experience combination;

[0023] if the number of screened experiences is less than the preset number of screened experiences, perform screened experience selection for the reference sequence.

[0024] Further, performing screened experience selection for the experience combination includes:

[0025] Perform combined analysis on each screened experience according to the reference order. When performing combined analysis on a single screened experience, record the screened experience as the target screened experience, and record the other screened experiences except the target screened experience as the reference screened experiences. Successively detect the reward evaluation deviation between each reference screened experience after the reference order of the target screened experience and the target screened experience until the reward evaluation deviation between the reference screened experience and the target screened experience is greater than the preset reward evaluation deviation, then stop the detection, and record the set of each screened experience before the reference order of the reference screened experience with the reward evaluation deviation greater than the preset reward evaluation deviation from the target screened experience as an experience combination, and continue to perform combined analysis on the screened experiences not recorded as experience combinations until each screened experience is recorded as an experience combination, then stop the combined analysis;

[0026] Perform screened experience selection for each experience combination. When performing screened experience selection for a single experience combination, randomly select a preset proportion of the screened experiences in the experience combination, and use each selected screened experience as the pre-updated experience.

[0027] Further, performing screened experience selection for the reference sequence includes:

[0028] Record the sequence obtained by sorting each screened experience according to the reference order as the reference sequence, divide the reference sequence into M equal parts, and select the screened experiences located at each equal division point as the pre-updated experiences.

[0029] Further, update the critic network;

[0030] Use the update formula (1) to update the critic network loss value L Q (θ), and the formula (1) is:

[0031]

[0032] Among them, N is the number of experiences randomly selected for update after determining the reward values of the pre-updated experiences in the experience pool, θ is the initial parameter of the critic network, i is 1, 2, 3, ……, N, s i is the current state of the i-th experience for update, and the state includes the actual position, actual speed, target position, and target speed of the robotic arm, a i is the current action of the i-th experience for update, and the action is the torque output by the robotic arm for the joint angle, Q θ (s i , a i ) is the predicted score of the critic network with the initial parameter θ for s i and a i of the i-th experience for update, r i is the reward value of the i-th experience for update, s′ i is the next state of the i-th experience for update, y(r i , s′ i ) is determined by formula (2), and formula (2) is:

[0033] y(r i , s′ i ) = r + γQ θ′ (s′ i , π Φ′ (s′ i )) (2),

[0034] Update the critic network using the method of stochastic gradient descent, and the update formula is formula (3),

[0035]

[0036] Among them, r is the reward value, γ is the discount factor, θ′ is the target network parameter of the critic network, π is the actor network to be trained, Ф′ is the target network parameter of the actor network, π Ф' (s′ i ) is the predicted action of the actor network with the target network parameter Ф′ for s′ i of the i-th experience for update, Q θ' (s′ i , π Ф′ (s′ i )) is the predicted score of the critic network with the target network parameter θ′ for s′ i and πФ' (s′ i )'s predicted score.

[0037] Furthermore, the actor network is updated using stochastic gradient descent, and the update formula is Formula (4),

[0038]

[0039] where a is equal to π Ф (s i ), Ф is the initial parameter of the actor network, and π Ф (s i ) is the predicted action of the actor network with the initial parameter Ф for the i-th experience used for update for s i .

[0040] Furthermore, the target network of the critic network is softly updated using Formula (5), and the target network of the actor network is softly updated using Formula (6),

[0041] θ′ = αθ+(1 - α)θ′ (5),

[0042] Ф′ = αФ+(1 - α)Ф′ (6),

[0043] where α is the soft update weight, α is greater than 0 and less than 1, and the soft update weights of the target networks of the critic network and the actor network are adjusted according to the reward evaluation value;

[0044] If the reward evaluation value is within the first preset reward evaluation value range, the soft update weight is adjusted to decrease;

[0045] If the reward evaluation value is within the second preset reward evaluation value range, the soft update weight is adjusted to increase;

[0046] where all values within the first preset reward evaluation value range are greater than all values within the second preset reward evaluation value range;

[0047] The reward evaluation value is determined according to the reward fluctuation coefficient and the reward mean coefficient. The reward fluctuation coefficient has a positive correlation with the reward evaluation value, and the reward mean coefficient has a positive correlation with the reward evaluation value.

[0048] Compared with the prior art, the beneficial effects of the present invention are as follows. In the technical solution of the present invention, the complex state of the action is determined according to the trajectory smoothness and the node covariation degree. The complexity of the manipulator action is effectively reflected by the trajectory smoothness and the node covariation degree. Then, different experience pool optimization methods are adaptively selected, making the selection of the experience pool optimization method more in line with the actual application scenario, avoiding the problem in the prior art that the reward method cannot be adaptively adjusted according to the actual state of the manipulator, and only setting the next state as the tracking target for reward, reducing the exploration of meaningless actions by the agent, and then being able to accurately describe the action when the complexity of the tracking task is high, improving the tracking effect and tracking stability of the trained model.

[0049] Further, in the present invention, the reward value of the pre-updated experience in the experience pool is determined according to the position error and the estimated position deviation, which is beneficial to giving positive rewards to correct actions, thus enhancing the logic between the reward feedback and the output action, reducing the exploration of meaningless actions by the agent, enabling the agent to automatically adapt to the changes in the environment during the learning process, and then being able to accurately describe the action when the complexity of the tracking task is high; determining the reward value according to the estimated position deviation is beneficial to reducing the computational burden while ensuring the tracking efficiency when the complexity of the tracking task is low, and then optimizing the accuracy and stability of the model for the manipulator trajectory tracking.

[0050] Further, in the present invention, the experience with the estimated position deviation greater than or equal to the standard estimated position deviation is selected as the screened experience. The estimated position deviation effectively reflects that the state transition tuple is beneficial for the agent to reach the end point. Then, the state transition tuples beneficial to achieving the goal are screened out for reward, avoiding the problem of inaccurate state description by the reward, and then making the state transition tuples with guiding significance have higher reward returns, avoiding the problem that the selected state transition tuple may not be able to guide the agent to reach the tracking target, resulting in the deep reinforcement learning control algorithm not being able to converge effectively for a long time or even having no positive convergence trend, and improving the accuracy of the model for the manipulator trajectory tracking.

[0051] Further, in the present invention, the experience selection method is determined to screen the experience for the experience combination or the reference sequence according to the number of screened experiences. The number of screened experiences effectively reflects the richness of the screened experiences for the off-policy experience replay. Then, different experience selection methods are adaptively selected, making the selection of the experience selection method more in line with the actual application scenario, being able to comprehensively consider the relevance and difference between the screened experiences, and then ensuring the diversity of the screened experiences, ensuring that the selected screened experiences are more in line with the actual needs, and then improving the training efficiency and training accuracy.

[0052] Furthermore, in the present invention, the soft update weights of the target network of the critic network and the target network of the actor network are adjusted according to the reward evaluation value. The reward evaluation value effectively reflects the beneficial degree of model learning, enabling the model to adaptively adjust the soft update weights of the target network of the critic network and the target network of the actor network according to the actual situation, guiding the model to adjust its learning strategy to adapt to the new environment, and further guiding the model to pay more attention to behaviors that are conducive to obtaining high rewards during parameter update, which helps the model gradually approach the optimal strategy during the training process, thereby improving the tracking effect and tracking stability of the trained model. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a schematic diagram of the robot tracking control learning method based on post hoc screening experience replay of the present invention;

[0054] Figure 2 is a flowchart of determining the action complexity state according to the trajectory smoothness and node covariation degree of the present invention;

[0055] Figure 3 is a flowchart of determining the experience pool optimization method according to the action complexity state of the present invention;

[0056] Figure 4 is a flowchart of determining the experience selection method according to the number of screened experiences of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] In order to make the objectives and advantages of the present invention clearer and more understandable, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0058] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.

[0059] It should be noted that in the description of the present invention, the terms indicating directions or positional relationships such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention.

[0060] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and defined, the terms "installation", "connection", and "linkage" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0061] Please refer to Figures 1 to 4 As shown, the present invention provides a robot tracking control learning method based on post-screening experience replay, including:

[0062] Initialize the target network parameters and the experience pool, and store the state transition tuples in the experience pool;

[0063] When the number of state transition tuples in the experience pool is greater than the preset number of state transition tuples, determine the action complexity state according to the trajectory smoothness and node covariation degree of the robot manipulator;

[0064] Determine the experience pool optimization method according to the action complexity state as determining the reward value of the pre-updated experience in the experience pool according to the position error and the estimated position deviation, or determining the reward value of the pre-updated experience in the experience pool according to the estimated position deviation;

[0065] Select the experiences with the estimated position deviation greater than or equal to the standard estimated position deviation as the screened experiences, and determine the experience selection method as screening experience selection for the experience combination or reference sequence according to the number of screened experiences, and record the screened experiences selected by the experience selection method as the pre-updated experiences;

[0066] Update the critic network and the actor network, and perform soft updates on the target networks of the critic network and the actor network respectively;

[0067] Adjust the soft update weights of the target networks of the critic network and the actor network according to the reward evaluation value.

[0068] The application scenario of the present invention is robotic arm trajectory tracking. The present invention uses the DDPG deep reinforcement learning algorithm for learning. The DDPG deep reinforcement learning algorithm includes an actor network, a target network of the actor network, a critic network, a target network of the critic network, and a reward function. Among them, the actor network is used to output the actions of the robotic arm, and the critic network is used to score the actions output by the actor network and predict the impact of the actions on the environment. The target networks of the actor network and the critic network are both updated in a soft update manner, which is easily understood by those skilled in the art and will not be elaborated herein.

[0069] Initialize the target network parameters and the experience pool, and store the state transition tuple in the experience pool, including: initialize the critic network Q with random network parameters θ and Ф θ and the actor network π Ф , initialize the target network parameters θ′ = θ, Ф′ = Ф. The initialization parameters Ф of the actor network and the initialization parameters θ of the critic network can sample the weight values from a normal distribution or a uniform distribution, and initialize the bias to zero or a small constant. These weights and biases will form the initial parameters of the critic network and be initialized for the experience pool. Use the actor network to output actions and input them into the environment to obtain the next state. Store the current state transition tuple (s, a, r, s′) as an experience in the experience pool. s is the current state of the experience, a is the current action of the experience, r is the reward value, and s′ is the next state of the experience. If s′ is the end state, reset the environment state. The end state is that the actual position of the robotic arm is the same as the target position and the actual speed is the same as the target speed, which is easily understood by those skilled in the art and will not be elaborated herein.

[0070] The number of state transition tuples is the total number of state transition tuples currently stored in the experience pool. The value of the preset number of state transition tuples can be determined by the user according to the actual application scenario. The greater the user's demand for improving the model accuracy, the greater the value of the preset number of state transition tuples. Provide a value for the preset number of state transition tuples, and record the average value of the number of state transition tuples corresponding to the historical records with qualified marks that can meet the user's needs and have qualified marks when selecting and screening experiences as the preset number of state transition tuples.

[0071] In the present invention, the historical records include but are not limited to the number of state transition tuples, trajectory smoothness, node covariance, the number of screened experiences, and reward evaluation values during the historical model training process, and each historical record corresponds to a qualified mark. The qualified mark records that the tracking effect and tracking stability of the trained model can meet the user's needs.

[0072] Specifically, the complex action states include:

[0073] A first complex action state where the trajectory smoothness is within a first preset trajectory smoothness range or the node covariation degree is within a first preset node covariation degree range;

[0074] A second complex action state where the trajectory smoothness is within a second preset trajectory smoothness range and the node covariation degree is within a second preset node covariation degree range;

[0075] Among them, each value within the first preset trajectory smoothness range is greater than each value within the second preset trajectory smoothness range, and each value within the first preset node covariation degree range is greater than each value within the second preset node covariation degree range.

[0076] Among them, the present invention is provided with a continuously circulating monitoring period. The determination of the complex action state is performed once at the end of each monitoring period. The duration of the monitoring period can be set according to the user's needs. The greater the user's requirement for the monitoring accuracy of the complex action state, the smaller the duration of the monitoring period. A value for the duration of the monitoring period is provided, and the duration of the monitoring period is 30 min.

[0077] The confirmation method of the trajectory smoothness is as follows: Detect the movement trajectory of any point at the end of the robotic arm of the robot currently performing tracking control learning within the current monitoring period. Denote this point as the target end point. At the end of the monitoring period, establish a three-dimensional rectangular coordinate system with the position of the target end point as the origin. Divide the movement trajectory into B equal parts. Denote each equal division point and the initial position and the termination position of the movement trajectory as reference points. The initial position is the position of the end effector of the mechanical part at the start of the current monitoring period, and the termination position is the position of the end effector of the mechanical part at the end of the current monitoring period. Draw a vector tangent to the movement trajectory at each reference point. The direction of the vector is the same as the movement direction. The vector corresponding to the j-th reference point is j = 1, 2, 3, ……, B + 2. Calculate the included angle corresponding to each trajectory segment. The trajectory segment is the movement trajectory between the j-th reference point and the (j + 1)-th reference point. The calculation formula for the included angle ε between the vector corresponding to the j-th reference point and the vector corresponding to the (j + 1)-th reference point is formula (7).

[0078]

[0079] Smoothing detection is performed for each included angle according to the moving direction of the moving trajectory. When performing smoothing detection for a single included angle, this included angle is denoted as the target included angle, and the angle change value between the included angle adjacent to the target included angle after the moving direction of the moving trajectory and the target included angle is calculated. If the angle change value is less than the preset angle change value, the target included angle is denoted as the smoothed included angle, and smoothing detection continues for the included angles that have not been smoothed until smoothing detection is completed for all included angles, at which point the smoothing detection stops; the number of smoothed included angles is denoted as the trajectory smoothness;

[0080] Regarding the value of the preset angle change value, the user can determine it according to the actual application scenario. It can be understood that the larger the value of the preset angle change value, the greater the degree of fluctuation of the moving trajectory of the robotic arm. A value of the preset angle change value is provided, and the preset angle change value is 30°.

[0081] The end of the robotic arm is the end that directly contacts the object. The end of the robotic arm includes, but is not limited to, grippers, suction cups, welding torches, and drills. The value of B can be determined by the user according to the actual application scenario. It can be understood that the larger the value of B, the greater the measurement accuracy of the trajectory smoothness. A value of B is provided, and B is 50.

[0082] The node co-variation degree is the number of nodes that change simultaneously within a single monitoring period; a node is a point in the robotic arm that connects various components or joints, which is easily understood by those skilled in the art and will not be elaborated here specifically.

[0083] All values within the first preset trajectory smoothness range are greater than or equal to the preset trajectory smoothness, all values within the second preset trajectory smoothness range are less than the preset trajectory smoothness, all values within the first preset node co-variation degree range are greater than or equal to the preset node co-variation degree, and all values within the second preset node co-variation degree range are less than the preset node co-variation degree.

[0084] Regarding the values of the preset trajectory smoothness and the preset node co-variation degree, the user can determine them according to the actual application scenario. It can be understood that the smaller the value of the preset trajectory smoothness and the larger the value of the preset node co-variation degree, the greater the complexity of the robotic arm's movement. A value of the preset trajectory smoothness and the preset node co-variation degree is provided. The optimization method for the detection experience pool is based on the historical records of determining the reward value according to the estimated position deviation. The average value of the trajectory smoothness corresponding to the historical records that can meet the user's needs is denoted as the preset trajectory smoothness, and the average value of the node co-variation degree corresponding to the historical records that can meet the user's needs is denoted as the preset node co-variation degree.

[0085] Specifically, the optimization method for the experience pool is determined according to the action complexity state, where,

[0086] In the first complex action state, the optimization method of the experience pool is to determine the reward value of the pre-updated experience in the experience pool according to the position error and the estimated position deviation;

[0087] In the second complex action state, the optimization method of the experience pool is to determine the reward value of the pre-updated experience in the experience pool according to the estimated position deviation.

[0088] Wherein, the t-th moment is the time corresponding to any state transition tuple. At the t-th moment, the position error at the t-th moment = the actual position at the t-th moment - the target position at the t-th moment;

[0089] In the first complex action state, the reward value at the t-th moment = the estimation coefficient × the estimated position deviation at the t-th moment + the position coefficient × the position error at the t-th moment; for the values of the estimation coefficient and the position coefficient, the user can obtain them through learning the historical records by means of a deep learning convolutional neural network. It can be understood that in the present invention, the reward value is used to reflect the feedback obtained by the intelligent communication from the environment after taking a certain action to judge the guiding degree of the current action on the tracking effect. The user can use the historical records for deep learning to determine the influence of the estimated position deviation and the position error on the guiding degree of the trajectory tracking effect respectively, and then correspondingly select the values of the estimation coefficient and the position coefficient. A set of values of the estimation coefficient and the position coefficient is provided, where the estimation coefficient is 0.5 and the position coefficient is 0.5.

[0090] In the second complex action state, the reward value at the t-th moment has a positive correlation with the estimated position deviation at the t-th moment.

[0091] Specifically, the state transition tuples with the estimated position deviation greater than or equal to the standard estimated position deviation are selected as the screened experiences;

[0092] The estimated position deviation at the t-th moment is determined according to the difference between the position error at the t-th moment and the position error at the (t + 1)-th moment.

[0093] Wherein, the standard estimated position deviation is 0, the t-th moment is the time corresponding to any state transition tuple, and the estimated position deviation at the t-th moment = the position error at the t-th moment - (the position error at the (t + 1)-th moment).

[0094] Specifically, the experience selection method is determined according to the number of screened experiences. Among them,

[0095] If the number of screened experiences is greater than or equal to the preset number of screened experiences, screened experience selection is performed for the experience combination;

[0096] If the number of screened experiences is less than the preset number of screened experiences, screened experience selection is performed for the reference sequence.

[0097] Among them, the number of screening experiences is the total amount of screening experiences. The value of the preset number of screening experiences can be determined according to the actual application scenario. The greater the user's demand for improving the model accuracy, the greater the value of the preset number of screening experiences. A value of the preset number of screening experiences is provided. The average value of the screening experiences corresponding to the historical records with qualified marks and whose selection accuracy can meet the user's needs when selecting screening experiences is recorded as the preset number of screening experiences.

[0098] Specifically, for the screening experience selection of the experience combination, it includes:

[0099] Perform combined analysis on each screening experience according to the reference order. When performing combined analysis on a single screening experience, record this screening experience as the target screening experience, and record the other screening experiences except the target screening experience as the reference screening experiences. Successively detect the reward evaluation deviation between each reference screening experience after the reference order of the target screening experience and the target screening experience until the reward evaluation deviation between the reference screening experience and the target screening experience is greater than the preset reward evaluation deviation, then stop the detection, and record the set of each screening experience before the reference order of the reference screening experience whose reward evaluation deviation from the target screening experience is greater than the preset reward evaluation deviation as an experience combination, and continue to perform combined analysis on the screening experiences not recorded as experience combinations until each screening experience is recorded as an experience combination, then stop the combined analysis;

[0100] For the screening experience selection of each experience combination, when performing screening experience selection on a single experience combination, randomly select a preset proportion of the screening experiences in the experience combination, and use each selected screening experience as the pre-updated experience.

[0101] Among them, the method for confirming the reference order is, it can be understood that the screening experience is the state transition tuple corresponding to a certain time. Each screening experience corresponds to a different time. The time corresponding to the screening experience corresponds to a position in the movement trajectory. The reference order is the time sequence of reaching each position in the movement trajectory;

[0102] The movement trajectory is the movement trajectory of any point at the end of the robotic arm of the robot currently performing tracking control learning within the current monitoring period;

[0103] The value of the preset proportion can be determined by the user according to the actual application scenario. The higher the user's accuracy requirement for training the model, the greater the value of the preset proportion. A value of the preset proportion is provided. The preset proportion is 20%;

[0104] The confirmation methods after and before the reference order are as follows: for a screening experience, record this screening experience as the target experience, detect the position of the time corresponding to the target experience in the movement trajectory, and record this position as the target position. The screening experiences corresponding to the times of each position after the movement trajectory reaches the target position are those after the reference order, and the screening experiences corresponding to the times of each position before the movement trajectory reaches the target position are those before the reference order.

[0105] The reward evaluation deviation is the absolute value of the difference between the reward values of two screening experiences. The value of the preset reward evaluation deviation can be determined by the user according to the actual application scenario. It can be understood that the larger the value of the preset reward evaluation deviation, the greater the difference degree between the two screening experiences. Provide a value of the preset reward evaluation deviation. The detection experience extraction method is the historical record of extracting screening experiences for an experience combination, and the average value of the reference deviations corresponding to each experience combination corresponding to the historical records that can meet the user's needs is recorded as the preset reward evaluation deviation. The reference deviation is the reward evaluation deviation corresponding to any two screening experiences in an experience combination.

[0106] Specifically, the selection of screening experiences for the reference sequence includes:

[0107] The sequence obtained by sorting each screening experience according to the reference order is recorded as the reference sequence. The reference sequence is divided into M equal parts, and the screening experiences located at each equal division point are selected as the pre-updated experiences.

[0108] Among them, the value of M can be determined by the user according to actual needs. Provide a value of M, and M is 50.

[0109] Specifically, update the critic network;

[0110] Use the update formula (1) to update the critic network loss value L Q (θ), and the formula (1) is:

[0111]

[0112] Among them, N is the number of experiences randomly selected for update after determining the reward values of the pre-updated experiences in the experience pool, θ is the initialization parameter of the critic network, i is 1, 2, 3,..., N, s i is the current state of the i-th experience for update. The state includes the actual position, actual speed, target position, and target speed of the robotic arm, a i is the current action of the i-th experience for update. The action is the torque output by the robotic arm for the joint angle, Q θ (s i , a i)The predicted score of the critic network with initialized parameter θ for the s of the i-th experience used for update i and a i , r i is the reward value of the i-th experience used for update, s′ i is the next state of the i-th experience used for update, y(r i , s′ i ) is determined by formula (2), and formula (2) is:

[0113] y(r i , s′ i ) = r + γQ θ′ (s′ i , π Φ′ (s′ i )) (2), Update the critic network using stochastic gradient descent, and the update formula is formula (3),

[0114]

[0115] where r is the reward value, γ is the discount factor, θ′ is the target network parameter of the critic network, π is the actor network to be trained, Ф′ is the target network parameter of the actor network, π Ф' (s′ i ) is the predicted action of the actor network with target network parameter Ф′ for the s′ of the i-th experience used for update i , Q θ' (s′ i , π Ф' (s′ i )) is the predicted score of the critic network with target network parameter θ′ for the s′ and π i and π Ф' (s′ i ) of the i-th experience used for update.

[0116] Among them, N is the number of experiences used for update randomly selected after determining the reward values of the pre-update experiences in the experience pool. It can be understood that after optimizing the reward values of the pre-update experiences according to different experience pool optimization methods, the state transition tuples of the experience pool are recorded as optimized experiences, and N optimized experiences are randomly selected from the experience pool. These N optimized experiences are the experiences used for update. The value of N can be determined by the user according to the actual application scenario. The greater the user's precision requirement for model training, the greater the value of N. Provide a value of N, and N is 40% of the number of experiences used for update. Q is the critic network to be trained; the loss value of the critic network is denoted as L QIt is represented by L(θ). Q L(θ) is used to measure the Q predicted by the critic network θ (s i , a i ), and the difference between the target Q value of the i-th experience used for update, y(r i , s′ i ), is the target Q value of the i-th experience used for update. is the gradient of the loss function L(θ) with respect to the parameter θ. The loss function L(θ) is the mean square error between Q θ (s i , a i ) and y(r i , s′ i ). Here, the calculation formula of L(θ) is the same as that of L Q (θ), which is easy for those skilled in the art to understand and will not be elaborated here.

[0117] Specifically, the actor network is updated.

[0118] Among them, the actor network is updated using the method of stochastic gradient descent, and the update formula is formula (4).

[0119]

[0120] where a is equal to π Ф (s i ), Ф is the initial parameter of the actor network, and π Ф (s i ) is the predicted action of the actor network with the initial parameter Ф for the s of the i-th experience used for update. i

[0121] Among them, is the policy gradient when the parameters of the actor network are updated. is the gradient of π Ф (s i ) with respect to the parameters. is the gradient of Q θ (s i , a) with respect to the parameter a.

[0122] Specifically, the target network of the critic network is softly updated using formula (5), and the target network of the actor network is softly updated using formula (6).

[0123] θ′ = αθ+(1 - α)θ′ (5),

[0124] Φ′ = αΦ+(1 - α)Φ′ (6).

[0125] Among them, α is the soft update weight, where α is greater than 0 and less than 1, and the soft update weights for the target network of the critic network and the target network of the actor network are adjusted according to the reward evaluation value;

[0126] If the reward evaluation value is within the first preset reward evaluation value range, the soft update weight is adjusted to decrease;

[0127] If the reward evaluation value is within the second preset reward evaluation value range, the soft update weight is adjusted to increase;

[0128] Among them, all values within the first preset reward evaluation value range are greater than all values within the second preset reward evaluation value range;

[0129] The reward evaluation value is determined according to the reward fluctuation coefficient and the reward mean coefficient. The reward fluctuation coefficient has a positive correlation with the reward evaluation value, and the reward mean coefficient has a positive correlation with the reward evaluation value.

[0130] Among them, the soft update weight is the proportion for updating the target network of the critic network or the target network of the actor network. The value of α directly affects the stability and performance of the reinforcement learning algorithm. Users can make a preliminary setting according to experience, and during the subsequent training process, adjust the soft update weights for the target network of the critic network and the target network of the actor network according to the reward evaluation value to ensure the stability of the reinforcement learning algorithm. α can be initially set to 0.05;

[0131] Reward evaluation value = reward fluctuation coefficient + reward mean coefficient. The reward fluctuation coefficient has a positive correlation with the reward fluctuation value, the reward mean coefficient has a positive correlation with the reward mean. The reward fluctuation value is the difference between the maximum reward value corresponding to each pre-update experience and the minimum reward value corresponding to each pre-update experience, and the reward mean is the average of the reward values corresponding to each pre-update experience;

[0132] All values within the first preset reward evaluation value range are greater than the first preset reward evaluation value, and all values within the second preset reward evaluation value range are less than the second preset reward evaluation value. Among them, the first preset reward evaluation value is greater than the second preset reward evaluation value;

[0133] The values of the first preset reward evaluation value and the second preset reward evaluation value can be determined according to the actual application scenario. It can be understood that the reward evaluation value reflects the fitness of the learning strategy of the model to the environment. The smaller the value of the first preset reward evaluation value and the larger the value of the second preset reward evaluation value, the greater the adjustment intensity of the reinforcement learning algorithm, and thus the greater the fitness of the learning strategy of the model to the environment. To improve, take the maximum value of the reward evaluation value corresponding to the historical record that can meet the user's needs as the first preset reward evaluation value, and take the minimum value of the reward evaluation value corresponding to the historical record that can meet the user's needs as the second preset reward evaluation value.

[0134] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

[0135] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A robot tracking control learning method based on post-screening experience playback, characterized in that: include: Initialize the target network parameters and experience pool, and store the state transition tuple in the experience pool; When the number of state transition tuples in the experience pool is greater than the preset number of state transition tuples, the complex state of the action is determined according to the trajectory smoothness and node covariance of the robot manipulator; The optimization method of determining the experience pool according to the complex state of the action is to determine the reward value of the pre-updated experience in the experience pool according to the position error and the estimated position deviation, or to determine the reward value of the pre-updated experience in the experience pool according to the estimated position deviation; Select the experience whose estimated position deviation is greater than or equal to the standard estimated position deviation as the screening experience, and determine the experience selection method to select the screening experience for the experience combination or reference sequence according to the number of screening experiences, and record the screening experience selected by the experience selection method as the pre-update experience; Update the critic network and actor network, and perform soft updates on the target network of the critic network and actor network respectively; Adjust the soft update weights of the target network of the critic network and the target network of the actor network according to the reward evaluation value; The experience selection method is determined according to the number of screening experiences, wherein: If the number of screening experiences is greater than or equal to the preset number of screening experiences, screening experiences are selected for the experience combination; If the number of screening experiences is less than the preset number of screening experiences, screening experiences are selected for the reference sequence; Empirical selection for screening against reference sequences, including: The sequence after sorting the screening experiences according to the reference order is recorded as the reference sequence, the reference sequence is divided into M equal parts, and the screening experiences at each equal part are selected as pre-update experiences.

2. The robot tracking control learning method based on post-screening experience playback according to claim 1 is characterized in that: Complex action states include: A first action complexity state in which the trajectory smoothness is within a first preset trajectory smoothness range or the node covariance is within a first preset node covariance range; A second action complexity state in which the trajectory smoothness is within a second preset trajectory smoothness range and the node covariance is within a second preset node covariance range; Among them, all values ​​within the first preset trajectory smoothness range are greater than all values ​​within the second preset trajectory smoothness range, and all values ​​within the first preset node covariance range are greater than all values ​​within the second preset node covariance range.

3. The robot tracking control learning method based on post-screening experience playback according to claim 2 is characterized in that: The experience pool optimization method is determined according to the complexity of the action, where: In the first complex action state, the experience pool optimization method is to determine the reward value of the pre-updated experience in the experience pool according to the position error and the estimated position deviation; In the second complex action state, the experience pool optimization method is to determine the reward value of the pre-updated experience in the experience pool according to the estimated position deviation.

4. The robot tracking control learning method based on post-screening experience playback according to claim 3 is characterized in that: Select the state transition tuple whose estimated position deviation is greater than or equal to the standard estimated position deviation as the screening experience; The estimated position deviation at time t is determined based on the difference between the position error at time t and the position error at time t+1.

5. The robot tracking control learning method based on post-screening experience playback according to claim 1 is characterized in that: Screening experience selection based on experience combination, including: Perform combination analysis on each screening experience according to the reference order. When performing combination analysis on a single screening experience, record the screening experience as the target screening experience, and record other screening experiences other than the target screening experience as reference screening experiences. Detect the reward evaluation deviation between each reference screening experience after the reference order of the target screening experience and the target screening experience in turn until the reward evaluation deviation between the reference screening experience and the target screening experience is greater than the preset reward evaluation deviation, then stop the detection, and record the set of each screening experience before the reference order of the reference screening experience whose reward evaluation deviation with the target screening experience is greater than the preset reward evaluation deviation as an experience combination, and continue to perform combination analysis on the screening experiences that are not recorded as experience combinations until each screening experience is recorded as an experience combination, then stop the combination analysis; Screening experience selection is performed for each experience combination. When screening experience selection is performed for a single experience combination, a preset proportion of screening experiences in the experience combination is randomly selected, and each selected screening experience is used as a pre-update experience.

6. The robot tracking control learning method based on post-screening experience playback according to claim 1 is characterized in that: Update the critic network; Use update formula (1) to update the critic network loss value L Q (θ) is updated, and formula (1) is: Where N is the number of experiences randomly selected for updating after determining the reward value of the pre-updated experience in the experience pool, θ is the initialization parameter of the critic network, and i is 1, 2, 3, ..., N, s i is the current state of the i-th experience used for updating, which includes the actual position, actual speed, target position and target speed of the robot arm, a i is the current action of the i-th experience used for updating, the action is the torque output by the robot arm on the joint angle, Q θ (s i , a i ) is the s of the critic network with initialization parameter θ for the i-th experience used for updating i and a i The prediction score, r i is the reward value of the i-th experience used for updating, S′ i is the next state of the i-th experience used for updating, y(r i , s′ i ) is determined by formula (2), which is: y(r i ,s′ i )=r+γQ θ′ (s′ i ,π Ф′ (s′ i )) (2), Use stochastic gradient descent to update the critic network. The update formula is formula (3): Among them, r is the reward value, γ is the discount factor, θ′ is the target network parameter of the critic network, π is the actor network to be trained, Ф′ is the target network parameter of the actor network, π Ф′ (s′ i ) is the target network parameter Φ′ of the actor network for the i-th experience used for updating s′ i The predicted action, Q θ′ (s′ i , π Ф′ (s′ i )) is the target network parameter θ′ of the critic network for the i-th experience s′ used for updating i and π Ф′ (s′ i )’s prediction score.

7. The robot tracking control learning method based on post-screening experience playback according to claim 6 is characterized in that: Use stochastic gradient descent to update the actor network. The update formula is formula (4): Where a is equal to π Ф (s i ), Ф is the initialization parameter of the actor network, π Ф (s i ) is the s of the actor network with initialization parameter Φ for the i-th experience used for updating i Predicted actions.

8. The robot tracking control learning method based on post-screening experience playback according to claim 7 is characterized in that: Formula (5) is used to perform soft update on the target network of the critic network, and formula (6) is used to perform soft update on the target network of the actor network. θ′=αθ+(1-α)θ′ (5), Ф′=αФ+(1-α)Ф′ (6), Among them, α is the soft update weight, α is greater than 0 and less than 1, and the soft update weights of the target network of the critic network and the target network of the actor network are adjusted according to the reward evaluation value; If the reward evaluation value is within the first preset reward evaluation value range, the soft update weight is adjusted to decrease; If the reward evaluation value is within the second preset reward evaluation value range, the soft update weight is increased and adjusted; Wherein, each value within the first preset reward evaluation value range is greater than each value within the second preset reward evaluation value range; The reward evaluation value is determined according to a reward fluctuation coefficient and a reward mean coefficient. The reward fluctuation coefficient is positively correlated with the reward evaluation value, and the reward mean coefficient is positively correlated with the reward evaluation value.

Citation Information

Patent Citations

  • Sparse reward-oriented deep reinforcement learning mechanical arm grabbing method

    CN118493388A

  • Batch limitation reinforcement learning method

    CN113112018A

  • Mechanical arm path planning method integrating reinforcement learning and fuzzy obstacle avoidance

    CN113232016A