Mechanical arm trajectory tracking control double-Q reinforcement learning method with disturbance

By applying dual Q reinforcement learning method and perturbation observer in the robotic arm system, the problems of external perturbation and model uncertainty in the robotic arm trajectory tracking control are solved, and high-precision tracking and good anti-interference performance are achieved.

CN119910656APending Publication Date: 2025-05-02QUFU NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510264355.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

In industrial manufacturing, robotic arm trajectory tracking control faces external perturbations and model uncertainty problems, resulting in low control accuracy and poor steady-state performance.

Method used

A tracking controller based on dual Q reinforcement learning is designed using a dual Q reinforcement learning method combined with perturbation observer. The controller optimizes the control strategy through a dual Q learning algorithm, using a perturbation observer to estimate and compensate for external perturbations and model uncertainties.

Benefits of technology

The track tracking accuracy and anti-interference ability of the robotic arm are improved, the control cost is reduced, the Q value overestimation problem in traditional Q learning is avoided, and the good dynamic performance of the system is ensured and the small steady-state error is small.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005300775530000021
    Figure BDA0005300775530000021
  • Figure BDA0005300775530000025
    Figure BDA0005300775530000025
  • Figure BDA0005300775530000035
    Figure BDA0005300775530000035
Patent Text Reader

Abstract

The invention relates to a disturbance-containing mechanical arm trajectory tracking control double-Q reinforcement learning method, and belongs to the technical field of mechanical arm intelligent control. According to the method, a double-Q reinforcement learning method is adopted to solve the problem of trajectory tracking control of the interfered mechanical arm. An interference observer is designed to estimate and counteract the influence of external interference and model inaccuracy, and the control precision and interference suppression capability of the mechanical arm system are improved; a tracking controller based on a double-Q reinforcement learning method is designed, and the control cost is kept while the dynamic and steady-state performance of the system is improved by utilizing the excellent performance of a double-Q learning algorithm; in addition, the problem that the Q value is too high in a traditional Q learning method is solved through the double-Q reinforcement learning method. According to the method, the advantages of the double-Q reinforcement learning method are effectively utilized, under the condition that the system contains disturbance and uncertain factors, the fast tracking performance of the mechanical arm system on the expected trajectory is achieved, and the good steady-state performance is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a control method, in particular to a double-Q reinforcement learning method for robot arm trajectory tracking control containing disturbance, and belongs to the field of robot arm intelligent control. Background Art

[0002] At present, robotic arms play an important role in industrial manufacturing. Realizing trajectory tracking control of robotic arms is one of the most basic requirements of robotic arm systems and has been widely used in the field of robotics. However, in actual industrial contexts, it is difficult to fully grasp and understand the kinematic and dynamic models of robotic arms. At the same time, external disturbances caused by working environment or uncertain factors always exist, which may significantly reduce the performance of robotic arm control systems. In order to ensure that the robotic arm can meet the stringent requirements of industrial manufacturing, it is a challenge for control theorists and engineers to achieve fast and high-precision tracking control of the robotic arm. Therefore, there is an urgent need to explore an effective control technology.

[0003] As a control method that does not require the precise model of the controlled object, reinforcement learning (RL) has emerged to broaden the range of options for designing control algorithms. One of the biggest advantages of RL is that it can calculate optimal control without a comprehensive understanding of the target system or the controlled object, especially when RL is directly applied to the target system, because RL can adapt to the uncertainty of the research object through data and use RL-based controllers to solve various complex problems.

[0004] As a typical reinforcement learning method, Q-learning has been successfully used to calculate the optimal energy loss in the system. However, the application of Q-learning does not always achieve the expected results. Especially in some random and uncertain environments, Q-learning will overestimate the Q value during the learning phase, thus failing to ensure the accuracy of selecting the optimal behavior. Therefore, the present invention is based on a dual-Q reinforcement learning method, introducing a disturbance observer to estimate and compensate for the influence of interference and uncertainty in the robotic arm system, while also avoiding the problem of overestimation of the Q value in traditional Q-learning, and well solving the problems of low system tracking accuracy and poor steady-state performance. Summary of the invention

[0005] The main purpose of the present invention is to solve the problem of disturbed robot trajectory tracking control by using a dual Q reinforcement learning method to address the shortcomings of the prior art. A disturbance observer is designed to estimate and offset the influence of external disturbances and model inaccuracies, thereby improving the control accuracy and disturbance suppression capability of the robot; a tracking controller based on the dual Q reinforcement learning method is designed to improve the tracking performance of the system by utilizing the advantages of the dual Q learning algorithm while maintaining the control cost. In addition, the dual Q reinforcement learning method also solves the problem of too high Q value encountered in the traditional Q learning method.

[0006] In order to achieve the above objectives, for an n-joint robotic arm system, the present invention provides a dual-Q reinforcement learning method for robotic arm trajectory tracking control with disturbance, comprising the following steps:

[0007] Step 1: Establish the dynamics model of the robotic arm system:

[0008]

[0009] In the formula, are joint position, joint velocity, and joint acceleration, respectively. is a symmetric positive definite inertia parameter matrix, is the centrifugal force and Coriolis force parameter matrix, is the gravity parameter matrix of the manipulator, is the control torque, is the control input variable, For external disturbance.

[0010]

[0011] In the formula, M0, C0, G0 are scalars of the system model parameter matrix, ΔM, ΔC, ΔG are uncertain terms of the system model parameter matrix. Therefore, the dynamic model (1) of the manipulator can be further expressed as:

[0012]

[0013] In the formula, represents the total unknown disturbance of the robot system, which consists of model uncertainty, load variation, and external disturbance.

[0014] Step 2, let x1 = q, The state space equation of the robotic arm system is obtained as follows:

[0015]

[0016] Step 3: Design a disturbance observer to estimate and compensate for the external disturbances and model parameter uncertainties of the system. The specific implementation steps are as follows:

[0017] 3-1, the disturbance observer of the designed system (4) is as follows:

[0018]

[0019] Where k(x1,x2) is the observation gain matrix, is the observed estimated value of the total disturbance to the system. Since the disturbance observer designed above has the acceleration signal of the state, it is impossible to directly realize the disturbance observation. In the manipulator system, using the velocity signal to obtain the acceleration signal may introduce noise and cause instability. Therefore, the observer needs to be modified as follows.

[0020] 3-2, construct the auxiliary function h(x1,x2), and define the internal state variables of the observer as:

[0021]

[0022] In the formula, is the internal state variable of the observer, and h(x1,x2) is the function vector to be designed. To avoid introducing acceleration signals, the perturbation observation gain matrix k(x1,x2) and h(x1,x2) have the following relationship:

[0023]

[0024] 3-3, design the disturbance observer structure. From equations (5)-(7), we get:

[0025]

[0026] In summary, after improving the disturbance observer (5), the obtained acceleration disturbance-free observer is:

[0027]

[0028] Step 4: Design a tracking controller based on the double-Q reinforcement learning method. The specific method is:

[0029] 4-1. First, according to the characteristics of robot tracking control, the desired tracking trajectory q is given d , define the tracking error as e d =q d -q.

[0030] In order to achieve tracking control of the desired trajectory, the control strategy of the tracking error is designed as follows:

[0031]

[0032] Where λ is a real number greater than 0, u1(t) is the feedback control strategy that can stabilize the robotic arm system, and u2(t) is the approximate optimal control strategy of the double Q reinforcement learning algorithm. The feedback control strategy can be expressed as u1(t)=q d +βe(t), where β is a constant.

[0033] 4-2, define the utility function of the system as:

[0034]

[0035] in, are the utility function constants of the state variables and the control variables, respectively. The utility function represents the instantaneous cost of the control variables at the current moment, and U(0,0)=0.

[0036] According to the previous discussion, we hope to find a suitable and feasible control strategy that can minimize the Q function as follows:

[0037]

[0038] Among them, ζ=t,t+1,t+2,... represents time t and any time thereafter, and the Q function is the sum of all time utility functions.

[0039] When the Q function in equation (12) is minimized, it is the optimal Q function Q * (e(t),u2(t)), the corresponding u2(t) is the optimal control strategy. This optimal control strategy can ensure that the error tends to zero while keeping the system stable, thereby achieving the tracking performance of the robot arm.

[0040] According to the Bellman optimality principle, the optimal Q function Q * (e(t),u2(t)) satisfies the following discrete-time HJB equation:

[0041]

[0042] Therefore, we can get the optimal tracking controller based on the double Q reinforcement learning method as:

[0043]

[0044] Double Q reinforcement learning strategy update steps:

[0045]

[0046] Double Q reinforcement learning is based on two Q functions: Q1 and Q2. Each Q function updates the Q value of the next state according to the other Q function. The action value a in row 6 of Table 1 * is the action with the greatest value in state s according to the action-value function Q1. However, unlike traditional Q-learning, traditional Q-learning uses the numerical value Q1(s t+1 ,a * )=max a Q1(s t+1 ,a) to update Q1(s,a), we use the value Q2(s t+1 ,a * )=max Q2(st+1 ,a) to update Q1(s,a). A similar update method is used for Q2, that is, to find t+1 The action b that maximizes Q2 * , then apply b * Get Q1(s t+1 ,b * )=max Q1(s t+1 ,a) to update Q2(s,a). Importantly, the two Q functions are learned from different sets of experience, but both value functions can be used when choosing an action to perform. Therefore, this algorithm is no less data efficient than traditional Q-learning.

[0047] The beneficial effects of the present invention are:

[0048] 1) The tracking controller based on the dual-Q reinforcement learning method proposed in the present invention can improve the tracking accuracy of the system and reduce the control cost.

[0049] 2) The proposed dual-Q reinforcement learning method can achieve real-time error tracking tasks and solve the poor performance problem caused by overestimation of Q values ​​in traditional Q learning, ensuring that the system has good dynamic performance and small steady-state error.

[0050] 3) The anti-interference framework of the robotic arm system constructed by the disturbance observer and the dual Q reinforcement learning method proposed in the present invention can effectively suppress the tracking error caused by interference and improve the tracking accuracy. At the same time, the observation value is fed back to the dual Q reinforcement learning tracking controller to improve the tracking performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is a topological diagram of the double-link robotic arm described in the present invention.

[0052] Figure 2 This is a structural block diagram of the control system of the present invention.

[0053] Figure 3 This is the training graph of the double-Q reinforcement learning algorithm of the present invention.

[0054] Figure 4 It is the trajectory tracking simulation curve of the angular position q of the robot arm links 1 and 2 under the dual Q reinforcement learning method of the present invention.

[0055] Figure 5 is the angular velocity of the robot arm links 1 and 2 under the dual Q reinforcement learning method of the present invention Trajectory tracking simulation curve.

[0056] Figure 6 is the disturbance τ of the robot arm links 1 and 2 under the dual Q reinforcement learning method of the present invention d1 ,τd2 And the simulation curve after observation compensation.

[0057] Figure 7 It is the simulation curve of the control torque τ1, τ2 of the robot arm links 1 and 2 under the double Q reinforcement learning method of the present invention.

[0058] Figure 8 Comparison curve of the dual Q reinforcement learning algorithm designed by the present invention and the traditional Q learning algorithm in terms of Q0 value.

[0059] Fig. 9 Comparison curve between the dual Q reinforcement learning algorithm designed for the present invention and the traditional Q learning algorithm on the learning curve.

[0060] Numbers in the figure: 1—first connecting rod, 2—second connecting rod, 3—first joint, 4—second joint. DETAILED DESCRIPTION

[0061] The present invention will be further described in detail below in conjunction with the accompanying drawings.

[0062] Take a robot with n=2 joints as an example. Figure 1 As shown, the robot arm includes two links: the first link 1 and the second link 2, and two constrained joints: the first joint 3 and the second joint 4. The mass and length of the first link 1 are m1 and l1 respectively, and the mass and length of the second link 2 are m2 and l2 respectively; the moments of inertia of the first link 1 and the second link 2 are J1 and J2 respectively; q1 is the angular position of the first link 1, and q2 is the angular position of the second link 2. The counterclockwise direction is specified as the positive rotation direction of the robot arm.

[0063] The present invention provides a dual-Q reinforcement learning method for robot arm trajectory tracking control containing disturbance, comprising the following steps:

[0064] Step 1: Establish the dynamics model of the robotic arm system:

[0065]

[0066] In the formula, are joint position, joint velocity, and joint acceleration, respectively. is a symmetric positive definite inertia parameter matrix, is the centrifugal force and Coriolis force parameter matrix, is the gravity parameter matrix of the manipulator, is the control torque, is the control input variable, For external disturbance.

[0067]

[0068] In the formula, M0, C0, G0 are scalars of the system model parameter matrix, ΔM, ΔC, ΔG are uncertain terms of the system model parameter matrix. Therefore, the dynamic model (1) of the manipulator can be further expressed as:

[0069]

[0070] In the formula, represents the total unknown disturbance of the robot system, which consists of model uncertainty, load variation, and external disturbance.

[0071] Step 2, let x1 = q, The state space equation of the robotic arm system is obtained as follows:

[0072]

[0073] Step 3: Design a disturbance observer to estimate and compensate for the external disturbances and model parameter uncertainties of the system. The specific implementation steps are as follows:

[0074] 3-1, the disturbance observer of the designed system (4) is as follows:

[0075]

[0076] Where k(x1,x2) is the observation gain matrix, is the observed estimated value of the total disturbance to the system. Since the disturbance observer designed above has the acceleration signal of the state, it is impossible to directly realize the disturbance observation. In the manipulator system, using the velocity signal to obtain the acceleration signal may introduce noise and cause instability. Therefore, the observer needs to be modified as follows.

[0077] 3-2, construct the auxiliary function h(x1,x2), and define the internal state variables of the observer as:

[0078]

[0079] In the formula, is the internal state variable of the observer, and h(x1,x2) is the function vector to be designed. To avoid introducing acceleration signals, the perturbation observation gain matrix k(x1,x2) and h(x1,x2) have the following relationship:

[0080]

[0081] 3-3, design the disturbance observer structure. From equations (5)-(7), we get:

[0082]

[0083] In summary, after improving the disturbance observer (5), the obtained acceleration disturbance-free observer is:

[0084]

[0085] 3-4, design the observer gain and the function to be designed. The observation error of the disturbance observer is defined as:

[0086]

[0087] According to equations (6)-(10), the observer error dynamic equation is:

[0088]

[0089] Since there is no a priori knowledge of the differential of the disturbance in practice, it is assumed that f varies slowly with respect to the disturbance observer, i.e. Then the error dynamic equation is:

[0090]

[0091] In the formula, the nonlinear gain matrix is ​​selected is a reversible matrix, and we can obtain the integral by substituting it into equation (7):

[0092] h(x1,x2)=ρ -1 x2 (13)

[0093] Step 4: Design a tracking controller based on the double-Q reinforcement learning method. The specific method is:

[0094] 4-1. First, according to the characteristics of robot tracking control, the desired tracking trajectory q is given d , define the tracking error as e d =q d -q.

[0095] In order to achieve tracking control of the desired trajectory, the control strategy of the tracking error is designed as follows:

[0096]

[0097] Where λ is a real number greater than 0, u1(t) is the feedback control strategy that can stabilize the robotic arm system, and u2(t) is the approximate optimal control strategy of the double Q reinforcement learning algorithm. The feedback control strategy can be expressed as u1(t)=q d +βe(t), where β is a constant.

[0098] 4-2, define the utility function of the system as

[0099]

[0100] in, are the utility function constants of the state variables and the control variables, respectively. The utility function represents the instantaneous cost of the control variables at the current moment, and U(0,0)=0.

[0101] According to the previous discussion, we hope to find a suitable and feasible control strategy that can minimize the Q function as follows:

[0102]

[0103] Among them, ζ=t,t+1,t+2,... represents time t and any time thereafter, and the Q function is the sum of all time utility functions.

[0104] When the Q function in equation (16) is minimized, it is the optimal Q function Q * (e(t),u2(t)), the corresponding u2(t) is the optimal control strategy. This optimal control strategy can ensure that the error tends to zero while keeping the system stable, thereby achieving the tracking performance of the robot arm.

[0105] According to the Bellman optimality principle, the optimal Q function Q * (e(t),u2(t)) satisfies the following discrete-time HJB equation:

[0106]

[0107] Therefore, we can get the optimal tracking controller based on the double Q reinforcement learning method as:

[0108]

[0109] Double Q reinforcement learning strategy update steps:

[0110]

[0111]

[0112] Double Q reinforcement learning is based on two Q functions: Q1 and Q2. Each Q function updates the Q value of the next state according to the other Q function. The action value a in row 6 of Table 1 * is the action with the greatest value in state s according to the action-value function Q1. However, unlike traditional Q-learning, traditional Q-learning uses the numerical value Q1(s t+1 ,a * )=max a Q1(s t+1 ,a) to update Q1(s,a), we use the value Q2(s t+1 ,a *)=max Q2(s t+1 ,a) to update Q1(s,a). A similar update method is used for Q2, that is, to find t+1 The action b that maximizes Q2 * , then apply b * Get Q1(s t+1 ,b * )=max Q1(s t+1 ,a) to update Q2(s,a). Importantly, the two Q functions are learned from different sets of experience, but two value functions can be used when selecting the action to be performed. Therefore, the data efficiency of this algorithm is not lower than that of traditional Q learning. The following is a further explanation of the present invention by giving the parameters of a two-link robot system.

[0113] Figure 1 The parameters of the robot system shown are as follows: m1 = 0.5 kg, m2 = 1.0 kg, l1 = 1.0 m, l2 = 0.8 m, J1 = 4 kg·m 2 , J2=4kg·m 2 , gravitational acceleration g = 9.807 m / s 2 .

[0114] Based on the above system parameters, other simulation conditions of the system are designed as follows:

[0115] The initial conditions of the robot are q1(0)=1.8, q2(0)=1.1, The given target tracking trajectory is: q d =[q 1d q 2d ] T =[1.25-1.4e -t +0.35e -4t 1.25+e -t +0.25e -4t ] T , where the simulation time is t∈(0,5s).

[0116] According to the above simulation conditions, the system is simulated to verify the trajectory tracking capability of the system.

[0117] The parameter ρ in formula (13) is diag{0.008, 0.0096}, and the parameter in formula (15) is The simulation results are as follows Figures 4 to 9 shown.

[0118] Figure 4The tracking curve of the target trajectory of the first link 1 angular position q1 and the second link 2 angular position q2 of the robot arm under the dual Q reinforcement learning tracking controller. The dotted curve in the figure represents the target angular position trajectory of the given link, and the solid curve represents the actual angular position trajectory of the robot arm link. Figure 4 It can be seen that the control method proposed in the present invention achieves satisfactory tracking performance, ensuring that the actual position q closely follows the target trajectory q d .

[0119] Figure 5 The angular velocity of the first link 1 of the robot arm under the dual Q reinforcement learning tracking controller and the angular velocity of the second link 2 The tracking curve of the target trajectory. The dotted curve in the figure represents the target angular velocity trajectory of the given link, and the solid curve represents the actual angular velocity trajectory of the robot arm link. Figure 5 It can be seen that the control method proposed in the present invention can effectively and accurately achieve the expected trajectory tracking target, and the tracking effect is good.

[0120] Figure 6 is the estimation and compensation curve of the disturbance observer for the uncertain parameters and external disturbances of the system. Figure 6 It can be seen that the disturbance observer designed in the present invention effectively estimates the external disturbance and uncertain parameters of the system and performs real-time compensation for the control input, which greatly improves the control accuracy and tracking effect of the system.

[0121] Figure 7 are the input torques τ1 and τ2 of the first link 1 and the second link 2 of the robot under double Q reinforcement learning control.

[0122] Figure 8 This is the comparison curve of Q0 value between traditional Q learning and double Q reinforcement learning method. Figure 8 As shown, for Double Q Learning, the discounted critical estimate of the long-term reward (2000 episodes) is lower than that of the Q Learning agent. This difference is because the Double Q RL algorithm takes a conservative approach when updating the target, using the minimum of the two Q functions. This difference is further amplified by the delay in the target update. Although Double Q Learning has a lower estimate over these 2000 episodes, the Episode Q0 value of the Double Q Learning agent shows a steady increase, which is different from the Q Learning agent.

[0123] Fig. 9 The learning curves of traditional Q learning and double Q reinforcement learning methods are given by Fig. 9As can be seen, the Q-learning agent seems to learn faster (on average around episode 600) but quickly gets stuck in a local optimum. Double-Q RL starts off slower but ultimately achieves higher rewards than Q-learning because it avoids overestimating the Q-values. The Double-Q RL agent shows a steady improvement in performance over its learning curve, indicating improved stability compared to the Q-learning agent.

[0124] The above results show that the dual-Q reinforcement learning method proposed in the present invention can effectively eliminate the influence of system interference and uncertain parameters under the premise of combining with the disturbance observer, has fast response speed, high tracking accuracy, ideal tracking performance and good control flexibility.

Claims

1. A dual-Q reinforcement learning method for trajectory tracking control of a robotic arm with disturbance, characterized in that: The following steps are involved: Step 1: Establish the dynamics model of the robotic arm system: In the formula, are joint position, joint velocity, and joint acceleration, respectively. is a symmetric positive definite inertia parameter matrix, is the centrifugal force and Coriolis force parameter matrix, is the gravity parameter matrix of the manipulator, is the control torque, is the control input variable, For external disturbance. Where M0, C0, G0 are scalars of the system model parameter matrix, ΔM, ΔC, ΔG are uncertain terms of the system model parameter matrix. Therefore, the dynamic model (1) of the manipulator can be further expressed as In the formula, represents the total unknown disturbance of the robot system, which consists of model uncertainty, load variation, and external disturbance. Step 2, let x1 = q, The state space equation of the robotic arm system is obtained as follows: Step 3: Design a disturbance observer to estimate and compensate for the external disturbances and model parameter uncertainties of the system. The specific implementation steps are as follows: 3-1, the disturbance observer of the designed system (4) is as follows: Where k(x1,x2) is the observation gain matrix, is the observed estimated value of the total disturbance to the system. Since the disturbance observer designed above has the acceleration signal of the state, it is impossible to directly realize the disturbance observation. In the manipulator system, using the velocity signal to obtain the acceleration signal may introduce noise and cause instability. Therefore, the observer needs to be modified as follows. 3-2, construct the auxiliary function h(x1,x2), and define the internal state variables of the observer as: In the formula, is the internal state variable of the observer, and h(x1,x2) is the function vector to be designed. To avoid introducing acceleration signals, the perturbation observation gain matrix k(x1,x2) and h(x1,x2) have the following relationship: 3-3, design the disturbance observer structure. From equations (5)-(7), we get: In summary, after improving the disturbance observer (5), the obtained acceleration disturbance observer is Step 4: Design a tracking controller based on the double-Q reinforcement learning method. The specific method is: 4-1. First, according to the characteristics of robot tracking control, the desired tracking trajectory q is given d , define the tracking error as e d =q d -q. In order to achieve tracking control of the desired trajectory, the control strategy of the tracking error is designed as follows: Where λ is a real number greater than 0, u1(t) is the feedback control strategy that can stabilize the robotic arm system, and u2(t) is the approximate optimal control strategy of the double Q reinforcement learning algorithm. The feedback control strategy can be expressed as u1(t)=q d +βe(t), where β is a constant. 4-2, define the utility function of the system as In the formula, Q and R are the utility function constants of the state variable and the control variable respectively. The utility function represents the instantaneous cost of the control variable at the current moment, and U(0,0)=0. According to the previous discussion, we hope to find a suitable and feasible control strategy that can minimize the Q function as follows: in, ζ=t,t+1,t+2,... represents time t and any time thereafter, and the Q function is the sum of all time utility functions. When the Q function in equation (12) is minimized, it is the optimal Q function Q * (e(t),u2(t)), the corresponding u2(t) is the optimal control strategy. This optimal control strategy can ensure that the error tends to zero while keeping the system stable, thereby achieving the tracking performance of the robot arm. According to the Bellman optimality principle, the optimal Q function Q * (e(t),u2(y)) satisfies the following discrete-time HJB equation: Therefore, we can get the optimal tracking controller based on the double Q reinforcement learning method as: Double Q reinforcement learning strategy update steps: Double Q reinforcement learning is based on two Q functions: Q1 and Q2. Each Q function updates the Q value of the next state according to the other Q function. The action value a in row 6 of Table 1 * is the action with the maximum value in state s according to the action-value function Q1. However, unlike traditional Q-learning, which uses the numerical value Q1(s t+1 ,a * )=max a Q1(s t+1 ,a) to update Q1(s,a), we use the value Q2(s t+1 ,a * )=max Q2(s t+1 ,a) to update Q1(s,a). A similar update method is used for Q2, that is, to find t+1 The action b that maximizes Q2 * , then apply b * Get Q1(s t+1 ,b * )=max Q1(s t+1 ,a) to update Q2(s,a). Importantly, the two Q functions are learned from different sets of experience, but both value functions can be used when choosing an action to perform. Therefore, this algorithm is no less data efficient than traditional Q-learning.

Citation Information

Cited By

  • Robot action decision-making method and device based on spatial relationship, equipment and medium

    CN120620218A

  • Six-degree-of-freedom mechanical arm disturbance compensation control method based on iterative learning observer

    CN120921362A