TD3 compensation control method for uncertain AUV trajectory tracking
By constructing the AUV dynamics model and the compensation control method of the TD3 algorithm, the modeling error and random interference problems in AUV trajectory tracking are solved, and the precise trajectory tracking is achieved in complex environments, improving the stability and performance of the controller.
Patent Information
- Application Number
- CN202510633248.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-26
AI Technical Summary
The existing AUV trajectory tracking control methods have shortcomings in dealing with problems such as underdrive, incompleteness, strong coupling, modeling errors and random interference, and it is difficult to achieve accurate trajectory tracking.
The dynamic model of the target AUV was constructed, and the control quantity of the nominal model was designed using the backward method and the Liyapunov function, and the modeling error and random interference were used to process the final control quantity, and the compensation control was performed through deep reinforcement learning technology.
Effectively reduce the impact of modeling errors and random perturbations on AUV precise trajectory tracking, improve the robustness and tracking accuracy of the controller, and is suitable for complex AUV physical characteristics environments.
Smart Images

Figure CN120540356A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot control technology, and in particular to a TD3 compensation control method for uncertain AUV trajectory tracking. Background Art
[0002] With the continuous development and utilization of marine resources, autonomous underwater vehicles (AUVs) have gained widespread attention and application in fields such as ocean mapping, energy exploration, military reconnaissance, and marine environmental monitoring. As an emerging ocean exploration tool, AUV trajectory tracking control technology is key to achieving these functions and is directly related to the quality and efficiency of mission completion.
[0003] In practical AUV trajectory tracking applications, numerous challenges and problems exist. For example, the underactuated nature of AUVs prevents them from achieving desired motion directly through control inputs; nonholonomic constraints limit the flexibility of control inputs; and strong coupling causes the motions of various degrees of freedom to interact, increasing control complexity. Furthermore, modeling errors and random interference are practical issues facing AUV control systems. These factors combine to create significant challenges for accurate AUV trajectory tracking.
[0004] To address these challenges, numerous scholars have conducted in-depth research. Early studies focused primarily on traditional control methods, such as sliding mode control and adaptive control. Sliding mode control offers strong robustness and can, to a certain extent, address inaccuracies in system model parameters and random disturbances. However, its control rate is not smooth enough, potentially causing problems such as chattering, which can affect practical applications. Adaptive control methods compensate for uncertainty through online estimation of system parameters, but they still have certain limitations when dealing with the complex dynamic characteristics of AUVs.
[0005] In recent years, deep reinforcement learning (DRL) technology has made significant progress in the field of intelligent robot control. Combining the strengths of deep learning and reinforcement learning, DRL can learn optimal control strategies through interaction with the environment, without requiring a precise system model. For example, neural network-based reinforcement learning algorithms have been developed for trajectory tracking of underactuated autonomous underwater vehicles (AUVs), and policy gradient-based RL has been successfully implemented for underwater cable tracking. While these methods have improved AUV trajectory tracking performance to some extent, they still face challenges such as high training costs and difficult convergence.
[0006] In summary, existing AUV trajectory tracking control methods still have shortcomings when dealing with complex practical problems and require further research and improvement. Therefore, developing a control method that can effectively deal with problems such as underactuation, non-holonomic issues, strong coupling, modeling errors, and random interference in AUV systems has important theoretical significance and practical application value. Summary of the Invention
[0007] The purpose of the present invention is to address the above technical problems and provide a TD3 compensation control method for uncertain AUV trajectory tracking, which can effectively reduce the impact of modeling errors and random disturbances on the precise trajectory tracking of AUV.
[0008] To achieve the above object, the present invention provides the following solutions:
[0009] A TD3 compensation control method for uncertain AUV trajectory tracking, comprising:
[0010] Construct a dynamic model of the target AUV;
[0011] Based on the dynamic model, the control variables of the nominal model are designed using the backward stepping method and Lyapunov function;
[0012] The TD3 algorithm is used to process modeling errors and random interference, obtain the final control quantity, and complete the compensation control of the target AUV trajectory tracking.
[0013] Optionally, the dynamic model of the target AUV is constructed as follows:
[0014]
[0015] where η = [x, y, ψ] T is the position information, v=[u,v,r] T is the speed information, is the derivative of position information with respect to time, is the time derivative of velocity information, J(η) is the rotation operator, M is the inertia matrix, C(v) is the Coriolis force matrix, D(v) is the water damping term matrix, τ is the control input, and d is the environmental disturbance force.
[0016] Optionally, based on the dynamic model, the control quantity of the nominal model is designed using the backward step method and the Lyapunov function, including:
[0017] Constructing a control target and a nominal model based on the dynamic model, and designing a virtual control variable for the nominal model;
[0018] A nominal control error is determined based on the virtual control variable, and a corresponding nominal control variable, ie, the control variable of the nominal model, is designed.
[0019] Optionally, the control target is:
[0020]
[0021] The nominal model is:
[0022]
[0023] The virtual control amount is:
[0024]
[0025] in, is the tracking error, t is the current moment, is the derivative of the nominal position information with respect to time, is the nominal model rotation operator matrix, is the nominal position information, is the nominal speed information, is the nominal model inertia matrix, is the time derivative of the nominal velocity information, is the nominal model Coriolis force matrix, is the nominal model water damping matrix, is the nominal control quantity, is the virtual control quantity, is the inverse matrix of the nominal model rotation operator, is the time derivative of the desired trajectory, is the nominal position information control error, k1 and k2 are diagonal matrices with positive diagonal elements.
[0026] Optionally, the nominal control error is:
[0027]
[0028] The nominal control quantity is:
[0029]
[0030] in, is the nominal position information control error, t is the current moment, is the nominal position information, η d is the expected trajectory, is the nominal speed information control error, is the nominal speed information, is the virtual control quantity, is the nominal control quantity, is the nominal model inertia matrix, is the time derivative of the virtual control quantity, is the nominal model Coriolis force matrix, is the nominal model water damping matrix.
[0031] Optionally, the TD3 algorithm is used to process modeling errors and random disturbances to obtain the final control quantity, including:
[0032] S1. Set the state value of the target AUV at different times, the input value and reward value of the TD3 algorithm, the action value function, the critic network and the action network;
[0033] S2, select any sequence from the experience replay and calculate the output of the target network;
[0034] S3. Define a temporal difference error, and update parameters of the critic network and the action network based on the temporal difference error;
[0035] S4. Update the parameters of the target network through a soft update mechanism;
[0036] S5. Repeat S1-S4 until the preset stop condition is reached.
[0037] Optionally, the timing differential error is:
[0038]
[0039] Where i represents the number of iterative learning, r(k) represents the reward value of the target AUV at time a(k), represents the output of the j-th target criticism network, s(k+1) represents the state value of the target AUV at time k+1, represents the output of the target network at time k+1, w′ j (i), represents the parameters of the j-th target critic network, represents the output of critic network 1, represents the output of the critic network 2, s(k) represents the state value of the target AUV at time k, represents the output of the target network at time k, w1(i) and w2(i) are the parameters of the critic network 1 and the critic network 2 respectively, δ 1,i , δ 2,i is the time series difference error, and γ is a positive constant.
[0040] Optionally, the parameters of the critic network and the action network are updated as follows:
[0041]
[0042] Among them, w 1,i+i represents the updated target critic network 1 parameters, w 1,i represents the target criticism network 1 parameter before update, w 2,i+1 represents the updated target critic network 2 parameters, w 2,i represents the target criticism network 2 parameters before update, θ i+1 represents the updated target action network parameters, θ i represents the target action network parameters before update, μ(s k ,θk ) represents the output of the action network at time k, w' i For smaller value or The value parameter, To find the sign of the partial derivative, α w , α θ are the learning rates of the target critic network and the target action network, respectively. They are the outputs of target-critic network 1 before update, target-critic network 2 before update, and target-action network before update, respectively.
[0043] Optionally, the parameters of the target network are updated as follows:
[0044]
[0045] Among them, w′ 1,i+1 represents the target criticism network 1 parameter after soft update, w′ 1,i represents the target criticism network 1 parameter before update, w′ 1,i represents the parameters of the critic network 1 before update, w′ 2,i+1 represents the target criticism network 2 parameters after soft update, w′ 2,i represents the target critic network 2 parameters before update, w′ 2,i represents the parameters of the critic network 2 before updating, θ′ i+1 represents the target action network parameters after soft update, θ i represents the target action network parameters before update, θ′ i represents the action network parameters before updating, and ξ is the soft update parameter.
[0046] The beneficial effects of the present invention are:
[0047] The present invention first establishes an ideal model, or nominal model, based on the dynamics of an incompletely underactuated, strongly coupled AUV. The control rate of the nominal model is designed using the backward step method and Lyapunov function. Deep reinforcement learning techniques are then used to address the modeling errors and random disturbances of the system model during actual operation, and a compensation method based on a double-delayed deep deterministic policy gradient (TD3) is proposed. The anti-interference control method designed in this invention can effectively reduce the impact of modeling errors and random disturbances on the AUV's precise trajectory tracking, while taking into account the complex physical characteristics of the AUV, such as underactuation, non-holonomy, and strong coupling, and has high practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] Figure 1 This is a flow chart of a TD3 compensation control method for uncertain AUV trajectory tracking according to an embodiment of the present invention;
[0050] Figure 2 Schematic diagram of the AUV horizontal coordinate system according to an embodiment of the present invention;
[0051] Figure 3 TD3 compensation control algorithm flow chart of an embodiment of the present invention;
[0052] Figure 4 This is an AUV trajectory tracking component diagram according to an embodiment of the present invention;
[0053] Figure 5 This is a diagram of the AUV trajectory tracking error according to an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0055] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0056] This embodiment provides a TD3 compensation control method for uncertain AUV trajectory tracking, including:
[0057] Construct a dynamic model of the target AUV;
[0058] Based on the dynamic model, the control variables of the nominal model are designed using the backward stepping method and Lyapunov function;
[0059] The TD3 algorithm is used to process modeling errors and random interference, obtain the final control quantity, and complete the compensation control of the target AUV trajectory tracking.
[0060] Specifically, this embodiment first establishes an ideal model, namely the nominal model, based on the dynamics of an incompletely underactuated, strongly coupled AUV. The control rate of the nominal model is designed based on the backward step method and the Lyapunov function. Deep reinforcement learning technology is then used to address the modeling errors and random disturbances of the system model in actual operation, and a compensation method based on double-delayed deep deterministic policy gradient (TD3) is proposed. The anti-interference control method designed in this invention can effectively reduce the impact of modeling errors and random disturbances on the AUV's precise trajectory tracking while taking into account the complex physical characteristics of the AUV, such as underactuation, non-holonomy, and strong coupling, and has high practical application value.
[0061] The following combination Figure 1 、 Figure 2 、 Figure 3 The TD3 compensation control method for uncertain AUV trajectory tracking provided in this embodiment is described in detail. Figure 1 As shown, the specific steps include:
[0062] Step 1: Establish a dynamic model of the underactuated AUV;
[0063] AUVs are usually considered as rigid bodies with six degrees of freedom moving in three-dimensional space with strong coupling characteristics. To simplify this coupling characteristic, the three-dimensional motion of AUVs is usually decomposed into two three-dimensional motions in the horizontal plane and the vertical plane. Research usually focuses on the horizontal motion of underactuated AUVs, such as Figure 2 As shown, the AUV dynamic model is established as follows:
[0064]
[0065] where η = [x, y, ψ] T is the position information, v=[u,v,r] T is the velocity information, η is the derivative of the position information with respect to time, is the time derivative of velocity information, J(η) is the rotation operator, M=diag(M x ,M y ,M ψ ) is the inertia matrix, C(v) is the Coriolis force matrix, D(v)=diag(X u ,Y v ,N r )+diag(D u |u|,D v |v|,D r |r) is the water damping matrix, τ=[F u ,F v ,F r ] T is the control input, d=[d u,d v ,d r ] T It is the environmental disturbance force. It includes:
[0066]
[0067] In actual scenarios, the precise dynamic characteristics of an AUV are usually difficult to obtain. Therefore, the parameter matrix in the AUV dynamics model can be decomposed into two parts:
[0068]
[0069] Among them, M, D(v), and C(v) are the precise parts of the parameters, and δM, δD(v), and δC(v) are the imprecise parts of the parameters.
[0070] Step 2: nominal model controller design and stability verification;
[0071] Based on the dynamic model, the control target and nominal model are constructed, and a virtual control variable is designed for the nominal model. Based on the virtual control variable, the nominal control error is determined, and the corresponding nominal control variable, i.e., the control variable of the nominal model, is designed. Specifically, the following are involved:
[0072] 2.1, assuming the desired trajectory is η d =[x d ,y d ,ψ d ] T , the actual trajectory is η, then the tracking error is defined as η(t) = η d (t)-η(t), the control target is:
[0073]
[0074] 2.2, in order to reduce the control complexity, the AUV dynamic model ignores uncertainty interference such as modeling error and random disturbance, and thus defines the nominal model:
[0075]
[0076] in, is the nominal position information, is the nominal speed information, is the derivative of the nominal position information with respect to time, is the time derivative of the nominal velocity information, is the nominal model rotation operator matrix, is the nominal model inertia matrix, is the time derivative of the nominal velocity information, is the nominal model Coriolis force matrix, is the nominal model water damping matrix, For control input.
[0077] 2.3, considering the strong coupling characteristics of its model, a virtual control rate is designed for Equation (6) to achieve hierarchical control of the speed and position of the AUV and reduce the control complexity. The virtual control quantity is defined as:
[0078]
[0079] in, is the virtual control quantity, is the inverse matrix of the nominal model rotation operator, is the time derivative of the desired trajectory, is the nominal position information control error, k1 and k2 are diagonal matrices with positive diagonal elements.
[0080] The nominal control error is defined as:
[0081]
[0082] in, is the nominal position information control error, is the nominal speed information control error.
[0083] 2.4, design the nominal AUV model control quantity as:
[0084]
[0085] 2.5, design the Liapunov function as:
[0086]
[0087] 2.6, to prove the stability of the controller, the Liapunov function is derived with respect to time as follows:
[0088]
[0089] Since the above equation shows For any small λ(T)>0, when t1>T, there exists For t>T. Therefore, for t>T, Rewritten as:
[0090]
[0091] By choosing large enough k1 and k2, and small enough λ, we can deduce Therefore, the control error of the nominal model converges asymptotically and the stability of the controller is proved.
[0092] Step 3, TD3 compensator design, see Figure 3 ;
[0093] 3.1, define s(k) and s(k+1) as the state values of AUV at time k and k+1, a(k) as the input value of TD3 compensator to AUV system at time k, r(k) as the reward value of AUV system under the action of time a(k), and define the difference between two adjacent moments as the algorithm sampling time T. Where s={xd,yd,ψ d ,x,y,ψ,u,v,r} and a=δτ and Defining the action-value function Used to determine the action strategy a(k) = μ(s). The critic network (CN) is used to estimate Q(s,a) to Q(s,a,w), and the actor network (AN) is used to estimate μ(s) to μ(s k ,θ k ).
[0094] 3.2, arbitrarily select a sequence (s(k), a(k), r(k), s(k+1)) from the experience replay and define the estimated output as where noise is the exploration noise of the action network (AN).
[0095] 3.3, let the outputs of target criticism network 1 and target criticism network 2 be and Similarly, the outputs of critic network 1 and critic network 2 are and The output of the target action network is μ′(s,θ). Therefore, the temporal difference (TD) error is defined as:
[0096]
[0097] Where i represents the number of iterative learning, r(k) represents the reward value of the target AUV at time a(k), represents the output of the j-th target criticism network, s(k+1) represents the state value of the target AUV at time k+1, represents the output of the target network at time k+1, represents the output of critic network 1, represents the output of the critic network 2, s(k) represents the state value of the target AUV at time k, represents the output of the target network at time k, δ 1,i , δ 2,i is the time difference (TD) error, γ is a positive constant, w′ j (i), w1(i), w2(i) are the parameters of the corresponding network.
[0098] 3.4, the parameters of the target criticism network CN1, CN2 and AN are updated as follows:
[0099]
[0100] Among them, w 1,i+1 represents the updated target critic network 1 parameters, w 1,i represents the target criticism network 1 parameter before update, w 2,i+1 represents the updated target critic network 2 parameters, w 2,i represents the target criticism network 2 parameters before update, θ i+1 represents the updated target action network parameters, θ i represents the target action network parameters before update, μ(s k ,θ k ) represents the output of the action network at time k. w' i corresponds to a smaller or The value parameter, To find the sign of the partial derivative, α w , α θ is the learning rate of the corresponding network, is the Q value of the corresponding network.
[0101] 3.5, the parameters of the target critic network (CN) and the target action network (AN) are updated through a soft update mechanism, which can be expressed mathematically as follows:
[0102]
[0103] Among them, w′ 1,i+1 represents the target criticism network 1 parameter after soft update, w 1,i represents the target criticism network 1 parameter before update, w′ 1,i represents the parameters of the critic network 1 before update, w′ 2,i+1 represents the target criticism network 2 parameters after soft update, w 2,i represents the target critic network 2 parameters before update, w′ 2,i represents the parameters of the critic network 2 before updating, θ′ i+1 represents the target action network parameters after soft update, θ i represents the target action network parameters before update, θ′ i represents the action network parameters before updating, and ξ is the soft update parameter.
[0104] 3.6, increment i by 1 and return to step 3.1 until the predetermined number of iterations is reached or the specified stopping condition is met.
[0105] The AUV dynamics model developed in this example comprehensively captures the inherently complex physical characteristics of AUVs, including insufficient maneuvers, system incompleteness, and strong nonlinear coupling. Furthermore, the model employs a robust framework that explicitly accounts for inaccuracies in system modeling and the random environmental perturbations encountered during actual AUV operation. This comprehensive modeling approach not only preserves the essential dynamic characteristics of the AUV system but also faithfully reproduces the challenges faced in real-world underwater environments, laying a solid foundation for subsequent controller design and performance evaluation.
[0106] The controller proposed in this example exhibits robust stability, enabling the system to effectively perform tracking tasks while inherently compensating for modeling inaccuracies and environmental disturbances. This stability characteristic not only ensures reliable performance under ideal conditions but also addresses fundamental control challenges associated with AUV dynamics, including insufficient action, system incompleteness, and strong nonlinear coupling. The controller's inherent robustness effectively mitigates the impact of these complex dynamic characteristics, providing a comprehensive solution to long-standing control challenges in AUV operation. This achievement simultaneously addresses multiple technical challenges while maintaining system stability and tracking accuracy, representing a significant advancement in the field of underwater vehicle control.
[0107] This example leverages the inherent convergence properties of the TD3 algorithm to develop a compensator that theoretically ensures higher control performance through systematic iterative improvements. In actual operation, the compensator continuously cycles through data acquisition, processing, and parameter optimization. The TD3 algorithm's dual-network architecture, coupled with its delayed policy update mechanism and target policy smoothing regularization, ensures robust learning from accumulated operational data while effectively mitigating overestimation bias in the value function approximation.
[0108] The following is a further explanation and verification of the TD3 compensation control method for uncertain AUV trajectory tracking proposed in this embodiment through experimental simulation:
[0109] In order to verify the effectiveness of the proposed AUV trajectory tracking control algorithm, numerical simulation was performed using Matlab. The parameters of the nominal model are set as: M x =216.6, M y =593.2, M ψ =28.7,X u =26.9, Y v =35.8, N r =3.5, D u =100.3, D v =503.8, D r=76.9. The model parameters are perturbed between 0.1 and 0.2, and the random perturbation is set to d = [200sin(t), 600cos(t), 10sin(t) + 10cos(t)] T The initial value of the AUV is set to η(0) = [0,0,0] T ,v(0)=[0,0,0] T , and set the AUV’s trajectory to track the desired trajectory as η d (t)=[20cos(0.1t),20sin(0.1t),0.1t] T Set the controller parameters to k1=diag([1,1,1]), k2=diag([0.01,0.01,0.01]), and the criticism network parameter learning rate to α w =0.001, the learning rate of the action network parameter is α θ =0.005, the experience pool size is 107, and the mini-batch sampling size is 256.
[0110] L represents the trajectory tracking without adding TD3 compensator, LRL represents the trajectory tracking with adding TD3 compensator, Figure 4 and Figure 5 It can be seen that after adding the TD3 compensator, the tracking effect of the nominal model controller is significantly better than that without the TD3 compensator. The specific performance is shown in Table 1.
[0111] Table 1
[0112]
[0113] After adding the TD3 compensator, the root mean square error in all directions is smaller than before. Simulation experiments show that the anti-interference control method designed in this paper can effectively reduce the impact of modeling errors and random disturbances on the precise trajectory tracking of autonomous underwater vehicles, fully considering the complex physical characteristics of AUVs, such as underactuation, incompleteness, and strong coupling.
[0114] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A TD3 compensation control method for uncertain AUV trajectory tracking, characterized by: include: Construct a dynamic model of the target AUV; Based on the dynamic model, the control variables of the nominal model are designed using the backward stepping method and Lyapunov function; The TD3 algorithm is used to process modeling errors and random interference, obtain the final control quantity, and complete the compensation control of the target AUV trajectory tracking.
2. The TD3 compensation control method for uncertain AUV trajectory tracking according to claim 1, characterized in that: The dynamic model of the target AUV is constructed as follows: where η = [x, y, ψ] T is the position information, v=[u,v,r] T is the speed information, is the derivative of position information with respect to time, is the time derivative of velocity information, J(η) is the rotation operator, M is the inertia matrix, C(v) is the Coriolis force matrix, D(v) is the water damping term matrix, τ is the control input, and d is the environmental disturbance force.
3. The TD3 compensation control method for uncertain AUV trajectory tracking according to claim 1, characterized in that: Based on the dynamic model, the control variables of the nominal model are designed using the backward step method and Lyapunov function, including: Constructing a control target and a nominal model based on the dynamic model, and designing a virtual control variable for the nominal model; A nominal control error is determined based on the virtual control variable, and a corresponding nominal control variable, ie, the control variable of the nominal model, is designed.
4. The TD3 compensation control method for uncertain AUV trajectory tracking according to claim 3, characterized in that: The control objectives are: The nominal model is: The virtual control amount is: in, is the tracking error, t is the current moment, is the derivative of the nominal position information with respect to time, is the nominal model rotation operator matrix, is the nominal position information, is the nominal speed information, is the nominal model inertia matrix, is the time derivative of the nominal velocity information, is the nominal model Coriolis force matrix, is the nominal model water damping matrix, is the nominal control quantity, is the virtual control quantity, is the inverse matrix of the nominal model rotation operator, is the time derivative of the desired trajectory, is the nominal position information control error, k1 and k2 are diagonal matrices with positive diagonal elements.
5. The TD3 compensation control method for uncertain AUV trajectory tracking according to claim 3, characterized in that: The nominal control error is: The nominal control quantity is: in, is the nominal position information control error, t is the current moment, is the nominal position information, η d is the expected trajectory, is the nominal speed information control error, is the nominal speed information, is the virtual control quantity, is the nominal control quantity, is the nominal model inertia matrix, is the time derivative of the virtual control quantity, is the nominal model Coriolis force matrix, is the nominal model water damping matrix.
6. The TD3 compensation control method for uncertain AUV trajectory tracking according to claim 1, characterized in that: The TD3 algorithm is used to process modeling errors and random disturbances to obtain the final control quantities, including: S1. Set the state value of the target AUV at different times, the input value and reward value of the TD3 algorithm, the action value function, the critic network and the action network; S2, select any sequence from the experience replay and calculate the output of the target network; S3. Define a temporal difference error, and update parameters of the critic network and the action network based on the temporal difference error; S4. Update the parameters of the target network through a soft update mechanism; S5. Repeat S1-S4 until the preset stop condition is reached.
7. The TD3 compensation control method for uncertain AUV trajectory tracking according to claim 6, characterized in that: The timing differential error is: Where i represents the number of iterative learning, r(k) represents the reward value of the target AUV at time a(k), represents the output of the j-th target criticism network, s(k+1) represents the state value of the target AUV at time k+1, represents the output of the target network at time k+1, w′ j (i), represents the parameters of the j-th target critic network, represents the output of critic network 1, represents the output of the critic network 2, s(k) represents the state value of the target AUV at time k, represents the output of the target network at time k, w1(i) and w2(i) are the parameters of the critic network 1 and the critic network 2 respectively, δ 1,i , δ 2,i is the time series difference error, and γ is a positive constant.
8. The TD3 compensation control method for uncertain AUV trajectory tracking according to claim 6, characterized in that: Update the parameters of the critic network and action network as follows: Among them, w 1,i+1 represents the updated target critic network 1 parameters, w 1,i represents the target criticism network 1 parameter before update, w 2,i+1 represents the updated target critic network 2 parameters, w 2,i represents the target critic network 2 parameters before update, θ i+1 represents the updated target action network parameters, θ i represents the target action network parameters before update, μ(s k ,θ k ) represents the output of the action network at time k, w' i For smaller value or The value parameter, To find the sign of the partial derivative, α w , α θ are the learning rates of the target critic network and the target action network, respectively. They are the outputs of target-critic network 1 before update, target-critic network 2 before update, and target-action network before update, respectively.
9. The TD3 compensation control method for uncertain AUV trajectory tracking according to claim 6, characterized in that: Update the parameters of the target network to: Among them, w′ 1,i+1 represents the target criticism network 1 parameter after soft update, w 1,i represents the target criticism network 1 parameter before update, w′ 1,i represents the parameters of the critic network 1 before update, w′ 2,i+1 represents the target criticism network 2 parameters after soft update, w 2,i represents the target critic network 2 parameters before update, w′ 2,i represents the parameters of the critic network 2 before updating, θ′ i+1 represents the target action network parameters after soft update, θ i represents the target action network parameters before update, θ′ i represents the action network parameters before updating, and ξ is the soft update parameter.