An unmanned aerial vehicle trajectory tracking method based on reinforcement learning and event triggering
By designing a state event triggering mechanism and neural network based on reinforcement learning and event triggering methods, the problem of high actuator movement frequency of unmanned aerial vehicles during long-term airborne stays was solved, and accurate trajectory tracking and life extension were achieved.
Patent Information
- Application Number
- CN202211475778.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-23
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2042-11-23
AI Technical Summary
When unmanned aerial vehicles are stationed in the air for a long time, the actuators operate at a high frequency, resulting in shortened lifespan and excessive communication burden, which is especially prominent when facing unknown disturbances and external interference.
A method based on reinforcement learning and event triggering is used to design a suitable state event triggering mechanism, which is combined with the actuator and evaluator neural networks to fit external disturbances and dynamic coupling to achieve optimal control.
It reduces the actuator action frequency, reduces the communication burden, extends the actuator life, and achieves accurate trajectory tracking under unknown disturbances.
Smart Images

Figure CN116300991B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application provides a trajectory tracking method for unmanned aerial vehicles based on reinforcement learning and event triggering, provides a control method for suppressing external disturbances to achieve optimal control and event triggering for unmanned aerial vehicle systems considering unknown disturbances, and belongs to the technical field of automatic control of unmanned aerial vehicles. BACKGROUND
[0002] Unmanned aerial vehicles are also permeated into all aspects of life, and are widely used in the fields of military and police for tasks such as aerial patrol, intelligence warning and battlefield reconnaissance, and also have outstanding performance in disaster relief and disaster prevention, surveying and mapping and other commercial fields. As a high-tech industry, unmanned aerial vehicles have great application potential in the military and civilian industries, and have become a hot research project in recent years and have developed rapidly. In practical applications, the communication bandwidth on board is often very limited, and it is necessary to consider reducing the communication amount as much as possible. One solution is to introduce an event-driven mechanism. Long-time hovering unmanned aerial vehicles have the need to reduce the action frequency of actuators, and this need is realized on the premise of completing flight tasks, so as to prolong the service life of the actuators.
[0003] The application "a trajectory tracking method for unmanned aerial vehicles based on reinforcement learning and event triggering" takes the above problems as the breakthrough point, and realizes optimal control based on the method of reinforcement learning and event triggering, solving the problems of actuator wear and tear and communication burden of long-time hovering unmanned aerial vehicles under parameter uncertainty and external disturbance. SUMMARY
[0004] The application aims to provide a trajectory tracking method for unmanned aerial vehicles based on reinforcement learning and event triggering, which realizes accurate tracking and reduces the action frequency of actuators.
[0005] Technical scheme: the application "a trajectory tracking method for unmanned aerial vehicles based on reinforcement learning and event triggering", the main content and steps are: determining the tracking error of the unmanned aerial vehicle according to the given desired trajectory, formulating the tracking problem, designing a suitable state-based event triggering mechanism condition, using a neural network to construct the actuator network and evaluator network in reinforcement learning, the actuator neural network is used to fit the external disturbance and dynamic coupling in the model, and the evaluator is used to fit the traditional long-term performance index function, and then the tracking control law of the system is obtained.
[0006] The application "a trajectory tracking method for unmanned aerial vehicles based on reinforcement learning and event triggering", the specific steps are as follows:
[0007] Step one: establish a six-degree-of-freedom model of the unmanned aerial vehicle, give a desired tracking trajectory, and obtain the desired position, attitude angle, speed and angular velocity according to the given trajectory.
[0008] Step two trajectory tracking error calculation: calculate the error between the actual trajectory and the desired trajectory, define the filter tracking error, and derive it.
[0009] Step three event-triggered mechanism design for input, design state quantity, and design reasonable event-triggered parameters.
[0010] Step four evaluator neural network design, design the tracking performance of the evaluation system, and use it for tracking performance improvement.
[0011] Step five actuator neural network design, approximation of uncertain terms in the model.
[0012] Step six model nonlinear control law calculation combined with neural network: calculate the control amount τ needed to eliminate the error between the desired trajectory and the actual trajectory.
[0013] In step one, the model establishment process is as follows:
[0014] First, establish the inertial coordinate system and the body coordinate system; thus, the six-degree-of-freedom nonlinear motion equation of the unmanned aerial vehicle can be obtained:
[0015]
[0016] The position coordinates of the unmanned aerial vehicle P = [x, y, z] T The coordinates of the origin of the unmanned aerial vehicle body coordinate system in the inertial coordinate system. Attitude angle Θ = [φ, θ, ψ] T The angle between the unmanned aerial vehicle body coordinate system and the inertial coordinate system. Velocity vector v = [u, v, w] T The velocity components of the unmanned aerial vehicle ground speed in the body coordinate system along the Ox, Oy, and Oz axes. Angular velocity vector Ω = [p, q, r] T The angular velocity components of the unmanned aerial vehicle in the body coordinate system around the Ox, Oy, and Oz axes. F v ,F ω ,B 12 ,B 11 ,B 22 ,B 21 f v and f ω represent the functions related to the state equation, τ = [τ v ,τ ω ] T = [τ u ,τ v ,τ w ,τ p ,τ q ,τ r ] T represent the input of the system. Define X = [P, Θ]T = [x, y, z, φ, θ, ψ] T represents the system output, X d = [P d , Θ d ] T = [x d , y d , z d , φ d , θ d , ψ d ] T represents the desired trajectory; is the derivative of the position coordinate, is the derivative of the velocity component, is the derivative of the angle component, is the derivative of the angular velocity, K is the rotation matrix of the body coordinate system to the inertial coordinate system, and R is the inverse angular velocity decomposition projection transformation matrix.
[0017] The trajectory curve is P d = [x d , y d , z d ] T is the position coordinate of the desired trajectory curve, is the derivative of the desired position coordinate, and the desired attitude angle calculation formula is
[0018] represents the derivative of the desired position coordinate in the x-axis direction, represents the derivative of the desired position coordinate in the y-axis direction, represents the derivative of the desired position coordinate in the z-axis direction.
[0019] The trajectory tracking error calculation method and derivative calculation of step two are as follows:
[0020] Define the tracking error e = X - X d , and define the filtered tracking error as: where K e is a symmetric positive definite matrix, is the derivative of the tracking error. Since the first and second derivatives of the desired trajectory are quite cumbersome and complex to solve, an instruction filter is introduced to estimate its value. A second-order instruction filter is used, which can quickly track the desired signal.
[0021] The following expressions can be obtained by differentiation:
[0022]
[0023] where, k vB is a positive definite constant diagonal matrix, k ω B is a positive definite constant diagonal matrix, B v0 = KB 11 B is a positive definite constant diagonal matrix, B ω0 = RB 22 B is a positive definite constant diagonal matrix, B are the second and first derivatives of the state estimation of the command filter with respect to position and attitude, and are the internal coupling, external disturbance and prediction error in position and attitude. and are the derivatives of the transformation matrices K and R.
[0024] wherein the event-triggered mechanism in step three is designed as follows:
[0025] The control quantity of the system depends on s, e, X d and wherein is the derivative of the desired trajectory X d . Define the comprehensive error state vector then the input can be converted from a function of time to a state-based representation, i.e. The controller is updated at non-periodic sampling time affected by the event-triggered condition, and the update will depend on the state ξ.
[0026] The measurement error is defined as follows:
[0027] e ξ = ξ(t j )-ξ,t∈(t j ,t j+1 ]
[0028] The designed event-triggered mechanism is as follows:
[0029]
[0030] t k+1 = inf{t>t k :||e ξ ||≥k s ‖s‖and‖s‖≥r s >0}
[0031] wherein u(t) is the real controller input, when the event-triggered error exceeds the event-triggering threshold, the current instantaneous state is the sampling state, which is passed to the controller, and the event-triggered error can be expressed as the event-triggering condition. When not triggered, the controller retains the state at the last time under the action of the zero-order holder t∈(t j ,t j+1 ]. wherein k s , rs Control parameters for the event-triggered design.
[0032] The evaluator neural network in step four is designed as follows:
[0033] The event-triggered adaptive controller only updates at t j , and (t j , t j+1 ] is the original value of the last time. The design function p(t j ) is as follows:
[0034]
[0035] Where c p is the preset threshold vector of s, s i (t j ) is the value of the i-th item of the filtered tracking error s at the t j -th time. p i (t j ) = 0 means good tracking performance, and p i (t j ) = 1 represents poor tracking performance. Unlike the time-based update, the p i (t j ) function also only updates at the jump time t j , rather than updating according to fixed time sampling. The long-term performance index function Q(t j ) is designed as follows:
[0036] Q(t j ) = a N p(t j+1 ) + a N-1 p(t j+2 ) + … + a j+1 p(t N )
[0037] Where 0 < a < 1 is the discount rate to be designed, and N is a preset positive integer large enough. The control objective is to make the tracking filtered error as small as possible. Therefore, from the expression, it can be known that the expected optimal performance is Q d = 0. By analogy with the standard Bellman equation, it can be obtained:
[0038] Q(t j ) = min{aQ(t j-1 ) - a N+1 p(t j )}
[0039] The input vector of the neural network is calculated as follows: The radial basis vector is calculated as follows: σ(X) = [σ(x1), σ(x2), …, σ(xN)]T m is the number of nodes of the neural network, b i and c i are the center and width of the radial basis function, respectively.
[0040] Because the long-term performance index function Q(t j ) contains the future tracking information, this part of information is unknown. Therefore, an evaluator neural network is introduced to fit it.
[0041]
[0042] where W c ∈R m×6 and σ c ∈R m×1 are the weight matrix and the neural network radial basis vector of the evaluator neural network, respectively, x c (t j ) is the input vector of the neural network, ε a (x c (t j )) is the fitting error of the neural network and is bounded.
[0043] Because the optimal evaluator neural network is unknown, the actual actuator neural network is used to fit it.
[0044]
[0045] where the actual actuator neural network is the current estimation value of W c . The difference error occurring at time t j can be expressed in the following form:
[0046]
[0047] The update purpose of the actuator neural network is to reduce e c T (t j )e c (t j ) as much as possible. Therefore, it is not updated during (t j , t j+1 ] and is updated at time t j using the gradient descent method and the σ-modification method. The σ-modification method is to increase a term in the original update law. The added term is which can improve the robustness of the gradient descent algorithm and make it not necessarily to satisfy the persistent excitation condition to guarantee that W c is bounded. The designed update law is:
[0048]
[0049]
[0050] where, is the updated actuator neural network weight matrix at time t j is the updated actuator neural network weight matrix at time t c is the learning rate to be designed, L c is the correction term coefficient to be designed. represents the error matrix of the optimal weight estimate.
[0051] where, the actuator neural network in step five is designed as follows:
[0052] The actuator neural network is used to fit the external disturbance:
[0053] where, and σ a (x)∈R m×1 are the actuator neural network weight matrix and the neural network radial basis vector, respectively. x is the input vector representing the neural network, ε a (x) is the fitting error of the actuator neural network and is bounded.
[0054] The update purpose of the actuator neural network is to make the system track the desired trajectory and ensure that Q(t j ) is as close to zero as possible. The actuator error can be defined as:
[0055] e a = s + Q(t j ) - Q d
[0056] Q d is the desired optimal performance.
[0057] Therefore, at time t j , the gradient descent method is used for updating, and the σ-correction method is adopted, and the update law is designed as follows:
[0058]
[0059] where, Γ a > 0 is the learning rate to be designed, σ a (x a ) is each Gaussian function of the radial basis vector, x a is the input vector representing the neural network, is the error matrix of the optimal weight estimate at time t j , L a > 0 is the correction term coefficient to be designed, is the error matrix of the optimal weight estimate at time t jThe actuator neural network weight matrix is updated in real time.
[0060] The controller in step six is designed as follows:
[0061] The control amount is calculated, and the control law of the system is designed as:
[0062]
[0063] Wherein, K is a symmetric positive definite matrix.
[0064] The beneficial effects of the present application are:
[0065] 1) The present method avoids model linearization and can be directly applied to nonlinear systems, having a certain universality;
[0066] 2) The present method builds a radial basis neural network to construct an actuator-evaluator reinforcement learning framework by building a framework in reinforcement learning, improves system performance, and realizes optimal control;
[0067] 3) The present method can introduce an event-triggered mechanism into the reinforcement learning controller design, reducing the actuator execution frequency, reducing the algorithm calculation amount and communication amount; BRIEF DESCRIPTION OF DRAWINGS
[0068] Figure 1 The present application is a schematic diagram of the coordinate system;
[0069] Figure 2 The present application is a controller design flowchart; DETAILED DESCRIPTION
[0070] The present application "a trajectory tracking method for unmanned aerial vehicles based on reinforcement learning and event triggering", the specific steps are as follows: step one: establish a six-degree-of-freedom model of an unmanned aerial vehicle, give a desired tracking trajectory, and obtain the desired position, attitude angle, speed and angular velocity according to the given trajectory.
[0071] First, the inertial coordinate system and the body coordinate system are established; so the six-degree-of-freedom nonlinear motion equation of the unmanned aerial vehicle can be obtained:
[0072]
[0073] The position coordinates of the unmanned aerial vehicle P = [x, y, z] T The coordinates of the origin of the body coordinate system of the unmanned aerial vehicle in the inertial coordinate system. Attitude angle Θ = [φ, θ, ψ] T The angle between the body coordinate system of the unmanned aerial vehicle and the inertial coordinate system. Velocity vector v = [u, v, w] T The velocity components of the ground speed of the unmanned aerial vehicle in the body coordinate system along the Ox, Oy and Oz axes. Angular velocity vector Ω = [p, q, r]T are the angular velocity components of the UAV in the body frame about the Ox, Oy, and Oz axes, respectively. v ,F ω ,B 12 ,B 11 ,B 22 ,B 21 are functions related to the state equation, f v and f ω represent external disturbances of the system, τ = [τ v , τ ω ] T = [τ u , τ v , τ w , τ p , τ q , τ r ] T represent the input of the system. Define X = [P, Θ] T = [x, y, z, φ, θ, ψ] T represent the output of the system, X d = [P d , Θ d ] T = [x d , y d , z d , φ d , θ d , ψ d ] T represents the desired trajectory.
[0074] The trajectory curve is P d = [x d , y d , z d ] T = [+xt2, 200 + 100 cos(0.05t), 0.1t + 19000] T m, and the desired attitude angle is calculated as Step 2: Trajectory tracking error calculation: Calculate the error between the actual trajectory and the desired trajectory, define the filtered tracking error, and derive it.
[0075] The trajectory tracking error calculation method and derivation calculation are as follows:
[0076] Define the tracking error e = X - X d , and the filtered tracking error is defined as: where K e is a symmetric positive definite matrix.
[0077] Wherein, since the first derivative and second derivative of the desired trajectory are quite cumbersome and complex to solve, it is necessary to introduce a command filter to estimate its value. A second-order command filter is adopted, which can quickly track the desired signal.
[0078] By derivation, the following expression is obtained:
[0079]
[0080] Wherein,
[0081] Step three: design the event-triggered mechanism for the input, design the state quantity, and design reasonable event-triggered parameters.
[0082] The event-triggered mechanism in step three is designed as follows:
[0083] The control quantity of the system depends on s, e, X d And Define the comprehensive error state vector Then the input can be converted from a function of time to a state representation, that is, The controller is updated at non-periodic sampling time affected by the event-triggered condition, and the update will depend on the state ξ.
[0084] Define the measurement error as follows:
[0085] e ξ = ξ(t j )- ξ, t ∈ (t j , t j+1 ]
[0086] The designed event-triggered mechanism is as follows:
[0087]
[0088] t k+1 = inf {t > t k : ||e ξ || ≥ k s ‖s‖and‖s‖≥r s > 0}
[0089] Wherein, u(t) is the real input of the controller, when the event-triggered error exceeds the event-triggered threshold, the current instantaneous state is the sampling state, which is passed to the controller, and the event-triggered error can be expressed as the event-triggered condition. When not triggered, the controller retains the state at the last time under the action of the zero-order holder t ∈ (t j , t j+1 ]. Wherein, k s , r s are the event-triggered control parameters to be designed.
[0090] Step 4: Design the evaluator neural network to evaluate the tracking performance of the system for tracking performance improvement.
[0091] The evaluator neural network is designed as follows:
[0092] The event triggers the adaptive controller only at t j Update occurs at (t j ,t j+1 ] is to maintain the original value of the previous moment. Define the function p(t j ) is as follows:
[0093]
[0094] Among them, c p The threshold vector preset for s. p i (t j )=0 means good tracking performance, p i (t j )=1 represents poor tracking performance. Different from time-based updating, p i (t j ) function also only at the jump time t j The update occurs at all times, rather than according to fixed time sampling. Therefore, the long-term performance indicator function Q(t j ) is defined in the following way:
[0095] Q(t j )=α N p(t j+1 )+α N-1 p(t j+2 )+…+α j+1 p(t N )
[0096] Where 0<α<1 is the discount rate to be designed, and N is a preset large enough positive integer. The control goal is to make the tracking and filtering error as small as possible. Therefore, from the expression, we can know that the expected optimal performance is Q d = 0. By analogy with the standard Bellman equation, we can obtain:
[0097]
[0098] The input vector of the neural network Calculate the σ(X) radial basis vector and calculate each Gaussian function
[0099] Because the long-term performance indicator function Q(t j) contains future tracking information, which is unknown. Therefore, an estimator neural network can be introduced to fit it.
[0100] Q(t j )=W c T σ c (x c (t j ))+ε c (x(t j ))
[0101] Since the optimal evaluator neural network is unknown, the actual actuator neural network is used for fitting.
[0102]
[0103] Among them, the actual actuator neural network It's W c The current estimated value of . j The differential error occurring at the moment can be expressed as follows:
[0104]
[0105] The goal of updating the actuator neural network is to minimize e c T (t j )e c (t j ), so in (t j ,t j+1 ] is not updated during t j The time update adopts the gradient descent method and the σ-correction method. The σ-correction method adds a term to the original update law. There is a conflict in the variable names here, so the added items are It can improve the robustness of the gradient descent algorithm and make it unnecessary to meet the continuous incentive condition to ensure W c Bounded. The designed update law is:
[0106]
[0107]
[0108] in, It is t j The updated actuator neural network weight matrix, Γ c >0 is the learning rate to be designed, L c >0 is the correction coefficient to be designed. The error matrix representing the best weight estimates.
[0109] Step five: The actuator neural network is designed to approximate the uncertain term in the model.
[0110] The actuator neural network is designed as follows:
[0111] The actuator neural network is used to fit the external disturbance:
[0112] The purpose of updating the actuator neural network is to enable the system to track the desired trajectory and ensure that Q(t j ) is as close to zero as possible. The actuator error can be defined as:
[0113] e a = s + Q(t j ) - Q d
[0114] Therefore, the gradient descent method is used to update at time t j , and the σ-modification method is used, and the update law is designed as follows:
[0115]
[0116] wherein, is the updated actuator neural network weight matrix at time t j , Γ c > 0 is the learning rate to be designed, and L c > 0 is the correction term coefficient to be designed.
[0117] Step six: The nonlinear control law of the model is calculated by combining the neural network: the control amount required to eliminate the error between the desired trajectory and the actual trajectory is calculated.
[0118] The control amount is calculated, and the control law of the system is designed as:
[0119]
[0120] wherein, K is a symmetric positive definite matrix.
Claims
1. A UAV trajectory tracking method based on reinforcement learning and event triggering, characterized in that: The specific steps are as follows: Step 1: Establish a six-degree-of-freedom model of the UAV, give the desired tracking trajectory, and obtain the desired position, attitude angle, speed, and angular velocity based on the given trajectory; Step 2: Calculate the trajectory tracking error: Calculate the error between the actual trajectory and the expected tracking trajectory, define the filtered tracking error, and calculate its derivative; Step 3: Design input event triggering mechanism, design state quantity, and design reasonable event triggering parameters; Step 4: Design an evaluator neural network to evaluate the tracking performance of the system. Step 5: Design the actuator neural network to approximate the uncertainty; Step 6: Combine the neural network to obtain the nonlinear control law calculation of the model; The controller design in step six is as follows: Calculate the control quantity and design the control law of the system as follows: Where K is a symmetric positive definite matrix.
2. The UAV trajectory tracking method based on reinforcement learning and event triggering according to claim 1, characterized in that: The process of establishing the six-degree-of-freedom model of the unmanned aerial vehicle described in step 1 is as follows: First, establish the inertial coordinate system and the body coordinate system; thus, the six-degree-of-freedom nonlinear motion equation of the UAV can be obtained: The position coordinates of the UAV P = [x, y, z] T is the coordinate of the origin of the UAV body coordinate system in the inertial coordinate system; attitude angle Θ=[φ,θ,ψ] T is the angle between the UAV body coordinate system and the inertial coordinate system; velocity vector v = [u, v, w] T is the velocity component of the UAV ground speed along the Ox, Oy, and Oz axes in the body coordinate system; angular velocity vector Ω = [p, q, r] T F is the angular velocity component of the UAV around the Ox, Oy, Oz axis in the body coordinate system; v ,F ω ,B 12 ,B 11 ,B 22 ,B 21 It represents the function related to the state equation, f v and f ω represents the external disturbance of the system, τ=[τ v ,τ ω ] T =[τ u ,τ v ,τ w ,τ p ,τ q ,τ r ] T Represents the input of the system; define X = [P, Θ] T =[x,y,z,φ,θ,ψ] T Indicates system output, X d =[P d ,Θ d ] T =[x d ,y d ,z d ,φ d ,θ d ,ψ d ] T represents the desired trajectory; is the derivative of the position coordinate, is the derivative of the velocity component, is the derivative of the angular component, is the derivative of the angular velocity, K is the rotation matrix from the hull coordinate system to the inertial coordinate system, and R is the transformation matrix for the inverse angular velocity decomposition projection; The trajectory curve is P d =[x d ,y d ,z d ] T is the position coordinate of the desired trajectory curve, is the derivative of the desired position coordinate, and the desired attitude angle calculation formula is Θ d =[φ d ,θ d ,ψ d ] T represents the desired position coordinate derivative of the x-axis, represents the desired position coordinate derivative in the y-axis direction, Represents the desired position coordinate derivative in the z-axis direction.
3. The UAV trajectory tracking method based on reinforcement learning and event triggering according to claim 1, characterized in that: The calculation method and derivative calculation of the trajectory tracking error in step 2 are as follows: Define tracking error e = XX d , the tracking error after filtering is defined as: Among them, K e is a symmetric positive definite matrix, is the derivative of the tracking error; since the solution of the first-order derivative and the second-order derivative of the desired trajectory is quite tedious and complicated, it is necessary to introduce a command filter to estimate its value. A second-order command filter is used to quickly track the desired signal; By taking the derivative, we can get the following expression: in, k v is a positive constant diagonal matrix, k ω is a set of positive constant diagonal matrices, B v0 =KB 11 is the correlation function, B ω0 =RB 22 is the related function, is the second-order derivative and first-order derivative of the state estimator position and attitude of the command filter, and is the internal coupling, external disturbance and estimation error in position and attitude; and are the derivatives of the transformation matrices K and R.
4. The UAV trajectory tracking method based on reinforcement learning and event triggering according to claim 1, characterized in that: The event triggering mechanism in step 3 is designed as follows: The control quantity of the system depends on s, e, X d and in is the expected trajectory X d The derivative of ; define the comprehensive error state vector Then the input can be transformed from a function of time to a state representation, that is, The controller is updated at the non-periodic sampling time affected by the event triggering condition, and its update will depend on the state ξ; The measurement error is defined as follows: e ξ =ξ(t j )-ξ,t∈(t j ,t j+1 ] The designed event triggering mechanism is as follows: u(t)=τ(t k ), t k+1 =inf{t>t k :||e ξ ||≥k s ‖s‖and‖s‖≥r s >0} Among them, u(t) is the input of the real controller. When the event trigger error exceeds the event trigger threshold, the current state is the sampling state and is passed to the controller. The event trigger error can be expressed as the event trigger condition. When not triggered, the controller retains the state t∈(t j ,t j+1 ]; where k s , r s Control parameters triggered by the event to be designed.
5. The UAV trajectory tracking method based on reinforcement learning and event triggering according to claim 1, characterized in that: The evaluator neural network in step 4 is designed as follows: The event triggers the adaptive controller only at t j Update occurs at (t j ,t j+1 ] is to maintain the original value of the previous moment; the design function p(t j ) is as follows: Among them, c p The threshold vector preset for s, s i (t j ) is the tth filter tracking error s j The value of the i-th item at the moment; p i (t j )=0 means good tracking performance, p i (t j )=1 represents the poor tracking performance; Unlike time-based updates, p i (t j ) function also only at the jump time t j Updates are made at all times, rather than based on fixed time sampling; long-term performance indicator function Q(t j ): Q(t j )=a N p(t j+1 )+a N-1 p(t j+2 )+…+a j+1 p(t N ) Where 0<α<1 is the discount rate to be designed, and N is a preset large enough positive integer; the control objective is to minimize the tracking and filtering error; therefore, from the expression, we can know that the expected optimal performance is Q d =0; By analogy with the standard Bellman equation, we can obtain: Q(t j )=min{αQ(t j-1 )-α N+1 p(t j ))} The input vector of the neural network Calculate the σ(X) radial basis vector and calculate each Gaussian function m is the number of nodes in the neural network, b i and c i are the center and width of the radial basis function respectively; Because the long-term performance indicator function Q(t j ) contains future tracking information, which is unknown; therefore, an evaluator neural network can be introduced to fit it; Among them, W c ∈R m×6 and σ c ∈R m×1 are the estimator neural network weight matrix and neural network radial basis vector, x c (t j ) is the input vector of the neural network, ε a (x c (t j )) is the fitting error of the neural network and is bounded; Since the optimal evaluator neural network is unknown, the actual actuator neural network is used for fitting; Among them, the actual actuator neural network It's W c The current estimated value of j The differential error occurring at the moment can be expressed as follows: The goal of updating the actuator neural network is to minimize e c T (t j )e c (t j ), so in (t j ,t j+1 ] is not updated during t j The time update adopts the gradient descent method and the σ-correction method; the σ-correction method adds a term to the original update law. There is a conflict in the variable names here, so the added items are It can improve the robustness of the gradient descent algorithm and make it unnecessary to meet the continuous incentive condition to ensure W c Bounded; the designed update law is: in, It is t j The updated actuator neural network weight matrix, Γ c >0 is the learning rate to be designed, L c >0 is the correction coefficient to be designed; The error matrix representing the best weight estimates.
6. The UAV trajectory tracking method based on reinforcement learning and event triggering according to claim 1, characterized in that: The actuator neural network design in step 5 is as follows: The actuator neural network is used to fit the external disturbance: in, and σ a (x)∈R m×1 are the actuator neural network weight matrix and the neural network radial basis vector respectively; x represents the input vector of the neural network, ε a (x) is the fitting error of the actuator neural network and is bounded; The purpose of updating the actuator neural network is to enable the system to track the desired trajectory and ensure Q(t j ) can be as close to zero as possible; the actuator error can be defined as: e a =s+Q(t j )-Q d Q d For the best performance desired; Therefore, at t j The gradient descent method is used for constant update, and the σ-correction method is adopted. The update law is designed as follows: Among them, Γ a >0 is the learning rate to be designed, σ a (x a ) Gaussian functions of radial basis vectors, x a is the input vector of the neural network, It is t j The error matrix of the optimal weight estimate at the moment, L a >0 is the correction coefficient to be designed, It is t j The updated executor neural network weight matrix.
Citation Information
Patent Citations
Water ship trajectory tracking control method for actuator asymmetric saturation
CN107065847A
Unmanned aerial vehicle hanging system online trajectory planning method based on event driving
CN113759979A