Reinforcement learning based roll missile output constraint adaptive control system and method
Patent Information
- Application Number
- CN202410026150.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2044-01-08
AI Technical Summary
传统的飞行器控制方法没有充分考虑多源不确定性的影响,采用基于智能算法的自适应的控制方法更加有效,具有重要的理论和实践意义
[0051]Compared with existing technologies, this invention has at least the following beneficial effects: This invention discloses an output constraint adaptive control method for a roll missile based on reinforcement learning, which includes attitude dynamics modeling of a missile subject to multi-source coupled uncertainties such as external disturbances and modeling installation errors; analysis and mathematical representation of multi-source uncertainties, simplification and specialization of the attitude dynamics model, obtaining a control-oriented system model, and dividing the control model into inner and outer loops; learning compensation using a reinforcement learning framework for unknown nonlinear dynamics; compensation using a disturbance boundary estimator for disturbances in the inner and outer loops; processing output constraints based on nonlinear coordinate transformation, and designing a reinforcement learning adaptive control law by combining reinforcement learning, disturbance estimation, and backstepping; finally, stability analysis based on Lyapunov functions, theoretically proving the effectiveness of the disclosed control method, showing that the system is stable and all signals eventually converge to bounded, realizing the missile's tracking of the desired attitude angle under complex uncertainties. This invention addresses the complex uncertainties in missile flight, extending the classical model to a more universally representative one that better aligns with practical engineering applications. It introduces reinforcement learning, interacting with the environment and constructing penalty functions to further improve control performance. The novel control method developed in this invention exhibits superior interference suppression capabilities and environmental adaptability.
Smart Images

Figure CN118131802B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of aircraft attitude control technology, specifically relating to an adaptive control system and method for output constraints of a roll missile based on reinforcement learning. Background Technology
[0002] A rolling missile is a type of missile that can continuously spin around its longitudinal axis during flight. Relying on the high-speed rotation of the missile axis, it can maintain the stability of the warhead's pointing direction and make it difficult for laser weapons to target it, thus achieving penetration against laser weapons. Attitude control technology is one of the key technologies of rolling missiles. During attitude control, rolling missiles are often affected by multi-source coupled uncertainties. These uncertainties mainly originate from thrust eccentricity, manufacturing errors, and center of gravity deviation, and are also affected by noise and complex external interference in the complex flight environment. Traditional aircraft control methods do not fully consider the impact of multi-source uncertainties. Adaptive control methods based on intelligent algorithms are more effective and have significant theoretical and practical implications. Summary of the Invention
[0003] To address the problems existing in the prior art, this invention provides a roll missile output constraint adaptive control system and method based on reinforcement learning. It fully considers the influence of multi-source uncertainties and adopts an adaptive control method based on intelligent algorithms to improve the effectiveness of the roll missile control method.
[0004] To achieve the above objectives, the technical solution adopted by this invention is: an adaptive control method for output constraints of a rolling missile based on reinforcement learning, comprising the following steps:
[0005] Establish the rotational kinematics and dynamics equations of the rolling missile body;
[0006] The rotational kinematics and dynamic equations of the rolling missile are integrated into a set of vector equations. The multi-source uncertainties of the missile are mathematically represented and modeled to construct a control-oriented attitude control model for the rolling missile, expressed in terms of a set of differential equations.
[0007] For the unknown nonlinear dynamics in the attitude control model of a rolling missile, an error-based penalty function is established, and a reinforcement learning actor-critic framework is constructed to approximate the unknown nonlinear dynamics. For the disturbances and network approximation errors existing in the inner and outer loop systems, an interference boundary estimator is used to estimate and compensate for them.
[0008] By transforming the coordinates, the output constraint problem of the inner and outer loop systems is transformed into a stability problem of new variables, resulting in an output constraint reinforcement learning adaptive control law. The desired attitude angle is then tracked using this constraint reinforcement learning adaptive control law.
[0009] Furthermore, the kinematics and dynamics equations of the rolling missile body are expressed as follows:
[0010]
[0011] Where α, β, and γ are the angle of attack, sideslip angle, and roll angle, respectively; ω x ω y ω z δ represents the angular velocity components along the missile's three axes. x δ y δ z These represent the aileron, yaw, and pitch deflection angles, respectively; q, S, L, and P represent the missile's dynamic pressure, reference area, reference length, and thrust, respectively; J x J y J z The moment of inertia is the rotational inertia of the three axes; The slope of the lift coefficient with respect to the angle of attack and pitch deflection; The slope of the lateral force coefficient with respect to the sideslip angle and yaw deflection angle; For the control moment coefficients of ailerons, yaw rudders, and pitch rudders; These are the derivatives of the damping moment coefficient caused by pitch downwash delay with respect to the angle of attack and the derivatives of the damping moment coefficient caused by yaw downwash delay with respect to the sideslip angle, respectively. Δf are the dimensionless rotational derivatives. z , Δf y , Δf x This indicates the parts of the system model that cannot be accurately modeled due to assembly errors, etc.; d M z(t), d My (t), d Mx (t) represents the disturbance torque experienced by the missile along the three axes, m is the missile mass, and V is the missile velocity.
[0012] Furthermore, the rotational kinematics and dynamic equations of the rolling missile are integrated into a set of vector equations, and the multi-source uncertainties experienced by the missile are mathematically represented. A control-oriented attitude control model for the rolling missile, expressed in vector form, is established, including two subsystems: an inner loop and an outer loop.
[0013]
[0014] in:
[0015]
[0016] x1(t)=[α β γ] T x2(t)=[ω z ω y ω x ] TThese are the attitude angle and angular velocity vector of the rolling missile, respectively; u(t) = [δ z δ y δ x ] T Δf(x1(t),x2(t)) represents the unknown nonlinear dynamics of the system; d1(t) and d2(t) are defined as the total disturbances experienced by the inner and outer loop systems, respectively.
[0017] Furthermore, for the unknown nonlinear dynamics in the attitude control model of a rolling missile, a reinforcement learning actor-critic framework is established to approximate the unknown nonlinear dynamics, including:
[0018] The unknown nonlinear Δf(x1(t),x2(t)) is approximated using an actor network:
[0019]
[0020] Where, Φ a (x1(t),x2(t)) are Gaussian functions. It is the approximation value of the execution network, ε. a W is an estimate of the approximation error of the actor network. a For ideal weights, It is W a The estimated value.
[0021] Furthermore, an error-based penalty function is established, and an observational compensation method is used to compensate for disturbances and approximation errors existing in the inner and outer loops of the system using a disturbance boundary estimator, including:
[0022] Design an integral penalty function for the control system:
[0023]
[0024] Where e is the systematic error and Q is the positive definite matrix;
[0025] The penalty function fitted by the critic network is:
[0026]
[0027] Where, ε c To evaluate the network fitting error, W c For ideal weights, For W c The estimated values are given by x1, x2, f1, f2, Δf, d, and Φ. a Φ c Substitute x1(t), x2(t), f1(x1(t)), f2(x1(t),x2(t)), Δf(x1(t),x2(t)), d(t), and Φ respectively.a (x1(t),x2(t)),Φ c (x1(t),x2(t));
[0028] The weight update laws for the actor network and critic network are designed as follows:
[0029]
[0030] Among them, τ, η, Γ, ω is a design parameter.
[0031] Furthermore, the estimation and compensation of the total disturbance in the inner and outer loops of the system using the disturbance boundary estimator includes:
[0032] Define the total disturbance of the inner loop system as Design an interference boundary estimator to estimate d1, and define the estimated value as... Regarding the interference present in the outer loop system, the total disturbance is defined as... set up This is an estimated value.
[0033] Furthermore, by transforming the coordinates, the output constraint problems of the inner and outer loop systems are converted into stability problems with new variables, resulting in the output constraint reinforcement learning adaptive control law, which includes:
[0034] The virtual control law for the inner-loop system is designed as follows:
[0035]
[0036] According to the Lyapunov function stability theorem, the adaptive parameter update law is:
[0037]
[0038] Where, k1, η D1 σ D1 ε D1 These are the parameters to be designed;
[0039] The outer loop system control signal is designed as follows:
[0040]
[0041] According to the Lyapunov function stability theorem, the adaptive parameter update law is:
[0042]
[0043] Where k2, η D2 σ D2 ε D2 These are the parameters to be designed.
[0044] This invention also provides a reinforcement learning-based adaptive control system for output constraints of a rolling missile, comprising a model building module, a reinforcement learning module, a solution module, and an execution module.
[0045] The model building module is used to establish the rotational kinematics and dynamic equations of the rolling missile body. The rotational kinematics and dynamic equations of the rolling missile are simplified. The missile is subject to multi-source coupling uncertainty of modeling assembly error and external disturbance torque. The above uncertainties are mathematically represented and model specialized to construct a control-oriented attitude control model of the rolling missile represented by differential equations, which consists of an inner loop and an outer loop.
[0046] The reinforcement learning module is used to establish an error-based integral penalty function for the unknown nonlinear dynamics in the attitude control model of the rolling missile, and to design a reinforcement learning actor-critic framework to approximate the unknown nonlinear dynamics; and to estimate and compensate for the disturbances and approximation errors in the inner and outer loop systems using a disturbance boundary estimator.
[0047] The solution module is used to transform the stability problem of the system under output constraints into a stability problem with new variables by introducing nonlinear coordinate transformation, and obtain the output constraint reinforcement learning adaptive control law. The inner loop control law consists of the conventional adaptive control law and the disturbance estimate, and the outer loop control law consists of the conventional adaptive control law, reinforcement learning, and the disturbance estimate.
[0048] The execution module is used to enable the rolling missile to track the desired attitude angle during flight by executing the constraint reinforcement learning adaptive control law.
[0049] The present invention provides a computer device, including a processor and a memory. The memory is used to store a computer executable program. The processor reads part or all of the computer executable program from the memory and executes it. When the processor executes part or all of the executable program, it can realize the adaptive control method for output constraints of a rolling missile based on reinforcement learning described in the present invention.
[0050] Simultaneously, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, it can implement the reinforcement learning-based adaptive control method for output constraints of a rolling missile as described in this invention.
[0051] Compared with existing technologies, this invention has at least the following beneficial effects: This invention discloses an output constraint adaptive control method for a roll missile based on reinforcement learning, which includes attitude dynamics modeling of a missile subject to multi-source coupled uncertainties such as external disturbances and modeling installation errors; analysis and mathematical representation of multi-source uncertainties, simplification and specialization of the attitude dynamics model, obtaining a control-oriented system model, and dividing the control model into inner and outer loops; learning compensation using a reinforcement learning framework for unknown nonlinear dynamics; compensation using a disturbance boundary estimator for disturbances in the inner and outer loops; processing output constraints based on nonlinear coordinate transformation, and designing a reinforcement learning adaptive control law by combining reinforcement learning, disturbance estimation, and backstepping; finally, stability analysis based on Lyapunov functions, theoretically proving the effectiveness of the disclosed control method, showing that the system is stable and all signals eventually converge to bounded, realizing the missile's tracking of the desired attitude angle under complex uncertainties. This invention addresses the complex uncertainties in missile flight, extending the classical model to a more universally representative one that better aligns with practical engineering applications. It introduces reinforcement learning, interacting with the environment and constructing penalty functions to further improve control performance. The novel control method developed in this invention exhibits superior interference suppression capabilities and environmental adaptability. Attached Figure Description
[0052] Figure 1 This is a schematic diagram of the controller loop. Detailed Implementation
[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0055] refer to Figure 1 The adaptive control method for output constraints of a rolling missile based on reinforcement learning provided by this invention includes the following steps:
[0056] Step 1. Establish the rotational kinematics and dynamics equations of the rolling missile body, as follows:
[0057]
[0058] Where α, β, and γ are the angle of attack, sideslip angle, and roll angle, respectively; ω x ω y ω z δ represents the angular velocity components along the missile's three axes. x δ y δ z These represent the aileron, yaw, and pitch deflection angles, respectively; q, S, L, and P represent the missile's dynamic pressure, reference area, reference length, and thrust, respectively; J x J y J z The moment of inertia is the rotational inertia of the three axes; The slope of the lift coefficient with respect to the angle of attack and pitch deflection; The slope of the lateral force coefficient with respect to the sideslip angle and yaw deflection angle; For the control moment coefficients of ailerons, yaw rudders, and pitch rudders; These are the derivatives of the damping moment coefficient caused by pitch downwash delay with respect to the angle of attack and the derivatives of the damping moment coefficient caused by yaw downwash delay with respect to the sideslip angle, respectively. Δf are the dimensionless rotational derivatives. z , Δf y , Δf x This indicates the parts of the system model that cannot be accurately modeled due to assembly errors, etc.; d Mz (t), d My (t), d Mx (t) represents the disturbance torque experienced by the missile along its three axes, m is the missile mass, and V is the missile velocity.
[0059] Step 2. Establish a control-oriented attitude control model for a roll missile, expressed as a system of differential equations, consisting of two subsystems: inner and outer loops.
[0060]
[0061] Where x1(t) = [α β γ] T x2(t)=[ω z ω y ω x ] T These are the attitude angle and angular velocity vector of the rolling missile, respectively; u(t) = [δ z δ y δ x ] T d1(t), d2(t) represents the control input of the system; Δf(x1(t), x2(t)) represents the unknown nonlinear dynamics of the system; d1(t) and d2(t) are defined as the total disturbances experienced by the inner and outer loop systems, respectively, in the following forms:
[0062]
[0063] Step 3. Construct a reinforcement learning framework. For the unknown dynamics in the roll missile attitude control model, i.e., Δf(x1(t),x2(t)) in the formula, establish an actor network for approximation:
[0064]
[0065] Where, Φ a (x1(t),x2(t)) are Gaussian functions. It is the approximation value of the execution network, ε. a W is an estimate of the approximation error of the execution network. a For ideal weights, It is W a The estimated value. Define the actor network weight error as... Define the actor network approximation error as:
[0066]
[0067] The actor network behavior error function is created as follows:
[0068]
[0069] The weight update law of the actor network is designed using the gradient descent method as follows:
[0070]
[0071] The integral penalty function for the control system is designed as follows:
[0072]
[0073] Design the penalty function for the critical network:
[0074]
[0075] Where, ε c To evaluate the network fitting error, W c For ideal weights, It is estimated as follows. The mean squared error function of the critic network is defined as:
[0076]
[0077] To make E c To minimize this, the weight update law for the critic network, designed using gradient descent, is:
[0078]
[0079] Here τ, η, Γ, ω is the parameter to be designed.
[0080] Step 4. Design the output constraint reinforcement learning adaptive control law for the rolling missile. To ensure the constraint on the tracking error, a nonlinear function is introduced:
[0081]
[0082] The desired angle is x d =[α d β d γ d ] T Error e1 = x1 - x d =[e α e β e γ ] T , Introduce the following transformations:
[0083] z α =g(σ α ,e α ),z β =g(σ β ,e β ),z γ =g(σ γ ,e γ (35)
[0084] Where σ α σ β σ γ Since the parameters are known, we have:
[0085]
[0086] K(x1)=Diag[k α k β k γ ],
[0087] definition Design an interference boundary estimator to estimate d1, and define the estimated value as... The virtual control law for the inner-loop system is designed as follows:
[0088]
[0089] According to the Lyapunov function stability theorem, the adaptive parameter update law for the inner-loop system is:
[0090]
[0091] Where k1, η D1 σD1 ε D1 For the parameters to be designed,
[0092] Similarly, regarding the interference present in the outer loop system, the total disturbance is defined as follows: set up For the estimated value, the outer loop system control signal is designed as follows:
[0093]
[0094] According to the Lyapunov function stability theorem, the adaptive parameter update law for the outer-loop system is:
[0095]
[0096] Where k2, η D2 σ D2 ε D2 These are the parameters to be designed.
[0097] Finally, this invention completes the stability analysis of the output constraint reinforcement learning adaptive control law for a rolling missile based on Lyapunov stability theory. First, the Lyapunov function is selected:
[0098]
[0099] Differentiating both sides of the equation, we get:
[0100]
[0101] Substituting the error state equation (36), control laws (37), (39), parameter update laws (38), (40), and reinforcement learning network update laws (29), (33) into (42) and simplifying, we get:
[0102]
[0103] Because of the inequality It holds true, and the basis functions are bounded, so ||Φ a ||≤Φ aM ,||Φ c ||≤Φ cM , According to Young's inequality, the above equation can be simplified to:
[0104]
[0105] in:
[0106]
[0107] According to Lyapunov's stability theorem, the system is stable, and all signals eventually converge to a bounded state. Given any initial angle, by appropriately selecting design parameters, the designed control law can ensure that the roll missile's angle of attack, sideslip angle, and roll angle stably track the desired attitude angle.
[0108] Based on the same inventive concept, this invention provides a roll missile output constraint adaptive control system based on reinforcement learning, including a model building module, a reinforcement learning module, a solution module, and an execution module;
[0109] The model building module is used to establish the rotational kinematics and dynamic equations of the rolling missile body. The rotational kinematics and dynamic equations of the rolling missile are simplified. The missile is subject to multi-source coupling uncertainty of modeling assembly error and external disturbance torque. The above uncertainties are mathematically represented and model specialized to construct a control-oriented attitude control model of the rolling missile represented by differential equations, which consists of an inner loop and an outer loop.
[0110] The reinforcement learning module is used to establish an error-based integral penalty function for the unknown nonlinear dynamics in the attitude control model of the rolling missile, and to design a reinforcement learning actor-critic framework to approximate the unknown nonlinear dynamics; and to estimate and compensate for the disturbances and approximation errors in the inner and outer loop systems using a disturbance boundary estimator.
[0111] The solution module is used to transform the stability problem of the system under output constraints into a stability problem with new variables by introducing nonlinear coordinate transformation, and obtain the output constraint reinforcement learning adaptive control law. The inner loop control law consists of the conventional adaptive control law and the disturbance estimate, and the outer loop control law consists of the conventional adaptive control law, reinforcement learning, and the disturbance estimate.
[0112] The execution module is used to enable the rolling missile to track the desired attitude angle during flight by executing the constraint reinforcement learning adaptive control law.
[0113] Optionally, the present invention also provides a computer device, including a processor and a memory, the memory being used to store a computer executable program, the processor reading part or all of the computer executable program from the memory and executing it, and the processor executing part or all of the compute executable program being able to implement the steps of the reinforcement learning-based adaptive control method for output constraints of a rolling missile as described in the present invention.
[0114] And a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps of the reinforcement learning-based adaptive control method for output constraints of a rolling missile as described in this invention.
[0115] A program that can be written in a computer programming language to perform the methods described in this application can be used. The computer program can be in the form of source code, object code, executable file or some intermediate form. The computer programming language can be C++, Java, Fortran, C# or Python.
[0116] The processor can be a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or an off-the-shelf programmable gate array (FPGA).
[0117] The memory described in this invention can be an internal storage unit of a laptop, tablet, desktop computer, mobile phone, or workstation, such as memory or hard disk; or it can be an external storage unit, such as a portable hard disk or flash memory card.
[0118] Computer-readable storage media can include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media can include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. Random access memory can include resistive random access memory (ReRAM).
Claims
1. A method for adaptive control of output constraints of a rolling missile based on reinforcement learning, characterized in that, Includes the following steps: Establish the rotational kinematics and dynamics equations of the rolling missile body; The rotational kinematics and dynamic equations of the rolling missile are simplified. The missile is subject to multi-source coupling uncertainty of modeling assembly errors and external disturbance torques. The above uncertainties are mathematically represented and modeled to construct a control-oriented attitude control model of the rolling missile represented by differential equations, consisting of an inner loop and an outer loop. To address the unknown nonlinear dynamics in the attitude control model of a rolling missile, an error-based integral penalty function is established, and a reinforcement learning actor-critic framework is designed to approximate the unknown nonlinear dynamics. A disturbance boundary estimator is used to estimate and compensate for disturbances and approximation errors in the inner and outer loop systems. For the unknown nonlinear dynamics in the rolling missile attitude control model, a reinforcement learning framework is introduced for compensation: first, an actor network is constructed to approximate the unknown nonlinear dynamics, including: (4) in, For Gaussian functions, It is the approximation value of the execution network. It is an estimate of the approximation error of the actor network. For ideal weights, yes The estimated value, Design an integral penalty function for the control system: (5) in, For systematic error, It is a positive definite matrix; The penalty function fitted by the critic network is: (6) in, To evaluate the network fitting error, For ideal weights, for The estimated value, , These are the attitude angle and angular velocity vector of the rolling missile, respectively; The system exhibits unknown nonlinear dynamics; , , These are angle of attack, sideslip angle, and roll angle, respectively. , , These are the angular velocity components along the missile's three axes; , , These are the aileron, yaw, and pitch deflection angles, respectively. The weight update laws for the actor network and critic network are designed as follows: (7) in, , , , , For design parameters; By introducing nonlinear coordinate transformation, the stability problem of the system under output constraints is transformed into a stability problem of new variables, resulting in an output-constrained reinforcement learning adaptive control law. The inner-loop control law consists of a conventional adaptive control law and disturbance estimates, while the outer-loop control law consists of a conventional adaptive control law, reinforcement learning, and disturbance estimates. During the flight of the rolling missile, the desired attitude angle is tracked by executing the constraint reinforcement learning adaptive control law.
2. The adaptive control method for output constraints of a rolling missile based on reinforcement learning according to claim 1, characterized in that, The kinematic and dynamic equations of the rolling missile body attitude are expressed as follows: (1) in, , , , These are the missile's dynamic pressure, reference area, reference length, and thrust, respectively. , , The moment of inertia is the rotational inertia of the three axes; , The slope of the lift coefficient with respect to the angle of attack and pitch deflection; , The slope of the lateral force coefficient with respect to the sideslip angle and yaw deflection angle; , , For the control moment coefficients of ailerons, yaw rudders, and pitch rudders; , These are the derivatives of the damping moment coefficient caused by pitch downwash delay with respect to the angle of attack and the derivatives of the damping moment coefficient caused by yaw downwash delay with respect to the sideslip angle, respectively. , , These are the dimensionless rotational derivatives. , , This indicates the parts of the system model that cannot be accurately modeled due to assembly errors, etc. , , This represents the disturbance torque experienced by the missile along its three axes. m For missile quality, V This refers to the missile's speed.
3. The adaptive control method for output constraints of a rolling missile based on reinforcement learning according to claim 1, characterized in that, The rotational kinematics and dynamics equations of the rolling missile are integrated into a system of differential equations. The multi-source uncertainties experienced by the missile are mathematically represented, and a control-oriented attitude control model for the rolling missile, expressed in differential equations, is established, including two subsystems: an inner loop and an outer loop. (2) in: (3) For system control input; The system exhibits unknown nonlinear dynamics; and Defined as the total disturbance experienced by the inner and outer loop systems.
4. The adaptive control method for output constraints of a rolling missile based on reinforcement learning according to claim 1, characterized in that, The use of a disturbance boundary estimator to estimate and compensate for disturbances and approximation errors existing in the inner and outer loops of the system includes: definition Design a disturbance boundary estimator for the inner loop system disturbance. Make an estimate and define the estimated value as follows: To address the interference in the outer loop system and the error in the reinforcement learning network approximation, the total perturbation is defined as... ,set up This is an estimated value.
5. The adaptive control method for output constraints of a rolling missile based on reinforcement learning according to claim 1, characterized in that, By transforming the coordinates, the output constraint problem of the inner and outer loop systems is converted into a stability problem with new variables, resulting in the output constraint reinforcement learning adaptive control law, which includes: The virtual control law for the inner-loop system is designed as follows: (8) According to the Lyapunov function stability theorem, the adaptive parameter update law is: (9) in, , , , These are the parameters to be designed; The outer loop system control signal is designed as follows: (10) According to the Lyapunov function stability theorem, the adaptive parameter update law is: (11) in , , , These are the parameters to be designed.
6. A roll missile output constraint adaptive control system based on reinforcement learning, characterized in that, It includes a model building module, a reinforcement learning module, a solution module, and an execution module; The model building module is used to establish the rotational kinematics and dynamic equations of the rolling missile body. The rotational kinematics and dynamic equations of the rolling missile are simplified. The missile is subject to multi-source coupling uncertainty of modeling assembly error and external disturbance torque. The above uncertainties are mathematically represented and model specialized to construct a control-oriented attitude control model of the rolling missile represented by differential equations, which consists of an inner loop and an outer loop. The reinforcement learning module is used to address the unknown nonlinear dynamics in the roll missile attitude control model. It establishes an error-based integral penalty function and designs a reinforcement learning actor-critic framework to approximate the unknown nonlinear dynamics. A disturbance boundary estimator is used to estimate and compensate for disturbances and approximation errors in the inner and outer loop systems. For the unknown nonlinear dynamics in the roll missile attitude control model, a reinforcement learning framework is introduced for compensation: first, an actor network is constructed to approximate the unknown nonlinear dynamics, including: (4) in, For Gaussian functions, It is the approximation value of the execution network. It is an estimate of the approximation error of the actor network. For ideal weights, yes The estimated value, Design an integral penalty function for the control system: (5) in, For systematic error, It is a positive definite matrix; The penalty function fitted by the critic network is: (6) in, To evaluate the network fitting error, For ideal weights, for The estimated value, , These are the attitude angle and angular velocity vector of the rolling missile, respectively; The system exhibits unknown nonlinear dynamics; , , These are angle of attack, sideslip angle, and roll angle, respectively. , , These are the angular velocity components along the missile's three axes; , , These are the aileron, yaw, and pitch deflection angles, respectively. The weight update laws for the actor network and critic network are designed as follows: (7) in, , , , , For design parameters; The solution module is used to transform the stability problem of the system under output constraints into a stability problem with new variables by introducing nonlinear coordinate transformation, and obtain the output constraint reinforcement learning adaptive control law. The inner loop control law consists of the conventional adaptive control law and the disturbance estimate, and the outer loop control law consists of the conventional adaptive control law, reinforcement learning, and the disturbance estimate. The execution module is used to enable the rolling missile to track the desired attitude angle during flight by executing the constraint reinforcement learning adaptive control law.
7. A computer device, characterized in that, It includes a processor and a memory, the memory being used to store a computer-executable program, the processor reading part or all of the computer-executable program from the memory and executing it, and the processor executing part or all of the computer-executable program is able to implement the reinforcement learning-based adaptive control method for output constraints of a rolling missile as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program that, when executed by a processor, can implement the reinforcement learning-based adaptive control method for output constraints of a rolling missile as described in any one of claims 1 to 5.