Non-affine reinforcement learning variable structure control system and method for rolling missile
Patent Information
- Application Number
- CN202410027331.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2044-01-08
AI Technical Summary
然而,滚转导弹飞行过程中姿态角和角速率变化快,三通道间耦合严重,且存在非仿射气动特征以及快速变化不确定性,给控制系统设计带来了困难
[0059]Compared with existing technologies, this invention has at least the following beneficial effects: This invention discloses a reinforcement learning adaptive variable structure attitude control method for a roll missile with non-affine control input, including kinematic and dynamic modeling of the pitch, yaw, and roll axes of the roll missile; simplification and specialization of the physical model based on the system's disturbance characteristics; separation of non-affine control quantities by adding an integral method to obtain an augmented affine system model oriented towards control, addressing the system's non-affine characteristics; fitting of unknown nonlinear dynamics in the system using a reinforcement learning network; observation compensation for bounded disturbances by developing a disturbance observer based on reinforcement learning; introduction of a sliding mode surface to eliminate chattering; and design of a reinforcement learning adaptive variable structure controller by combining reinforcement learning, disturbance observation, and adaptive backstepping; finally, the stability of the disclosed control method is proved by designing a Lyapunov function, demonstrating that the missile can still achieve attitude angle stability or track the desired attitude angle under complex disturbance conditions. This invention studies the dynamic characteristics of a rolling missile, considering the multi-source uncertainties encountered during flight and the non-affine characteristics of the control input. The model is extended to an augmented affine differential equation system, facilitating further controller design. This invention introduces reinforcement learning into function approximation and disturbance observer design, further improving anti-interference performance by evaluating control performance and designing update laws. The model considered in this invention involves uncertainties with more diverse characteristics, making it more universally representative.
Smart Images

Figure CN117970801B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of aircraft attitude control technology, specifically relating to a non-affine reinforcement learning variable structure control system and method for a rolling missile. Background Technology
[0002] The rapid roll of a missile can avoid being burned by laser interceptor weapons. Therefore, the design of the control system for a rolling missile is an important research topic. However, the rapid changes in attitude angle and angular rate during the flight of a rolling missile, the severe coupling between the three channels, and the existence of non-affine aerodynamic characteristics and rapidly changing uncertainties bring difficulties to the design of the control system. Variable structure control is a class of adaptive nonlinear control methods that can automatically switch the controller structure when the system state is in different spatial regions. This invention proposes a non-affine reinforcement learning variable structure control system for a rolling missile by combining the advantages of reinforcement learning algorithms, variable structure control, and non-affine control. This system can effectively improve the controlled object's ability to suppress non-affine characteristics and internal and external disturbances. Summary of the Invention
[0003] To address the problems existing in the prior art, this invention provides a non-affine reinforcement learning variable structure control method for rolling missiles. By designing a reinforcement learning-based variable structure control method, the design difficulties caused by non-affine structures are overcome. This method not only eliminates chattering but also improves the robustness and response speed of the rolling missile system.
[0004] To achieve the above objectives, the technical solution adopted by this invention is: a non-affine reinforcement learning variable structure control method for a rolling missile, comprising the following steps:
[0005] Establish the rotational kinematics and dynamics equations of the rolling missile body;
[0006] Based on the rotational kinematics and dynamics equations of the rolling missile body, and considering the non-affine characteristics of the rolling missile control signal and the internal and external disturbances it is subjected to, an augmented affine attitude model of the rolling missile under control is established by adding an integral method.
[0007] A reinforcement learning framework is used to learn and compensate for the unknown nonlinear dynamics in the augmented affine system model. The reinforcement signal is generated using error information, and a reinforcement learning-based interference observer is used to observe and compensate for internal and external interference.
[0008] By introducing the sliding mode variable structure method into adaptive control to eliminate chattering, and combining disturbance observation and reinforcement learning, a reinforcement learning variable structure control law for roll missiles with non-affine structure is obtained.
[0009] During the flight of the rolling missile, the reinforcement learning variable structure control law for rolling missiles with non-affine structures is executed to track the desired attitude angle.
[0010] Furthermore, the kinematics and dynamics equations of the rolling missile body are expressed as follows:
[0011]
[0012]
[0013]
[0014]
[0015]
[0016]
[0017] Among them, γ, ψ, These are roll angle, yaw angle, and pitch angle, respectively; ω x ω y ω z These are the components of the missile's angular velocity along the three axes; J x J y J z M is the moment of inertia of the three axes. x M y M z d represents the control torque on the three axes. M x(t), d My (t), d Mz (t) represents the disturbance torque on the three axes.
[0018] Furthermore, the attitude kinematics and dynamics of the rolling missile are integrated into a control-oriented non-affine system model, incorporating external disturbances and unknown nonlinear dynamics. This model includes two-order subsystems.
[0019]
[0020]
[0021] Constructing an augmented affine system model:
[0022]
[0023]
[0024]
[0025] in: ω=[ω x ω y ω z ] T,
[0026]
[0027] in, ω=[ω x ω y ω z ] T These represent the missile's three attitude angles and the components of its attitude angular velocity along the three axes; u is the missile's control input; M(θ,ω,u,t)=[M x (θ,ω,u,t)M y (θ,ω,u,t)M z (θ,ω,u,t)] T The system represents a non-affine function containing control inputs; Δf(θ(t),ω(t)) is the unknown nonlinear dynamics of the system; d ω (t)=[d Mx (t) d My (t) d Mz (t)] T This represents the attitude disturbance torque.
[0028] Furthermore, the execution network in the reinforcement learning network is used to approximate the unknown nonlinear function Δf existing in the system:
[0029]
[0030] Where, Φ a For Gaussian functions, To approximate the value of the network, the integral penalty function of the control system is designed as J(t), and the evaluation network fitting penalty function is designed as follows:
[0031]
[0032] Where, ε c To evaluate the network fitting error, To approximate the value of the execution network, the weight update laws for the execution network and the evaluation network are designed as follows:
[0033]
[0034]
[0035] Among them, σ, η, τ, Let the parameters to be designed satisfy the boundedness of the basis functions.
[0036] Furthermore, in the augmented affine system, the first two stages employ an adaptive back-reasoning method, while the last stage subsystem introduces the concept of variable structure to construct an adaptive sliding mode variable structure control law.
[0037] Furthermore, combining adaptive control and sliding mode variable structure methods, a reinforcement learning variable structure adaptive control law is designed, including:
[0038] Design the virtual control signal ω for the first-order subsystem. c for:
[0039]
[0040] Design a reinforcement learning-based interference observer:
[0041]
[0042]
[0043] Where ξ is the intermediate variable of the disturbance observer, L is the gain to be designed, P(ω) = Lω is the function to be designed, and the virtual control law of the second-order subsystem is u. c Designed as follows:
[0044]
[0045] Furthermore, the system sliding surface is designed in the extended subsystem as follows:
[0046]
[0047] Where c i Let p be the undetermined coefficients of the sliding surface. n-1 +c n-1 p n-2 +…+c2p+c1, for Hurwitz stability; the adaptive sliding mode variable structure control law is designed as follows:
[0048] u f =u f1 +u f2
[0049]
[0050]
[0051] Where b1 and b2 > 0 are the parameters to be designed.
[0052] With the same inventive concept as the above method, the present invention also provides a non-affine reinforcement learning variable structure control system for a rolling missile, including a model building module, a reinforcement learning module, an adaptive control law acquisition module, and an attitude stabilization module.
[0053] The model building module is used to establish the kinematics and dynamics equations of the rolling missile attitude; considering the non-affine characteristics of the rolling missile control signal and the characteristics of the disturbances it receives, an augmented affine attitude model of the rolling missile oriented towards control is established by adding the integral method.
[0054] The reinforcement learning module is used to learn and compensate for the unknown nonlinear dynamics in the augmented affine system model using a reinforcement learning framework, generate reinforcement signals using error information, and compensate for internal and external interference observations using a reinforcement learning-based interference observer.
[0055] The adaptive control law acquisition module is used to combine adaptive control, reinforcement learning, interference observation and sliding mode variable structure method to design reinforcement learning variable structure control laws for roll missiles with non-affine structure.
[0056] The attitude stabilization module is used to enable the rolling missile to perform the control law tracking or stabilize to the desired attitude angle during flight.
[0057] Another computer device is provided, including a processor and a memory. The memory is used to store a computer-executable program. The processor reads the computer-executable program from the memory and executes it. When the processor executes the program, it can implement the non-affine reinforcement learning variable structure control method for rolling missiles described in this invention.
[0058] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the non-affine reinforcement learning variable structure control method for rolling missiles described in the present invention.
[0059] Compared with existing technologies, this invention has at least the following beneficial effects: This invention discloses a reinforcement learning adaptive variable structure attitude control method for a roll missile with non-affine control input, including kinematic and dynamic modeling of the pitch, yaw, and roll axes of the roll missile; simplification and specialization of the physical model based on the system's disturbance characteristics; separation of non-affine control quantities by adding an integral method to obtain an augmented affine system model oriented towards control, addressing the system's non-affine characteristics; fitting of unknown nonlinear dynamics in the system using a reinforcement learning network; observation compensation for bounded disturbances by developing a disturbance observer based on reinforcement learning; introduction of a sliding mode surface to eliminate chattering; and design of a reinforcement learning adaptive variable structure controller by combining reinforcement learning, disturbance observation, and adaptive backstepping; finally, the stability of the disclosed control method is proved by designing a Lyapunov function, demonstrating that the missile can still achieve attitude angle stability or track the desired attitude angle under complex disturbance conditions. This invention studies the dynamic characteristics of a rolling missile, considering the multi-source uncertainties encountered during flight and the non-affine characteristics of the control input. The model is extended to an augmented affine differential equation system, facilitating further controller design. This invention introduces reinforcement learning into function approximation and disturbance observer design, further improving anti-interference performance by evaluating control performance and designing update laws. The model considered in this invention involves uncertainties with more diverse characteristics, making it more universally representative. Attached Figure Description
[0060] Figure 1 This is a schematic diagram of the controller loop. Detailed Implementation
[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0062] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0063] The non-affine reinforcement learning variable structure control method for a rolling missile provided by this invention can be implemented as follows:
[0064] Step 1. Establish the kinematic dynamic equations of the rolling missile attitude, as follows:
[0065]
[0066]
[0067]
[0068]
[0069]
[0070]
[0071] Among them γ, ψ, These are roll angle, yaw angle, and pitch angle, respectively; ω x ω y ω z These are the components of the missile's angular velocity along the three axes; J x J y J z M is the moment of inertia of the three axes. x M y M z d represents the control torque on the three axes. Mx (t), d My (t), d Mz (t) represents the disturbance torque on the three axes.
[0072] Step 2. Analyze the disturbance characteristics of the system, perform mathematical characterization, and establish a control-oriented non-affine system model after simplifying and specializing the attitude kinematics equations. The non-affine system model is divided into two-order subsystems:
[0073]
[0074]
[0075] in: ω=[ω x ω y ω z ] T ,
[0076]
[0077] in, ω=[ω x ω y ω z ] T These represent the missile's three attitude angles and the components of its attitude angular velocity along the three axes; u is the missile's control input; M(θ,ω,u,t)=[M x (θ,ω,u,t) M y(θ,ω,u,t) M z (θ,ω,u,t)] T The system represents a non-affine function containing control inputs; Δf(θ(t),ω(t)) is the unknown nonlinear dynamics of the system; d ω (t)=[d Mx (t) d My (t) d Mz (t)] T This represents the attitude disturbance torque. Clearly, the control signal is non-affine, requiring further analysis using an augmented system. An augmented affine system model is constructed by adding an integral method, containing a third-order subsystem:
[0078]
[0079]
[0080]
[0081] To simplify the analysis process, the functions are represented in a simplified manner below, such as Δf(θ,ω,t) being replaced by Δf.
[0082] Step 3. Construct a reinforcement learning framework. For the unknown nonlinear function Δf in the system, approximate it using an execution network:
[0083]
[0084] Where, Φ a Let Gaussian function be defined. To obtain the approximate value of the network, the integral penalty function of the control system is designed as follows:
[0085]
[0086] Where e is the systematic error and Q is the positive definite matrix. Design the evaluation network fitting penalty function:
[0087]
[0088] Here ε c To evaluate the network fitting error, we define: To approximate the value of the execution network, the weight update laws for the execution network and the evaluation network are designed as follows:
[0089]
[0090]
[0091] Here σ, η, τ, Let be the parameters to be designed. They satisfy the boundedness of the basis functions, therefore ||Φ a||≤Φ aM ,||Φ c ||≤Φ cM ,
[0092] Step 4. Design a reinforcement learning variable structure control law for a roll missile with a non-affine structure. The first two stages use adaptive backpropagation technology, and the last stage designs an adaptive sliding mode variable structure control law.
[0093] Define θ d For the desired angle, z1 = θ - θ d z2=ω-ω c , where ω c Since this is a virtual control signal for a first-order system, we have:
[0094]
[0095] The virtual control signal for the first-order subsystem is designed as follows:
[0096]
[0097] Design the Lyapunov function as follows:
[0098]
[0099] Differentiating both sides of the equation, we get:
[0100]
[0101] According to z2=ω-ω c z3 = Mu c u c For a second-order virtual control law, the system error function is:
[0102]
[0103] Design an interference observer:
[0104]
[0105]
[0106] Where ξ is the intermediate variable of the interference observer, L is the gain to be designed, and P(ω)=Lω is the function to be designed. Let the observation error be... The virtual control law is designed as follows:
[0107]
[0108] Define the Lyapunov function as:
[0109]
[0110] Taking the derivative, substituting the virtual control law, the interference observer, and the network weight update law, we obtain the following according to Young's inequality:
[0111]
[0112] Because z3 = Mu c ,have:
[0113]
[0114] The system sliding surface is defined as follows:
[0115]
[0116] Where c i Let p be the undetermined coefficients of the sliding surface. n-1 +c n-1 p n-2 +...+c2p+c1, which is Hurwitz stable.
[0117] The sliding mode variable structure control law is designed as follows:
[0118] u f =u f1 +u f2
[0119]
[0120]
[0121] Where b1 and b2 > 0 are the parameters to be designed.
[0122] Choose the Lyapunov function:
[0123]
[0124] Differentiating the Lyapunov function yields:
[0125]
[0126] Lyapunov's derivative can be rewritten as:
[0127]
[0128] Step 5. Proof of stability:
[0129] Define the Lyapunov function of the system as:
[0130] V = V1 + V2 + V3 (23)
[0131] Differentiation yields:
[0132]
[0133] From the above formula, it can be seen that by reasonably selecting the design parameters k1, k2, k3, σ, τ, L, b1, and b2, the following can be achieved: And when the sliding surface s = 0, we have z i =0, at this time the system attitude angle stably tracks the desired attitude angle, and the system is asymptotically stable.
[0134] Based on the same inventive concept, the present invention also provides a non-affine reinforcement learning variable structure control system for a rolling missile, including a model building module, a reinforcement learning module, an adaptive control law acquisition module, and an attitude stabilization module.
[0135] The model building module is used to establish the kinematics and dynamics equations of the rolling missile attitude; considering the non-affine characteristics of the rolling missile control signal and the characteristics of the disturbances it receives, an augmented affine attitude model of the rolling missile oriented towards control is established by adding the integral method.
[0136] The reinforcement learning module is used to learn and compensate for the unknown nonlinear dynamics in the augmented affine system model using a reinforcement learning framework, generate reinforcement signals using error information, and compensate for internal and external interference observations using a reinforcement learning-based interference observer.
[0137] The adaptive control law acquisition module is used to combine adaptive control, reinforcement learning, interference observation and sliding mode variable structure method to design reinforcement learning variable structure control laws for roll missiles with non-affine structure.
[0138] The attitude stabilization module is used to enable the rolling missile to perform the control law tracking or stabilize to the desired attitude angle during flight.
[0139] On the other hand, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the non-affine reinforcement learning variable structure control method for rolling missiles described in the present invention.
[0140] The computer equipment may be a laptop, desktop computer, workstation, or vehicle-mounted computer.
[0141] The processor described in this invention may be a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or an off-the-shelf programmable gate array (FPGA).
[0142] The memory described in this invention can be an internal storage unit of a laptop, desktop computer, workstation, or vehicle-mounted computer, such as memory or hard disk; or it can be an external storage unit, such as a portable hard disk or flash memory card.
[0143] The present invention can also provide a computer device, including a processor and a memory, wherein the memory is used to store a computer executable program, the processor reads the computer executable program from the memory and executes it, and the processor can implement the non-affine reinforcement learning variable structure control method for rolling missiles described in the present invention when executing the computer executable program.
[0144] Computer-readable storage media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media can include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. Random access memory can include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).
Claims
1. A non-affine reinforcement learning variable structure control method for a rolling missile, characterized in that, Includes the following steps: Establish the rotational kinematics and dynamics equations of the rolling missile body; Based on the rotational kinematics and dynamics equations of the rolling missile, and considering the non-affine characteristics of the rolling missile's control signals as well as the internal and external disturbances it experiences, an augmented affine attitude model of the rolling missile under disturbance is established for control purposes by adding an integral method. The attitude kinematics and dynamics of the rolling missile are integrated into a control-oriented non-affine system model, incorporating external disturbances and unknown nonlinear dynamics. This model includes two-order subsystems. Constructing an augmented affine system model: in: , , , , ; in, , These represent the missile's three attitude angles and the components of its attitude angular velocity along the three axes, respectively. It is the control input for the missile; This represents a system containing non-affine functions for control inputs; It is an unknown nonlinear dynamic of the system; Represents attitude disturbance torque; , , These are roll angle, yaw angle, and pitch angle, respectively. , , These are the components of the missile's angular velocity along the three axes; , , The moment of inertia is the rotational inertia of the three axes; , , The control torques are on the three axes. , , The disturbance torque on the three axes; A reinforcement learning framework is used to learn and compensate for the unknown nonlinear dynamics in the augmented affine system model. The reinforcement signal is generated using error information, and a reinforcement learning-based interference observer is used to observe and compensate for internal and external interference. By introducing the sliding mode variable structure method into adaptive control to eliminate chattering, and combining disturbance observation and reinforcement learning, a reinforcement learning variable structure control law for roll missiles with non-affine structure is obtained. During the flight of the rolling missile, the reinforcement learning variable structure control law for rolling missiles with non-affine structures is executed to track the desired attitude angle.
2. The non-affine reinforcement learning variable structure control method for rolling missiles according to claim 1, characterized in that, The kinematic and dynamic equations of the rotating missile body are expressed as follows: 。 3. The non-affine reinforcement learning variable structure control method for rolling missiles according to claim 2, characterized in that, Using the execution network in reinforcement learning networks to approximate unknown nonlinear functions in the system : in, For Gaussian functions, To obtain the approximate value of the network, the integral penalty function of the control system is designed as follows: The design of the evaluation network fitting penalty function is as follows: in, To evaluate the network fitting error, To approximate the value of the execution network, the weight update laws for the execution network and the evaluation network are designed as follows: in, , , , Let the parameters to be designed satisfy the boundedness of the basis functions. , , , .
4. The non-affine reinforcement learning variable structure control method for rolling missiles according to claim 1, characterized in that, In the augmented affine system, the first two stages employ an adaptive back-inference method, while the last stage subsystem introduces the concept of variable structure to construct an adaptive sliding mode variable structure control law.
5. The non-affine reinforcement learning variable structure control method for rolling missiles according to claim 1, characterized in that, Combining adaptive control and sliding mode variable structure methods, the reinforcement learning variable structure adaptive control law is designed as follows: Design the virtual control signal for the first-order subsystem. for: Design a reinforcement learning-based interference observer: in For the intermediate variables of the interference observer, For the gain to be designed, For the function to be designed, the virtual control law of the second-order subsystem Designed as follows: 。 6. The non-affine reinforcement learning variable structure control method for a rolling missile according to claim 1, characterized in that, The system sliding surface is designed in the extended subsystem as follows: in Let the coefficients of the sliding surface be undetermined, satisfying For Hurwitz stability, the adaptive sliding mode variable structure control law is designed as follows: in, , These are the parameters to be designed.
7. A non-affine reinforcement learning variable structure control system for a rolling missile, characterized in that, It includes a model building module, a reinforcement learning module, an adaptive control law acquisition module, and an attitude stabilization module; The model building module is used to establish the attitude kinematics and dynamics equations of the rolling missile. Considering the non-affine characteristics of the rolling missile's control signals and the characteristics of the disturbances it experiences, an augmented affine attitude model of the rolling missile oriented towards control is established by adding an integral method. The attitude kinematics and dynamics of the rolling missile are integrated into a control-oriented non-affine system model, including external disturbances and unknown nonlinear dynamics. This model includes a two-order subsystem. Constructing an augmented affine system model: in: , , , , ; in, , These represent the missile's three attitude angles and the components of its attitude angular velocity along the three axes, respectively. It is the control input for the missile; This represents a system containing non-affine functions for control inputs; It is an unknown nonlinear dynamic of the system; Represents attitude disturbance torque; , , These are roll angle, yaw angle, and pitch angle, respectively. , , These are the components of the missile's angular velocity along the three axes; , , The moment of inertia is the rotational inertia of the three axes; , , The control torques are on the three axes. , , The disturbance torque on the three axes; The reinforcement learning module is used to learn and compensate for the unknown nonlinear dynamics in the augmented affine system model using a reinforcement learning framework, generate reinforcement signals using error information, and compensate for internal and external interference observations using a reinforcement learning-based interference observer. The adaptive control law acquisition module is used to combine adaptive control, reinforcement learning, interference observation and sliding mode variable structure method to design reinforcement learning variable structure control laws for roll missiles with non-affine structure. The attitude stabilization module is used to enable the rolling missile to perform the control law tracking or stabilize to the desired attitude angle during flight.
8. A computer device, characterized in that, It includes a processor and a memory, the memory being used to store a computer-executable program, the processor reading the computer-executable program from the memory and executing it, and the processor executing the program being able to implement the non-affine reinforcement learning variable structure control method for a rolling missile as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program that, when executed by a processor, enables the implementation of the non-affine reinforcement learning variable structure control method for a rolling missile as described in any one of claims 1-6.