A method for decision of variable aircraft steady flight deceleration maneuver based on reinforcement learning

By using reinforcement learning-based methods, a nonlinear model and aerodynamic moment model of the variant aircraft were established, the overall control structure of the variant aircraft was designed, and the agent was trained to optimize the variant strategy. This solved the problems of aerodynamics and control efficiency of the variant aircraft in level flight deceleration maneuvers, and achieved the best performance of the variant aircraft in specific maneuvers.

CN119670575BActive Publication Date: 2025-11-28NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411880125.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-11-28
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

Existing technologies struggle to guarantee optimal aerodynamics, flight efficiency, and control efficiency when variant aircraft perform maneuvers, especially during level flight deceleration maneuvers, where the potential of variant aircraft cannot be effectively utilized.

Method used

A reinforcement learning-based approach is used to establish a nonlinear model and aerodynamic moment model for the variant aircraft, design the overall control structure of the variant aircraft, train the agent using the reinforcement learning algorithm DQN, optimize the variant strategy of the variant aircraft in level flight deceleration maneuvers, including speed, attitude and altitude control, and use a reward and penalty function to guide the variant decision-making.

Benefits of technology

It achieves optimal aerodynamics and control efficiency for the variant aircraft during level flight deceleration maneuvers, enhancing handling and maneuverability. In particular, the asymmetric variable sweep wing structure improves lateral maneuverability, while the variable dihedral tail balances stability, ensuring the aircraft achieves optimal performance during specific maneuvers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119670575B_ABST
    Figure CN119670575B_ABST
Patent Text Reader

Abstract

The application relates to a variant airplane steady flight deceleration maneuver decision method based on reinforcement learning and belongs to the technical field of aviation. The application comprises the following steps: establishing a kinematic model and a dynamic model of a variant airplane; establishing an aerodynamic force and an aerodynamic moment model of the variant airplane and considering the aerodynamic force and the aerodynamic moment change caused by the variant; designing an overall control structure of the steady flight deceleration maneuver, using PI control on the speed, using backstepping control on the attitude angle to keep the attitude stable, using PID control on the height in the outer loop of the pitch angle to keep the height unchanged; designing a reward and punishment function of the steady flight deceleration maneuver; and designing an intelligent agent based on a reinforcement learning DQN algorithm, discretizing the possible degrees of the left and right sweep angles and the V-tail upper anti-angle to adapt to the discrete action space of the DQN algorithm. The application uses the reinforcement learning control method to ensure that the airplane can achieve optimal aerodynamic force, flight efficiency and control efficiency when performing a specific maneuver.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of aviation technology and relates to a variable aircraft steady flight deceleration maneuver decision-making method based on reinforcement learning. BACKGROUND

[0002] An aircraft with high efficient aerodynamic performance and the ability to quickly change shape to obtain maneuverability is an important development trend of future new-type fighter aircraft, and intelligent adaptive variable strategy and adaptive control method of the variable aircraft are important measures to improve the maneuverability of the new-type variable aircraft. Although certain progress has been made in the research on the variable aircraft in recent years, the research is still limited by the development of related technologies such as aerodynamic characteristics, variable characteristics and maneuver performance requirements.

[0003] There are many methods for studying the variable control mechanism. The variable structure can be taken as a system input, the variable aircraft model input can be expanded, and the traditional control method can be directly used for control. Meanwhile, the characteristics of the variable aircraft in different configurations can be analyzed separately, and then different control laws with different parameters can be designed for different variable configurations. However, no matter which control method is used, the variable aircraft is controlled as a conventional aircraft, and the maximum performance of the variable aircraft cannot be guaranteed.

[0004] As one of the most popular control methods at present, the reinforcement learning method has made breakthroughs in many fields after the improvement of computer computing power. Reinforcement learning is a method of continuous trial and error. Most of them determine the change gradient of the variable controller output through the size of the current reward and punishment, and finally guarantee that the system gets the maximum reward. Therefore, as long as the reward and punishment function is designed well, after a large number of sample training, the variable control mechanism can achieve the expected effect. SUMMARY

[0005] The technical problem solved by the application is:

[0006] In order to avoid the shortcomings of the prior art, the application provides a variable aircraft steady flight deceleration maneuver decision-making method based on reinforcement learning, which is used to solve the problem of optimal variable configuration of the variable aircraft in executing a maneuvering task. For the steady flight deceleration maneuver, the variable strategy method is studied through reinforcement learning to ensure that the aircraft can achieve optimal aerodynamic force, flight efficiency and control efficiency when executing the steady flight deceleration maneuver.

[0007] In order to solve the above technical problems, the technical scheme adopted by the application is:

[0008] A variable aircraft steady flight deceleration maneuver decision-making method based on reinforcement learning, characterized in that it comprises:

[0009] S1, establishing a nonlinear model of the variable aircraft, an aerodynamic force and an aerodynamic moment model;

[0010] S2, the nonlinear model established in step S1, the aerodynamic force and moment model variant aircraft model design variant aircraft overall control structure, including speed control method, height control method and attitude control method;

[0011] S3, on the basis of step S2, establish the environment of the steady flight deceleration maneuver task, to obtain the state information of the variant aircraft and the reward and punishment information of the steady flight deceleration maneuver in reinforcement learning;

[0012] S4, design the reward and punishment function of the steady flight deceleration maneuver task;

[0013] S5, design the reinforcement learning intelligent agent based on DQN algorithm, train the intelligent agent to obtain the final variant decision algorithm.

[0014] The further technical scheme of the application: in step S1, the nonlinear model of the variant aircraft and the aerodynamic force and moment model are established, specifically:

[0015] (a) the establishment of the kinematic model and the dynamic model of the variant aircraft

[0016] The force equation of the variant aircraft is:

[0017]

[0018] The above formula can obtain the dynamic model:

[0019]

[0020] Wherein, V is the speed of the aircraft, alpha is the angle of attack of the aircraft, beta is the side slip angle of the aircraft, T is the engine thrust, [p q r] T respectively, [D Y L] T are the drag, side force and lift, [F Ix F Iy F Iz ] T is the component of the inertia force caused by the variant process in the airflow coordinate system, and has

[0021]

[0022] Wherein, S x is the component of the aircraft static moment on the x-axis of the machine system;

[0023] The angular velocity equation of the variant aircraft is:

[0024]

[0025] Wherein, [phi theta psi] T is the attitude angle of the aircraft, Rolling moment, pitching moment and yawing moment, respectively, Ix M Iy M Iz ] T is the component of the moment of inertia force caused by morphing process in the body coordinate system, g is the acceleration of gravity,

[0026] The attitude angle equations and the navigation equations of the morphing aircraft are the same as those of the conventional aircraft, and are expressed as follows,

[0027]

[0028] (b) Establishment of the aerodynamic force and moment model of the morphing aircraft

[0029] The aerodynamic force and moment are expressed as:

[0030] Lift force:

[0031] L = C L Q S w

[0032] Drag force:

[0033] D = C D Q S w

[0034] Side force:

[0035] Y = C Y Q S w

[0036] Rolling moment:

[0037]

[0038] Pitching moment:

[0039] M A = C m Q S w c A

[0040] Yawing moment:

[0041] N A = C n Q S w b

[0042] wherein C L , C D and C Y are the lift coefficient, the drag coefficient and the side force coefficient, respectively; C l , C m and C nare the roll moment coefficient, the pitch moment coefficient and the yaw moment coefficient, respectively; is the dynamic pressure, and p is the air density;

[0043] Let the left and right elevon deflection angles be δ el and δ er , and the downward deflection is positive, which produces a pitch moment M A that is negative, i.e. a nose-down moment, and at the same time, since it is a V-tail configuration, the left and right elevons simultaneously produce a yaw moment, and the yaw moment produced by the right elevon is negative when the right elevon is deflected, so when the left elevon is deflected negatively and the right elevon is deflected positively, a negative yaw moment N A is produced; let the left and right aileron deflection angles be δ al and δ ar , and the downward deflection is positive, the left aileron is deflected negatively, and the right aileron is deflected positively, which produces a roll moment that is negative; in addition, when the sweepback angle Λ is 60°, it is considered that the force and moment changes produced by the sweepback angle are 0, so the influence of the sweepback angle on the aerodynamic force and moment can be calculated;

[0044] The aerodynamic force coefficients and the aerodynamic moment coefficients are expressed as:

[0045] The lift coefficient:

[0046]

[0047] The drag coefficient:

[0048]

[0049] The side force coefficient:

[0050]

[0051] The roll moment coefficient:

[0052]

[0053] The pitch moment coefficient:

[0054]

[0055] The yaw moment coefficient:

[0056]

[0057] wherein C L0 is the zero angle of attack lift coefficient, C Lα is the lift curve slope; C D0 is the zero angle of attack drag coefficient, C Dα , C Dβ are the derivatives of the drag coefficient with respect to the angle of attack and the sideslip angle, respectively; C m0 is the zero angle of attack pitch moment coefficient, and Cmα Cp is the derivative of the pitching moment coefficient with respect to the angle of attack mq Cp is the derivative of the pitching moment coefficient with respect to the pitch rate Yβ Cp is the derivative of the pitching moment coefficient with respect to the pitch rate lβ Cp is the derivative of the pitching moment coefficient with respect to the pitch rate nδ Cp is the derivative of the pitching moment coefficient with respect to the pitch rate lp Cp is the derivative of the pitching moment coefficient with respect to the pitch rate np Cp is the derivative of the pitching moment coefficient with respect to the pitch rate lr Cp is the derivative of the pitching moment coefficient with respect to the pitch rate nr Cp is the derivative of the pitching moment coefficient with respect to the pitch rate Lδe Cp is the derivative of the pitching moment coefficient with respect to the pitch rate Dδe Cp is the derivative of the pitching moment coefficient with respect to the pitch rate Yδe Cp is the derivative of the pitching moment coefficient with respect to the pitch rate lδe Cp is the derivative of the pitching moment coefficient with respect to the pitch rate mδe Cp is the derivative of the pitching moment coefficient with respect to the pitch rate nδe Cp is the derivative of the pitching moment coefficient with respect to the pitch rate Lδa Cp is the derivative of the pitching moment coefficient with respect to the pitch rate Dδa Cp is the derivative of the pitching moment coefficient with respect to the pitch rate Yδa Cp is the derivative of the pitching moment coefficient with respect to the pitch rate lδa Cp is the derivative of the pitching moment coefficient with respect to the pitch rate mδa Cp is the derivative of the pitching moment coefficient with respect to the pitch rate nδa Cp is the derivative of the pitching moment coefficient with respect to the pitch rate LΛ Cp is the derivative of the pitching moment coefficient with respect to the pitch rate DΛ Cp is the derivative of the pitching moment coefficient with respect to the pitch rate YΛ Cp is the derivative of the pitching moment coefficient with respect to the pitch rate lΛ Cp is the derivative of the pitching moment coefficient with respect to the pitch rate mΛ Cp is the derivative of the pitching moment coefficient with respect to the pitch rate nΛ Cp is the derivative of the pitching moment coefficient with respect to the pitch rate

[0058] The further technical scheme of the present application is that in step S2, the overall control structure of the variable aircraft is designed based on the variable aircraft model established in step S1, and specifically:

[0059] Firstly, PI control is used for the speed, and the control law is as follows:

[0060] T=k p,V (V ref -V)+k i,V ∫(V ref -V)dt

[0061] Wherein, V ref is the speed command, k p,V is the proportional coefficient, and k i,V is the integral coefficient;

[0062] The height control law includes two parts, which are the height to the pitch angle and the pitch angle to the elevator, wherein the pitch angle to the elevator control is included in the attitude control, and the PID control law of the height to the pitch angle is given first, and the form is as follows:

[0063]

[0064] where, θ ref is the pitch angle command, h ref is the height command, k p,h , k i,h , k d,h are proportional, integral, and derivative coefficients, respectively.

[0065] The lateral direction uses attitude control and sideslip angle control, where the attitude control uses backstepping method as follows:

[0066]

[0067] The sideslip angle control uses PI control, and when controlling the sideslip angle, the yaw angle velocity command in the attitude control is replaced by the sideslip angle as follows:

[0068] r ref = k p,β (β ref - β) + k i,β ∫(β ref - β)dt

[0069] where, β ref is the sideslip angle command, k p,β , k i,β are proportional and integral coefficients, respectively.

[0070] A further technical solution of the present application: in step S3, the environment of the steady flight deceleration maneuver task is established on the basis of step S2, to obtain the state information of the morphing aircraft and the reward and punishment information of the steady flight deceleration maneuver in reinforcement learning, specifically:

[0071] The morphing structure intelligent agent agent is trained using the steady flight deceleration maneuver;

[0072] By setting the corresponding reward reward, the intelligent agent agent is trained to converge in the direction of the maximum return, and finally the control of the desired morphing structure in the steady flight deceleration maneuver is obtained. In order to ensure that the reinforcement learning control application range is wider, some randomness needs to be added to each training task, including the randomness of the command and the randomness of the initial value.

[0073] The observation state of the environment has 9 quantities, which are [V, α, β, φ, θ, ψ, p, q, r] T These 9 state quantities; and not only the current state needs to be obtained, but also the next state needs to be calculated; for these 9 state quantities, the morphing aircraft nonlinear model is used, and the Runge-Kutta algorithm is used to update and calculate the next state.

[0074] The input of the environment, i.e., the output of the agent, is three angle quantities of the variant structure, which are respectively a left sweep angle Λ l , a right sweep angle Λ r and an upper dihedral angle Λ d of the tail wing; before starting each simulation, an initialization operation is set, i.e., in the speed control loop, a speed instruction signal is randomized;

[0075] In the reinforcement learning training of the steady flight deceleration maneuver, the reward r t of the current moment is calculated according to the formula of the reward and punishment function, and then the return u t size can be calculated from r t ; the current return u t represents the sum of rewards in all rounds at t moment and after t moment; in order to emphasize that the future reward is not as important as the current reward, the discount rate γ is defined to define the discounted return; if the return is larger, it means that the speed of the aircraft can track the instruction signal faster; otherwise, the smaller the return is, the slower the aircraft tracks the instruction signal; in addition, a certain randomness is added in the final control output, so as to ensure that the strategy will not fall into local optimization.

[0076] The whole training process is as follows: when the simulation starts, the aircraft gets a random speed instruction, and performs the corresponding deceleration maneuver, in which the potential energy of the aircraft does not change; then the single-step update is performed through the time punishment of each step, and when the deceleration maneuver task is completed, a larger reward is obtained, at this time, the larger reward is obtained but not updated; in this way, the dynamic performance of the aircraft executing the deceleration maneuver task can be indirectly obtained, and the update of the strategy network Actor and the value network Critic is performed according to the dynamic performance; the update of the strategy network changes the output of the aircraft variant structure in real time, if the reward is larger, the network is updated in this direction larger, otherwise, the smaller the reward is, the smaller the network is updated; in this way, after a large number of rounds of training, it can be ensured that the strategy network can control the variant structure well according to the current state.

[0077] The further technical scheme of the present application is that the reward and punishment function of the steady flight deceleration maneuver task is designed in step S4, and specifically:

[0078] r t =r s (t)+r d (t)+r tp (t)

[0079]

[0080] r tp (t)=-1

[0081] Wherein, r d (t) represents the dynamic performance reward, r s(t) represents a steady-state performance reward and penalty, r tp (t) represents a time penalty, and c1 and c2 are normal numbers.

[0082] A further technical solution of the present application is that in step S5, a reinforcement learning agent based on the DQN algorithm is designed, and the agent is trained to obtain a final variant decision algorithm, specifically:

[0083] In the network structure of the DQN, the input layer observation has 9 nodes, which are 9 state quantities of the aircraft [V, a, b, f, q, y, p, q, r]. T Then, it sequentially passes through a full connection layer, a relu activation function, a full connection layer, a relu activation function, and an output layer.

[0084] Since the DQN algorithm can only be used for discrete action space, and the variant output of the variant aircraft uses continuous action space, it is necessary to discretize the possible actions, i.e., to discretize the left and right sweep angles and the V tail up angle.

[0085] In addition, as an off-policy method, the DQN is designed to store past experiences in a buffer:

[0086] D←D∪(s t ,a t ,r t ,s t+1 )

[0087] Each time the training is performed, a group of data packets (s t ,a t ,r t ,s t+1 ) is selected from the buffer, and then the DQN network is updated. The update of the DQN network parameters is:

[0088]

[0089] wherein a represents a learning rate, represents the gradient of the optimization function J Q (0);

[0090] The TD algorithm is used to update the network, so

[0091]

[0092] The TD error of the DQN network update is:

[0093]

[0094] wherein Q(s t ,at ;θ t ) represents the neural network structure.

[0095] A computer system is characterized by comprising: one or more processors, and a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method described above.

[0096] A computer-readable storage medium is characterized by storing computer-executable instructions, which, when executed, are used to implement the above-described method.

[0097] A computer program product is characterized by including computer-executable instructions, which, when executed, are used to implement the above-described method.

[0098] The beneficial effects of this invention are as follows:

[0099] This invention provides a reinforcement learning-based decision-making method for level flight deceleration maneuvers of vari-plane aircraft, designing novel vari-wing strategies and configurations to improve aircraft stability and enhance its handling and maneuverability. The invention utilizes reinforcement learning control methods to ensure optimal aerodynamics, flight efficiency, and control efficiency when performing specific maneuvers. Furthermore, the asymmetric variable-sweep wing structure further increases the aircraft's lateral maneuverability and offers advantages such as a smaller turning radius; the dihedral tail balances longitudinal and lateral stability. The invention focuses on vari-plane control methods for the level flight deceleration maneuver. Using the reinforcement learning-based DQN algorithm, the optimal vari-plane configuration for level flight deceleration maneuvers is determined. Attached Figure Description

[0100] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.

[0101] Figure 1 A framework diagram for a reinforcement learning-based decision-making algorithm for deceleration maneuvers of variant aircraft is shown.

[0102] Figure 2 Variant aircraft object graph for reinforcement learning applications.

[0103] Figure 3 This is a diagram of the overall control structure for level flight deceleration maneuvers.

[0104] Figure 4 This is a block diagram of the reward and punishment function.

[0105] Figure 5 Environmental design diagram for level flight deceleration maneuvers.

[0106] Figure 6 is a DQN network structure diagram.

[0107] Figure 7 is a reward and punishment diagram during training.

[0108] Figure 8 is a reinforcement learning agent variant structure output diagram during a steady flight deceleration maneuver.

[0109] Figure 9 is a speed change diagram.

[0110] Figure 10 is an attitude angle change diagram. DETAILED DESCRIPTION

[0111] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. In addition, the technical features involved in the various embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.

[0112] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged as appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0113] In the reinforcement learning process, it is necessary to amplify the influence of the variant structure on the flight performance as much as possible, which requires the aircraft to fly steadily; when the aircraft is in an unstable flight state, the change of the flight state may easily make the reinforcement learning not easy to converge. Therefore, the content of the present application controls the aircraft to fly at a constant height, ensures that the gravitational potential energy of the aircraft remains unchanged, prevents the mutual conversion of potential energy and kinetic energy from affecting the convergence of reinforcement learning, and makes reinforcement learning only learn and update the aircraft's variant strategy.

[0114] Since the height needs to be kept unchanged and the attitude needs to be kept unchanged as far as possible during the deceleration maneuver, a speed control law, a height control law and an attitude control law need to be designed respectively, then an environment and a reward and punishment function of the flat flight deceleration maneuver are designed, a reinforcement learning agent based on DQN is designed to learn the morphing mechanism (wing sweep angle, V-tail up angle) of the morphing aircraft during the deceleration maneuver, and the optimal morphing configuration under the maneuver is obtained, finally, the morphing aircraft flat flight deceleration maneuver decision algorithm based on reinforcement learning is simulated.

[0115] As shown in Figure 1 The application provides a morphing aircraft flat flight deceleration maneuver decision method based on reinforcement learning, which comprises the following steps:

[0116] (1) a nonlinear model, an aerodynamic force and an aerodynamic moment model of the morphing aircraft are established;

[0117] (2) a speed control method, a height control method and an attitude control method of the morphing aircraft are designed;

[0118] (3) an environment of the flat flight deceleration task is established to obtain the state information of the morphing aircraft and the reward and punishment information of the flat flight deceleration maneuver in reinforcement learning;

[0119] (4) a reward and punishment function of the flat flight deceleration maneuver task is designed;

[0120] (5) a reinforcement learning agent based on the DQN algorithm is designed, and the agent is trained to obtain a final morphing decision algorithm.

[0121] In order to enable those skilled in the art to better understand the application, the application will be described in detail below with reference to specific embodiments.

[0122] Embodiment 1

[0123] The object to which the reinforcement learning is applied in this embodiment is a morphing aircraft as shown in the figure, the morphing aircraft has a variable structure of left and right wing sweep angles and V-tail up angles, and the sweep angles of the left and right wings can be asymmetrically changed. Figure 2

[0124] (a) establishment of the kinematic model and the dynamic model of the morphing aircraft

[0125] The force equation of the morphing aircraft is:

[0126]

[0127] The above formula can be obtained as:

[0128]

[0129] ​where V is the aircraft velocity, a is the angle of attack, β is the sideslip angle, T is the engine thrust, [p q r] T are the three angular velocities, [D Y L] T are the drag, side force and lift, [F Ix F Iy F Iz ] T are the components of the inertia force in the airflow axes, and

[0130]

[0131] where S x is the component of the aircraft static moment in the body x axis.

[0132] The angular velocity equations of the morphing aircraft are:

[0133]

[0134] where [φ θ ψ] T are the aircraft attitude angles, are the roll, pitch and yaw moments, [M Ix M Iy M Iz ] T are the components of the inertia moment in the body axes, and g is the gravitational acceleration,

[0135] The attitude angle equations and the navigation equations of the morphing aircraft are the same as those of the ordinary aircraft, and are expressed as follows:

[0136]

[0137] (b) Establishment of the morphing aircraft aerodynamic force and aerodynamic moment model

[0138] The aerodynamic force and aerodynamic moment can be expressed as:

[0139] Lift:

[0140] L = C L QS w

[0141] Drag:

[0142] D = C D QS w

[0143] Side force:

[0144] Y = C Y QS w

[0145] Rolling moment:

[0146]

[0147] Pitching moment:

[0148] M A = C m QS w c A

[0149] Yawing moment:

[0150] N A = C n QS w b

[0151] where C L , C D and C Y are lift, drag and side force coefficients respectively; C l , C m and C n are rolling, pitching and yawing moment coefficients respectively. is dynamic pressure, and p is air density.

[0152] Let the left and right elevator deflection angles be δ el and δ er , and the downward deflection is positive, and the generated pitching moment M A is negative, that is, it generates a nose-down moment, and at the same time, since it is a V-tail configuration, the left and right elevators simultaneously generate a yawing moment, and the yawing moment generated by the right deflection is negative, so when the left elevator is negative and the right elevator is positive, a negative yawing moment N A is generated. Let the left and right aileron deflection angles be δ al and δ ar , and the downward deflection is positive, and the left aileron is negative and the right aileron is positive, and the generated rolling moment L A is negative. In addition, when the sweepback angle Λ is 60°, it is considered that the force and moment changes generated by the sweepback angle are 0, and the influence of the sweepback angle on the aerodynamic force and moment can be calculated.

[0153] The aerodynamic force coefficients and the aerodynamic moment coefficients can be expressed as:

[0154] Lift coefficient:

[0155]

[0156] Drag coefficient:

[0157]

[0158] Side force coefficient:

[0159]

[0160] Rolling moment coefficient:

[0161]

[0162] Pitching moment coefficient:

[0163]

[0164] Yawing moment coefficient:

[0165]

[0166] where C L0 is the zero-lift coefficient, C Lα is the lift curve slope; C D0 is the zero-lift drag coefficient, C Dα , C Dβ are the derivatives of the drag coefficient with respect to the angle of attack and the sideslip angle, respectively; C m0 is the zero-lift pitching moment coefficient, C mα is the derivative of the pitching moment coefficient with respect to the angle of attack, C mq is the derivative of the pitching moment coefficient with respect to the pitch rate; C Yβ , C lβ , C nβ are the derivatives of the side force coefficient, the rolling moment coefficient, and the yawing moment coefficient with respect to the sideslip angle, respectively; C lp , C np are the derivatives of the rolling moment coefficient and the yawing moment coefficient with respect to the roll rate; C lr , C nr are the derivatives of the rolling moment coefficient and the yawing moment coefficient with respect to the yaw rate; C Lδe , C Dδe , C Yδe , C lδe , C mδe , C nδe are the control derivatives of the elevators; C Lδa , C Dδa , C Yδa , C lδa , C mδa , C nδa are the control derivatives of the ailerons; C LΛ Λ, C DΛ Λ, C YΛ Λ, C lΛ Λ, C mΛ Λ, C nΛ Λ are the changes in the aerodynamic force coefficients and the aerodynamic moment coefficients due to the wing sweepback angle and the tail dihedral angle.

[0167] (c) Control structure of the steady flight deceleration maneuver

[0168] The present application is mainly to study the performance of the variable aircraft in different configurations in the steady flight deceleration maneuver. Therefore, in the process of reinforcement learning, it is necessary to maximize the influence of the variable structure on the flight performance, which requires the aircraft to fly steadily. When the aircraft is in unstable flight, the change of flight state may easily make the reinforcement learning not easy to converge. The overall control structure as shown in Figure 3 The present application will control the aircraft to fly at a constant height, ensuring that the gravitational potential energy of the aircraft remains unchanged, preventing the mutual conversion of potential energy and kinetic energy from affecting the convergence of reinforcement learning. In order to achieve steady flight of the aircraft, it is necessary to ensure that the aircraft has good control performance, in the longitudinal direction, the aircraft is controlled to fly at a constant height, ensuring that the gravitational potential energy of the aircraft changes little; in the lateral direction, the aircraft is initially controlled to fly without sideslip and roll. However, the performance changes embodied by the variable structure cannot be fully reflected in the case of only steady flight, so in the process of reinforcement learning training, the aircraft needs to be made to perform a steady flight deceleration maneuver, and then the best control of the variable structure under this maneuver is obtained.

[0169] First, PI control is used for speed, and the control law is as follows:

[0170] T=k p,V (V ref -V)+k i,V ∫(V ref -V)dt

[0171] Where V ref is the speed command, k p,V is the proportional coefficient, and k i,V is the integral coefficient.

[0172] The height control law includes two parts, height to pitch angle and pitch angle to elevator, respectively, and the pitch angle to elevator control is included in the attitude control. Here, the PID control law of height to pitch angle is given, which is as follows:

[0173]

[0174] Where θ ref is the pitch angle command, h ref is the height command, k p,h , k i,h , and k d,h are the proportional, integral, and differential coefficients, respectively.

[0175] In the lateral direction, attitude control and sideslip angle control are used, and the attitude control uses backstepping method as follows:

[0176]

[0177] The sideslip angle control uses PI control, and when the sideslip angle is controlled, the yaw angle velocity instruction in the attitude control is replaced by the sideslip angle, as follows:

[0178] r ref = k p,β (β ref - β) + k i,β ∫(β ref - β)dt.

[0179] Where β ref is the sideslip angle instruction, k p,β and k i,β are proportional and integral coefficients respectively.

[0180] (d) Cruise deceleration maneuver reward-punishment function design

[0181] The main purpose of the reinforcement learning algorithm is to enable the agent to obtain the maximum cumulative reward, and the reward and punishment of each step will guide the convergence of the action network. Generally speaking, there is no big rule in the design of the reward and punishment function, and almost all of them are designed and adjusted according to experience. In some complex behaviors, reinforcement learning may cause some "strange consequences" due to the design of the reward and punishment function, which maximize the benefits as much as possible, but deviate from our goal. Therefore, when designing the reward and punishment function, we will adjust the parameters or change the structure according to the results.

[0182] In the present application, the reward and punishment function design process of the cruise deceleration maneuver is as follows, and the structure diagram of the reward and punishment function is as shown in Figure 4 .

[0183] For the control task of acceleration and deceleration, we expect it to be able to track the instruction speed quickly and efficiently. According to the foregoing content, the speed control has a corresponding control law, that is, by controlling the thrust, the tracking of the speed instruction is ensured. This part of the content is to use the control of the variable structure to change the aerodynamic parameters of the fuselage, so that the original control can be more efficient. Therefore, the design standard of the reward and punishment function is the same as that of the traditional control, including two parts: dynamic performance reward and steady-state performance reward.

[0184] For dynamic performance, we mainly hope that the speed can reach the instruction as soon as possible. Therefore, a "penalty of -1" can be given to each step reward, and then a positive reward is given when the speed approaches the instruction.

[0185] For the steady-state performance, we hope that the speed can be stabilized near the target as soon as possible and for a period of time, and when the condition is met, a positive reward is given and the round is ended.

[0186] Therefore, according to the above requirements, the reward and punishment function can be expressed as follows:

[0187] r(t) = r s (t)+r d (t)+r tp (t)

[0188]

[0189] r tp (t)=-1

[0190] Where, r d (t) represents the dynamic performance reward / penalty, r s (t) represents the steady-state performance reward / penalty, r tp (t) represents the time penalty, and c1 and c2 are positive constants.

[0191] (e) Level Flight Deceleration Maneuvering Environment Design

[0192] like Figure 5 As shown, in this invention, a variant structure agent is trained using level flight deceleration maneuvers. By setting appropriate rewards, the agent is trained to converge in the direction of maximizing the reward, ultimately obtaining the desired control of the variant mechanism during level flight deceleration maneuvers. During training, it is necessary to maintain a fixed altitude in the longitudinal direction. To ensure a wider range of applications for reinforcement learning control, some randomness needs to be added to each training task, including randomness in commands and randomness in initial values.

[0193] The observed state of the environment consists of nine quantities: V, α, β, φ, θ, ψ, p, q, and r. It is necessary not only to obtain the current state but also to calculate the state at the next moment. For these nine state quantities, a variant aircraft kinematic model is used, updated and the state for the next time step is calculated via the Runge-Kutta algorithm. The input to the environment, i.e., the agent's output, consists of three angular quantities of the variant structure: the left sweep angle Λ. l Right sweep angle Λ r and the dihedral angle Λ of the tail fin d Finally, before each simulation begins, an initialization operation is set up, which randomizes the speed command signal in the speed control loop. This allows the trained results to adapt to different speed targets, thereby making the level flight deceleration task tend towards the optimal.

[0194] In reinforcement learning training for level flight deceleration, the reward / penalty r at the current moment is calculated based on the formula of the reward / penalty function. t And then by r t The return u can be calculated. t Size. Current return u tThe sum of the rewards in all rounds at time t and after is represented. In order to emphasize that the future reward is not as important as the current reward, the discounted return can be defined by a discount rate γ. If the return is larger, it means that the speed of the aircraft can track the command signal faster; on the contrary, the smaller the return is, the slower the aircraft tracks the command signal. Therefore, the training goal of reinforcement learning can be understood as maximizing the return. Since the return as a random variable is related to the current behavior and the current state, in order to evaluate the pros and cons of the current behavior, the action value can be defined as the expectation of the return with respect to the current behavior. The larger the action value is, the more effective the control strategy of the variant structure can improve the performance of the aircraft. The action value can help update the policy network of the variant structure, and ultimately obtain a satisfactory control strategy.

[0195] The task of the constant altitude speed reduction maneuver is to find a variant strategy so that the aircraft can complete the speed reduction maneuver better and faster. Therefore, in the initialization of the simulation, the speed command will be randomly processed to ensure that the data during training is in different degrees of speed reduction state of the aircraft. In order to highlight the influence of the variant structure on the flight state, the constant altitude flight is guaranteed to exclude the conversion between potential energy and kinetic energy of the aircraft due to changes in altitude, which interferes with the influence of the variant structure on the speed reduction of the aircraft. In addition, a certain randomness will be added to the final control output to ensure that the strategy will not fall into a local optimum.

[0196] The whole training process is as follows: when the simulation starts, the aircraft gets a random speed command and performs the corresponding speed reduction maneuver, during which the potential energy of the aircraft does not change. Then the single-step update is performed through the time penalty of each step, and a larger reward is obtained when the speed reduction maneuver task is completed, at which time the larger reward is not updated. In this way, the dynamic performance of the aircraft executing the speed reduction maneuver task can be indirectly obtained, and the update of the policy network Actor and the value network Critic is performed according to this. The update of the policy network will change the output of the aircraft variant structure in real time. If the reward is larger, the network will be updated in this direction, otherwise it will be updated smaller. After a large number of rounds of training, it can be ensured that the policy network can control the variant structure well according to the current state.

[0197] (f) Reinforcement learning agent design based on DQN

[0198] The network structure of DQN is shown in Figure 6 In the DQN algorithm, there is only one deep neural network, namely the optimal action value network. In practical application of the DQN algorithm, in order to prevent the adverse effects of bootstrapping, a target network with the same structure as the value network will also be designed.

[0199] In the network structure of the DQN of the application, the input layer observation has 9 nodes, which are 9 state quantities V, a, b, f, q, y, p, q, r of the aircraft, then passes through a fully connected layer with 882 nodes, and then passes through a relu activation function, and then passes through a fully connected layer with 882 nodes, and then passes through a relu activation function, and then passes through an output layer with 294 nodes. Since the DQN algorithm can only be used for discrete action space, and the output of the variable aircraft uses continuous action space, it is necessary to discretize the possible actions. Considering the problem of computing power, the left and right sweep angles and the V tail up angle are discretized every 6°, as follows:

[0200] Λ l = 24:6:60

[0201] Λ r = 24:6:60

[0202] Λ d = 30:6:60

[0203] The code is as follows:

[0204]

[0205] Therefore, the number of action spaces obtained by this discretization method is:

[0206] N action = 7*7*6 = 294

[0207] This is also the reason why the output layer of this DQN has 294 nodes. In addition, as an off-policy method, DQN can be designed to store past experiences in a buffer.

[0208] D <- D U (s t , a t , r t , s t+1 )

[0209] Each time the training is performed, a group of data packets (s t , a t , r t , s t+1 ) is selected from the buffer, and then the DQN network is updated. The update of the DQN network parameters is as follows:

[0210]

[0211] where a represents the learning rate, represents the gradient of the optimization function J Q (0).

[0212] Here the TD algorithm is used to update the network, so

[0213]

[0214] The TD error for DQN network update is:

[0215]

[0216] where Q(s t ,a t ; θ t ) represents the neural network structure.

[0217] The designed reinforcement learning-based variable aircraft deceleration maneuver decision algorithm is simulated and verified as follows:

[0218] The initial condition of the aircraft is a steady flight state, the height is 1000 m, the speed is 60 m / s, and the angle of attack is 4.23°.

[0219] First, the single-step reward and the average reward of 5 steps after training 1000 times are given, as shown in Figure 7 From Figure 7 , it can be seen that the round reward gradually converges to a larger value after training, and the effect is not very good when continuing to train, and it will start to deteriorate. Finally, when the training times are adjusted to 1000, the performance of DQN is better. The trained agent will be used for simulation verification. The simulation results are shown in Figure 8 .

[0220] The trained agent is used for simulation, and the response of the speed and attitude is obtained, as shown in Figure 9 and Figure 10 .

[0221] According to the simulation results of the steady flight deceleration, it can be seen that when the aircraft decelerates, the external thrust is 0 at this time, the reinforcement learning method is trained, and the asymmetric swept wing makes the aircraft roll, and then the attitude control of the aircraft at this time makes the aircraft roll back to the middle, and at the same time, these control surfaces also increase the drag of the aircraft, so that the aircraft has a large deceleration. Specifically, in 10 seconds, the speed of the aircraft decreases from 60 m / s to 43.3 m / s.

[0222] The above simulation results prove the effectiveness of the variable aircraft steady flight deceleration maneuver decision algorithm designed based on reinforcement learning, which can successfully train and optimize the variable strategy of the variable structure, and finally ensure that the variable aircraft can realize the rapid reduction of flight speed only by changing the variable structure under the condition that the height and attitude of the variable aircraft are basically unchanged.

[0223] The above merely illustrates the specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any skilled person in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present application, and these modifications or replacements shall be covered within the protection scope of the present application.

Claims

1. A method for decision-making on deceleration maneuvers during level flight of a variant aircraft based on reinforcement learning, characterized in that, Comprise: S1, the nonlinear model of the variable aircraft, the aerodynamic force and the aerodynamic moment model are established; Specifically: (a) The establishment of the kinematic model and the dynamic model of the variable aircraft The force equation of the variable aircraft: The above formula can obtain the dynamic model: where is the aircraft velocity, is the aircraft angle of attack, is the aircraft sideslip angle, is the engine thrust, are the three angular velocities, are the drag, side force and lift forces, respectively, is the component of the inertia force caused by the morphing process in the airflow coordinate system, and has wherein is the component of the aircraft static moment on the system of axes on the axis The angular velocity equation of the variable aircraft: wherein is the aircraft attitude angle, are the roll, pitch and yaw moments, respectively, is the component of the moment of inertia due to the morphing process in the body axis system, is the gravitational acceleration, , , , , , , , , , , , , , , ; The attitude angle equation and the navigation equation of the variable aircraft are the same as those of the ordinary aircraft, which are expressed as follows, (b) The establishment of the aerodynamic force and the aerodynamic moment model of the variable aircraft The aerodynamic force and the aerodynamic moment are expressed as: Lift: Drag: Lateral force: Rolling moment: Pitching moment: Yaw moment: wherein, , and are the lift, drag and side force coefficients, respectively; , and are the roll, pitch and yaw moment coefficients, respectively; is the dynamic pressure, is the air density; Let the left and right elevators deflection angle be and , the downward deflection is positive, the generated pitching moment is negative, that is, it generates a nose-down moment, at the same time, since it is a V-tail configuration, the left and right elevators generate a yawing moment at the same time, the yawing moment generated by the right deflection is negative, so when the left elevator is negative and the right elevator is positive, a negative yawing moment is generated ; let the left and right aileron deflection angle be and , the downward deflection is positive, the left aileron is negative deflection, the right aileron is positive deflection, the generated rolling moment is negative; in addition, when the rear sweep angle is 60°, it is considered that the force and moment change generated by the rear sweep angle is 0, the influence of the rear sweep angle on the aerodynamic force and the aerodynamic moment can be calculated. The aerodynamic force coefficients and the aerodynamic moment coefficients are expressed as: Lift coefficient: Drag coefficient: Lateral force coefficient: Rolling moment coefficient: Pitching moment coefficient: Yaw moment coefficient: wherein C L0 is the zero-lift coefficient, C L' is the slope of the lift curve; C D0 is the zero-lift drag coefficient, C D' is the derivative of the drag coefficient with respect to the angle of attack, C D" is the derivative of the drag coefficient with respect to the sideslip angle; C M0 is the zero-lift pitching moment coefficient, C M' is the derivative of the pitching moment coefficient with respect to the angle of attack, C M" is the derivative of the pitching moment coefficient with respect to the pitch rate; C Y is the side force coefficient, C L is the lift coefficient, C L' is the derivative of the side force coefficient with respect to the sideslip angle, C L" is the derivative of the roll moment coefficient with respect to the roll rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; C L'" is the derivative of the yaw moment coefficient with respect to the yaw rate; S2, the nonlinear model, the aerodynamic force and the aerodynamic moment model of the variable aircraft model established in step S1 are designed The overall control structure of the variable aircraft, including the speed control method, the height control method and the attitude control method; S3, on the basis of step S2, the environment of the constant flight deceleration maneuver task is established, which is used to obtain the state information of the variable aircraft and the reward and punishment information of the constant flight deceleration maneuver in reinforcement learning; S4, the reward and punishment function of the constant flight deceleration maneuver task is designed; S5, the reinforcement learning agent based on DQN algorithm is designed, and the final variable decision algorithm is obtained by training the agent.

2. The method of claim 1, wherein, In step S3, on the basis of step S2, the environment of the constant flight deceleration maneuver task is established, which is used to obtain the state information of the variable aircraft and the reward and punishment information of the constant flight deceleration maneuver in reinforcement learning, specifically: The variable structure agent agent is trained using the constant flight deceleration maneuver; By setting the corresponding reward reward, the agent agent converges to the direction of the maximum return, and finally the control of the variable structure in the constant flight deceleration maneuver is obtained; In order to ensure that the application range of reinforcement learning control is wider, some randomness needs to be added to each training task, including the randomness of the instruction and the randomness of the initial value; The observed states of the environment have 9 quantities, which are These 9 state quantities; and not only the current state needs to be obtained, but also the next state needs to be calculated; for these 9 state quantities, the variant aircraft nonlinear model is used, and the Runge-Kutta algorithm is used to update and calculate the next state. The input of the environment, i.e. the output of the agent, is the 3 angles of the variant configuration, respectively the left sweep angle , the right sweep angle and the up dihedral angle of the tail ; at the beginning of each simulation, an initialization operation is set, i.e. in the speed control loop, the speed command signal is randomized; In the intensive learning training of level flight deceleration maneuvers, the reward and penalty at the current moment are calculated according to the formula of the reward and penalty function. And then from The return can be calculated. Size; Current Return express The sum of rewards at the current moment and in all subsequent rounds; to emphasize that future rewards are less important than current rewards, a discount rate can be used. Define the discount reward; a larger reward indicates that the aircraft can track the command signal faster; conversely, a smaller reward indicates that the aircraft tracks the command signal slower. In addition, a certain degree of randomness will be added to the final control output to ensure that the strategy does not get stuck in local optima. The whole training process is: when the simulation starts, the aircraft gets a random speed instruction, and performs the corresponding deceleration maneuver. The potential energy of the aircraft does not change in this process; Then the single step is updated through the time penalty of each step, and there is a larger reward when the deceleration maneuver task is completed, which is not updated at this time; In this way, the dynamic performance of the aircraft executing the deceleration maneuver task can be indirectly obtained, and the update of the policy network Actor and the value network Critic is carried out according to this; The update of the policy network will change the output of the variable structure of the aircraft in real time. If the reward is larger, the network will update in this direction larger, and vice versa. After a large number of rounds of training, the policy network can ensure that the variable structure can be well controlled according to the current state.

3. The method of claim 1, wherein, In step S4, the reward and punishment function of the constant flight deceleration maneuver task is designed, specifically: wherein, represents a dynamic performance reward, represents a steady state performance reward, represents a time penalty, and is a positive constant.

4. A computer system, characterized by Comprise: One or more processors, a computer readable storage medium, for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors realize the method of claim 1.

5. A computer-readable storage medium, characterized in that Computer executable instructions are stored, which when executed, implement the method of claim 1.

6. A computer program product, characterised in that Computer executable instructions are included, which when executed, implement the method of claim 1.

Citation Information

Patent Citations

  • High-speed aircraft attitude control method based on reinforcement learning and application

    CN118394099A

  • Stratospheric airship trajectory tracking method based on reinforcement learning optimal control

    WO2024216870A1