Reinforcement learning control method and system for attitude of spinning aircraft under saturated nonlinearity

Through the dynamic model of the autogyro and the reinforcement learning control method, the attitude limit cycle oscillation problem caused by the saturation nonlinearity of the autogyro servo is solved, and the automatic optimization and high-precision control of the attitude control parameters are achieved.

CN119512159BActive Publication Date: 2025-09-12CIVIL AVIATION UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411674460.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2025-09-12
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

The saturation nonlinearity of the servo of a spinning aircraft under high-speed rolling causes attitude limit cycle oscillation and instability. Existing control methods fail to effectively solve the internal nonlinearity of the servo, and parameter optimization relies on manual adjustment, which is cumbersome and not applicable to nonlinear models.

Method used

The reinforcement learning control method is adopted to obtain the mathematical rudder deflection angle command value through the spinning aircraft dynamics model, decompose it into the physical rudder deflection angle command value, consider the saturation nonlinearity of the servo, use the expanded observer to estimate and compensate the total disturbance, optimize the attitude control parameters, and automatically adjust the attitude control in combination with the reinforcement learning network.

Benefits of technology

The acquisition efficiency and accuracy of attitude control parameters are improved, higher control precision is achieved, attitude limit cycle oscillation is suppressed, and system stability is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119512159B_ABST
    Figure CN119512159B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for attitude reinforcement learning control of a spinning aircraft under saturated nonlinearity, comprising: obtaining command values ​​for mathematical rudder angles for yaw and pitch channels; obtaining command values ​​for physical rudder angles for the yaw and pitch channels based on the command values ​​for the mathematical rudder angles; obtaining the physical rudder angles for the yaw and pitch channels based on the command values ​​for the physical rudder angles; generating mathematical rudder angles for the yaw and pitch channels based on the physical rudder angles; and generating a state of a controlled object based on the command values ​​for the quasi-angle of attack and the quasi-sideslip angle at the current moment and an output state, and transmitting the state to a trained reinforcement learning network model. The present invention takes saturated nonlinearity into account when obtaining the physical rudder angles, can automatically adjust attitude control parameters, improve the efficiency and accuracy of acquiring attitude control parameters, and thereby achieve higher control accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of aircraft attitude control, and in particular to a method and system for controlling the attitude of a spinning aircraft under saturated nonlinearity through reinforcement learning. Background Art

[0002] For autorotating vehicles, the roll channel maintains high-speed, uncontrolled rotation, necessitating attitude or overload control of the pitch and yaw channels to achieve the desired flight trajectory. For example, this involves controlling the pitch channel's angle of attack and the yaw channel's sideslip angle, or controlling the overload of the pitch and yaw channels. Autorotating vehicles are often controlled using a servo system, where deflection of the control surfaces changes the flight attitude. Unlike non-rotating vehicles, the high-speed roll of autorotating vehicles results in complex coupling effects between the dynamics of the pitch and yaw channels. These effects primarily include inertial coupling caused by the gyroscopic effect, aerodynamic coupling caused by the Magnus effect, and control coupling caused by servo response delays. Furthermore, high-speed roll increases the frequency of servo commands. The saturation nonlinearity within the servo can lead to large errors in tracking high-frequency commands, resulting in severe attitude limit cycle oscillations and even instability within the outer-loop attitude control system.

[0003] Currently, the main approach to controlling autogyro vehicles is to model the servo as a first- or second-order inertial link. However, this approach ignores the saturation nonlinearity within the servo. However, servo saturation can lead to severe attitude limit cycle oscillations and even instability. Therefore, servo saturation must be considered in controller design. While some control schemes have considered the internal backlash nonlinearity of the servo, these have only been addressed through stability analysis and not through controller design and parameter optimization. Furthermore, the control parameters of autogyro vehicles are manually adjusted and fixed. Manual parameter adjustment relies primarily on trial and error based on personal experience, a tedious and time-consuming process. Furthermore, existing parameter optimization methods rely on linear models of the object, such as pole placement and loop shaping in overload autopilots. These methods are also inapplicable when servo nonlinearities are present in the controlled object. Summary of the Invention

[0004] In view of the above technical problems, the technical solution adopted by the present invention is:

[0005] According to a first aspect of the present invention, a method for attitude reinforcement learning control of a spinning aircraft under saturated nonlinearity is provided, the method comprising the following steps:

[0006] S100, based on current attitude control parameters, the output state of the controlled object, and the command values ​​of the quasi-angle of attack and quasi-sideslip angle of the controlled object received at the current moment, obtain the command value of the mathematical rudder angle of the current yaw channel and the command value of the mathematical rudder angle of the current pitch channel; wherein the controlled object is a dynamic model of a spinning aircraft, and the output state of the controlled object includes the quasi-angle of attack, the quasi-sideslip angle, the yaw angular velocity, and the pitch angular velocity.

[0007] S200 , decompose the mathematical rudder angle instruction value of the current yaw channel and the mathematical rudder angle instruction value of the current pitch channel respectively to obtain the physical rudder angle instruction value of the current yaw channel and the physical rudder angle instruction value of the current pitch channel.

[0008] S300 , obtaining the physical rudder deflection angle of the current yaw channel and the physical rudder deflection angle of the current pitch channel based on the instruction value of the physical rudder deflection angle of the current yaw channel and the instruction value of the physical rudder deflection angle of the current pitch channel.

[0009] S400: Generate a mathematical rudder angle for the current yaw channel and a mathematical rudder angle for the current pitch channel based on the physical rudder angle for the current yaw channel and the physical rudder angle for the current pitch channel, and send them to the controlled object to obtain the output state of the controlled object at the current moment.

[0010] S500. Based on the command value of the quasi-angle of attack, the command value of the quasi-sideslip angle, and the output state at the current moment, generate a state of the controlled object at the current moment as the current state of the controlled object. The state of the controlled object includes a control error of the quasi-angle of attack, a control error of the quasi-yaw angle, a yaw angular velocity, and a pitch angular velocity. The control error of the quasi-angle of attack is the difference between the command value of the quasi-angle of attack and the quasi-angle of attack. The control error of the quasi-sideslip angle is the difference between the command value of the quasi-sideslip angle and the quasi-sideslip angle.

[0011] S600: Send the current state of the controlled object to the trained reinforcement learning network model to obtain the current posture control parameters; and execute S100.

[0012] According to a second aspect of the present invention, a saturated nonlinear attitude reinforcement learning control system for a spinning aircraft is provided, the system comprising: a control parameter optimization module, a mathematical rudder angle command value generation module, a mathematical rudder angle command value decomposition module, a physical servo module, a physical rudder angle synthesis module, a controlled object, and a state generation module; the controlled object is a dynamic model of the spinning aircraft. The mathematical rudder angle command value generation module is used to obtain the command value of the mathematical rudder angle of the current yaw channel and the command value of the mathematical rudder angle of the current pitch channel based on the current attitude control parameters, the output state of the controlled object, and the command values ​​of the quasi-attack angle and quasi-sideslip angle of the controlled object received at the current moment, wherein the output state includes the quasi-attack angle, the quasi-sideslip angle, the yaw angular velocity, and the pitch angular velocity; the mathematical rudder angle command value decomposition module is used to decompose the command value of the mathematical rudder angle of the current yaw channel and the command value of the mathematical rudder angle of the current pitch channel. The physical rudder angle of the current yaw channel and the physical rudder angle of the current pitch channel are decomposed respectively to obtain the instruction value of the physical rudder deflection angle of the current yaw channel and the instruction value of the physical rudder deflection angle of the current pitch channel, and send them to the physical rudder deflection angle synthesis module; the physical rudder deflection angle synthesis module is used to obtain the physical rudder deflection angle of the current yaw channel and the physical rudder deflection angle of the current pitch channel based on ... physical The rudder angle is used to generate a mathematical rudder angle for the current yaw channel and a mathematical rudder angle for the current pitch channel, and send them to the controlled object; the controlled object is used to obtain an output state at a current moment based on the received mathematical rudder angle for the current yaw channel and the mathematical rudder angle for the current pitch channel, and send the output state to the state generation module and the mathematical rudder angle command value generation module; the state generation module is used to generate the state of the controlled object at a current moment based on the command value of the quasi-angle of attack, the command value of the quasi-sideslip angle, and the output state at a current moment, as the current state of the controlled object, and send it to the control parameter optimization module; the state of the controlled object includes a control error of the quasi-angle of attack, a control error of the quasi-yaw angle, a yaw angular velocity, and a pitch angular velocity; the control error of the quasi-angle of attack is the difference between the command value of the quasi-angle of attack and the quasi-angle of attack, and the control error of the quasi-sideslip angle is the difference between the command value of the quasi-sideslip angle and the quasi-sideslip angle; the control parameter optimization module is used to generate a current attitude control parameter based on the received current state of the controlled object, and send the parameter to the mathematical rudder angle command value generation module.

[0013] The present invention has at least the following beneficial effects:

[0014] The embodiment of the present invention provides a method for attitude reinforcement learning control of a spinning aircraft under saturated nonlinearity, comprising: obtaining an instruction value of a mathematical rudder deflection angle of a current yaw channel and an instruction value of a mathematical rudder deflection angle of a current pitch channel based on current attitude control parameters, an output state of a controlled object, and instruction values ​​of a quasi-attack angle and a quasi-sideslip angle of the controlled object received at a current moment, wherein the current attitude control parameters are obtained by inputting the current state of the controlled object into a trained reinforcement learning network model; decomposing the instruction value of the mathematical rudder deflection angle of the current yaw channel and the instruction value of the mathematical rudder deflection angle of the current pitch channel respectively to obtain an instruction value of a physical rudder deflection angle of the current yaw channel and an instruction value of a physical rudder deflection angle of the current pitch channel. The present invention takes into account saturation nonlinearity when obtaining the physical rudder angle, can automatically adjust the attitude control parameters, improve the acquisition efficiency and accuracy of the attitude control parameters, and thus can achieve higher control accuracy.

[0015] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 A flowchart of a method for reinforcement learning control of a spinning aircraft attitude under saturated nonlinearity provided by an embodiment of the present invention;

[0018] Figure 2 A structural block diagram of a saturated nonlinear attitude reinforcement learning control system for a spinning aircraft provided by an embodiment of the present invention;

[0019] Figure 3 Schematic diagram of the evaluation network structure;

[0020] Figure 4 It is a structural diagram of the action network;

[0021] Figures 5 to 10 Schematic diagram of the experimental results. DETAILED DESCRIPTION

[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0024] It should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be performed in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. A process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. A process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0025] The embodiment of the present invention provides a method for attitude reinforcement learning control of a spinning aircraft under saturated nonlinearity, which may include: Figure 1 The following steps are shown:

[0026] S100, based on current attitude control parameters, the output state of the controlled object, and the command values ​​of the quasi-angle of attack and quasi-sideslip angle of the controlled object received at the current moment, obtain the command value of the mathematical rudder angle of the current yaw channel and the command value of the mathematical rudder angle of the current pitch channel; wherein the controlled object is a dynamic model of a spinning aircraft, and the output state of the controlled object includes the quasi-angle of attack, the quasi-sideslip angle, the yaw angular velocity, and the pitch angular velocity.

[0027] In an embodiment of the present invention, the dynamic model of the spinning aircraft satisfies the following conditions:

[0028]

[0029]

[0030]

[0031]

[0032] Among them, m, V, and ψ V are mass, velocity, ballistic inclination and ballistic deviation respectively, and They are V, and ψ V The first-order derivative of , γ, ψ and θ represent the roll angle, yaw angle and pitch angle respectively; α, β and γ V are respectively the quasi-attack angle, quasi-sideslip angle and quasi-speed roll angle, is the first-order derivative of θ, is the first-order derivative of γ; ω x 、ω y and ω z are the roll angular velocity, yaw angular velocity and pitch angular velocity respectively, and are ω x 、ω y and ω z The first derivative of J x 、J y and J z are the moments of inertia of the roll channel, yaw channel, and pitch channel respectively; P, X, Y, Z are thrust, drag, lift, and lateral force respectively; M x 、M y and M z They are rolling moment, yaw moment and pitching moment respectively, q is the flight pressure, q = ρV 2 / 2, ρ is the atmospheric density. S is the reference area of ​​the spinning vehicle; c1 and c2 are the longitudinal and lateral reference lengths of the spinning vehicle respectively; c x (γ), c y (γ) and c z (γ) are the drag coefficient, lift coefficient and lateral force coefficient respectively. x (γ), c y (γ) and c z (γ) is also referred to as c x 、c y and c z . m x (γ), m y (γ) and m z (γ) are the rolling moment coefficient, yaw moment coefficient and pitch moment coefficient respectively, below, m x (γ), m y (γ) and m z (γ) is also referred to as mx 、m y and m z The variable corresponding to the “·” in the coefficient brackets represents the variable that affects the coefficient. Ma is the Mach number, Ma=V / V s , V s is the speed of sound. yaw and δ pitch They are the mathematical rudder deflection angle of the yaw channel and the mathematical rudder deflection angle of the pitch channel respectively.

[0033] In the embodiment of the present invention, since the autogyro is unpowered during attitude control, P=0.

[0034] Based on equations (1)-(4), the small perturbation model of the pitch and yaw directions of the spinning aircraft can be derived as follows:

[0035]

[0036] in, d α ,d β Respectively express the impact The error factor of , the expressions of the other parameters are:

[0037]

[0038] in, represents the partial derivative of the aerodynamic coefficient γ with respect to the variable *.

[0039] In an embodiment of the present invention, the attitude control parameters may include a control gain of a quasi-attack angle, a control gain of a quasi-sideslip angle, a pitch angular velocity feedback gain, and a yaw angular velocity feedback gain.

[0040] Based on Equation (5), the dynamic equations of α and β can be rewritten as:

[0041]

[0042] in:

[0043]

[0044] Similarly, based on formula (5), ω can be derived z and ω y The dynamic equation is:

[0045]

[0046] in, and The control instructions for the mathematical rudder angles of the pitch channel and the yaw channel are represented by the command values, respectively, z and f yare the total disturbance factors of the pitch channel and yaw channel respectively, and their expressions are:

[0047]

[0048] In the embodiment of the present invention, the total disturbance factor includes coupling effects such as inertial coupling, aerodynamic coupling caused by the Magnus effect, and the control error of the mathematical rudder angle. Combining equations (9) and (10), the time derivative of equation (7) can be obtained:

[0049]

[0050] Because ω z and ω y It can be directly measured by the rate gyro, so it is used for inner loop feedback to increase the damping coefficient. and Can be designed as:

[0051]

[0052] in, and are the pitch angular rate feedback gain and the yaw angular rate feedback gain, and is the virtual control signal that needs further design. At this time, formula (11) becomes:

[0053]

[0054] After the damping is enhanced, the virtual control signal ( and ) to α and β can be approximated as a first-order object, so Equation (13) can be written as:

[0055]

[0056] in, and represents the input gain after the model is reduced in order, that is, is the input gain of the pitch channel, is the input gain of the yaw channel. α and t β is the total disturbance factor including model uncertainty, coupling effect and model reduction, which are the total disturbance factor of α and the total disturbance factor of β respectively. and The calculation process is the same as Take as an example to illustrate. Ignoring the coupling effect, error factor, and control error of the mathematical rudder angle in equation (5), the small perturbation model in the pitch direction can be obtained as:

[0057]

[0058] Substituting equation (12) into the above equation and performing Laplace transform, we can get Transfer function to α for:

[0059]

[0060] Where e is a complex variable. According to formula (15), we can calculate The two extreme points p1 and p2 (p1<p2) of In order to compensate for the total disturbance factor, the following nonlinear extended state observer is designed based on formula (14):

[0061]

[0062] Among them, z α,1 and z α,2 are for α and t respectively α Estimates of z β,1 and z β,2 are β and t respectively. β Estimates, l α,o and l β,o (o=1,2) represents the gain of the observer, which can be designed as:

[0063] l α,1 =2ω α ,l α,2 =ω α 2 ,l β,1 =2ω β ,l β,2 =ω β 2

[0064] Among them, ω α and ω β represents the bandwidth of the observer and can be an empirical value. The purpose of introducing the function fal(γ) is to achieve "small gain for large error and large gain for small error" to avoid saturation of the control signal. Its specific form is:

[0065]

[0066]

[0067] Among them, λ α and λ β are the switching thresholds of the quasi-angle of attack and quasi-sideslip angle estimation errors, which can be empirical values. α and δ βare the exponents of the estimation errors of the quasi-angle of attack and quasi-sideslip angle, respectively, which can be empirical values. Then, the virtual control signal can be written as follows:

[0068]

[0069] Among them, α c and β c is the control instruction of α and β, that is, the instruction value of α and β, K α and K β are the control gains of α and β respectively. Combining equations (12) and (17), the final mathematical command of the rudder angle can be obtained as:

[0070]

[0071] Among them, K α , K β and The optimization is performed using a reinforcement learning algorithm. That is, in S100, the command values ​​of the mathematical steering angle of the yaw channel and the mathematical steering angle of the pitch channel satisfy equation (18).

[0072] In an embodiment of the present invention, an extended observer estimates and compensates for the total disturbance including the servo saturation nonlinearity. When optimizing the control parameters, attitude limit cycle suppression is reflected in the objective function of the model, so that the optimized control parameters can suppress the attitude limit cycle to the greatest extent.

[0073] S200 , decompose the mathematical rudder angle instruction value of the current yaw channel and the mathematical rudder angle instruction value of the current pitch channel respectively to obtain the physical rudder angle instruction value of the current yaw channel and the physical rudder angle instruction value of the current pitch channel.

[0074] In S200, the command values ​​of the physical rudder deflection angle of the yaw channel and the physical rudder deflection angle of the pitch channel meet the following conditions:

[0075]

[0076] in, is the command value of the physical rudder deflection angle of the yaw channel, is the command value of the physical rudder angle of the pitch channel, σ is the intermediate value,

[0077] Furthermore, in the embodiment of the present invention, if the amplitude of the command value of the physical rudder deflection angle is Exceeding the maximum amplitude u of the servo δ , then the command value of the mathematical rudder angle needs to be readjusted as follows:

[0078]

[0079] S300 , obtaining the physical rudder deflection angle of the current yaw channel and the physical rudder deflection angle of the current pitch channel based on the instruction value of the physical rudder deflection angle of the current yaw channel and the instruction value of the physical rudder deflection angle of the current pitch channel.

[0080] Furthermore, S300 specifically includes:

[0081] S301, obtain rudder angle tracking error And obtain the rudder angle tracking error of the pitch channel δ y is the physical rudder deflection angle of the yaw channel, δ z It is the physical rudder deflection angle of the pitch channel.

[0082] In the embodiment of the present invention, the physical rudder deflection angle used to calculate the tracking error in S301 is the physical rudder deflection angle obtained at the previous moment.

[0083] S302, based on and the maximum speed of the servo Get the rate limit output value u i , where i is x or y, u i The following conditions must be met:

[0084] That is, if set up if set up if set up

[0085] S303, based on u i and linear models Get the linear output value r i ,in, and r i The first and second derivatives of a i is the time constant, b i The time constant and steady-state gain can be obtained by conducting experiments on the servo, measuring the input and output data of the servo, and fitting the data.

[0086] S304, based on r i and the maximum amplitude u of the servo δ , get δ i , δ i The following conditions must be met:

[0087]

[0088] S400: Generate a mathematical rudder angle for the current yaw channel and a mathematical rudder angle for the current pitch channel based on the physical rudder angle for the current yaw channel and the physical rudder angle for the current pitch channel, and send them to the controlled object to obtain the output state of the controlled object at the current moment.

[0089] Furthermore, in S400, the mathematical rudder deflection angle of the yaw channel and the mathematical rudder deflection angle of the pitch channel meet the following conditions:

[0090] δ yaw =δ y cosγ-δ z sinγ,

[0091] δ pitch =δ y sinγ+δ z cosγ.

[0092] It should be noted that for actual flight tests, both the autorotor and the physical servos are actual hardware devices. The physical servos directly change the aircraft's attitude by deflecting the control surfaces, so there's no need to synthesize the physical rudder angles. For mathematical simulations, both the autorotor and the physical servos are represented by mathematical models. Because the input to the autorotor's dynamics model is the mathematical rudder angle, synthesizing the physical rudder angles is necessary according to the above formula.

[0093] S500. Based on the command value of the quasi-angle of attack, the command value of the quasi-sideslip angle, and the output state at the current moment, generate a state of the controlled object at the current moment as the current state of the controlled object. The state of the controlled object includes a control error of the quasi-angle of attack, a control error of the quasi-yaw angle, a yaw angular velocity, and a pitch angular velocity. The control error of the quasi-angle of attack is the difference between the command value of the quasi-angle of attack and the quasi-angle of attack. The control error of the quasi-sideslip angle is the difference between the command value of the quasi-sideslip angle and the quasi-sideslip angle.

[0094] S600: Send the current state of the controlled object to the trained reinforcement learning network model to obtain the current posture control parameters; and execute S100.

[0095] In an embodiment of the present invention, when the trained reinforcement learning network model receives the current state of the controlled object, it obtains the corresponding action, that is, the posture control parameter, through the action network.

[0096] Furthermore, in an embodiment of the present invention, the reinforcement learning network model may include a first evaluation network, a second evaluation network, a first target evaluation network, a second target evaluation network, an action network, and a target action network.

[0097] Furthermore, the trained reinforcement learning network model can be obtained through the following steps:

[0098] S1, initialize the parameters of the first evaluation network, the second evaluation network, the first target evaluation network, the second target evaluation network, the action network and the target action network, and initialize the experience replay pool; set the training round counter p=1; the experience replay pool stores multiple sample data, each sample data is a four-tuple information (s, a, r, s'), where s represents the state, s=(e α ,e β ,ω z ,ω y ), e α is the control error of the quasi-angle of attack, e α =α c -α, α c is the quasi-attack angle command value, α is the quasi-attack angle, e β is the control error of the quasi-sideslip angle, e β =β c -β,β c is the quasi-sideslip angle command value, β is the quasi-sideslip angle, ω y is the yaw angular velocity, ω z is the pitch angular velocity, a represents the action, a=(K,K d ), K and K d are posture control parameters, r represents the reward obtained by selecting action a, and s′ represents the state obtained after taking action a at time s.

[0099] In the embodiment of the present invention, the network parameters and the experience replay pool may be generated in a random manner.

[0100] In the embodiment of the present invention, since the dynamics of the pitch channel and the yaw channel are similar, the control parameters of the two channels are set to be the same, that is, K α =K β =K, Setting the dimension of the action space of a to 2 can improve the convergence speed of the model.

[0101] In the embodiment of the present invention, based on the following considerations: (1) the control accuracy of the angle of attack and sideslip angle is reflected by the angle control error; (2) the limit cycle oscillation of the attitude angle is suppressed, which is reflected by the angular velocity; (3) the control signal saturation is avoided, which is reflected by the control instruction of the mathematical rudder deflection angle, i.e., the command value, the set return r satisfies the following conditions:

[0102]

[0103]

[0104]

[0105] Among them, c e 、cω and c u are the return coefficients, is the mathematical rudder angle command value of the yaw channel, is the mathematical rudder angle command value for the pitch channel, tanh() is the hyperbolic tangent function, R1 is the penalty value when the control error is greater than the first set error threshold e1, and R2 is the reward value when the control error is less than the second set error threshold e2. The smaller the attitude angle control error, the smaller the angular velocity, and the smaller the mathematical rudder angle command value, the greater the reward. In this embodiment of the present invention, e1 and e2 can be empirical values.

[0106] S2: If p>p0 or the average value of the rewards of p1 consecutive training rounds in p training rounds is R>R0, the current reinforcement learning network model is used as the trained reinforcement learning network model; if p≤p0 and R≤R0, set the training step counter t=1 and execute S3; p0 is the preset training round threshold, R0 is the preset reward threshold, and p0 and R0 can be set based on actual needs.

[0107] In the embodiment of the present invention, the last round number in p1 consecutive training rounds is p, that is, the p1 consecutive training rounds are the p-p1+1th round, the p-p1+2th round, ..., the pth round. p1 can be an empirical value.

[0108] S3, if t≤t0, execute S4, otherwise, set p=p+1 and execute S2; t0 is a preset time step threshold, which can be an empirical value.

[0109] In the embodiment of the present invention, the time period of each time step may be the same, which is ΔT, and may be an empirical value.

[0110] S4, the current state of the controlled object s c Input into the current action network to get the selected action a c .

[0111] S5, based on a c Obtain the mathematical rudder deflection angle of the yaw channel and the mathematical rudder deflection angle of the pitch channel, and send them to the controlled object to obtain the output state of the controlled object, and obtain the corresponding feedback r based on the output state of the controlled object c and the new state s c ′, forming the corresponding four-tuple information (s c , a c , r c , s c ′) and added to the current experience replay pool.

[0112] S6, sample N from the current experience replay pool using the priority experience replay method sFour-tuple information is used as training sample data, and each training sample data (s m ,a m ,r m ,s′ m ) is input into the current target action network to obtain the corresponding optimal action a′ m , m ranges from 1 to N s .

[0113] In the embodiment of the present invention, a′ m The following conditions must be met:

[0114]

[0115] Among them, ε′ represents random noise, which obeys normal distribution, has a mean of 0, and a variance of σ′. clip() is a clip function used to convert the input value π φ′ (s′ m )+ε′ is limited to the interval [a min ,a max ], a min and a max They represent the lower and upper bounds of the action, π φ′ (s′ m ) means to convert s′ m The output result obtained by inputting the target network.

[0116] S7, the state s of the current node in each training sample data m and the corresponding action a m Input them into the current first evaluation network respectively, obtain the corresponding first evaluation network output results and second evaluation network output results, and convert the s′ corresponding to each quadruple information into m and a′ m The results are input into the current first target evaluation network and the second target evaluation network respectively to obtain the corresponding output results of the first target evaluation network and the second target evaluation network.

[0117] S8, based on the first evaluation network output result, the second evaluation network output result, the first target evaluation network output result and the second target evaluation network output result of each training sample data, update the parameters of the current first evaluation network and the second evaluation network; if (t / d) is an integer, use the parameters of the current first evaluation network to update the parameters of the current action network, and use the parameters of the current evaluation network to update the parameters of the current target evaluation network and use the parameters of the updated action network to update the parameters of the current target action network, and execute S9; if (t / d) is not an integer, execute S9.

[0118] Since accurate estimation of the value function is a prerequisite for reasonable updating of control parameters, updating the action network after the evaluation network is fully trained can improve model stability.

[0119] S9, set t=t+1, and execute S3.

[0120] Furthermore, in an embodiment of the present invention, the parameters of the evaluation network j are updated as follows:

[0121]

[0122] Where j = 1 or 2, θ j (k+1) is the parameter of evaluation network j at the k+1th iteration, θ j (k) is the parameter of the evaluation network at the kth iteration, and the evaluation network j is the first evaluation network or the second evaluation network; H k The Hessian matrix corresponding to the kth iteration is specifically p θ ×p θ The Hessian matrix, p θ is the size of the parameters of the evaluation network. m is the intermediate variable, γ d is the discount factor, is the output result of the first evaluation network, is the output result of the first target evaluation network, is the output result of the second evaluation network, is the output result of the second target evaluation network, J c To evaluate the objective function of the network, For J c Regarding the parameter θ of the evaluation network j j The gradient of , λ is the damping factor. for About θ i The gradient, It can be calculated by existing techniques such as back propagation.

[0123] The parameters of the target evaluation network j are updated as follows:

[0124] θ j ′=τθ j +(1-τ)θ j ';

[0125] θ j ′ is the parameter of the target evaluation network j, the target evaluation network j is the first target evaluation network or the second target evaluation network, and τ is the weighting coefficient.

[0126] Furthermore, in an embodiment of the present invention, the parameters of the action network are updated as follows:

[0127] φ=φ+χ a ▽ φ J a ;

[0128] Among them, φ is the parameter of the action network, χ a represents the learning rate of the action network, ▽ φ J a Represents the objective function J of the action network a With respect to the gradient of φ, is the output result of the first evaluation network.

[0129] in, express With respect to the gradient of a, φ π φ (s m ) represents π φ (s m ) The gradient of φ can be calculated using existing techniques such as back propagation.

[0130] In this embodiment of the present invention, the parameters of the target action network are updated as follows:

[0131] φ′=τφ+(1-τ)φ′;

[0132] Among them, φ′ is the parameter of the target action network and τ is the weight coefficient.

[0133] Based on the same inventive concept, an embodiment of the present invention provides a saturated nonlinear spinning aircraft attitude reinforcement learning control system, such as Figure 2 As shown, the system includes: a mathematical rudder deflection angle command value generation module 1, a mathematical rudder deflection angle command value decomposition module 2, a physical servo module 3, a physical rudder deflection angle synthesis module 4, a controlled object 5, a state generation module 6, and a control parameter optimization module 7. Among them, the controlled object 5 is a dynamic model of the spinning aircraft.

[0134] In an embodiment of the present invention, the mathematical rudder angle command value generation module 1 is used to obtain the command value of the mathematical rudder angle of the current yaw channel and the command value of the mathematical rudder angle of the current pitch channel based on the current attitude control parameters, the output state of the controlled object, and the command values ​​of the quasi-attack angle and quasi-sideslip angle of the controlled object received at the current moment.

[0135] In the embodiment of the present invention, the instruction value of the mathematical rudder angle of the current yaw channel and the instruction value of the mathematical rudder angle of the current pitch channel can be obtained by the aforementioned formula (18).

[0136] The mathematical rudder angle command value decomposition module 2 is used to decompose the command value of the mathematical rudder angle of the current yaw channel and the command value of the mathematical rudder angle of the current pitch channel respectively, obtain the command value of the physical rudder angle of the current yaw channel and the command value of the physical rudder angle of the current pitch channel, and send them to the physical servo module.

[0137] In the embodiment of the present invention, the command values ​​of the physical rudder deflection angle of the yaw channel and the physical rudder deflection angle of the pitch channel meet the following conditions:

[0138]

[0139] in, is the command value of the physical rudder deflection angle of the yaw channel, is the command value of the physical rudder angle of the pitch channel, σ is the intermediate value,

[0140] Furthermore, in the embodiment of the present invention, if the amplitude of the command value of the physical rudder deflection angle is Exceeding the maximum amplitude u of the servo δ , then the command value of the mathematical rudder deflection angle needs to be readjusted as follows:

[0141]

[0142] The physical servo module 3 is used to obtain the physical rudder deflection angle of the current yaw channel and the physical rudder deflection angle of the current pitch channel based on the instruction value of the physical rudder deflection angle of the current yaw channel and the instruction value of the physical rudder deflection angle of the current pitch channel, and send them to the physical rudder deflection angle synthesis module.

[0143] In this embodiment of the present invention, the physical servo module 3 may include a tracking error calculation module, a rate saturation module, a linearization module, and an amplitude saturation module. The tracking error calculation module is configured to calculate the rudder angle tracking error based on the currently received physical rudder angle command value and the physical rudder angle, and send the calculated value to the rate saturation module, thereby implementing the process defined in S301. The rate saturation module is configured to limit the rudder angle tracking error based on the rudder angle tracking error and the maximum rate of the servo, obtain a corresponding rate-limited output value, and send the calculated value to the linearization module, thereby implementing the process defined in S302. The linearization module is configured to perform linear processing on the received rate-limited output value to obtain a corresponding linear output value, and send the calculated value to the amplitude saturation module, thereby implementing the process defined in S303. The amplitude saturation module is configured to limit the linear output value based on the received linear output value and the maximum amplitude of the servo, obtain a corresponding amplitude-limited output value as the physical rudder angle, and send the calculated value to the tracking error calculation module and the physical rudder angle synthesis module 4.

[0144] The physical rudder angle synthesis module 4 is used to generate the mathematical rudder angle of the current yaw channel and the mathematical rudder angle of the current pitch channel based on the physical rudder angle of the current yaw channel and the physical rudder angle of the current pitch channel, and send them to the controlled object.

[0145] In an embodiment of the present invention, the mathematical rudder deflection angle of the yaw channel and the mathematical rudder deflection angle of the pitch channel meet the following conditions:

[0146] δ yaw =δ y cosγ-δ z sinγ,

[0147] δ pitch =δ y sinγ+δ z cosγ.

[0148] The controlled object 5 is used to obtain the output state at the current moment based on the received mathematical rudder deflection angle of the current yaw channel and the mathematical rudder deflection angle of the current pitch channel, and send it to the state generation module and the mathematical rudder deflection angle command value generation module; the output state includes the quasi-angle of attack, the quasi-sideslip angle, the yaw angular velocity and the pitch angular velocity.

[0149] By inputting the mathematical rudder deflection angle of the current yaw channel and the mathematical rudder deflection angle of the current pitch channel into the controlled object, the corresponding output state will be obtained through calculation.

[0150] The state generation module 6 is used to generate the state of the controlled object at the current moment based on the command value of the quasi-angle of attack, the command value of the quasi-sideslip angle and the output state at the current moment, as the current state of the controlled object, and send it to the control parameter optimization module 7.

[0151] Among them, the state of the controlled object includes the control error of the quasi-attack angle, the control error of the quasi-yaw angle, the yaw angular velocity and the pitch angular velocity. The control error of the quasi-attack angle is the difference between the command value of the quasi-attack angle and the quasi-attack angle, and the control error of the quasi-sideslip angle is the difference between the command value of the quasi-sideslip angle and the quasi-sideslip angle.

[0152] The control parameter optimization module 7 is used to generate current attitude control parameters based on the received current state of the controlled object and send them to the mathematical rudder deflection instruction value generation module.

[0153] In this embodiment of the present invention, the control parameter optimization module 7 is equipped with a trained reinforcement learning network model. In this embodiment of the present invention, upon receiving the current state of the controlled object, the trained reinforcement learning network model obtains the corresponding action, i.e., posture control parameters, through the action network. The structure and training process of the reinforcement learning network model can be found in the description of the previous embodiment.

[0154] (Example)

[0155] (1) Model and algorithm parameters

[0156] (1) Aircraft model parameters:

[0157]

[0158]

[0159]

[0160] (2) Obtain relevant parameters of the physical rudder deflection angle:

[0161] u eδ =5,u δ =20,a z =a y =8882.6, b z =b y =180.

[0162] (3) Control parameters that do not need to be optimized in the attitude controller:

[0163] ω α =ω β =5,δ α =δ β =0.2,λ α =λ β =0.8,δ max =20,α c =1,β c =0.

[0164] (4) Reinforcement learning network model parameters:

[0165] In the embodiment of the present invention, the structures of the four evaluation networks are the same, such as Figure 3 As shown, the structures of the two action networks are the same, such as Figure 4 shown. Figure 3 and Figure 4 The detailed explanation of the parameters in is given in Tables 1 and 2 below. Each evaluation network has 2369 learning parameters, and each action network has 142 parameters. All the parameters of the reinforcement learning algorithm are as follows:

[0166] c e =100, c ω =10, c u =10, e1=1, e2=0.5, a min =[0,0], a max =[50,10], Ns =128,σ′=[1,0.5],

[0167] γ d =0.997,λ=0.99,χ a =0.0001, d=3, τ=0.005, p0=3000, p1=20, R0=2000, ΔT=0.1, t0=200, σ=[1,0.5].

[0168] Table 1 Detailed information of the evaluation network structure

[0169]

[0170] Table 2 Detailed information of the action network structure

[0171]

[0172]

[0173] (2) Algorithm Results

[0174] The rewards during training are as follows Figure 5 As shown in Figure 2, the average return tends to stabilize after the 250th training round, indicating that the algorithm has achieved rapid convergence. Subsequently, the trained action network is used to control the attitude of the spinning aircraft. The simulation results are shown in Figure 2. Figures 6 to 10 shown. Figure 6 The optimized control parameters during the simulation are displayed. It can be seen that the control parameters can quickly jump to a stable value, ensuring the rapid response of the attitude angle and stability in steady state. Figure 7 and Figure 8 The command values ​​and actual response values ​​of the mathematical rudder deflection angle and the physical rudder deflection angle are shown. It can be seen from the figure that both the mathematical rudder deflection angle and the physical rudder deflection angle have large response errors, which are mainly caused by the saturation nonlinearity inside the servo. However, if Figure 8 As shown in , the control parameters optimized by reinforcement learning can still effectively track the command values ​​of the angle of attack and sideslip angle, and no attitude limit cycle phenomenon occurs. Figure 9 As shown in Figure 2, the pitch and yaw angular rates fluctuate rapidly only in the first second and then quickly stabilize.

[0175] An embodiment of the present invention also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method described in the embodiment of the present invention.

[0176] An embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer instructions are used to execute the method described in the embodiment of the present invention.

[0177] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. This is not limited herein.

[0178] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for attitude reinforcement learning control of a spinning aircraft under saturated nonlinearity, characterized in that: The method comprises the following steps: S100, based on current attitude control parameters, an output state of a controlled object, and command values ​​of a quasi-angle of attack and a quasi-sideslip angle of the controlled object currently received, obtaining a command value of a mathematical rudder angle for a current yaw channel and a command value of a mathematical rudder angle for a current pitch channel; wherein the controlled object is a dynamic model of a spinning vehicle, and the output state of the controlled object includes the quasi-angle of attack, the quasi-sideslip angle, the yaw angular velocity, and the pitch angular velocity; S200, decomposing the mathematical rudder angle command value of the current yaw channel and the mathematical rudder angle command value of the current pitch channel respectively to obtain the physical rudder angle command value of the current yaw channel and the physical rudder angle command value of the current pitch channel; S300, acquiring the physical rudder angle of the current yaw channel and the physical rudder angle of the current pitch channel based on the instruction value of the physical rudder angle of the current yaw channel and the instruction value of the physical rudder angle of the current pitch channel; S400: Generate a mathematical rudder angle for the current yaw channel and a mathematical rudder angle for the current pitch channel based on the physical rudder angle for the current yaw channel and the physical rudder angle for the current pitch channel, and send the generated mathematical rudder angle to the controlled object to obtain an output state of the controlled object at the current moment. S500: Based on the command value of the quasi-angle of attack, the command value of the quasi-sideslip angle, and the output state at the current moment, a state of the controlled object at the current moment is generated as the current state of the controlled object; the state of the controlled object includes a control error of the quasi-angle of attack, a control error of the quasi-yaw angle, a yaw angular velocity, and a pitch angular velocity. The control error of the quasi-angle of attack is the difference between the command value of the quasi-angle of attack and the quasi-angle of attack, and the control error of the quasi-sideslip angle is the difference between the command value of the quasi-sideslip angle and the quasi-sideslip angle. S600: Send the current state of the controlled object to the trained reinforcement learning network model to obtain the current posture control parameters; and execute S100.

2. The method according to claim 1, characterized in that The reinforcement learning network model includes a first evaluation network, a second evaluation network, a first target evaluation network, a second target evaluation network, an action network, and a target action network; wherein the trained reinforcement learning network model is obtained by the following steps: S1, initialize the parameters of the first evaluation network, the second evaluation network, the first target evaluation network, the second target evaluation network, the action network and the target action network, and initialize the experience replay pool; set the training round counter p=1; the experience replay pool stores multiple sample data, each sample data is a four-tuple information (s, a, r, s'), where s represents the state, s=(e α ,e β ,ω z ,ω y ), e α is the control error of the quasi-angle of attack, e α =α c -α, α c is the quasi-attack angle command value, α is the quasi-attack angle, e β is the control error of the quasi-sideslip angle, e β =β c -β,β c is the quasi-sideslip angle command value, β is the quasi-sideslip angle, ω y is the yaw angular velocity, ω z is the pitch angular velocity, a represents the action, a=(K,K d ), K and K d are all posture control parameters, r represents the reward obtained by selecting action a, and s′ represents the state obtained after taking action a at time s; S2: If p>p0 or the average value of the rewards of p1 consecutive training rounds in p training rounds is R>R0, the current reinforcement learning network model is used as the trained reinforcement learning network model; if p≤p0 and R≤R0, set the training step counter t=1 and execute S3; p0 is the preset training round threshold, R0 is the preset reward threshold; the last round number in the p1 consecutive training rounds is p; S3, if t≤t0, execute S4, otherwise, set p=p+1 and execute S2; t0 is the preset time step threshold; S4, the current state of the controlled object s c Input into the current action network to get the selected action a c ; S5, based on a c Obtain the mathematical rudder deflection angle of the yaw channel and the mathematical rudder deflection angle of the pitch channel, and send them to the controlled object to obtain the output state of the controlled object, and obtain the corresponding feedback r based on the output state of the controlled object c and the new state s c ′, forming the corresponding four-tuple information (s c , a c , r c , s c ') and added to the current experience replay pool; S6, sample N from the current experience replay pool using the priority experience replay method s Four-tuple information is used as training sample data, and each training sample data (s m ,a m ,r m ,s′ m ) is in the state s′ of the next node m Input into the current target action network to obtain the corresponding optimal action a′ m , m ranges from 1 to N s ; S7, the state s of the current node in each training sample data m and the corresponding action a m Input them into the current first and second evaluation networks respectively, obtain the corresponding output results of the first evaluation network and the second evaluation network, and convert the s′ corresponding to each quadruple information into m and a′ m Input them into the current first target evaluation network and the second target evaluation network respectively to obtain the corresponding first target evaluation network output result and the second target evaluation network output result; S8, based on the output result of the first evaluation network, the output result of the second evaluation network, the output result of the first target evaluation network and the output result of the second target evaluation network of each training sample data, update the parameters of the current first evaluation network and the second evaluation network; if (t / d) is an integer, use the parameters of the current first evaluation network to update the parameters of the current action network, and use the parameters of the current evaluation network to update the parameters of the current target evaluation network and use the parameters of the updated action network to update the parameters of the current target action network, and execute S9; if (t / d) is not an integer, execute S9; S9, set t=t+1, and execute S3.

3. The method according to claim 2, characterized in that The return r satisfies the following conditions: Among them, c e 、c ω and c u are the return coefficients, is the mathematical rudder angle command value of the yaw channel, is the mathematical rudder angle command value of the pitch channel, tanh() is the hyperbolic tangent function, R1 is the penalty value when the control error is greater than the first set error threshold e1, and R2 is the reward value when the control error is less than the second set error threshold e2.

4. The method according to claim 2, characterized in that The parameters of the evaluation network j are updated as follows: Where j = 1 or 2, θ j (k+1) is the parameter of evaluation network j at the k+1th iteration, θ j (k) is the parameter of the evaluation network at the kth iteration, and the evaluation network j is the first evaluation network or the second evaluation network; H k is the Hessian matrix corresponding to the kth iteration, ψ m is the intermediate variable, γ d is the discount factor, is the output result of the first evaluation network, is the output result of the first target evaluation network, is the output result of the second evaluation network, is the output result of the second target evaluation network, J c To evaluate the objective function of the network, For J c Regarding the parameter θ of the evaluation network j j The gradient of , λ is the damping factor; The parameters of the target evaluation network j are updated as follows: i j ′=τθ j +(1-τ)θ j ′; θ j ′ is the parameter of the target evaluation network j, the target evaluation network j is the first target evaluation network or the second target evaluation network, and τ is the weighting coefficient.

5. The method according to claim 2, characterized in that The parameters of the action network are updated as follows: Among them, φ is the parameter of the action network, χ a represents the learning rate of the action network, Represents the objective function J of the action network a With respect to the gradient of φ, is the output result of the first evaluation network; The parameters of the target action network are updated as follows: φ′=τφ+(1-τ)φ′; Among them, φ′ is the parameter of the target action network and τ is the weight coefficient.

6. The method according to claim 1, wherein The dynamic model of the spinning aircraft meets the following conditions: Among them, m, V, and ψ V are mass, velocity, ballistic inclination and ballistic deviation respectively, and They are V, and ψ V The first derivative of , γ, ψ and Denote the roll angle, yaw angle, and pitch angle respectively; α, β, and γ V are respectively the quasi-attack angle, quasi-sideslip angle and quasi-speed roll angle, for The first derivative of is the first-order derivative of ψ, is the first-order derivative of γ; ω x 、ω y and ω z are the roll angular velocity, yaw angular velocity and pitch angular velocity respectively, and are ω x 、ω y and ω z The first derivative of J x 、J y and J z are the moments of inertia of the roll channel, yaw channel, and pitch channel respectively; P, X, Y, Z are thrust, drag, lift, and lateral force respectively; M x 、M y and M z are the rolling moment, yaw moment and pitching moment respectively, q is the flight pressure; S is the reference area of ​​the spinning vehicle; c1 and c2 are the longitudinal and lateral reference lengths of the spinning vehicle respectively; c x (γ), c y (γ) and c z (γ) are drag coefficient, lift coefficient and lateral force coefficient, respectively, m x (γ), m y (γ) and m z (γ) are the rolling moment coefficient, yaw moment coefficient and pitching moment coefficient respectively, Ma is the Mach number, δ yaw and δ pitch They are the mathematical rudder deflection angle of the yaw channel and the mathematical rudder deflection angle of the pitch channel respectively.

7. The method according to claim 6, characterized in that The attitude control parameters include a control gain of a quasi-attack angle, a control gain of a quasi-sideslip angle, a pitch angular velocity feedback gain, and a yaw angular velocity feedback gain; In S100, the command values ​​of the mathematical rudder angle of the yaw channel and the mathematical rudder angle of the pitch channel meet the following conditions: in, K is the mathematical rudder angle command value of the pitch channel, α is the control gain of α, α c is the command value of α, z α,2 is the total perturbation factor t on α α Estimates, is the pitch angular velocity feedback gain, is the input gain of the pitch channel, K is the command value of the mathematical rudder angle of the yaw channel, β is the control gain of β, β c is the command value of β, z β,2 is the total perturbation factor t on β β Estimates, is the yaw rate feedback gain, is the input gain of the yaw channel.

8. The method according to claim 7, characterized in that In S200, the command values ​​of the physical rudder deflection angle of the yaw channel and the physical rudder deflection angle of the pitch channel meet the following conditions: in, is the command value of the physical rudder deflection angle of the yaw channel, is the command value of the physical rudder angle of the pitch channel, σ is the intermediate value, S300 specifically includes: S301, obtaining the physical rudder angle tracking error of the yaw channel And obtain the physical rudder angle tracking error of the pitch channel δ y is the physical rudder deflection angle of the yaw channel, δ z is the physical rudder deflection angle of the pitch channel; S302, based on and the maximum speed of the servo Get the rate limit output value u i , where i is x or y, u i The following conditions must be met: S303, based on u i and linear models Get the linear output value r i ,in, and r i The first and second derivatives of a i is the time constant, b i is the steady-state gain; S304, based on r i and the maximum amplitude u of the servo δ , get δ i , δ i The following conditions must be met:

9. The method according to claim 7, characterized in that In S400, the mathematical rudder deflection angle of the yaw channel and the mathematical rudder deflection angle of the pitch channel meet the following conditions: d yaw =d y cosγ-δ z sing, d pitch =d y sinγ+δ z cosγ.

10. A saturated nonlinear attitude reinforcement learning control system for a spinning aircraft, characterized in that: The system includes: a control parameter optimization module, a mathematical rudder deflection angle command value generation module, a mathematical rudder deflection angle command value decomposition module, a physical servo module, a physical rudder deflection angle synthesis module, a controlled object and a state generation module; the controlled object is a dynamic model of a spinning aircraft; The mathematical rudder angle command value generation module is used to obtain the mathematical rudder angle command value for the current yaw channel and the mathematical rudder angle command value for the current pitch channel based on the current attitude control parameters, the output state of the controlled object, and the command values ​​of the quasi-angle of attack and quasi-sideslip angle of the controlled object received at the current moment; the output state includes the quasi-angle of attack, the quasi-sideslip angle, the yaw angular velocity, and the pitch angular velocity; The mathematical rudder angle command value decomposition module is used to decompose the mathematical rudder angle command value of the current yaw channel and the mathematical rudder angle command value of the current pitch channel respectively, obtain the physical rudder angle command value of the current yaw channel and the physical rudder angle command value of the current pitch channel, and send them to the physical servo module; The physical servo module is used to obtain the physical rudder deflection angle of the current yaw channel and the physical rudder deflection angle of the current pitch channel based on the instruction value of the physical rudder deflection angle of the current yaw channel and the instruction value of the physical rudder deflection angle of the current pitch channel, and send them to the physical rudder deflection angle synthesis module; The physical rudder deflection angle synthesis module is used to generate a mathematical rudder deflection angle of the current yaw channel and a mathematical rudder deflection angle of the current pitch channel based on the physical rudder deflection angle of the current yaw channel and the physical rudder deflection angle of the current pitch channel, and send the generated mathematical rudder deflection angles to the controlled object; The controlled object is used to obtain an output state at a current moment based on the received mathematical rudder deflection angle of the current yaw channel and the mathematical rudder deflection angle of the current pitch channel, and send the output state to the state generation module and the mathematical rudder deflection angle command value generation module; The state generation module is configured to generate a state of the controlled object at a current moment based on the command value of the quasi-angle of attack, the command value of the quasi-sideslip angle, and the output state at the current moment, as the current state of the controlled object, and send the generated state to the control parameter optimization module; the state of the controlled object includes a control error of the quasi-angle of attack, a control error of the quasi-yaw angle, a yaw angular velocity, and a pitch angular velocity; the control error of the quasi-angle of attack is the difference between the command value of the quasi-angle of attack and the quasi-angle of attack; and the control error of the quasi-sideslip angle is the difference between the command value of the quasi-sideslip angle and the quasi-sideslip angle; The control parameter optimization module is used to generate current attitude control parameters based on the received current state of the controlled object and send the parameters to the mathematical rudder deflection instruction value generation module.

Citation Information

Patent Citations

  • High-precision overload control method

    CN112180965A

  • Flight attitude control method

    CN114200950A