Nonlinear system actuator fault ppb-siadp fault-tolerant control method

By using a single-network incremental adaptive dynamic programming method, the optimization control problem under actuator failure in a nonlinear system is solved, achieving efficient fault-tolerant control, reducing computational burden and improving system performance.

CN116661307BActive Publication Date: 2025-12-19NANJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310551647.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-16
Publication Date
2025-12-19
Estimated Expiration
2043-05-16

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle actuator failures in nonlinear systems, leading to performance degradation. Furthermore, traditional fault-tolerant control methods are computationally burdensome and fail to achieve optimal control results.

Method used

The single-network incremental adaptive dynamic programming (SIADP) method is adopted. By transforming the predetermined performance boundary and recursive least squares identification, an incremental neural network observer is designed. Combined with the anti-saturation performance index, efficient estimation and control of actuator faults can be achieved.

Benefits of technology

It reduces the learning parameters and online computation burden, improves the efficiency of actuator fault handling, and realizes optimized control under hardware redundancy and actuator grouping schemes, which is suitable for unknown nonlinear dynamic systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116661307B_ABST
    Figure CN116661307B_ABST
Patent Text Reader

Abstract

The application discloses a nonlinear system actuator fault PPB-SIADP fault-tolerant control method. For a nominal nonlinear system, a new error tracking system is obtained through predetermined performance boundary transformation, and an approximate linear time-varying system is obtained by using incremental nonlinear technology; according to the redundancy characteristics and functions of the actuator, an actuator grouping scheme is proposed, and an incremental neural network observer is designed to approximate multiple actuator faults; a functional performance index is designed to process the actuator saturation problem, and a corresponding Hamilton-Jacobi-Bellman equation is derived; a single network incremental adaptive dynamic programming algorithm is used to realize optimal control. The new single network incremental adaptive dynamic programming scheme is composed of an optimized evaluation network, and the application can shorten the learning time and reduce the calculation burden in the control process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of fault-tolerant control, in particular to a nonlinear system actuator fault PPB-SIADP fault-tolerant control method. BACKGROUND

[0002] Engineering systems are becoming more and more complex, and various faults may occur during operation. Due to long-term and frequent execution of control tasks, actuators are prone to failure. After the actuator fails, not only the original control law cannot be executed, but also the output of the controlled object is directly affected, thereby reducing the performance of the entire system. However, the fault-tolerant control (FTC) method provides an effective way to improve the reliability and safety of complex nonlinear systems. It is of great significance to develop an FTC scheme to deal with such faults and maintain acceptable system performance.

[0003] Generally, the FTC method can be divided into two categories: passive fault-tolerant control (PFTC) and active fault-tolerant control (AFTC). PFTC mainly relies on the robustness of the controller itself to reduce the impact of a certain type of assumed fault, thereby achieving fault-tolerant control. The feature of the PFTC controller is that all possible fault types need to be known in advance, and its conservative design makes it difficult to achieve better control performance. From the functional point of view of handling faults, other researchers have also classified active fault-tolerant control schemes (AFTCS) into fault detection, identification (diagnosis), and adjustment schemes. AFTC mainly utilizes a real-time fault detection and diagnosis (FDD) module to detect and obtain fault information online, and then reconstructs the control law to make the system achieve a more satisfactory control effect. AFTC mainly includes three aspects: the FDD module, the reconfigurable controller module, and the controller reconfiguration design mechanism. Compared with passive fault-tolerant control, the control effect of active fault-tolerant control will be better.

[0004] In the case of considering optimization, the approximate dynamic programming (ADP) is introduced into the design of the FTC scheme, so as to obtain better control effect. The ADP overcomes the shortcomings of the traditional DP by combining the theories of reinforcement learning and neural network with the classic dynamic programming, and can obtain the approximately optimal control law, and is an effective method for solving the optimal control of the nonlinear system. The traditional dynamic programming (DP) is difficult to implement due to the curse of dimensionality, that is, when the state signal and the input signal dimension of the nonlinear system increase, the data storage and the calculation amount of the DP will increase. In order to overcome these shortcomings, as an important branch of reinforcement learning, the ADP framework effectively solves the problem by estimating the cost function through the neural network, on-line or off-line approximation iteration. The core challenge of deriving the solution of the nonlinear optimal control problem is usually attributed to solving certain Hamilton-Jacobi-Bellman (HJB) equations. The HJB equation is nonlinear, and it is difficult to solve for a general nonlinear dynamic system. In fact, except for very special problems, there is no closed-form solution for such equations. Therefore, engineers have developed numerical solutions for the HJB equation. In order to obtain such a numerical solution, the efficient algorithm of the ADP can be used to obtain it. The incremental control is often combined with the adaptive control to reduce the dependence on the dynamics knowledge, and the algorithm can handle the system nonlinearities and uncertainties without identifying the global system. In the related literature, the incremental adaptive dynamic programming (IADP) algorithm is introduced to solve the adaptive tracking control problem of the nonlinear system. However, the stability analysis of the system is not given in the related work. And in these algorithms, a large number of tuning parameters are common shortcomings. According to the standard gradient descent algorithm, the estimated value of the expected weight of the neural network should be updated, which inevitably needs to calculate a large amount of data on-line, and greatly increases the calculation burden. It is worth noting that some researchers have studied the control strategy based on fewer learning parameters, but these control schemes do not consider the optimized FTC. Therefore, how to develop a new single-network incremental adaptive dynamic programming (SIADP) with fewer learning parameters and less on-line calculation burden and generalize it to the FTC strategy is an important problem. SUMMARY

[0005] The application discloses a nonlinear system actuator fault PPB-SIADP fault-tolerant control method, which comprises the following steps.

[0006] To achieve the above technical purposes, the technical scheme adopted by the application is as follows:

[0007] A nonlinear system actuator fault PPB-SIADP fault-tolerant control method comprises the following steps:

[0008] S1, an actuator fault model is established, and m actuators are divided into q groups according to the redundancy characteristics and functions of the actuators; a continuous MIMO nonlinear system with unknown actuator faults is considered as:

[0009]

[0010] In the formula, is a system state vector, f(.) is: an unknown nonlinear function, and sat(·) represents a saturation nonlinear function; and ξ G (t) is an additional fault after the actuator grouping, μ G (t) = [μ1(t) μ2(t) … μ q (t)] T = Bμ(t), is an ideal virtual control law, is an ideal input vector, is a control distribution matrix, p i is the number of actuators in the i-th group;

[0011] S2, setting constraint boundary for state tracking error, making equivalent transformation for system according to predetermined performance boundary, transforming state tracking problem with error constraint into state regulation problem without constraint condition; transformed tracking error dynamic model is:

[0012] z(t) = Γ - (z(t), p l (t), p u (t));

[0013]

[0014] wherein, p u (t) and p l (t) are upper and lower boundaries of predetermined dynamic performance function;

[0015] S3, making incremental modeling for error system with actuator fault and saturation characteristic according to transformed nonlinear system full quantity model, obtained discrete linear error incremental model is:

[0016]

[0017] wherein, Δz t =z t -z t-1 , Δsat(μ Gξ,t )=Δsat(μ G,t )-Δsat(ξ G,t ), Δsat(μ G,t )=sat(μ G,t )-sat(μ G,t-1 ); expressing p u (t) and p l (t) as p u,t , p l,t , then Δp u,t =p u,t -p u,t-1 , Δp l,t =p l,t -p l,t-1 ; parameter matrix F t-1 , G t-1 , and P t-1 are F°(x t-1 , sat(μ Gξ,t-1 ), p l,t-1 , p u,t-1 ), G°(x t-1 , sat(μ Gξ,t-1 ), p l,t-1 , p u,t-1 ) and F°(x t-1 , sat(μ Gξ,t-1 ), p l,t-1 , p u,t-1 ), G°(x t-1 , sat(μ Gξ,t-1 ), p l,t-1 , p u,t-1 ) at t-1 time respectively; and F t-1 , G t-1 , and P t-1 are F(x t-1 , sat(μ Gξ,t-1 ), p l,t-1 , p u,t-1 ), G(x t-1 , sat(μ Gξ,t-1 ), p l,t-1 , p u,t-1 ) and F(x t-1 , sat(μ Gξ,t-1 ), p l,t-1 , p u,t-1 ), G(x t-1 , sat(μ Gξ,t-1 ), p l,t-1 , p u,t-1 ) respectively. and P°(x t-1 , sat(μ Gξ,t-1 ), p l,t-1 , p u,t-1 );

[0018] S4, the system matrix of the linear error increment model is identified by using recursive least square identification;

[0019] S5, the actuator fault is estimated by using the adaptive increment fault observer based on the RBF neural network; wherein, the adaptive increment fault neural network is defined as:

[0020]

[0021] In the formula, is the approximation of Δz(t), is the approximation of W o , is a positive definite matrix; F o , G o , and P° respectively represent the parameter matrices F°[z(t0), sat(μ Gξ (t0)),p l (t0), p u (t0)]、G°[z(t0), sat(μ Gξ (t0)),p l (t0), p u (t0)]、 and P °[z(t0), sat(μ Gξ (t0)),p l (t0), p u (t0)] at t0;

[0022] The estimated actuator fault is:

[0023]

[0024] In the formula, η o > 0 is the learning rate, is the approximation error of Δz(t), is an activation function, l o > 1 is the number of hidden layer neurons; v s is a sampling period;

[0025] S6, the optimal performance index of anti-saturation is set, and a single network increment adaptive dynamic programming control method is proposed to obtain an approximate optimal fault-tolerant control strategy, and the optimal increment control strategy obtained is:

[0026]

[0027] In the formula, It is a saturator model, a bounded monotonically increasing function is selected according to the actual actuator output range, and satisfies || phi ( · ) || <= phi max , phi max Is a normal number; It is a diagonal positive definite matrix, It is a diagonal matrix of actuator saturation bound, Saturation bound of the i-th actuator; Discount factor gamma is in (0, 1), which is used to control the attention degree of short-term cost or long-term cost; Mu G (t-1) is the control strategy at t-1 time; It is the estimated value of weight W c (t),

[0028] Compared with the prior art, the beneficial effects of the present application are as follows:

[0029] First, the nonlinear system actuator fault PPB-SIADP fault-tolerant control method of the present application uses a single neural network structure to approximate the solution of HJB, instead of using an execution network-evaluation network structure, compared with the traditional IADP control algorithm, and the control signal can be directly obtained by using the information of the evaluation network. More importantly, only the Euclidean norm of the weight estimate is updated, instead of directly updating the weight estimate, which greatly reduces the adaptive learning parameters and reduces the online calculation burden;

[0030] Second, the nonlinear system actuator fault PPB-SIADP fault-tolerant control method of the present application is based on a discrete linear incremental system, and an incremental neural network observer is proposed to estimate multiple actuator faults. Considering the hardware redundancy and actuator grouping scheme, the processing efficiency of actuator faults is improved;

[0031] Third, the nonlinear system actuator fault PPB-SIADP fault-tolerant control method of the present application obtains a new anti-saturation scheme based on SIADP by redesigning the anti-saturation performance index;

[0032] Fourth, the nonlinear system actuator fault PPB-SIADP fault-tolerant control method of the present application first combines PPB and SIADP to propose a new type of predictive performance model-free optimal tracking control scheme;

[0033] Fifth, the nonlinear system actuator fault PPB-SIADP fault-tolerant control method of the present application realizes the proposed adaptive optimal FTC scheme without any prior knowledge of nonlinear dynamics, which improves the applicability of the control method. BRIEF DESCRIPTION OF DRAWINGS

[0034] Figure 1 Block diagram for grouping actuators

[0035] Figure 2 Structure diagram for tailless flying wing aircraft

[0036] Figure 3 Attitude control double loop control block diagram for flying wing aircraft

[0037] Figure 4 Nonlinear system PPB-SIADP fault-tolerant control schematic diagram with actuator failure and saturation characteristics of the present application. DETAILED DESCRIPTION

[0038] The embodiments of the present application are further described in detail below with reference to the accompanying drawings.

[0039] The present application discloses a nonlinear system actuator failure PPB-SIADP fault-tolerant control method, comprising the following steps:

[0040] Step 1: First, an actuator failure model is established, and an actuator grouping scheme is proposed according to the redundancy characteristics and functions of the actuator;

[0041] Consider a continuous MIMO nonlinear system with unknown actuator failure:

[0042]

[0043] where is the system state vector, is the actual input vector with actuator failure, is the ideal input vector, is the additional actuator failure, f(·): is an unknown nonlinear function, sat(·) represents a saturation nonlinear function, i.e.

[0044]

[0045] Assumption 1: All system states are controllable and observable, and all system states are directly measurable.

[0046] Assumption 2: The additional actuator failure ξ(t) is unknown and satisfies ||ξ(t)||≤δ1, where δ1 is a positive integer.

[0047] The present application considers unknown actuator failure (stuck and failure). The failure model is described as:

[0048]

[0049]

[0050]

[0051] where λ i ∈ [0, 1] represents the effectiveness value of μ i (t), denotes the i-th actuator of the controlled system is stuck. Detailed faults are shown in Table 1, where t i denotes the different time instants when the corresponding fault occurs.

[0052] Table 1 Actuator faults

[0053]

[0054] Then the additional actuator faults can be written as:

[0055]

[0056] Assumption 3. When the nonlinear system of equation (1) has actuator faults in the form of equation (6), the control system can still complete the task using the faultless or still available actuators. Assumption 3 is a necessary condition for the FTC scheme of the nonlinear system (1). In order to deal with actuator redundancy, the m actuators are divided into q groups according to their physical meaning:

[0057]

[0058] where pi + p2 +... p q = m, and p j represents the number of the j-th (0 < j < q) group of actuators. The corresponding diagram of actuator grouping is shown in Figure 1 Figure 1, and equation (7) can be rewritten as:

[0059] μ G (t) = [μ1(t) μ2(t)... μ q (t)] T = Bμ(t) (8) ;

[0060] where is the ideal virtual control law, is the control allocation matrix. b ij is the element of B, which can be expressed as:

[0061]

[0062] According to equation (8), equation (1) can be written as:

[0063]

[0064] where and ξ G (t) is an additional fault after the actuator grouping.

[0065] Step 2: Set the constraint boundary for the state tracking error, make an equivalent transformation to the system according to the predetermined performance boundary, and convert the state tracking problem with error constraints into a state regulation problem without constraints;

[0066] Define the desired state signal as The tracking error of the nonlinear system (10) can be obtained as:

[0067] e(t) = x(t) - x d (t) (11) ;

[0068] In the proposed control objective, the predetermined performance indicates that the tracking error e i (t) is limited in the preset range area, which can be described as:

[0069]

[0070] where ρ ui (t) and ρ li (t) are the upper and lower bounds of the predetermined dynamic performance function, and their specific definitions are:

[0071] ρ li (t) = -σ li ρ i (t), ρ ui (t) = σ ui ρ i (t) (13) ;

[0072] In equation (13), positive numbers σ ui and σ li are used to adjust the boundary, which also indicates the relationship between the upper and lower bounds. ρ i (t) : is a continuous and sufficiently smooth, constant positive and monotonically decreasing function, which satisfies In this application, ρ i (t) is chosen as an exponential function shown in equation (14):

[0073]

[0074] where ρ i0 , ρ i∞ and k i are positive numbers used to adjust the predetermined boundary. Specifically, ρ i0 limits the positive and negative overshoots of e i (t). ρ i∞ is the ei The upper bound of the steady state of (t) is allowed to vary. i The convergence rate of (t) is constrained by i The decreasing rate of (t) is determined by k i Adjustments are made. Subsequent control law design aims to force the tracking error within a predetermined domain.

[0075] Define a continuous variable represents the equivalent system variable of the original nonlinear system (10) after a predetermined performance transformation, which satisfies the following equation:

[0076]

[0077] where is a strictly increasing function to be designed, which satisfies all the conditions of equation (16):

[0078]

[0079] To design a S(z i ) that satisfies equation (16), in this application, S(z i ) is designed in the form of equations (17)-(18):

[0080]

[0081]

[0082] Combining equations (17)-(18), we can derive:

[0083]

[0084] Taking the first derivative of equation (19), we get:

[0085]

[0086] After a simple transformation of , we get:

[0087]

[0088] Then equation (20) can be written as:

[0089]

[0090] Therefore, according to equation (19) and equation (22), the tracking error dynamics model after transformation can be rewritten as:

[0091]

[0092] According to equation (15), equation (20) can be rewritten as:

[0093]

[0094] According to equation (15), equation (22) can be rewritten as:

[0095]

[0096] The new tracking error dynamics model is:

[0097]

[0098] Step 3: Incremental modeling of the error system with actuator faults and saturation characteristics for the transformed nonlinear system full state model;

[0099] In order to obtain the incremental form of the error system dynamics model (26) with actuator faults and actuator saturation characteristics, consider the first-order Taylor series expansion around z(t0) and μ Gξ (t0):

[0100]

[0101] where H.O.T = O[(z(t)-z(t0) 2 , (sat(μ Gξ (t)-μ Gξ (t0))) 2 , (ρ l (t)-ρ l (t0)) 2 , (ρ u (t)-ρ u (t0)) 2 ] represents the remaining high-order terms. The state transition matrix F(·), the control energy efficiency matrix G(·), the predetermined performance upper bound parameter matrix and the predetermined performance lower bound parameter matrix P (·) of system (26) are:

[0102]

[0103] Therefore, by rewriting (27), a continuous linear incremental model can be obtained at time t0:

[0104]

[0105] Although the system is continuous, in practice, computers use digital signals for data acquisition and processing, so it is necessary to discretize the continuous system. By assuming a sampling frequency of 1 / v shigh enough, the state derivative can be expressed as:

[0106]

[0107] Equation (29) can then be written as:

[0108] z t+1 = z t + v s Ξ - (z(t), p l (t), p u (t), sat(μ Gξ (t))) (31) ;

[0109] By the same procedure of Equations (27)-(29), we have:

[0110]

[0111] According to Equations (28) and (31), F°[z(t0), μ G ξ(t0), p l (t0), p u (t0)], G°[z(t0), μ Gξ (t0), p l (t0), p u (t0)], and P °[z(t0), μ Gξ (t0), p l (t0), p u (t0)] can be expressed as:

[0112]

[0113] where F°[z(t0), sat(μ Gξ (t0)), p l (t0), p u (t0)], G°[z(t0), sat(μ Gξ (t0)), p l (t0), p u (t0)], , p u (t0)] and P °[z(t0), sat(μ Gξ (t0)), p l (t0), p u (t0)] can be simplified as F o , G o , and P.

[0114] The system is processed by this incremental method at each step, so that z t-1 Gξ,t-1 l,t-1 u,t-1 instead of z(t0), sat(μ Gξ (t0)), ρ l (t0) and ρ u (t0). The discrete linear incremental model is then shown as:

[0115]

[0116] where Δz t = z t - z t-1 , Δsat(μ Gξ,t ) = Δsat(μ G,t ) - Δsat(ζ G,t ), Δsat(μ G,t ) = sat(μ G,t ) - sat(μ G,t-1 ), Δρ l,t = ρ l,t - ρ l,t-1 and Δρ u,t = ρ u,t - ρ u,t-1 . The parameter matrices F°(x t-1 , sat(μ Gξ,t-1 ), ρ l,t-1 , ρ u,t-1 ), G°(x t-1 , sat(μ Gξ,t-1 ), ρ l,t-1 , ρ u,t-1 ), and P °(x t-1 , sat(μ Gξ,t-1 ), ρ l,t-1 , ρ u,t-1 ) can be simplified as F t-1 , G t-1 , and P t-1 in this section. Since the states and control signals of the system are bounded, the state transition matrix F t-1 and the input distribution matrix G t-1 of the system are also bounded. Therefore, it is assumed that ||F t-1 || ≤ F max and ||G t-1 || ≤ G max . Also because the predetermined performance upper bound function ρ​​​u,t and lower bound function p l,t which are both bounded, so the restructured and P t-1 is also bounded.

[0117] Step 4: The system matrix of the linear error increment model is identified by using recursive least square identification;

[0118] RLS is an improved method of LS, which is an online identification algorithm that does not rely on a large amount of data and has high identification accuracy.

[0119] The principle of RLS is to modify the identification parameters at the last time after obtaining new sampling data, and obtain new parameter estimates by recursion, so as to realize the function of online identification.

[0120] First, define the input information matrix and the parameter matrix

[0121]

[0122]

[0123] Second, the core formula of RLS is:

[0124]

[0125]

[0126]

[0127]

[0128] where is the prediction error obtained from equation (34) and equation (36), is the gain matrix, is the estimate of the covariance matrix. At each sampling step, the identification result of the model can be obtained by sequentially executing equations (37)-(40).

[0129] Step 5: The actuator fault is estimated by the adaptive incremental fault observer based on RBF neural network;

[0130] Since the additional actuator fault value is unknown, an adaptive incremental fault observer based on RBF neural network is proposed. Since the system (29) is incremental, the fault observer is designed in incremental form. The corresponding incremental actuator fault is defined as:

[0131]

[0132] where is the desired weight vector, is the activation function, l o > 1 is the number of hidden layer neurons, and is the approximation error of the RBF neural network.

[0133] Then Δμ Gξ (t) can be expressed as:

[0134]

[0135] Substituting (42) into the continuous incremental system (29), we have:

[0136]

[0137] Then the adaptive incremental fault neural network can be defined as:

[0138]

[0139] where is the approximation of Δz(t), is the approximation of W o , and is a positive definite matrix.

[0140] The weight vector can be updated as:

[0141]

[0142] where η o > 0 is the learning rate, is the approximation error of Δz(t).

[0143] To match the actuator fault estimation result with the discrete RLS identification and optimal control law design, (45) should be rewritten as:

[0144]

[0145] Then the approximate discrete actuator fault is obtained in incremental form:

[0146]

[0147] Then the corresponding actuator fault can be constructed as:

[0148]

[0149] Step 6: Set the optimal performance index of anti-saturation, and propose a single network incremental adaptive dynamic programming control method to obtain the approximate optimal fault-tolerant control strategy.

[0150] Consider the infinite horizon quadratic cost function as

[0151]

[0152] where R > 0 and Q ≥ 0. is the cost function at each sampling step. The discount factor γ ∈ (0, 1) ensures that the cost (49) is finite for any state. By adjusting γ, it can control the degree of attention to short-term cost or long-term cost.

[0153] If s(μ G,t-1 + Δ μG,t ) = (μ G,t-1 + Δμ G,t ) T R(μ G,t-1 + Δμ G,t ) is taken as the performance index function, then the optimal control law obtained for the control system design with actuator saturation is difficult to guarantee that the system reaches the optimal, and even can lead to instability of the system. Therefore, after considering the actuator saturation problem, the cost function (49) can be rewritten as

[0154]

[0155] where is the anti-saturation performance index function. Let be a diagonal positive definite matrix, be a diagonal actuator saturation boundary matrix, be the saturation boundary of the i-th actuator. is the saturation model, which can be selected according to the actual actuator output range. It is a bounded monotonic increasing function and satisfies ||φ(·)||≤φ max , φ max is a normal number. φ -1 (Δμ G (i)+μ G (t-1)) = [φ -1 (Δμ G,1 (i)+μ G,1 (i-1)), …, φ -1 (Δμ G,m (i)+μ G,m (i-1)) T .

[0156] Equation (50) can be rewritten as

[0157]

[0158] where Then, the discrete form of Hamilton function:

[0159]

[0160] Optimal cost function J * (z t It can be defined as:

[0161]

[0162] According to Bellman's optimality principle, J * (z t Satisfies the HJB equation:

[0163]

[0164] By solving the partial derivative of equation (54), that is... Then the closed-loop incremental optimal control strategy can be obtained.

[0165]

[0166] in Then, a corresponding optimal control strategy can be constructed:

[0167]

[0168] Substituting equation (56) into equation (54), we can obtain the discrete form of the HJB equation:

[0169]

[0170] Directly designing the overall control strategy This differs from traditional ADP. This chapter first designs a closed-loop incremental control strategy. Then in μ G,t-1 and An overall control strategy was built on this basis.

[0171] Since the solution to the HJB equation (57) is difficult to find directly, an adaptive online policy iteration (PI) control algorithm, as shown in Table 2, is proposed to approximate the solution to the HJB equation.

[0172] Table 2

[0173]

[0174] Based on the adaptive online PI control algorithm, a novel SIADP optimal control algorithm is proposed to approximate the incremental optimal control law. The optimal cost function J... * (z t It can be approximated as:

[0175]

[0176] where is the desired weight vector, is the activation function, and l > 1 is the number of neurons in the hidden layer, is the approximation error of the evaluation network.

[0177] A novel ANN is used to estimate the optimal cost function as follows:

[0178]

[0179] where and is the approximation error of the evaluation network.

[0180] Then the discrete form of the Hamiltonian function can be expressed as:

[0181]

[0182] where and Since the approximation error of the evaluation network is bounded, the is also bounded, that is,

[0183] Since the desired weight W c (t) is unknown, the estimate of the cost function can be defined as:

[0184]

[0185] Then the approximate Hamiltonian function can be expressed as:

[0186]

[0187] According to the gradient descent algorithm and the chain rule, the weight update is achieved by minimizing the objective function designed in the form of squared residuals:

[0188]

[0189] The weight update rule of the evaluation network is updated as equations (64) and (65):

[0190]

[0191]

[0192] where η c > 0 is the learning rate of the evaluation network.

[0193] According to equations (60) and (62), we have:

[0194]

[0195] where and

[0196] According to the formula (55) to formula (66), can obtain:

[0197]

[0198] Therefore, according to formula (55) and formula (59), the expected incremental optimal control strategy can be described as:

[0199]

[0200] where and Therefore, assuming and

[0201] The corresponding actual optimal incremental control strategy is:

[0202]

[0203] Examples

[0204] Step 1: Take the flying wing aircraft nonlinear system as an example, as shown in Figure 2 The flying wing aircraft model and actuator fault model are established, and an actuator grouping scheme is proposed according to the redundancy characteristics and functions of the actuator;

[0205] According to the principle of time scale separation, the state variables can be divided into four groups with different response speeds, and the inner two rings are the angular rate ring and the attitude ring. The mathematical model expression is:

[0206]

[0207]

[0208] where is the inertia matrix. u = [μ a μ e μ r ] T is the control input vector. μ a , μ e and μ r represent the total deflection angle of aileron, elevator and drag rudder respectively. ω = [p q r] T is the angular rate vector in the body reference coordinate system, p is the pitch rate, q is the roll rate, and r is the yaw rate. is the dynamic pressure, b is the wingspan, S ω is the wing area, c A is the mean aerodynamic chord, C nβ , C np , C nr , C nμa and C nμr is the aerodynamic derivative of the yawing moment, C lβ , C lp , C lr , C lμa and C lμr is the aerodynamic derivative of the rolling moment. C m0 , C mα , C mq and C mμe is the aerodynamic derivative of the pitching moment. M T is the pitching moment provided by the engine. Let and . [μ, α, β] T are the attitude angle vector, μ is the roll angle, α is the angle of attack, and β is the sideslip angle. χ is the flight azimuth angle, and γ is the flight path angle.

[0209] Let then the two-loop model of the flying wing aircraft can be obtained as shown in equation (72):

[0210]

[0211] where x1(t) = [μ, α, β] T , x2(t) = [p, q, r] T , μ(t) = [μ a (t), μ e (t), μ r (t)] T . The two-loop control block diagram of the flying wing aircraft attitude control is shown in Figure 2 . Figure 3 The reference signals and the desired command signals generated by each control loop are shown in

[0212] Considering the unknown actuator faults and saturation characteristics, equation (72) can be rewritten as:

[0213]

[0214] where is the system state vector, is the actual input vector with actuator faults, is the ideal input vector, is the additional actuator fault, f(·): is an unknown nonlinear function, sat(·) denotes a saturation nonlinear function, i.e.,

[0215]

[0216] Assumption 1 All system states are controllable and observable, and all system states are directly measurable.

[0217] Assumption 2 The additional actuator fault ξ(t) is unknown and satisfies ||ξ(t)||≤δ1, where δ1 is a positive integer.

[0218] This chapter considers unknown actuator faults (stuck-at and loss-of-control). The fault model is described as:

[0219]

[0220]

[0221]

[0222] where λ i ∈ [0, 1] represents the effectiveness of μ i (t), denotes that the i-th actuator of the controlled system is stuck-at. Detailed faults are shown in Table 1, where t i denotes the different time instants when the corresponding fault occurs.

[0223] Then the additional actuator fault can be written as:

[0224]

[0225] Assumption 3 When the nonlinear system (73) has actuator faults of the form (78), the control system can still complete the task using the fault-free or post-fault available actuators.

[0226] Assumption 3 is a necessary condition for the FTC scheme of the nonlinear system (73). To handle actuator redundancy, the m actuators are divided into q groups according to their physical meaning.

[0227]

[0228] where pi + p2 + … p q = m, and p i represents the number of actuators in the j-th (0 < j ≤ q) group. The corresponding diagram of actuator grouping is shown in Figure 1 (79) can be rewritten as:

[0229] μ G (t) = [μ1(t) μ2(t) … μ q (t)]T = Bμ(t) (80);

[0230] where is the ideal virtual control law, is the control allocation matrix. b ij is an element of B, which can be expressed as:

[0231]

[0232] According to equation (80), equation (73) can be written as:

[0233]

[0234] where and ξ G (t) is the additional fault after actuator grouping.

[0235] Step 2: Set the constraint boundary for the state tracking error, make an equivalent transformation to the system according to the predetermined performance boundary, and convert the state tracking problem with error constraints into a state regulation problem without constraint conditions;

[0236] Define the desired state signal as The tracking error of the nonlinear system (73) can be obtained as:

[0237] e(t) = x(t) - x d (t) (83);

[0238] In the proposed control objective, the predetermined performance indicates that the tracking error e i (t) is limited within the preset range area, which can be described as:

[0239]

[0240] where ρ ui (t) and ρ li (t) are the upper and lower bounds of the predetermined dynamic performance function, and their specific definitions are:

[0241] ρ li (t) = -σ li ρ i (t), ρ ui (t) = σ ui ρ i (t) (85);

[0242] In equation (85), positive numbers σ ui and σ li are used to adjust the boundary, which also indicates the relationship between the upper and lower bounds. ρ i (t): It is a continuous, sufficiently smooth, always positive, and monotonically decreasing function that satisfies In this article, we choose ρ i (t) is the exponential function shown in equation (86):

[0243]

[0244] Where ρ i0 ρ i∞ and k i It is a positive number used to adjust predetermined boundaries. Specifically, ρ i0 Restricted e i Positive and negative overshoot of (t). ρ i∞ It is e i The range of the steady-state upper limit allowed for (t). i The convergence rate constraint of (t) depends on ρ i The rate of decrease of (t), which is determined by k i Adjustments were made. The subsequent control law was designed to force the tracking error within a predetermined range.

[0245] Define a continuous variable The system variables that represent the equivalent system variables of the original nonlinear system (73) after a predetermined performance transformation satisfy the following equation:

[0246]

[0247] in It is a strictly increasing function to be designed, which satisfies all the conditions of equation (88):

[0248]

[0249] In order to design a Satisfying equation (88), in this example, the design For the form of equations (89)-(90):

[0250]

[0251]

[0252] Combining equations (89) and (90), we can deduce that:

[0253]

[0254] Taking the first derivative of equation (91), we can obtain:

[0255]

[0256] right With simple transformation, we have

[0257]

[0258] Then, equation (92) can be written as

[0259]

[0260] Therefore, according to equation (91) and equation (94), the transformed tracking error dynamics model can be rewritten as

[0261]

[0262] According to equation (87), equation (91) can be rewritten as

[0263]

[0264] Similarly, according to equation (87), equation (94) can be rewritten as

[0265]

[0266] The new tracking error dynamics model is

[0267]

[0268] Step 3: Incremental modeling for the error system with actuator faults and saturation characteristics is performed for the transformed nonlinear system full-state model;

[0269] In order to obtain the incremental form of the error system dynamics model (82) with actuator faults and actuator saturation characteristics, consider the first-order Taylor series expansion around z(t0) and μ Gξ (t0):

[0270]

[0271] where H.O.T = o[(z(t)-z(t0) 2 , (sat(μ Gξ (t)-μ Gξ (t0))) 2 , (ρ l (t)-ρ l (t0)) 2 , (ρ u (t)-ρ u (t0)) 2 ] represent the remaining high-order terms. The state transition matrix F(·) of system (82), the control energy efficiency matrix G(·), the predetermined performance upper bound parameter matrix and the predetermined performance lower bound parameter matrix areP (·)for:

[0272]

[0273] Therefore, by rewriting (99), a continuous linear increment model can be obtained at time t0:

[0274]

[0275] Although the system is continuous, computers actually use digital signals for data acquisition and processing, therefore it is necessary to discretize the continuous system. This is achieved by assuming a sampling frequency of 1 / v. s If it is high enough, the state derivative can be expressed as:

[0276]

[0277] Then equation (98) can be written as:

[0278] z t+1 =z t +v s Ξ - (z(t), ρ l (t), ρ u (t),sat(μ Gξ (t))) (103);

[0279] By following the same steps as in equations (99) to (101), we can obtain:

[0280]

[0281] According to equations (100) and (104), F°[z(t0), μ Gξ (t0), ρ l (t0), ρ u (t0)]、G°[z(t0), μ Gξ (t0), ρ l (t0), ρ u (t0)]、 and P °[z(t0), μ Gξ (t0), ρ l (t0), ρ u [t0] can be represented as:

[0282]

[0283] Where F°[z(t0),sat(μ) Gξ (t0)), ρ l (t0),ρ u(t0)], G° [z(t0), sat(μ Gξ (t0)), p l (t0), p u (t0)], and P° [z(t0), sat(μ Gξ (t0)), p l (t0), p u (t0)] can be simplified in this section to F o , G o , and P °.

[0284] The system is processed at each step by this incremental method, so that z t-1 , sat(μ Gξ,t-1 ), p l,t-1 and p u,t-1 can be used instead of z(t0), sat(μ Gξ (t0)), p l (t0) and p u (t0). The discrete linear incremental model is then shown as:

[0285]

[0286] where Δz t = z t - z t-1 , Δsat(μ Gξ,t ) = Δsat(μ G,t ) - Δsat(ζ G,t ), Δsat(μ G,t ) = sat(μ G,t ) - sat(μ G,t-1 ), Δp l,t = p l,t - p l,t-1 and Δp u,t = p u,t - p u,t-1 . The parameter matrices F°(x t-1 , sat(μ Gξ,t-1 ), p l,t-1 , p u,t-1 ), G°(x t-1 , sat(μ Gξ,t-1 ), p l,t-1 , p u,t-1 ), and P°(x t-1 , sat(μ Gξ,t-1 ), p l,t-1 , p u,t-1In this chapter, it can be simplified to F. t-1 G t-1 , and P t-1 Because the system's state and control signals are bounded, the system's state transition matrix F t-1 and input assignment matrix G t-1 It is also bounded. Therefore, assume ||F t-1 ||≤F max and ||G t-1 ||≤G max Also because of the predetermined upper bound function ρ. u,t and lower bound function ρ l,t They are all inherently bounded, so through reconstruction... and P t-1 It is also bounded.

[0287] Step 4: Identify the system matrix of the linear error increment model using recursive least squares identification;

[0288] RLS is an improved version of LS, and it is an online identification algorithm that does not rely on a large amount of data and has high identification accuracy.

[0289] The principle of RLS is to modify the identification parameters from the previous time step after obtaining new sampled data, and then obtain new parameter estimates through recursion, thereby achieving the function of online identification.

[0290] First, define the input information matrix. and parameter matrix

[0291]

[0292]

[0293] Secondly, the core formula of RLS is:

[0294]

[0295]

[0296]

[0297]

[0298] in The prediction error is obtained from equations (106) and (108). It is the gain matrix. is the estimate of the covariance matrix. At each sampling step, the identification result of the model can be obtained by sequentially executing equations (109)-(112).

[0299] Step 5: The actuator fault is estimated by the adaptive incremental fault observer based on RBF neural network.

[0300] Since the additional actuator fault value is unknown, an adaptive incremental fault observer based on RBF neural network is proposed. Since the system (101) is incremental, the fault observer is designed in incremental form. The corresponding incremental actuator fault is defined as:

[0301]

[0302] where is the desired weight vector, is the activation function, l o >1 is the number of hidden layer neurons, and is the approximation error of the RBF neural network;

[0303] Then Δμ Gξ (t) can be expressed as:

[0304]

[0305] Substituting equation (114) into the continuous incremental system (101), we can get:

[0306]

[0307] Then the adaptive incremental fault neural network can be defined as:

[0308]

[0309] where is the approximation of Δz(t), is the approximation of W o , and is a positive definite matrix. The weight vector can be updated as:

[0310]

[0311] where η o >0 is the learning rate, is the approximation error of Δz(t).

[0312] In order to match the actuator fault estimation result with the discrete RLS identification and optimal control law design, (117) should be rewritten as:

[0313]

[0314] Then the approximate discrete actuator fault is obtained in incremental form:

[0315]

[0316] The corresponding actuator fault can be constructed as:

[0317]

[0318] Step 6: Set the optimal performance index of anti-saturation, and propose a single network incremental adaptive dynamic programming control method to obtain an approximate optimal fault-tolerant control strategy.

[0319] Consider the infinite time domain quadratic cost function:

[0320]

[0321] where R > 0 and Q ≥ 0. is the cost function at each sampling step. The discount factor γ ∈ (0, 1) ensures that the cost (49) at any state is finite. By adjusting γ, it can control the degree of attention to short-term cost or long-term cost.

[0322] If s(μ G,t-1 + Δμ G,t ) = (μ G,t-1 + Δμ G,t ) T R(μ G,t-1 + Δμ G,t ) is taken as a performance index function, then the optimal control law obtained by the control system design with saturated actuators is difficult to guarantee that the system reaches the optimal, and it may even lead to instability of the system. Therefore, after considering the actuator saturation problem, the cost function (121) can be rewritten as:

[0323]

[0324] where is the anti-saturation performance index function. Let be a diagonal positive definite matrix, be a diagonal actuator saturation bound matrix, be the saturation bound of the i-th actuator. is the saturation model, which can be selected according to the output range of the actual actuator. It is a bounded monotonic increasing function and satisfies ||φ(·)|| ≤ φ max , φ max is a constant. φ -1 (Δμ G (i) + μ G (i-1)) = [φ -1 (ΔμG,1 (i)+μ G,1 (i-1)),..., φ -1 (Δμ G,m (i)+μ G,m (i-1)) T .

[0325] Equation (122) can be rewritten as:

[0326]

[0327] where Then, the Hamiltonian function in discrete form is:

[0328]

[0329] The optimal cost function J * (z t ) can be defined as:

[0330]

[0331] According to the Bellman optimality principle, J * (z t ) satisfies the HJB equation:

[0332]

[0333] By solving the partial derivative of equation (126), that is, then the closed-form incremental optimal control strategy

[0334]

[0335] where then the corresponding optimal control strategy can be constructed:

[0336]

[0337] Substituting equation (128) into equation (126), the discrete form of the HJB equation can be obtained:

[0338]

[0339] Unlike the traditional ADP that directly designs the overall control strategy This section first designs the closed-loop incremental control strategy and then constructs the overall control strategy G,t-1

[0340] ​​Since it is difficult to find the solution of the HJB equation (129) directly, an adaptive online policy iteration (PI) control algorithm as shown in Table 2 is proposed to approximate the solution of the HJB equation.

[0341]

[0342]

[0343] Based on the adaptive online PI control algorithm, a new SIADP optimal control algorithm is proposed to approximate the incremental optimal control law. The optimal cost function J * (z t ) can be approximated as:

[0344]

[0345] where W is the desired weight vector, is the activation function, l > 1 is the number of neurons in the hidden layer, is the approximation error of the critic network.

[0346] A new type of ANN is used to estimate the optimal cost function as follows:

[0347]

[0348] where and are the approximation errors of the critic network, respectively.

[0349] Then the discrete form of the Hamiltonian function can be expressed as:

[0350]

[0351] where and Since the approximation error of the critic network is bounded, the is also bounded, that is,

[0352] Since the desired weight W c (t) is unknown, the estimate of the cost function can be defined as:

[0353]

[0354] Then the approximate Hamiltonian function can be expressed as:

[0355]

[0356] According to the gradient descent algorithm and the chain rule, the weight update is achieved by minimizing the objective function designed in the form of squared residuals:

[0357]

[0358] The weight update rule of the critic network is updated as Equations (136) and (137):

[0359]

[0360]

[0361] where η c > 0 is the learning rate of the critic network.

[0362] According to Equations (132) and (134), we have:

[0363]

[0364] where and

[0365] According to Equations (137)-(138), we have:

[0366]

[0367] Therefore, according to Equations (127) and (131), the desired incremental optimal control policy can be described as:

[0368]

[0369] where and Therefore, assuming and

[0370] the corresponding actual optimal incremental control policy is:

[0371]

[0372] Those skilled in the art will appreciate that embodiments of the application can be devised for a method, a system, or a computer program product. Accordingly, the present application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code. Embodiments of the present application can be implemented with various computer program languages such as the object-oriented programming language Java and the interpreted scripting language JavaScript, etc.

[0373] The present application is described in reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0374] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0375] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0376] While preferred embodiments of the application have been described, modifications and variations can be apparent to those skilled in the art once aware of the general underlying concepts. Accordingly, the appended claims are intended to embrace all such modifications and variations as fall within the scope of the application.

[0377] It will be apparent to those skilled in the art that various modifications and variations can be made to the present application without departing from the spirit or scope of the application. Thus, it is intended that the present application cover modifications and variations of this application provided they come within the scope of the appended claims and their equivalents.

Claims

1. A nonlinear system actuator fault PPB-SIADP fault-tolerant control method, characterized in that, The nonlinear system PPB-SIADP fault-tolerant control method The method comprises the following steps: S1, an actuator fault model is established, and m actuators are divided into q groups according to the redundancy characteristics and functions of the actuators; considering a continuous MIMO nonlinear system with unknown actuator faults as: where is the system state vector, f(·) is an unknown nonlinear function, and sat(·) denotes a saturation nonlinear function; and ξ G (t) is an additional fault after actuator grouping, μ G (t) = [μ1(t) μ2(t) … μ q (t)] T = Bμ(t), is an ideal virtual control law, is an ideal input vector, is a control allocation matrix, p i is the number of actuators in the ith group;​ S2, a constraint boundary is set for a state tracking error, the system is equivalently transformed according to a predetermined performance boundary, a state tracking problem with an error constraint is converted into a state regulation problem without a constraint condition; a transformed tracking error dynamic model is: z(t) = Γ - (z(t), p l (t), p u (t)); where p u (t) and p l (t) are upper and lower bounds of a predetermined dynamic performance function. S3, an incremental modeling is performed on an error system containing actuator faults and having a saturation characteristic for a transformed nonlinear system full quantity model, and a discrete linear error incremental model obtained is: where Δz t = z t - z t+1 , Δsat(μ Gξ,t ) = sat(μ G,t ) - sat(μ G,t ), Δsat(μ c,t ) = sat(μ G,t ) - sat(μ G,t-1 ); ρ u (t) and ρ l (t) are expressed as ρ u,t , ρ l,t , then Δρ u,t = ρ u,t - ρ u,t-1 , Δρ l,t = ρ l,t - ρ l,t-1 ; the parameter matrices F t-1 , G t-1 , and P t-1 are F°(x t-1 , sat(μ Gξ,t-1 ), ρ l,t-1 , ρ u,t-1 ), G°(x t-1 , sat(μ Gξ,t-1 ), ρ l,t-1 , ρ u,t-1 ) and P °(x t-1 , sat(μ Gξ,t-1 ), ρ l,t-1 , ρ u,t-1 ) at t-1 time, respectively. S4, a recursive least square identification is used to identify a system matrix of the linear error incremental model; S5, an adaptive incremental fault observer based on a RBF neural network is used to estimate the actuator faults; wherein the adaptive incremental fault neural network is defined as: wherein is an approximation of Δz(t), is an approximation of W o , is a positive definite matrix; F o , G o , and P ° respectively denote the parameter matrices F° [z(t0), sat(μ Gξ (t0)), p l (t0), p u (t0)], G° [z(t0), sat(μ Gξ (t0)), p l (t0), p u (t0)], and P ° [z(t0), sat(μ Gξ (t0)), p l (t0), p u (t0)] at the instant t0. The estimated actuator faults are: wherein η o >0 is the learning rate, is the approximation error of Δz(t), is the activation function, l o >1 is the number of hidden layer neurons; v s is the sampling period; S6, an optimal performance index of anti-saturation is set, a single-network incremental adaptive dynamic programming control method is used to obtain an approximate optimal fault-tolerant control strategy, and an optimal incremental control strategy obtained is: wherein is a saturator model, a bounded monotonically increasing function is selected according to the actual actuator output range, and satisfies max , φ max is a normal number; is a diagonal positive definite matrix, is a diagonal matrix of actuator saturation bounds, is the saturation bound of the i-th actuator; the discount factor γ∈(0, 1) is used to control the degree of attention to short-term cost or long-term cost; μ G (t-1) is the control strategy at t-1 time; is the estimated value of the weight W c (t), 2. The nonlinear system actuator fault PPB-SIADP fault-tolerant control method according to claim 1, wherein, In step S1, the step of establishing the actuator fault model comprises the following steps: S11, a continuous MIMO nonlinear system with unknown actuator faults is constructed as: where is the system state vector, is the actual input vector with actuator faults, is the ideal input vector, is an additional actuator fault, f(·): is an unknown nonlinear function, sat(·) denotes a saturation nonlinear function: S12, the following assumptions are made: all system states are controllable and observable, and all system states are directly measured; an additional actuator fault ξ(t) is unknown and satisfies ||ξ(t)||≤δ1, wherein δ1 is a positive integer; The fault model is described as follows: where λ i ∈ [0, 1] represents μ i (t) of the performance value, represents the i-th actuator of the controlled system occurs dead, actuator fault as shown in Table 1, wherein t i represents the different time when the corresponding fault occurs: Table 1 Then the additional actuator fault is: S13, m actuators are divided into q groups according to the physical meaning of the actuators: where p1+p2+...p q = m, and p j represents the number of the jth group of actuators, 0 < j < q; The actuator grouping is rewritten as: μ G (t) = [μ1(t) μ2(t) … μ q (t)] T = Bμ(t); wherein is the ideal virtual control law, is the control allocation matrix; b ij is an element of B, denoted as: The continuous MIMO nonlinear system in step S11 is rewritten as: wherein and ξ G (t) is an additional fault after the actuator grouping.

3. The nonlinear system actuator fault PPB-SIADP fault-tolerant control method according to claim 1, wherein, In step S2, the process of setting a constraint boundary for a state tracking error, equivalently transforming the system according to a predetermined performance boundary, and converting a state tracking problem with an error constraint into a state regulation problem without a constraint condition comprises the following steps: S21, defining the desired state signal as The tracking error of the nonlinear system is obtained as e(t) = x(t) - x d (t); Among the proposed control objectives, the predetermined performance indicates that the tracking error e i (t) is limited in a region of a pre-set range, i.e.: where p ui (t) and p li (t) are upper and lower bounds of a predetermined dynamic performance function, which are defined as follows: p li (t) = -σ li p i (t), p ui (t) = σ ui p i (t); where the positive number σ ui and σ li for adjusting the boundary, indicating the relationship between the upper and lower bounds; p i (t): is a continuous and sufficiently smooth, constant positive and monotonically decreasing function, satisfying p i (t) is an exponential function: where ρ i0 , ρ i∞ , and k i are positive numbers for adjusting the predetermined boundary; ρ i0 limits the positive and negative overshoots of e i (t); ρ i∞ is the range allowed for the upper limit of e i (t) in the steady state; the constraint on the convergence speed of e i (t) depends on the decrement rate of e i (t), which is adjusted by k i ; S22, define a continuous variable represents the equivalent system variable of the nonlinear system after a predetermined performance transformation, which satisfies the following equation: wherein is a strictly increasing function to be designed, which satisfies the following conditions: S23, S(z i ) is designed: It is derived that: For z i Taking the first derivative, we get: right Perform the transformation: Then z i the first derivative of z transforms to: S24, the tracking error dynamic model after transformation is rewritten as: z(t) = Γ(e(t), p l (t), p u (t)) z is obtained i is: is obtained is: The transformed tracking error dynamic model is: z(t) = Γ - (z(t), p l (t), p u (t)); 4. The nonlinear system actuator fault PPB-SIADP fault-tolerant control method according to claim 3, wherein, In step S3, the process of performing incremental modeling on an error system containing actuator faults and having a saturation characteristic for a transformed nonlinear system full quantity model comprises the following steps: A first order Taylor series expansion around z(t0) and μ Gξ (t0) is performed: where H.O.T = o[(z(t) - z(t0)) 2 , (sat(μ Gξ (t) - μ Gξ (t0))) 2 , (ρ l (t) - ρ l (t0)) 2 , (ρ u (t) - ρ u (t0)) 2 ] represents the remaining higher order terms; the state transition matrix F(·), the control energy efficiency matrix G(·), the predetermined performance upper bound parameter matrix and the predetermined performance lower bound parameter matrix P (·) are given by: A continuous linear incremental model at t0 is obtained: The continuous system is discretized, and a state derivative is represented as: It is derived that: F°[z(t0), μ Gξ (t0), p l (t0), p u (t0)] and G°[z(t0), μ Gξ (t0), p l (t0), p u (t0)] are given by: and P °[z(t0), μ Gξ (t0), p l (t0), p u (t0)] are given by: with z t-1 , sat(μ Gξ,t-1 ), p l,t-1 and p u,t-1 instead of z(t0), sat(μ cξ (t0)), p l (t0) and p u (t0), the discrete linear increment model reads:

5. The nonlinear system actuator fault PPB-SIADP fault-tolerant control method according to claim 4, wherein, In step S4, the process of identifying a system matrix of a linear error incremental model by using a recursive least square identification comprises the following steps: defining an input information matrix and a parameter matrix A core formula of the recursive least square identification is obtained as: wherein is the prediction error, is the gain matrix, is an estimate of the covariance matrix.

6. The nonlinear system actuator fault PPB-SIADP fault-tolerant control method according to claim 5, wherein, In step S5, the process of estimating the actuator faults by using an adaptive incremental fault observer based on a RBF neural network comprises the following steps: S51, a fault observer is designed in an incremental form, and a corresponding incremental actuator fault is defined as: wherein is the desired weight vector, is an activation function, l o >1 is the number of hidden layer neurons, and is the approximation error of the RBF neural network; S52, Δμ Gξ (t) is represented by: Substituting the continuous incremental system obtains: Then the adaptive incremental fault neural network is defined as: wherein is an approximation of Δz(t), is an approximation of W o , is a positive definite matrix; the weight vector is updated to: where η o >0 is the learning rate, is the approximation error of Δz(t); S53, in order to make the actuator fault estimation results and discrete RLS identification and optimal control law design matching, the weight vector is rewritten as: The approximate discrete actuator fault is obtained in incremental form: The corresponding actuator fault is constructed as:

7. The nonlinear system actuator fault PPB-SIADP fault-tolerant control method according to claim 6, wherein, In step S6, the optimal performance index of anti-windup is set, and the process of obtaining the approximate optimal fault-tolerant control strategy by the single network incremental adaptive dynamic programming control method includes the following steps: S61, considering the infinite time domain quadratic cost function: where R > 0 and Q > 0; is the cost function at each sample step; the discount factor γ e (0, 1) ensures that the cost at any state is finite; the degree of focus on short-term or long-term cost is controlled by adjusting γ; S62, after considering the actuator saturation problem, the quadratic cost function is rewritten as: wherein is an anti-windup performance index function; let is a diagonal positive definite matrix, is a diagonal matrix of actuator saturation bounds, is the saturation bound of the i-th actuator; is a saturator model, which is chosen according to the actual actuator output range by a bounded monotonic increasing function and satisfies ||φ(·)||≤φ max , φ max is a positive constant; S63, the quadratic cost function is further rewritten as: wherein S64, define the discrete form of Hamilton function as: The optimal cost function J * (z t ) is defined as: According to the Bellman optimality principle, J * (z t ) satisfies the HJB equation: through the partial derivative of the HJB equation obtaining a closed-form incremental optimal control policy wherein The corresponding optimal control strategy is constructed as follows: The discrete form of HJB equation is obtained: An adaptive online policy iteration control algorithm is used to approximate the solution of HJB equation; S65, on the basis of adaptive online PI control algorithm, using the optimized SIADP optimal control algorithm, to approximate the incremental optimal control law; optimal cost function J * (z t ) is approximately: wherein is the desired weight vector, is the activation function, l > 1 is the number of neurons in the hidden layer, is the approximation error of the network; The optimal cost function is estimated as: wherein and is the approximation error of the corresponding evaluation network; The discrete form of Hamilton function is expressed as: wherein and is bounded; S66, define the cost function estimation as: The approximate Hamilton function is expressed as: According to the gradient descent algorithm and chain rule, the weight update is realized by minimizing the objective function designed in the form of square residual: The weight update method of the evaluation network is updated as: where η c >0 is the learning rate of the evaluation network; Get: wherein and Pushed: The desired incremental optimal control strategy is: wherein and The corresponding actual optimal incremental control strategy is:

8. The nonlinear system actuator fault PPB-SIADP fault-tolerant control method according to claim 7, wherein, In step S64, the process of using an adaptive online policy iteration control algorithm to approximate the solution of HJB equation includes: S641, select an initial allowable increment control strategy the control strategy μ at time t-1 G,t-1 and a positive number ε; S642, policy evaluation is performed according to the following equation to solve J i (z t ): S643, update the incremental control strategy to: S644, if ||J i (z t )-J i-1 (z t )|| < ε, stop iteration, and obtain the approximate increment optimal control policy; otherwise, let i = i + 1, and return to S642.