An Online Transfer Heuristic Dynamic Programming Method Based on a Disturbance Observer

Through an online transfer heuristic dynamic programming method based on perturbation observer, combined with transfer learning and adaptive criticism control, the optimal control problem of uncertain nonlinear systems is solved, rapid learning and robust control are achieved, and the stability and efficiency of the system are improved.

CN118467895BActive Publication Date: 2025-07-04BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410491460.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-23
Publication Date
2025-07-04
Estimated Expiration
2044-04-23

AI Technical Summary

Technical Problem

The existing optimal control method has limitations when dealing with uncertain nonlinear systems, especially the difficulty of initial strategy dependence and interference processing, which is difficult to effectively apply in actual engineering.

Method used

The online delivery heuristic dynamic programming method based on perturbation observer is adopted, combining transfer learning and adaptive criticism control, a robust optimal control strategy is designed, interference is estimated through perturbation observer and disturbance compensation control is designed, and the truncation mechanism and neural network update control strategy are used to realize online control.

Benefits of technology

It speeds up learning speed, reduces computing resource consumption, improves the control performance of uncertain nonlinear systems, ensures robustness and stability, and the neural network weight error is consistently bounded under certain conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118467895B_ABST
    Figure CN118467895B_ABST
Patent Text Reader

Abstract

The present invention discloses an online transfer heuristic dynamic programming method based on a disturbance observer, belonging to the field of nonlinear time-varying delay system control, and adopting a method combining transfer learning and adaptive critic control; using sample data collected from source tasks to obtain prior knowledge, designing a robust optimal control strategy for nonlinear discrete systems, and using it to guide the online control process of the nonlinear system of the target task; proposing a new type of attenuation function with a truncation mechanism to avoid negative transfer effects and save computational resources; proposing a disturbance compensation control mechanism to solve uncertainties; proving the characteristics of uncertain nonlinear systems under robust optimal control and that the weight errors of neural networks are ultimately uniformly bounded under certain given conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of optimal control, and particularly to an online transfer heuristic dynamic programming method based on a disturbance observer. Background Art

[0002] Optimal control is an important research field in control theory and control engineering, aiming to find control strategies to optimize system performance indices and guide a dynamic system to an equilibrium point. However, optimal control is restricted in practical engineering by limitations such as the curse of dimensionality and the requirement for an accurate model. Traditional dynamic programming methods, such as the Hamilton-Jacobi-Bellman (HJB) equation, are difficult to solve. Therefore, the application of traditional methods in nonlinear systems is limited.

[0003] In response to the challenges of optimal control, researchers have turned to adaptive dynamic programming (ADP), using reinforcement learning and neural networks to alleviate the limitations of traditional dynamic programming. Traditional ADP methods include policy iteration and value iteration, but in practical applications, obtaining an acceptable initial policy is challenging. Therefore, some research has focused on improving value iteration methods, such as multi-step schemes and the utilization of historical data, to accelerate the learning speed and improve control performance. In addition, methods such as guaranteed cost control and robust optimal control have also been proposed to handle the control problems of uncertain systems without an initial acceptable policy.

[0004] However, despite the certain progress made in existing research, there are still some problems. For example, when dealing with disturbances, existing methods still have limitations, especially for the design of general optimal control strategies for uncertain nonlinear systems. Additionally, the selection of a specific upper bound function and the dependence on the initial policy remain challenges. Therefore, although there are some solutions, how to better handle these problems still requires further research and exploration. Summary of the Invention

[0005] The object of the present invention is to overcome the defects existing in the prior art and solve the optimal control problem of a class of uncertain nonlinear discrete-time systems.

[0006] To achieve the above object, the present invention proposes an online transfer heuristic dynamic programming (THDP) method based on a disturbance observer, including the following steps:

[0007] S1. Construct a general nonlinear system model with disturbances, and the specific form is:

[0008] x k+1 =f(x k )+g u (x k )u k +g d (x k )dk ,

[0009] where x k ∈R n is the state variable, u k ∈R m is the control input, d k ∈R l is the disturbance, satisfying ξ1, ξ2 are known constants, and ξ1, ξ2 > 0. f(x k ) ∈ R n and g u (x k ) ∈ R n×m are known stationary functions. The disturbance gain of d k is g d (x k ) ∈ Rn×1. fxk is a Lipschitz continuous function on the set Ω ∈ Rn containing the origin.

[0010] And introduce Assumption 1: There exists a disturbance gain γ(x k ) ∈ R m×1 , such that g d (x k ) = g u (x k )γ(x k ).

[0011] S2. Construct the performance index function and design two parts of the control input. First, consider the case where the system d k = 0, and define the following performance index function as:

[0012]

[0013] The utility function is where Q and R are symmetric positive definite matrices. According to the Bellman optimal principle, the optimal cost function can be described as:

[0014]

[0015] The corresponding optimal control is:

[0016]

[0017] Using the optimal control the discrete-time HJB equation becomes:

[0018]

[0019] When the discrete-time HJB equation is linear in the system and the performance index is in quadratic form, it can be simplified to a solvable Riccati equation. However, for non-linear systems, solving the HJB equation is usually challenging. The present invention will solve the optimal control problem through the designed THDP algorithm.

[0020] S3. Construct the source task and the target task, and design a time-decreasing transfer function with a truncation mechanism. First, define the tasks involved in the present invention: Task T consists of a state space S, an action space C, a system function H(·), and a utility function N(·). It can be expressed as T = {S, C, H(·), N(·)}. The goal of the task is to minimize the cost function, that is, to minimize the sum of the utility functions.

[0021] The online control process includes training on current data and continuously updating the control strategy, which requires a learning period. In addition, since only one data point is learned each time, there is no guarantee that the control strategy will be updated in a positive direction. By utilizing transfer knowledge, the control strategy can be assisted in updating at the early stage of online control, thereby accelerating the learning speed and greatly reducing the number of samples required in the learning process.

[0022] To achieve transfer learning, the present invention defines the source task and the target task as T s = {S′, C′, H′(·), N(·)}, T t = {S, C, H(·), N(·)}. The evaluation strategy trained to maturity in T s is used to facilitate the learning process of T t . In T t , the knowledge initially imparted from T s plays a positive role in the learning of the optimal control strategy. However, due to certain differences in the systems between T s and T t , there may be a situation where the knowledge or experience in the source domain has a negative impact on the learning in the target domain, that is, the so-called negative transfer. In this case, the knowledge in the source domain may interfere with the learning task in the target domain, resulting in a performance decline. To prevent the above situation, the present invention designs a time-decreasing transfer function with a truncation mechanism:

[0023]

[0024] where a, b, c, and m are adjustable parameters, which determine the influence of the transferred knowledge on T t . Among them, a and b are positive numbers to ensure that the function is monotonically decreasing. It should be noted that the selection of a, b, and c should satisfy the condition of 0 < η k < 1. m is a positive number representing the truncation threshold. If η kIf it is less than m, it is assumed that the influence of transferred knowledge can be ignored. Therefore, the new cost function incorporating transferred knowledge can be expressed as:

[0025] V(x k ) = η k J tr (x k ) + (1 - η k )J d (x k ),

[0026] where J tr (x k ) is the evaluation strategy as the prior knowledge of T s , and J d (x k ) is the evaluation strategy of T t . Therefore, the optimal strategy u * (x k ) conforms to

[0027]

[0028] Transfer learning is a method of transferring knowledge from a set of source tasks to a target task. Therefore, the selection of source tasks directly affects the transfer effect. When there are significant differences between the source task and the target task, it will lead to serious negative transfer problems. The present invention selects a task with a certain similarity to T t as T s .

[0029] S4. Construct a disturbance observer to estimate the disturbance d k , and design observer-based disturbance compensation control. Since the disturbance d k is unknown, to estimate the disturbance d k , design a disturbance observer as:

[0030]

[0031] where is the intermediate variable of g d , p k ∈R w is the intermediate variable, is the estimated disturbance. The dynamic equation of the disturbance error can be obtained from the following formula:

[0032]

[0033] The following theorem can be used to estimate the upper bounds of and . Theorem 1: For the system If the gain matrix Z is a Schur matrix, then and The upper bounds of satisfy the following inequality:

[0034]

[0035]

[0036] Research has shown that the disturbance observer error and the disturbance estimate have upper bounds under certain conditions. Then, an observer-based disturbance compensation control was designed. Based on the disturbance observer, the disturbance compensation control The robust optimal control can be expressed as follows:

[0037]

[0038] Then, the system with the control input can be expressed as:

[0039]

[0040] Considering the input as u * (x k ) and the closed-loop system defined by , the state x k and the observation error both have asymptotic stability.

[0041] S5. Construct a neural network for the source task and design the weight update rate. For easy distinction, the symbols in T s are marked with a superscript '.

[0042] According to the definition of the cost function J(x k , u(x k ))), the calculation of the cost function of the current state J d (x k ) involves the control strategy and the future state. This cannot be achieved in real-time online control. Therefore, the present invention uses the critic network to obtain an approximation expressed as:

[0043]

[0044] where and are the estimated values of the ideal weights of the critic network. μ c represents the number of neurons in the hidden layer. The activation function is selected as ψ(q) = (e q - e -q ) / (e q + e -q ). The error of the critic network is defined as:

[0045]

[0046] The training objective of the critic network is to minimize the following objective function:

[0047]

[0048] By the gradient descent method, has the following update rule:

[0049]

[0050] where α c > 0 is the learning rate. By constructing the actor network, an approximate optimal control u′ a (x k ) can be obtained, which can be expressed as:

[0051]

[0052] where and are the weight vectors of the actor network. The error function is defined as:

[0053] ε′ a,k = u′ r (x k ) - u′(x k ),

[0054] where the target control vector u′(x k ) can be obtained by the following formula:

[0055]

[0056] The training objective of the actor network is to minimize the following objective function:

[0057]

[0058] Update by the gradient descent method

[0059]

[0060] where α a > 0 is the learning rate. After random initialization, only and will be updated, while and remain unchanged.

[0061] S6. Construct a neural network for the target task and design a weight update rate. During the execution of T t , the prior knowledge buffer is from T sTransfer it, and the critic network collaborates to evaluate u r (x k ). Therefore, the update rules of the critic networks of T s and T t are different. By constructing the critic network, the approximation can be expressed as:

[0062]

[0063] where, and are the estimated ideal weights of the critic network. The policy evaluation transferred from the source task can be expressed as:

[0064]

[0065] According to V(x k ) = η k J tr (x k )+(1 - η k )J d (x k ), then the new cost function of the target task is:

[0066]

[0067] The error of the critic network is defined as:

[0068]

[0069] By training the critic network, minimize the performance metric Then update rate of:

[0070]

[0071] The approximate optimal control policy u r (x k ) is expressed as:

[0072]

[0073] where, and are the weight vectors of the actor network. The error function is defined as:

[0074] ε a,k = u r (x k ) - u(x k ),

[0075] where, The training objective of the executor network is to minimize the following objective function Then The update rate of is:

[0076]

[0077] For the online execution process of the described THDP scheme, it can be proved that the weight error is uniformly ultimately bounded.

[0078] S7. Calculate the robust optimal control law and update the system state. The specific process is as follows: Calculate through the following formula

[0079]

[0080] For a nonlinear system with unknown disturbances, the control input is expected to be designed as:

[0081] u(x k ) = u r (x k ) + u d (x k ),

[0082] where, u r (x k ) is the optimal control strategy without considering disturbances, and u d (x k ) is the disturbance compensation part. Update x k , k → k + 1 through the following formula:

[0083] x k+1 = f(x k ) + g u (x k )u k + g d (x k )d k .

[0084] Furthermore, the characteristics of the uncertain nonlinear system under robust optimal control and the weight error of the neural network are ultimately uniformly bounded under certain given conditions are proved.

[0085] Furthermore, a computer-readable storage device, the storage device stores a computer program, characterized in that when the computer program is executed, it implements an optimal control method for a nonlinear time-varying time-delay system.

[0086] Compared with the prior art, the present invention has the following technical effects: for nonlinear discrete systems, an online transfer heuristic dynamic programming method based on a disturbance observer is proposed, which adopts a method combining transfer learning with adaptive critical control; sample data collected from source tasks are used to obtain prior knowledge, and a robust optimal control strategy is designed for the nonlinear discrete system, and used to guide the online control process of the nonlinear system of the target task; a new attenuation function with a truncation mechanism is proposed to avoid negative transfer effects and save computing resources; a disturbance compensation control mechanism is proposed to solve uncertainty; it is proved that the characteristics of uncertain nonlinear systems under robust optimal control and the weight errors of neural networks are ultimately uniformly bounded under certain given conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Figure 1 The present invention is a flow chart of an online transfer heuristic dynamic programming method based on a disturbance observer.

[0088] Figure 2 This is a graph showing weight changes during the training process of Example 1 of the present invention.

[0089] Figure 3 This is a disturbance observer curve diagram of Example 1 of the present invention.

[0090] Figure 4 This is a closed-loop system state and control input curve diagram of Example 1 of the present invention.

[0091] Figure 5 This is a graph showing the changes in weights during the training process of Example 2 of the present invention.

[0092] Figure 6 This is a disturbance observer curve diagram of Example 2 of the present invention.

[0093] Figure 7 This is a closed-loop system state and control input curve diagram of Example 2 of the present invention. DETAILED DESCRIPTION

[0094] To further illustrate the features of the present invention, please refer to the following detailed description of the present invention and the attached Figure 1-7 The accompanying drawings are for reference and illustration purposes only and are not intended to limit the scope of protection of the present invention.

[0095] The specific embodiments of the present invention are described in detail below in conjunction with the accompanying drawings:

[0096] like Figure 1 As shown, this embodiment discloses an online transfer heuristic dynamic programming method based on a disturbance observer, comprising the following steps:

[0097] S1. In aerospace, missile, satellite, and industrial control systems, construct a general nonlinear system model with disturbances and initialize the parameters. The specific form is:

[0098] x k+1 =f(x k )+g u (x k )u k +g d (x k )d k ,

[0099] where x k ∈R n is the state variable, u k ∈R m is the control input, d k ∈R l is the disturbance, satisfying f(x k )∈R n and g u (x k )∈R n ×m are known stationary functions. The disturbance gain of d k is g d (x k )∈R n×1 . f(x k ) is a Lipschitz continuous function containing the origin on the set Ω∈R n .

[0100] Initialize the following parameters Q, R, a, b, c, m, x0, x′0, α a 、α c 、the maximum number of sampling times k max and all network weights. Among them, Q and R are symmetric positive definite matrices in the utility function; a, b, c, and m are adjustable parameters in the time-decreasing transfer function, which determine the influence of the transferred knowledge on T t ; α a 、α c are the learning rates of the neural network.

[0101] S2. Construct the performance index function and design two parts of the control input. First, consider the case where the system d k =0, and define the following performance index function as:

[0102]

[0103] The utility function is where Q and R are symmetric positive definite matrices. According to the Bellman optimal principle, the optimal cost function can be described as:

[0104]

[0105] The corresponding optimal control is:

[0106]

[0107] Using the optimal control The discrete-time HJB equation becomes:

[0108]

[0109] The discrete-time HJB equation can be simplified to a solvable Riccati equation when the system is linear and the performance index is in quadratic form. However, for non-linear systems, solving the HJB equation is usually challenging. The present invention will use the THDP algorithm to solve the optimal control problem.

[0110] S3. Construct the source task and the target task, and design a time-decreasing transfer function with a truncation mechanism. First, define the tasks involved in the present invention: Task T consists of a state space S, an action space C, a system function H(·), and a utility function N(·). It can be expressed as T = {S, C, H(·), N(·)}. The goal of the task is to minimize the cost function, that is, to minimize the sum of the utility functions.

[0111] The online control process includes training on current data and continuously updating the control strategy, which requires a learning period. In addition, since only one data point is learned each time, it cannot be guaranteed that the control strategy will be updated in a positive direction. By utilizing the transfer knowledge, it is possible to assist in updating the control strategy in the early stage of online control, thereby accelerating the learning speed and significantly reducing the number of samples required in the learning process.

[0112] Adopt a time-decreasing transfer function with a truncation mechanism:

[0113]

[0114] where a, b, c, and m are adjustable parameters that determine the influence of the transferred knowledge on T t . Among them, a and b are positive numbers to ensure that the function is monotonically decreasing. It should be noted that the choices of a, b, and c should satisfy the condition 0 < η k < 1. m is a positive number representing the truncation threshold. If η k is less than m, it is assumed that the influence of the transferred knowledge can be ignored. Therefore, the new cost function including the transferred knowledge can be expressed as:

[0115] V(x k ) = η k Jtr (x k )+(1 - η k )J d (x k ),

[0116] where J tr (x k ) is the evaluation strategy of T s as prior knowledge, and J d (x k ) is the evaluation strategy of T t . Therefore, the optimal strategy u * (x k ) satisfies.

[0117]

[0118] The present invention selects a task with a certain similarity to T t as T s .

[0119] S4. Construct a disturbance observer to estimate the disturbance d k , and design observer - based disturbance compensation control. Since the disturbance d k is unknown, to estimate the disturbance d k , design the disturbance observer as:

[0120]

[0121] where is the intermediate variable of g d , p k ∈R w is the intermediate variable, and is the estimated disturbance. The dynamic equation of the disturbance error can be obtained from the following formula:

[0122]

[0123] The disturbance observer error and the disturbance estimated value have upper bounds under certain conditions. Then, observer - based disturbance compensation control is designed. According to the disturbance observer, design the disturbance compensation control The robust optimal control can be expressed as follows:

[0124]

[0125] Then, the system with the control input can be expressed as:

[0126]

[0127] Considering the input is u * (x k ) and the closed-loop system defined by both the state x k and the observation error are asymptotically stable.

[0128] S5. Train the neural network for the source task. The specific process is as follows: Let the number of sampling times k = 1, and train the neural network constructed with the initial parameters in S1 according to the following formula. First, calculate

[0129]

[0130] where and are the estimated values of the ideal weights of the critic network. μ c represents the number of neurons in the hidden layer. The activation function is selected as ψ(q)=(e q -e -q ) / (e q +e -q ). Update The specific process is as follows: Define the error of the critic network as:

[0131]

[0132] The training objective of the critic network is to minimize the following objective function:

[0133]

[0134] Through the gradient descent method, the update rule of

[0135]

[0136] is as follows: c where α > 0 is the learning rate.

[0137] Calculate u′ r (x k ):An approximate optimal control u′ a (x k ) can be obtained by constructing the actor network, which can be expressed as:

[0138]

[0139] where and are the weight vectors of the actor network. Update The error function is defined as:

[0140] ε′a,k = u' r (x k ) - u'(x k ),

[0141] where the target control vector u'(x k ) can be obtained by the following formula:

[0142]

[0143] The training objective of the executor network is to minimize the following objective function:

[0144]

[0145] Update by gradient descent

[0146]

[0147] where α a > 0 is the learning rate. After random initialization, only and will be updated, while and remain unchanged. After is updated once, let k = k + 1, and stop updating when k ≥ k max .

[0148] S6. Train the target task neural network. The specific process is as follows: Let k = 1. During the execution of T t , the prior knowledge buffer is transferred from T s . The critic network collaborates to evaluate u r (x k ). Therefore, the update rules of the critic networks of T s and T t are different. By constructing the critic network, the approximation value is calculated as:

[0149]

[0150] where and are the estimated ideal weights of the critic network. The policy evaluation transferred from the source task can be expressed as:

[0151]

[0152] And based on the approximation value calculate

[0153]

[0154] Update The specific method is as follows. First, the error of the critic network is defined as:

[0155]

[0156] By training the critic network, the following performance index is minimized to the greatest extent:

[0157]

[0158] The update rate is:

[0159]

[0160] The approximate optimal control strategy u r (x k ) is calculated by the following formula:

[0161]

[0162] where and are the weight vectors of the actor network. Update The specific method is as follows. First, the error function of the actor network is defined as:

[0163] ε a,k = u r (x k ) - u(x k ),

[0164] where The training objective of the actor network is to minimize the following objective function:

[0165]

[0166] The update rate is obtained by the gradient descent method as:

[0167]

[0168] S7. Calculate the robust optimal control law and update the system state. The specific process is: Calculate by the following formula

[0169]

[0170] For a nonlinear system with unknown disturbances, the control input is expected to be designed as:

[0171] u(x k ) = u r (x k ) + ud (x k ),

[0172] where u r (x k ) is the optimal control strategy without considering interference, and u d (x k ) is the interference compensation part. Update x k , k→k + 1 through the following formula:

[0173] x k+1 = f(x k ) + g u (x k )u k + g d (x k )d k .

[0174] Through the above steps, precise control and adaptive adjustment can be achieved in the missile guidance system, thereby improving the hit accuracy and combat effectiveness of the missile. This method can adapt to various missile types and combat scenarios and provides new ideas and methods for the further development of missile guidance technology.

[0175] Furthermore, the characteristics of the uncertain nonlinear system under robust optimal control and the weight error of the neural network are proven to be ultimately uniformly bounded under certain given conditions. The specific process is as follows:

[0176] First, introduce Assumption 2 and Lemma 2. Assumption 2: The ideal weights of the executor network and the critic network (denoted by and respectively) exist and are bounded, that is where and are both positive constants.

[0177] Lemma 2: Define When updating the weight of the critic network, the difference ΔΛ c,k = Λ c,k+1 - Λ c,k satisfies the following inequality:

[0178]

[0179] where Considering the closed-loop system x k+1 = f(x k ) + g u (x k )u k + g d (x k )d k, when Hypotheses 1 and 2 hold, the weights are updated according to the derived update rate. Then the state x k and the weight error are uniformly ultimately bounded. Consider the Lyapunov function as:

[0180] Λ = Λ1 + Λ2 + Λ3 + Λ4 + Λ5 + Λ6

[0181] where

[0182] the difference of Λ1 satisfies:

[0183]

[0184] the difference of Λ2 is:

[0185]

[0186] the difference of Λ3 is:

[0187]

[0188] According to Lemma 2, the difference of Λ4 satisfies the following inequality:

[0189]

[0190] According to it can be obtained that

[0191]

[0192] where ψ′ c,k is the derivative of ψ c,k We denote where denotes the Hadamard product. Due to the boundedness of the activation function, the inequality holds, where ψ′ m is the upper bound of ψ′ c,k .

[0193] the difference of Λ5 is:

[0194]

[0195] According to the Cauchy - Schwarz inequality, the following inequality holds:

[0196]

[0197] Then there is:

[0198]

[0199] Define χ 2 as:

[0200]

[0201] where ψ cm 、 ψ am 、U m respectively represent ψ c 、 ψ a 、U k upper limits. Select the learning rate and hyperparameters as l2>3l1||g d || 2 . Then for any there is:

[0202]

[0203] Since it can be guaranteed that ΔΛ < 0, according to the above formula, it can be seen that both the state and the weight error are uniformly ultimately bounded. Q.E.D.

[0204] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.

[0205] The following is a simulation example. The present invention will give two examples to illustrate the effect of the proposed robust optimization control scheme. Both of these cases show the advantages of the proposed THDP algorithm.

[0206] Consider a nonlinear system whose disturbance converges to a constant over time:

[0207]

[0208] where the disturbance d k = 0.5 arctan(0.04k + 1). To implement the proposed THDP algorithm, the system of Ts is selected as:

[0209] As shown in the specific implementation manner, first learn the prior knowledge from T s . Train the critic network of the nonlinear system given in Example 1 to obtain the approximate optimal critic weights. Subsequently, the obtained approximate optimal critic weights are used as prior knowledge to guide T tThe online control process. For the target task, the fixed matrices in the utility function are selected as follows: Q = [1, 0; 0, 1, 0], R = 0.1. The critic network and the actor network use a three-layer backpropagation neural network with a structure of 2-8-1. The state variable x k is used as the input to both networks. The critic network generates the cost function, while the actor network generates the control input. The learning rates of the critic network and the actor network are set to α a = α c = 0.05. To reduce negative transfer and save computational resources, the parameters of the truncation decay functions a, b, c, m are set to 5, 0.08, 0.24, 0.45 respectively. Figure 2 The weight convergence trajectories of the critic network and the actor network during the training process are given.

[0210] The disturbance observer gain and the disturbance compensation control gain are set to L i = 0.3. The observation results and the observation errors of the disturbance observer are as Figure 3 shown, indicating that the designed disturbance observer has good performance. The disturbance compensation control is designed according to the observed disturbance.

[0211] Combined with the robust optimal control strategy can be obtained. Figure 4 shows the state trajectory controlled by the robust optimal control, x0 = [2, 1] T .

[0212] The second example considers the following nonlinear system with oscillatory disturbances:

[0213] x k+1 = f(x k ) + g u (x k )u k + g d (x k )d k

[0214] where, d k = 0.5sin(0.04k + 1). We set the initial state x0 = [0.5, 1] T . At the same time, to implement THDP, the following system is selected as the source task system:

[0215]

[0216] The parameters of the utility function are selected as follows Q = [0.1, 0; 0, 0.1], R = 0.1. The structures of the two networks are the same as in Example 1. The weights of the critic network obtained through training will be used as prior knowledge to guide the policy evaluation of the target task. Among them, α a= α c = 0.05, and the parameters of a, b, c, and m are set to 3, 0.01, 0.37, and 0.3 respectively. Figure 5 It shows the weight convergence trajectories of the critic network and the actor network during the training process. Obviously, the proposed THDP algorithm effectively accelerates the learning speed. The structures of these two networks are the same as those in Example 1. The weights of the critic network obtained through training are used as prior knowledge to guide the policy evaluation of the target task. Figure 5 It shows the weight convergence trajectories of the critic network and the actor network during the training process.

[0217] To handle unknown disturbances, the disturbance observer gain and the disturbance compensation control gain are set to L i = 0.1 to obtain satisfactory estimation performance. From Figure 6 it can be seen that the designed disturbance observer has a relatively accurate observation effect. According to the robust optimal control composed of the policy generated by the actor network and the disturbance compensation part, the system state and input curves are as Figure 7 shown.

[0218] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An online transfer heuristic dynamic programming method based on a disturbance observer, characterized in that: It includes the following steps: S1. Construct a general non - linear system model with perturbations; S2. Construct a performance index function and design two parts of control inputs; S3. Construct a source task and a target task, and design a time - decreasing transfer function with a truncation mechanism; S4. Construct a disturbance observer to estimate the disturbance d k and design an observer-based disturbance compensation control; S5. Construct a neural network for the source task and design a weight update rate; S6. Construct a neural network for the target task and design a weight update rate; S7. Calculate the robust optimal control law and update the system state; Among them, in step S3, the task T consists of a state space S, an action space C, a system function H(·) and a utility function N(·); expressed as T = {S, C, H(·), N(·)}; the goal of the task is to minimize the cost function, that is, to minimize the sum of the utility functions; Define the source task and the target task as T s ={S′, C′, H′(·), N(·)}, T t ={S, C, H(·), N(·)}; In T s the trained and mature evaluation strategy is used to facilitate the learning process of T t ; Let the time - decreasing transfer function with a truncation mechanism be: where a, b, c, and m are adjustable parameters that determine the impact of the transferred knowledge on T t ; where a and b are positive numbers to ensure that the function is monotonically decreasing; the choices of a, b, and c should satisfy the condition 0 < η k < 1; m is a positive number representing the truncation threshold; if η k is less than m, the impact of the transferred knowledge is considered negligible; thus, the new cost function incorporating the transferred knowledge is expressed as: V(x k ) = η k J tr (x k ) + (1 - η k )J d (x k ), Among them, J tr (x k ) is an evaluation strategy for the T s prior knowledge, and J d (x k ) is the evaluation strategy for T t ; therefore, the optimal strategy u * (x k ) conforms to it; Among them, in step S4, to estimate the disturbance d k , a disturbance observer is designed as follows: Among them, |z i | < 1, i ∈ [1, w]; is the intermediate variable g d , and p k ∈R w is the intermediate variable, is the estimated disturbance; the disturbance error is obtained as the dynamic equation of: For the system If the gain matrix Z is a Schur matrix, then and The upper bound of satisfies the following inequality: Design disturbance compensation control The robust optimal control is expressed as follows: The system with control input is expressed as: Considering the input is u * (x k ) and the closed-loop system defined by both the state x k and the observation error are stable.

2. The online transfer heuristic dynamic programming method based on a disturbance observer according to claim 1, characterized in that: In step S1, the specific form is: x k+1 = f(x k ) + g u (x k )u k + g d (x k )d k , where x k ∈ R n is the state variable, u k ∈ R m is the control input, d k ∈ R l is the disturbance, satisfying ||d k+1 - d k || ≤ ξ2}; ξ1, ξ2 are known constants, and ξ1, ξ2 > 0; f(x k ) ∈ R n and g u (x k ) ∈ R n×m are known smooth functions; the disturbance gain of d k is g d (x k ) ∈ R n×1 ; f(x k ) is a Lipschitz continuous function on the set Ω ∈ R n containing the origin; Suppose there exists a perturbation gain γ(x k ) ∈ R m×1 , such that g d (x k ) = g u (x k )γ(x k ).

3. An online transfer heuristic dynamic programming method based on a disturbance observer according to claim 1 or 2, characterized in that: In step S2, consider the case where the system d k = 0, and define the following performance metric function as: The utility function is where Q and R are symmetric positive definite matrices; according to the Bellman optimality principle, the optimal cost function is described as: The corresponding optimal control is: Using optimal control The discrete-time HJB equation becomes:

4. An online transfer heuristic dynamic programming method based on a disturbance observer according to claim 1, characterized in that: In step S5, an approximate value is obtained using the critic network Expressed as: Among them, and are the estimated values of the ideal weights of the judge network; μ c represents the number of neurons in the hidden layer; the activation function is selected as ψ(q) = (e q -e -q ) / (e q +e -q ); the error of the judge network is defined as:

5. An online transfer heuristic dynamic programming method based on a disturbance observer according to claim 4, characterized in that: The training objective of the critic network is to minimize the following objective function: By the gradient descent method, The update rule is as follows: where α c > 0 is the learning rate; an approximate optimal control u′ a (x k ) is obtained by constructing an executor network, expressed as: Among them, and are the executor network weight vectors; the error function is defined as: ε′ a,k = u′ r (x k ) - u′(x k ), wherein, the target control vector u′(x k ) is obtained by the following formula: The training objective of the actor network is to minimize the following objective function: Update by gradient descent method Among them, α a > 0 is the learning rate. After random initialization, only and will be updated, while and remain unchanged.

6. A method for online transfer heuristic dynamic programming based on a disturbance observer according to claim 1 or 4 or 5, characterized in that: In step S6, by constructing a critic network, the approximation is expressed as: Among them, and are the estimated ideal weights of the critic network; the policy evaluation transferred from the source task is expressed as: According to V(x k ) = η k J tr (x k )+(1 - η k )J d (x k ), the new cost function of the target task is:

7. An online transfer heuristic dynamic programming method based on a disturbance observer according to claim 6, characterized in that: The error of the critic network is defined as: Minimize the performance metric by training a critic network then update rate of: Approximate optimal control strategy u r (x k ) is expressed as: Among them, and are the weight vectors of the executor network; the error function is defined as: ε a,k = u r (x k ) - u(x k ), Among them, The training objective of the executor network is to minimize the following objective function Then The update rate of is:

8. An online transfer heuristic dynamic programming method based on a disturbance observer according to claim 1, characterized in that: In step S7, calculate using the following formula For a non - linear system with unknown perturbations, the control input is expected to be designed as: u(x k ) = u r (x k ) + u d (x k ), where, u r (x k ) is the optimal control strategy without considering interference, and u d (x k ) is the interference compensation part; update x k , k → k + 1 through the following formula: x k+1 = f(x k ) + g u (x k )u k + g d (x k )d k 。

Citation Information

Patent Citations

  • Controllers, observers, and applications thereof

    CN102354104A

  • Online converter parameter identification method based on observer

    CN106096298A