An optimal tracking control method for time-delay systems based on inverse reinforcement Q-learning

By reconstructing performance indicators and learning control strategies in time-delay systems, combining Smith predictor and inverse reinforcement Q-learning, the problem of unknown performance indicators in time-delay systems is solved, and efficient optimal tracking control effects are achieved. It is suitable for linear discrete-time systems with unmeasurable states.

CN119960306BActive Publication Date: 2025-09-23THE 54TH RESEARCH INSTITUTE OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510096883.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-09-23
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Existing technologies have difficulty effectively solving the problem of unknown performance indicators in time-delay systems, resulting in poor control effects. In addition, traditional inverse reinforcement learning methods are computationally complex and inefficient, and fail to effectively deal with the problems of unmeasurable state data and control input time delays.

Method used

By observing the input and output data of the target system, reconstructing the performance indicators and learning the control strategy, a Smith predictor is designed to predict the state at future moments. Combined with the inverse reinforcement Q-learning method, only using historical output and control input data, a single-loop iterative output feedback method is proposed to solve the optimal tracking control problem of time-delay systems.

Benefits of technology

In the presence of time delay, it can effectively reconstruct performance indicators and learn optimal control strategies, realize trajectory tracking of the controlled system and the target system, improve the stability and control performance of the control system, reduce the number of iterations, and improve learning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960306B_ABST
    Figure CN119960306B_ABST
Patent Text Reader

Abstract

The present invention discloses an optimal tracking control method for a time-delay system based on inverse reinforcement Q learning, belonging to the technical field of optimal tracking control. The method first sets performance indicator input weights, initial performance indicator state weights, and a threshold, and collects trajectory data of a target system. The method then obtains control strategy parameters by collecting the target system's elevator actuator voltage and the aircraft's offset. The method then obtains the control strategy by collecting the controlled system's elevator actuator voltage, the aircraft's offset, and control strategy parameters. The method then obtains performance indicator state weights by collecting the controlled system's elevator actuator voltage, the aircraft's offset, and control strategy parameters. Finally, the iteration termination condition is determined to obtain the controlled system's performance indicator state weights and the optimal control strategy. The method enables the controlled system and the target system to have the same trajectory, thereby tracking the target system's trajectory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optimal tracking control of linear discrete-time systems, and in particular to an optimal tracking control method for a time-delay system based on inverse reinforcement Q learning. Background Art

[0002] Optimal tracking control is an important research area in modern control theory. Its goal is to find the optimal control strategy for a dynamic system by minimizing predefined performance indicators, while ensuring the stability of the closed-loop system and ultimately tracking the target trajectory. Reinforcement Q-learning methods solve the optimal tracking control problem for unknown system models by solving the optimal control strategy. However, these methods require fixed or manually pre-given performance indicators of the controlled system, lacking systematic theoretical guidance and scientific calculation rules, and often have a certain degree of randomness. In actual applications, they are often affected by various factors, such as human interference, equipment aging, and bad weather. It is difficult to manually give appropriate performance indicators based on experience, which makes it impossible to apply them to optimal tracking control methods. If the performance indicators are not given appropriately, satisfactory control effects may not be achieved.

[0003] Inverse reinforcement learning achieves optimal tracking control by reconstructing performance indicators. It typically considers two homogeneous systems: a target system and a controlled system. The target system already has a satisfactory trajectory under the optimal control policy, known as the target trajectory. To track the target trajectory, the controlled system reconstructs performance indicators by observing the target trajectory and simultaneously learns the control policy. Traditional inverse reinforcement learning methods do not consider the stability of the closed-loop system, which is essential in the control field. Inverse reinforcement learning is also known as inverse optimal control in the field of machinery, but the two differ in structure. While inverse optimal control aims to reconstruct performance indicators based on the stable state or trajectory of the controlled system, inverse reinforcement learning aims to reconstruct performance indicators based on the target trajectory of the target system. Therefore, combining inverse reinforcement learning and inverse optimal control in the control field can both learn unknown performance indicators and find the optimal control policy. However, the learning process is a double-loop iterative structure, where the control strategy calculation is nested within the performance indicator calculation. This results in computational complexity and low efficiency. Combining this with reinforced Q-learning methods is an important approach to solving optimal tracking control problems. However, existing technologies haven't addressed the issues of unmeasurable state data and control input time delays. In various time-delay systems, the pure delay of the controlled object not only reduces system stability and degrades transient characteristics, but also affects the overall system dynamics, significantly impacting the control performance of the control system. In the field of inverse reinforcement learning, time-delay systems have yet to be studied. Summary of the Invention

[0004] In light of this, the present invention provides an optimal tracking control method for time-delay systems based on inverse reinforcement Q-learning. This method uses only historical output and control input trajectories to reconstruct a performance indicator. The optimal control strategy corresponding to this performance indicator enables the controlled system to track the target trajectory.

[0005] The present invention is achieved through the following technical solutions:

[0006] An optimal tracking control method for a time-delay system based on inverse reinforcement Q-learning is disclosed. The method determines the performance indicator weights of a controlled system by observing the input and output data generated by the target system and learns the control strategy of the target system. The state variables x(k) of the target system are defined as the angle of attack, pitch rate, and elevator angle. The control input u(k) is defined as the elevator actuator voltage that controls the elevator change. The system output y(k) is defined as the aircraft's offset relative to the flight directions x, y, and z. The method comprises the following steps:

[0007] S1, set the number of iterations i = 0, give a fixed performance indicator input weight R> 0, and select the initial performance indicator state weight and a positive constant ε as a threshold, and collect the trajectory data y of the target system d (k) and u d (k);

[0008] S2, by collecting the elevator actuator voltage of the target system and the aircraft's offset relative to the flight direction x, y, z, the control strategy parameters are solved by formula (40)

[0009]

[0010] Among them, w d (k) is the actual control input u of the target system d (k) and output y d (k), which is the matrix composed of the elevator actuator voltage that controls the elevator change and the aircraft's offset relative to the flight direction x, y, and z; z d (k) is the actual control input u of the target system d (k) and output y d (k) is a matrix; (k), (kt), and (k-t+1) represent the data collected at different times; the performance index input weight R is fixed;

[0011] S3, by collecting the voltage of the elevator actuator that controls the elevator change of the controlled system, the offset of the aircraft relative to the flight direction x, y, z, and the control strategy parameters solved in step S2 The control strategy is solved by equation (30)

[0012]

[0013] Then u(kt) is calculated by equation (32);

[0014]

[0015] Where u(k) is the actual control input of the controlled system, that is, the elevator actuator voltage that controls the elevator change; z(k) is the matrix formed by the actual control input u(k) and output y(t) of the controlled system.

[0016] S4, by collecting the voltage of the elevator actuator that controls the elevator change of the controlled system, the offset of the aircraft relative to the flight direction, and the control strategy parameters obtained in step S2 The performance index state weight is solved by formula (43):

[0017]

[0018] Among them, the inverse reinforcement learning learning rate α∈[0,1] is artificially given, w(k) is the matrix formed by actually collecting the control input u(k) and output y(k) of the controlled system, The control strategy parameters Obtained by subtracting the performance indicator state weight and the input weight matrix, the performance indicator input weight R is fixed;

[0019] S5, determine the end condition of the iteration, if Then stop the iteration and let Obtain the performance indicator state weights and optimal control strategy of the controlled system; otherwise, set i=i+1 and return to step S2 for the next iteration.

[0020] The beneficial effects of the present invention are:

[0021] 1. The present invention can solve the problem of a controlled system tracking the behavior trajectory of a target system with a demonstration effect but unknown performance indicators when there is a time delay between the controlled system and the target system.

[0022] 2. The present invention solves the problem of unknown performance indicators or inappropriate manual pre-setting in reinforcement learning.

[0023] 3. The present invention does not require system model information and can reconstruct appropriate performance indicators and learn optimal control strategies using only measured historical output and control input data.

[0024] 4. This invention proposes for the first time the design idea and method of an output feedback controller for a time-delay system with unmeasurable state, and ensures a satisfactory tracking control effect with a more efficient single-loop iterative method. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a flow chart of an optimal tracking control method for a time-delay system based on inverse reinforcement Q learning in an embodiment of the present invention;

[0026] Figure 2 It is a schematic diagram of the aircraft flight;

[0027] Figure 3 In the learning process of the present invention, there is no delay. and The convergence curve diagram of

[0028] Figure 4 In the learning process of the present invention, there is no delay. and and and Error curve diagram;

[0029] Figure 5 In the case of time delay, the learning process of the present invention method and The convergence curve diagram of

[0030] Figure 6 In the case of time delay, the learning process of the present invention method and and and Error curve diagram;

[0031] Figure 7 This is the control input trajectory and output tracking performance curve of the controlled system of comparison method 1;

[0032] Figure 8 It is a graph showing the control input trajectory and output tracking performance of the controlled system according to the method of the present invention;

[0033] Figure 9 This is the flow chart of the traditional inverse reinforcement learning double-loop iterative method;

[0034] Figure 10 Compared with method 2, and Error curve diagram;

[0035] Figure 11 The present invention is a method of learning process and Error curve diagram. DETAILED DESCRIPTION

[0036] In order to further illustrate the solution and effects of the present invention, the present invention is described in detail below with reference to the accompanying drawings and simulation examples, but it should not be understood as limiting the scope of protection of the present invention.

[0037] An optimal tracking control method for time-delay systems based on inverse reinforcement Q-learning determines the performance indicator weights of the controlled system by observing the input and output data generated by the target system and learns the control strategy of the target system.

[0038] When an aircraft is flying steadily over a short period of time, the primary considerations are the impact of the angle of attack, pitch rate, and elevator angle on the aircraft's flight attitude. The angle of attack refers to the angle between the aircraft's velocity vector and the wing chord line, the pitch rate refers to the angular velocity of the aircraft's rotation about its lateral axis, and the elevator angle refers to the angle between the aircraft's elevator and horizontal stabilizer. The angle of attack and pitch rate directly measure changes in the aircraft's attitude. Changes in the elevator angle cause changes in the aerodynamic torque acting on the elevator, thereby affecting the aircraft's flight attitude. Therefore, this method defines the system's state variables x(k) as the angle of attack, pitch rate, and elevator angle, the control input u(k) as the elevator actuator voltage that controls the elevator's movement, and the system's output y(k) as the aircraft's offset relative to the flight directions x, y, and z.

[0039] The purpose of this method is to demonstrate the target trajectory of a target system with demonstration effect but unknown performance indicators through the target system. At the same time, a controlled system attempts to construct performance indicators by observing the target trajectory and complete the tracking of the target system trajectory.

[0040] like Figure 1 As shown, the method includes the following steps:

[0041] S1, set the number of iterations i = 0, give a fixed performance indicator input weight R> 0, and select the initial performance indicator state weight and a positive constant ε as a threshold, and collect the trajectory data y of the target system d (k) and u d (k);

[0042] S2, by collecting the elevator actuator voltage of the target system and the aircraft's offset relative to the flight direction x, y, z, the control strategy parameters are solved by formula (40)

[0043]

[0044] Among them, w d (k) is the actual control input u of the target system d(k) and output y d (k), which is the matrix composed of the elevator actuator voltage that controls the elevator change and the aircraft's offset relative to the flight direction x, y, and z; z d (k) is the actual control input u of the target system d (k) and output y d (k) is a matrix; (k), (kt), and (k-t+1) represent the data collected at different times; the performance index input weight R is fixed;

[0045] S3, by collecting the elevator actuator voltage of the controlled system, the aircraft's offset relative to the flight direction x, y, z, and the control strategy parameters obtained in step S2 The control strategy is solved by equation (30)

[0046]

[0047] Then u(kt) is calculated by equation (32);

[0048]

[0049] Where u(k) is the actual control input of the controlled system, that is, the elevator actuator voltage that controls the elevator change; z(k) is the matrix formed by the actual control input u(k) and output y(k) of the controlled system.

[0050] S4, by collecting the voltage of the elevator actuator that controls the elevator change of the controlled system, the offset of the aircraft relative to the flight direction, and the control strategy parameters obtained in step S2 The performance index state weight is solved by formula (43):

[0051]

[0052] Among them, the inverse reinforcement learning learning rate α∈[0,1] is artificially given, w(k) is the matrix formed by actually collecting the control input u(k) and output y(k) of the controlled system, The control strategy parameters Obtained by subtracting the performance indicator state weight and the input weight matrix, the performance indicator input weight R is fixed;

[0053] S5, determine the end condition of the iteration, if Then stop the iteration and let Obtain the performance indicator state weights and optimal control strategy of the controlled system; otherwise, set i=i+1 and return to step S2 for the next iteration.

[0054] This method has the following characteristics:

[0055] (1) A controlled system with time delay in control input is described, a Smith predictor is designed to predict the state at future moments, and the optimal tracking control problem of inverse reinforcement learning for time-delayed systems is proposed.

[0056] (2) The conditions for the solution of the optimal tracking control problem of inverse reinforcement learning were designed. Two correction equations for the control strategy parameters and performance index weights were proposed using the system model information, and the optimal control strategy was solved.

[0057] (3) Combining the idea of ​​reinforced Q learning, a data-driven inverse reinforcement Q learning method for time-delay systems is proposed. This method is an output feedback method that only uses historical output and control input data to simultaneously reconstruct performance indicators and learn control strategies.

[0058] The principle of this method is as follows:

[0059] The state space equation of a linear discrete-time controlled system with control input delay is as follows:

[0060]

[0061] in, and are the state, output and control input of the controlled system respectively, and t≥0 is the time delay unit. and is the system model information, assuming that the matrix (AB) is controllable and the matrix (AC) is observable.

[0062] The desired target trajectory of the controlled system (1) is generated by the following target system:

[0063]

[0064] in, and are the target system state, output and control input respectively, and t≥0 is the time delay unit. The controlled system (1) and the target system (2) are homogeneous systems, and both have the same system model information. The target system (2) has been d The optimal control effect is achieved under the action of .

[0065] The performance indicators that the target system can minimize are defined as:

[0066]

[0067] Among them, the target performance indicator state weight and the target performance index input weight Rd >0 is a given matrix, the target control strategy parameter Satisfies the Bellman equation:

[0068]

[0069] Or satisfy the algebraic Riccati equation:

[0070] 0=Q d -P d +A T P d AA T P d B(R d +B T P d B) -1 B T P d A (5)

[0071] Therefore, the target control input u can be obtained d (kt) and target control strategy K d for:

[0072] u d (kt)=-K d x d (k)=-(R d +B T P d B) -1 B T P d Ax d (k) (6)

[0073] Among them, the target control strategy K d =(R d +B T P d B) -1 B T P d A.

[0074] But when there is a time delay, that is, t≠0, the control input u at the current time k is calculated by formula (6): d (kt) is meaningless. Therefore, this method uses the concept of prediction to rewrite Equation (6) as:

[0075] u d (k)=-K d x d (k+t) (7)

[0076] Then, a Smith predictor is designed to predict the state at future moments. The target system state can be rewritten using historical output and control input data as follows:

[0077] x d (k+t)=Lz d (k) (8)

[0078] Among them, the prediction matrix L is:

[0079]

[0080] in,

[0081]

[0082] N=[B AB…A n+t+q-1 ]

[0083] O=[(CA n+q-1 ) T …(CA) T C T ]

[0084]

[0085] Substituting equation (8) into equation (7), we can get the new target control input:

[0086]

[0087] or

[0088]

[0089] in, It is the target control strategy of the time-delay system. The above solves the problem of using historical output and control input data to predict the state x at the future moment. d (k+t) problem.

[0090] The weight Q in the target system performance index (3) d and R d 、P d And the control strategy in formula (11) is unknown, and the system model information A, B and C are also unknown. Only the input and output data generated by the target system (2) are measurable.

[0091] The optimal tracking control problem of inverse reinforcement learning studied in this method is: by observing the input u generated by the target system d (kt) and output y d (k) data to determine the performance indicators of the controlled system and learn the control strategy so that it can have the same trajectory as the target system.

[0092] According to the optimal control theory, the performance index of the controlled system (1) is defined as:

[0093]

[0094] Among them, the performance index weight Q = Q T ≥0 and R>0 are any given matrices, and the control strategy parameter P=P T >0 corresponding to the optimal control input u(kt) * and the optimal control strategy K * They are:

[0095] u(kt) * =-K * x(k)=-(R+B T PB) -1 B T PAx(k) (14)

[0096] Among them, the optimal control strategy K * =(R+B T PB) -1 B T PA satisfies the Bellman equation:

[0097] x(k) T Px(k)=x(k) T Qx(k)+u(kt) T Ru(kt)+x(k+1) T Px(k+1) (15)

[0098] Or satisfy the algebraic Riccati equation:

[0099] 0=Q-P+A T PA-A T PB(R+B T PB) -1 B T PA (16)

[0100] Similar to the problem encountered in step (1), when there is a time delay, that is, when t≠0, it is meaningless to calculate the control input u(kt) at the current time k using equation (14). Therefore, this method uses the concept of prediction to rewrite equation (14) as:

[0101] u(k) * =-K * x(k+t) (17)

[0102] Then, a Smith predictor is designed to predict the state at future time. The system state can be rewritten using historical output and control input data as:

[0103] x(k+t)=Lz(k) (18)

[0104] Among them, the prediction matrix L is shown in formula (9), and z(k) is:

[0105]

[0106] Since the controlled system (1) is controllable and observable, the prediction matrix L is full rank. Substituting Equation (18) into Equation (17), the optimal control input of the time-delay system can be obtained as:

[0107]

[0108] or

[0109]

[0110] in, This is the optimal control strategy for the time-delay system. The above solves the problem of using historical output and control input data to predict the state x(k+t) at the future moment.

[0111] In order to solve the optimal tracking control problem of inverse reinforcement learning in step (1), this method gives R>0, and then needs to find a pair of weights Q equivalent to the target performance index d and R d Q and R, further solve the control strategy parameter P, and then we can make Equation (20) Equal to formula (11)

[0112] This method gives the conditions for the solution of the optimal tracking control problem of inverse reinforcement learning: if the control strategy parameters P simultaneously satisfy the algebraic Riccati equation (16) and the following parameter correction equation:

[0113]

[0114] Then formula (20) Equal to formula (11) Rewrite Equation (22) into iterative form:

[0115]

[0116] Formula (23) can be used to solve P i+1 The above can be regarded as the correction process of P, which is to ensure that P after each update i+1 Can make Closer

[0117] Based on the concept of inverse optimal control, the inverse reinforcement learning learning rate α∈(0, 1] is introduced on the basis of formula (16) to obtain Q i+1 The iterative equation is:

[0118]

[0119] Substituting formula (24) into formula (25) we can get the new iterative form:

[0120]

[0121] Formula (26) can be used to solve Q i+1 The above is the inverse optimal control process, which is to ensure that Q after each update i+1 Both correspond optimally to P i+1 .

[0122] Rewrite equation (21) into an iterative form:

[0123]

[0124] Formula (27) can be used to solve

[0125] Therefore, based on equations (23), (26), and (27), the following steps are performed:

[0126] 1) Set the iteration step i = 0, give the performance indicator input weight R> 0, select the initial performance indicator state weight Q0 ≥ 0 and a small positive constant ε, and collect the target system trajectory data y d (k) and u d (k);

[0127] 2) Use formula (23) to solve P i+1 ;

[0128] 3) Use equation (26) to solve Q i+1 ;

[0129] 4) Use formula (27) to solve K i+1 ;

[0130] 5) Final iteration end condition judgment: If || P i+1 -P i ||<ε, then stop the iteration and let Otherwise, i=i+1, and the process returns to step 2) to proceed to the next iteration.

[0131] In the following process, step 2) requires the system model information to be known when solving the inverse reinforcement learning optimal tracking control problem. Considering that it is difficult to establish an accurate system model in practical applications, step 3) proposes a data-driven inverse reinforcement Q learning optimal tracking control method based on step 2), so that only the behavioral trajectory data of the target system (2) and the controlled system (1) are used in the learning process without the need for system model information.

[0132] First, define the Q function form through formula (23):

[0133]

[0134] in,

[0135]

[0136] Define the delay system core matrix for:

[0137]

[0138] Define a new augmented vector w d (k)=[z d (k) T u d (k) T ] T , substituting formula (30) into formula (28), we can get the Q function as follows:

[0139]

[0140] Then the control strategy satisfy:

[0141]

[0142] in, Redefine the state weight of delay system performance indicators for:

[0143]

[0144] Substituting formula (18) into the Bellman equation, we can obtain:

[0145]

[0146] Substituting equation (18) into the controlled system (1) yields:

[0147] Lz(k-t+1)=ALz(kt)+Bu(kt) (35)

[0148] Substitute equation (35) into z(k-t+1) T L T P i+1 Lz(k-t+1) can be obtained:

[0149]

[0150] in,

[0151]

[0152] So z(kt) T L T P i+1 Lz(kt) can be rewritten as:

[0153]

[0154] Combining equations (29), (33) and (37), we can obtain:

[0155]

[0156] The above defines the form of the time-delay system Q function. It can be seen from Equation (32) that no system model information is required when updating the control strategy, which provides the possibility of implementing a model-free method.

[0157] Substituting equations (28) and (33) into equation (23), we can obtain:

[0158]

[0159] Formula (40) can be used to solve Substituting Equation (27) into Equation (26) can be rewritten as:

[0160]

[0161] Multiplying both sides of formula (41) by x(k) and combining it with formula (18) yields:

[0162]

[0163] Define a new augmented vector ω(k)=[z(k) T u(k) T ] T , combining equations (36), (38) and (39) we can get:

[0164]

[0165] Formula (43) can be used to solve

[0166] Therefore, based on Equations (32), (40), and (43), this method proposes an optimal tracking control method for time-delay systems based on inverse reinforcement Q learning. The specific steps are as follows:

[0167] 1) Initialization, set the iteration step i = 0, give the performance indicator input weight R> 0, and select the initial performance indicator state weight and a positive constant ε as a threshold, and collect the target trajectory data y d (k) and u d (k);

[0168] 2) Update the control strategy parameters and solve using equation (40)

[0169] 3) Control strategy update, using equation (32) to solve And calculate u(kt);

[0170] 4) Update the performance index weights and use formula (43) to solve

[0171] 5) Iteration end condition determination, if Then stop the iteration and let Otherwise, i=i+1, and the process returns to step 2) for the next iteration.

[0172] This method does not require any system model information during the solution process, and only uses the target trajectory y of the target system d (k) and u d (k) and the trajectories y(k) and u(k) of the controlled system. In addition, this method does not require system state data during the solution process, so it can be applied to linear discrete-time systems whose states are unmeasurable.

[0173] The following is a specific simulation example:

[0174] Simulation Example 1:

[0175] In order to verify the effectiveness of the present method, this example provides simulation results of the present method in the absence and presence of time delay.

[0176] Taking the short-time period aircraft flight attitude stability control model as an example, the aircraft flight diagram is as follows Figure 2As shown. When an aircraft is flying steadily for a short period of time, the main factors affecting the aircraft's flight attitude are the angle of attack, pitch rate, and elevator angle. The angle of attack refers to the angle between the aircraft's velocity vector and the wing chord line, the pitch rate refers to the angular velocity of the aircraft's rotation around the horizontal axis, and the elevator angle refers to the angle between the aircraft's elevator and horizontal tail. Changes in the angle of attack and pitch rate can directly measure changes in the aircraft's attitude. Changes in the elevator angle will cause changes in the aerodynamic torque acting on the elevator, thereby affecting the aircraft's flight attitude. Therefore, the aircraft's state variables are set to angle of attack, pitch rate, and elevator angle, that is, x(t) = [α q δ] T The aircraft satisfies the following dynamic relationship:

[0177]

[0178] The weights of target system performance indicators (3) are selected as follows:

[0179]

[0180] (1) Simulation results when there is no delay

[0181] When there is no time delay, that is, when t=0, Figure 3 Describes the system kernel matrix learned by this method Performance indicator weight and control strategies From the curve convergence diagram, we can see that all three finally converge to a constant value. Figure 4 Describes what this method has learned and Respectively with the target parameters and The error curve between and The converged solution of and but Converged This is a typical multi-solution characteristic of the inverse optimal control problem.

[0182] Simulation results show that in the absence of time delay, this method finds a performance indicator equivalent to the target performance indicator for the controlled system. Satisfactory tracking control effects can be achieved using only historical output and control input data without the need for state information, further demonstrating the effectiveness of the method of the present invention.

[0183] (2) Simulation results with time delay

[0184] When there is a time delay, select t=2. Figure 5 Describes the kernel matrix of the time-delay system learned by this method Performance indicator weight and control strategies From the convergence curve diagram, we can see that all three finally converge to a constant value. Figure 6 Describes what this method has learned and Respectively and The error curve between and The converged solution of and but Converged This is a typical multi-solution characteristic of the inverse optimal control problem.

[0185] Simulation results show that in the presence of time delay, this method finds a performance indicator equivalent to the target performance indicator for the controlled system, and a satisfactory tracking control effect can be achieved using only historical output and control input data, further illustrating the effectiveness of the method of the present invention.

[0186] Simulation Example 2:

[0187] To further verify the effectiveness of this method, this example presents two sets of comparative experiments in the presence of time delay: 1) Comparative Experiment 1 is conducted with the reinforcement learning method to illustrate the advantages of this method in control effect; 2) Comparative Experiment 2 is conducted with the dual-loop iterative inverse reinforcement learning method to illustrate the advantages of this method in the number of iterations.

[0188] (1) Comparative Experiment 1

[0189] Table 1 Evaluation indexes of the method of the present invention and comparative method 1

[0190]

[0191] Figure 7 The control input trajectory and output tracking performance of the controlled system of comparison method 1 are described. It can be seen that the tracking effect is not ideal and the tracking error is obvious. This is caused by the inappropriate weight of the performance index given by the human. This is also the control risk of this type of optimal control method. Figure 8 The control input trajectory and output tracking performance of the controlled system using the proposed method are described. Clearly, this method achieves significantly better tracking results. To quantify the difference in tracking performance between the two methods, the integral absolute error (IAE) and mean square error (MSE) are introduced to evaluate the control effectiveness of the two methods. The evaluation metrics are shown in Table 1, demonstrating that the proposed method achieves superior tracking control results.

[0192] (2) Comparative Experiment 2

[0193] The calculation of performance index weights in the dual-loop iterative inverse reinforcement learning method is outside the control strategy calculation. The method flow is as follows: Figure 9 shown. Figure 10 Describes the simulation results of the outer loop control strategy iteration of the comparison method 2, Figure 11 The control strategy iteration simulation results of the method of the present invention are described. Table 2 shows the number of iterations of the two methods.

[0194] Table 2 The number of iterations of the method of the present invention and the comparative method 2

[0195]

[0196] Combine Figure 10 、 Figure 11 As can be seen from Table 2, the outer loop of method 2 iterated 651 times, while the corresponding inner loop iterated 1041 times, for a total of 1692 iterations. The method of the present invention achieved the same parameter convergence effect with a total of 748 iterations. Therefore, the number of iterations of the method of the present invention is significantly reduced, and the learning efficiency is improved.

[0197] In summary, the present invention compensates for the impact of time delay on system performance by designing a Smith predictor, and at the same time finds the solution conditions for the optimal tracking control problem of inverse reinforcement learning. Then, combined with the Q-learning idea, a data-driven inverse reinforcement Q-learning method is proposed. It only uses historical output and control input data to simultaneously reconstruct performance indicators and learn optimal control strategies, avoiding the use of mathematical models and state information. It is an output feedback method. In the case of a time delay in the control input, for a target system with a demonstration effect but unknown performance indicators, the controlled system reconstructs the performance indicators and stabilizes the system only by observing the operating trajectory of the target system. The optimal control strategy corresponding to the performance indicator can make the controlled system and the target system have the same trajectory, thereby tracking the trajectory of the target system.

Claims

1. An optimal tracking control method for a time-delay system based on inverse reinforcement Q learning, characterized in that: The performance index weights of the controlled system are determined by observing the input and output data generated by the target system, and the control strategy of the target system is learned. The state variables x(k) of the target system are defined as the angle of attack, pitch rate, and elevator angle, the control input u(k) is defined as the elevator actuator voltage that controls the elevator change, and the system output y(k) is defined as the aircraft's offset relative to the flight directions x, y, and z. The method includes the following steps: S1, set the number of iterations i = 0, give a fixed performance indicator input weight R> 0, and select the initial performance indicator state weight and a positive constant ε as a threshold, and collect the trajectory data yd(k) and u of the target system d (k); S2, by collecting the elevator actuator voltage of the target system and the aircraft's offset relative to the flight direction x, y, z, the control strategy parameters are solved by formula (40) Among them, ω d (k) is the actual control input u of the target system d (k) and output y d (k), which is the matrix composed of the elevator actuator voltage that controls the elevator change and the aircraft's offset relative to the flight direction x, y, and z; z d (k) is the actual control input u of the target system d (k) and output y d (k) is a matrix; (k), (kt), and (k-t+1) represent the data collected at different times; the performance index input weight R is fixed; S3, by collecting the voltage of the elevator actuator that controls the elevator change of the controlled system, the offset of the aircraft relative to the flight direction x, y, z, and the control strategy parameters solved in step S2 The control strategy is solved by Then u(kt) is calculated by equation (32); Where u(k) is the actual control input of the controlled system, that is, the elevator actuator voltage that controls the elevator change; z(k) is the matrix formed by the actual control input u(k) and output y(k) of the controlled system. S4, by collecting the voltage of the elevator actuator that controls the elevator change of the controlled system, the offset of the aircraft relative to the flight direction, and the control strategy parameters obtained in step S2 The performance index state weight is solved by formula (43): Among them, the inverse reinforcement learning learning rate α∈[0,1] is artificially given, w(k) is the matrix formed by actually collecting the control input u(k) and output y(k) of the controlled system, The control strategy parameters Obtained by subtracting the performance indicator state weight and the input weight matrix, the performance indicator input weight R is fixed; S5, determine the end condition of the iteration, if Then stop the iteration and let Obtain the performance indicator state weights and optimal control strategy of the controlled system; otherwise, set i=i+1 and return to step S2 for the next iteration.