An optimal tracking control method and system with a specified convergence rate

By establishing a multi-input multi-output control system model under linear discrete time, setting an initial regulator and controller, adding a virtual control strategy with a specified convergence rate, and optimizing the controller, the problem of unreliable convergence rate in dynamic unknown systems is solved, and stable optimal tracking control is achieved.

CN116382080BActive Publication Date: 2025-11-14GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310325060.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-29
Publication Date
2025-11-14
Estimated Expiration
2043-03-29

AI Technical Summary

Technical Problem

Existing technologies are not applicable to dynamic and unknown systems, and cannot guarantee that the convergence speed of the optimal tracking controller will meet the preset standard.

Method used

A model of a multi-input multi-output control system in linear discrete time is established, an initial regulator and an initial controller are set, the Sylvester map is obtained, and a virtual control strategy with a specified convergence rate is added. The initial regulator and controller are optimized by collecting historical data to obtain the optimal tracking controller.

Benefits of technology

It is suitable for dynamic unknown systems, has a wide range of applications, reduces the consumption of computing resources, and ensures the convergence and uniqueness of the optimal controller. It can stably track the specified reference signal and ensure that the system output meets the convergence speed requirements during transient response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116382080B_ABST
    Figure CN116382080B_ABST
Patent Text Reader

Abstract

This invention provides an optimal tracking control method and system with a specified convergence rate. The method includes establishing a linear discrete-time multi-input multi-output (MIMO) control system model, setting an initial regulator and an initial controller, and obtaining the Sylvester map of the initial regulator; adding a virtual control strategy with a specified convergence rate to the established model; collecting historical data of the MIMO control system model with the added virtual control strategy; optimizing the initial regulator and initial controller based on the acquired historical data; obtaining the optimal tracking controller based on the optimized regulator and controller; and performing tracking control on the MIMO control system model. This invention designs the optimal controller based on data-driven principles, is applicable to dynamically unknown systems, and has a wider range of applications. Furthermore, this invention can stably track a specified signal and ensure that the convergence rate of the system output meets requirements during transient response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of control technology for reinforcement learning and multiple-input multiple-output systems, and more specifically, to an optimal tracking control method and system with a specified convergence rate. Background Technology

[0002] In recent years, with the continuous development of industrial technology, the connection methods and integration levels of systems have become increasingly complex, while the requirements for their control precision are also constantly increasing. Therefore, theoretical research on high-precision control of complex systems is particularly important. In the field of control science, reinforcement learning algorithms have been applied to the design of adaptive control and optimal control strategies. By learning from the system's operating data, optimal control strategies for unknown dynamic systems can be designed, thereby maximizing the system's economic benefits.

[0003] Reinforcement learning is a technique that utilizes feedback from the interaction between an agent and its unknown environment to optimize decision-making problems. This technology is inspired by the biological world, where organisms can improve their behavior to survive and thrive through interaction with their environment. Reinforcement learning algorithms are based on the idea that successful control decisions need to be remembered, and that reinforcement signals make them easier to use a second time. This theory originates from animal experiments, where the neurotransmitter dopamine was observed to act as a reinforcement signal, promoting learning at the neuronal level. From a theoretical perspective, reinforcement learning algorithms are closely related to direct and indirect adaptive optimal control methods.

[0004] Current technologies disclose an optimal tracking control method for multi-timescale systems based on reinforcement learning. First, it decomposes the multi-timescale tracking problem into a linear quadratic tracking problem for a slow subsystem and a dynamic game problem for a fast subsystem using singular perturbation theory. Then, based on this, it proposes a non-policy reinforcement learning algorithm that uses only real-time system measurement data to find a feedforward controller based on output regulation theory. The operating index can track its specified target value through an approximately optimal method, achieving optimal tracking control for the multi-timescale system. However, existing methods require knowledge of the system's internal dynamics when dealing with the tracking controller, making them unsuitable for systems with unknown dynamics. Furthermore, existing methods do not limit the dynamic performance of the optimal tracking controller, failing to guarantee that the system's tracking convergence speed reaches a preset standard when the controlled object tracks a specified reference signal. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, such as unsuitability for systems with unknown dynamics and inability to guarantee convergence speed, this invention provides an optimal tracking control method and system with a specified convergence speed. This method has a wide range of applications and can avoid the need to understand the internal dynamics of complex system controllers when designing them. In addition, this invention also limits the dynamic performance of the optimal tracking controller, enabling the control system to converge strictly according to the specified speed.

[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0007] An optimal tracking control method with a specified convergence rate includes the following steps:

[0008] S1: Establish a multi-input multi-output control system model under linear discrete time;

[0009] S2: Set the initial regulator and initial controller of the multi-input multi-output control system model, and obtain the Sylvester mapping of the initial regulator;

[0010] S3: Add a virtual control strategy with a specified convergence rate to the multi-input multi-output control system model;

[0011] S4: Collect historical data of the multi-input multi-output control system model with virtual control strategy added, optimize the initial regulator and initial controller based on the acquired historical data, obtain the optimal tracking controller based on the optimized regulator and controller, and perform tracking control on the multi-input multi-output control system model.

[0012] Preferably, the linear discrete-time multi-input multi-output control system model established in step S1 is specifically as follows:

[0013] x(k+1)=Ax(k)+Bu(k)+Dv(k)

[0014] v(k+1)=Ev(k)

[0015] y(k)=Cx(k)

[0016] y d (k)=Fv(k)

[0017] e(k)=y(k)-y d (k)

[0018] Where x(k) is the state variable of the multi-input multi-output control system at time k, u(k) is the input data of the multi-input multi-output control system at time k, v(k) is the external state variable of the multi-input multi-output control system at time k, and y(k) and y d(k) represents the output data and reference signal of the multi-input multi-output control system at time k, and e(k) represents the output error variable of the multi-input multi-output control system at time k; A, B, C, D, E, F are the first, second, third, fourth, fifth and sixth constant coefficient matrices of the multi-input multi-output control system, respectively.

[0019] The established linear discrete-time multi-input multi-output control system model satisfies the following constraints:

[0020] The constant coefficient matrix pair (A, B) is stable; the constant coefficient matrix pair (A, C) is observable; the eigenvalues ​​of the fifth constant coefficient matrix E are all outside the unit circle; the matrix Full rank, where λ i Let I be the i-th eigenvalue of the first constant coefficient matrix A, and let I be the identity matrix.

[0021] The state variable x(k), external state variable v(k), output error variable e(k), input data u(k), and output data y(k) of a multi-input multi-output control system at time k can be detected and recorded.

[0022] Preferably, the initial regulator set in step S2 is specifically:

[0023]

[0024]

[0025] Where (X,U) is the initial regulator matrix set calculated based on the constant coefficient matrix, satisfying the following equation:

[0026] XE-AX=BU+D

[0027] CX = F

[0028] Where (X,U) is the initial regulator matrix group.

[0029] Preferably, the initial controller set in step S2 is specifically:

[0030] u(k) = -K x x(k)-K v v(k)

[0031] Among them, K x and K v These are the first and second optimal control gain matrices of the initial controller, respectively.

[0032] Preferably, the Sylvester mapping of the initial regulator in step S2 is specifically as follows:

[0033]

[0034] Among them, X i The definition is as follows:

[0035] Define X0 as a matrix with all zero elements and the same dimension as matrix X;

[0036] Define X1 as the seventh constant coefficient matrix satisfying CX1 = F, I q ={X2,X3,…,X d+1 To satisfy CX i The set of constant coefficient matrices with constant coefficients equal to 0, where X, X1, X2, ..., X d+1 Satisfy the following expression:

[0037]

[0038] According to the Sylvester mapping of the initial regulator, we have:

[0039]

[0040] Where d is the number of elements in the set of constant coefficient matrices.

[0041] Preferably, the virtual control strategy with a specified convergence rate in step S3 is specifically as follows:

[0042]

[0043] in, K 0 Given a known initially stable control strategy. A =A, B =B, where γ is the convergence rate parameter;

[0044] The virtual control strategy with a specified convergence rate satisfies the following Bellman equation:

[0045]

[0046] in,

[0047]

[0048] in, The parameters of the Bellman equations are given by a virtual control strategy with a specified convergence rate.

[0049] matrix The expressions for each part of the matrix block are as follows:

[0050]

[0051]

[0052]

[0053]

[0054]

[0055]

[0056] The virtual control strategy with a specified convergence rate satisfies the following Bellman equation and can be expressed in the following form:

[0057]

[0058] in, Specifically:

[0059]

[0060] Specifically:

[0061]

[0062] Where γ is the convergence rate parameter.

[0063] Preferably, in step S4, historical data of a multi-input multi-output control system model with added virtual control strategy is collected, specifically by:

[0064] The specific amount of historical data collected is as follows:

[0065]

[0066] Where n, m, and q are the dimensions of the state variables, input variables, and external state variables of the multi-input multi-output control system model with added virtual control strategy, respectively.

[0067] The historical data collected for the multi-input multi-output control system model with added virtual control strategy are as follows:

[0068]

[0069] in, and The expression is as follows:

[0070]

[0071] in, and These are the first and second parameters from the collected historical data, respectively.

[0072] Preferably, in step S4, the specific method for optimizing the initial regulator and the initial controller based on the acquired historical data is as follows:

[0073] The initial regulator matrix set (X, U) is optimized based on the acquired historical data to obtain the optimized regulator matrix set (X). * U * The initial regulator is optimized using the following method:

[0074] Using the acquired historical data, the optimized regulator matrix set (X) is obtained by solving the following formula. * U * ):

[0075]

[0076] in,

[0077] Solve according to the following formula. and Specifically:

[0078]

[0079]

[0080] in,

[0081] The acquired historical data is used for data learning, and the initial controller is optimized. The optimized controller is as follows:

[0082]

[0083] Among them, u * (k) represents the optimized controller. and satisfy:

[0084]

[0085]

[0086] in, and This is the control gain matrix of the optimized controller.

[0087] Preferably, in step S4, the optimal tracking controller obtained based on the optimized regulator and controller specifically includes:

[0088] The optimal tracking controller is specifically:

[0089]

[0090] in, and This is the control gain matrix of the optimized controller.

[0091] The present invention also provides an optimal tracking control system with a specified convergence rate, which, when applying the above-described optimal tracking control method with a specified convergence rate, includes:

[0092] Model building unit: used to build a model of a multi-input multi-output control system in linear discrete time;

[0093] Initialization unit: used to set the initial regulator and initial controller of the multi-input multi-output control system model, and to obtain the Sylvester mapping of the initial regulator;

[0094] Virtual control unit: used to add a virtual control strategy with a specified convergence rate to the multi-input multi-output control system model;

[0095] Data acquisition and optimization unit: used to acquire historical data of the multi-input multi-output control system model with virtual control strategy added, optimize the initial regulator and initial controller based on the acquired historical data, obtain the optimal tracking controller based on the optimized regulator and controller, and perform tracking control on the multi-input multi-output control system model.

[0096] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0097] This invention provides an optimal tracking control method and system with a specified convergence rate. The method involves: establishing a linear discrete-time multi-input multi-output (MIMO) control system model; setting an initial regulator and an initial controller for the MIMO control system model and obtaining the Sylvester map of the initial regulator; adding a virtual control strategy with a specified convergence rate to the MIMO control system model; collecting historical data of the MIMO control system model with the added virtual control strategy; optimizing the initial regulator and initial controller based on the acquired historical data; finally, obtaining the optimal tracking controller based on the optimized regulator and controller, and performing tracking control on the MIMO control system model.

[0098] This invention utilizes reinforcement learning and data-driven design for optimal controllers, making it applicable to dynamically unknown systems with a wider range of applications and lower computational resource consumption. Furthermore, this invention ensures the convergence and uniqueness of the designed optimal controller, independent of specific system internal dynamics. When the optimal controller, regulator, and their parameters are applied to a multiple-input multiple-output (MIMO) control system, it can stably track a specified reference signal and ensure that the convergence speed of the system output meets requirements during transient response. Attached Figure Description

[0099] Figure 1 The flowchart is for an optimal tracking control method with a specified convergence speed provided in Example 1.

[0100] Figure 2 This is a structural diagram of the LCL-coupled inverter-type distributed generation system provided in Example 2.

[0101] Figure 3 The diagram shows the results of the optimal tracking control method provided in Example 2.

[0102] Figure 4 This is a comparison chart of the system output error and the reference error with a specified convergence rate provided in Example 2.

[0103] Figure 5 This is a schematic diagram showing the number of iterations for the optimal tracking control method provided in Example 2.

[0104] Figure 6 This is a structural diagram of an optimal tracking control system with a specified convergence speed provided in Example 3. Detailed Implementation

[0105] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent.

[0106] To better illustrate this embodiment, some parts in the accompanying drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions;

[0107] It will be understood by those skilled in the art that certain well-known structures and their descriptions may be omitted in the accompanying drawings.

[0108] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0109] Example 1

[0110] like Figure 1 As shown, this embodiment provides an optimal tracking control method with a specified convergence rate, including the following steps:

[0111] S1: Establish a multi-input multi-output control system model under linear discrete time;

[0112] S2: Set the initial regulator and initial controller of the multi-input multi-output control system model, and obtain the Sylvester mapping of the initial regulator;

[0113] S3: Add a virtual control strategy with a specified convergence rate to the multi-input multi-output control system model;

[0114] S4: Collect historical data of the multi-input multi-output control system model with virtual control strategy added, optimize the initial regulator and initial controller based on the acquired historical data, obtain the optimal tracking controller based on the optimized regulator and controller, and perform tracking control on the multi-input multi-output control system model.

[0115] In the specific implementation process, firstly, a multi-input multi-output (MIMO) control system model under linear discrete time is established; then, an initial regulator and an initial controller are set for the MIMO control system model, and the Sylvester mapping of the initial regulator is obtained; a virtual control strategy with a specified convergence rate is added to the MIMO control system model; historical data of the MIMO control system model with the added virtual control strategy is collected, and the initial regulator and initial controller are optimized based on the acquired historical data; finally, the optimal tracking controller is obtained based on the optimized regulator and controller, and tracking control is performed on the MIMO control system model.

[0116] This method utilizes reinforcement learning and data-driven design of optimal controllers, making it suitable for systems with unknown dynamics and a wider range of applications. Furthermore, this method ensures the convergence and uniqueness of the designed optimal controller, independent of specific system internal dynamics. When the optimal controller, regulator, and their parameters are applied to a multiple-input multiple-output (MIMO) control system, it can stably track a specified reference signal and ensure that the convergence speed of the system output meets the requirements during transient response.

[0117] Example 2

[0118] This embodiment provides an optimal tracking control method with a specified convergence rate, including the following steps:

[0119] S1: Establish a multi-input multi-output control system model under linear discrete time;

[0120] S2: Set the initial regulator and initial controller of the multi-input multi-output control system model, and obtain the Sylvester mapping of the initial regulator;

[0121] S3: Add a virtual control strategy with a specified convergence rate to the multi-input multi-output control system model;

[0122] S4: Collect historical data of the multi-input multi-output control system model with virtual control strategy added, optimize the initial regulator and initial controller based on the acquired historical data, obtain the optimal tracking controller based on the optimized regulator and controller, and perform tracking control on the multi-input multi-output control system model;

[0123] The linear discrete-time multi-input multi-output control system model established in step S1 is specifically as follows:

[0124] x(k+1)=Ax(k)+Bu(k)+Dv(k)

[0125] v(k+1)=Ev(k)

[0126] y(k)=Cx(k)

[0127] y d (k)=Fv(k)

[0128] e(k)=y(k)-y d (k)

[0129] Where x(k) is the state variable of the multi-input multi-output control system at time k, u(k) is the input data of the multi-input multi-output control system at time k, v(k) is the external state variable of the multi-input multi-output control system at time k, and y(k) and y d (k) represents the output data and reference signal of the multi-input multi-output control system at time k, and e(k) represents the output error variable of the multi-input multi-output control system at time k; A, B, C, D, E, F are the first, second, third, fourth, fifth and sixth constant coefficient matrices of the multi-input multi-output control system, respectively.

[0130] The established linear discrete-time multi-input multi-output control system model satisfies the following constraints:

[0131] The constant coefficient matrix pair (A, B) is stable; the constant coefficient matrix pair (A, C) is observable; the eigenvalues ​​of the fifth constant coefficient matrix E are all outside the unit circle; the matrix Full rank, where λ i Let I be the i-th eigenvalue of the first constant coefficient matrix A, and let I be the identity matrix.

[0132] The state variable x(k), external state variable v(k), output error variable e(k), input data u(k), and output data y(k) of the multi-input multi-output control system at time k can be detected and recorded;

[0133] The initial regulator set in step S2 is specifically:

[0134]

[0135]

[0136] Where (X,U) is the initial regulator matrix set calculated based on the constant coefficient matrix, satisfying the following equation:

[0137] XE-AX=BU+D

[0138] CX = F

[0139] Where (X,U) is the initial regulator matrix group;

[0140] The initial controller set in step S2 is specifically as follows:

[0141] u(k) = -K x x(k)-K v v(k)

[0142] Among them, K x and K v These are the first and second optimal control gain matrices of the initial controller, respectively;

[0143] The Sylvester mapping of the initial regulator in step S2 is specifically as follows:

[0144]

[0145] Among them, X i The definition is as follows:

[0146] Define X0 as a matrix with all zero elements and the same dimension as matrix X;

[0147] Define X1 as the seventh constant coefficient matrix satisfying CX1 = F, I q ={X2X3,…,X d+1 To satisfy CX i The set of constant coefficient matrices with constant coefficients equal to 0, where X, X1, X2, ..., X d+1 Satisfy the following expression:

[0148]

[0149] According to the Sylvester mapping of the initial regulator, we have:

[0150]

[0151] Where d is the number of elements in the set of constant coefficient matrices;

[0152] The virtual control strategy with a specified convergence rate in step S3 is specifically as follows:

[0153]

[0154] in, K 0 Given a known initially stable control strategy. A =γA, B =γB, where γ is the convergence rate parameter;

[0155] The virtual control strategy with a specified convergence rate satisfies the following Bellman equation:

[0156]

[0157] in,

[0158]

[0159] in, The parameters of the Bellman equations are given by a virtual control strategy with a specified convergence rate.

[0160] matrix The expressions for each part of the matrix block are as follows:

[0161]

[0162]

[0163]

[0164]

[0165]

[0166]

[0167] The virtual control strategy with a specified convergence rate satisfies the following Bellman equation and can be expressed in the following form:

[0168]

[0169] in, Specifically:

[0170]

[0171] Specifically:

[0172]

[0173] Where γ is the convergence rate parameter;

[0174] In step S4, historical data of a multi-input multi-output control system model with added virtual control strategy is collected. The specific method is as follows:

[0175] The specific amount of historical data collected is as follows:

[0176]

[0177] Where n, m, and q are the dimensions of the state variables, input variables, and external state variables of the multi-input multi-output control system model with added virtual control strategy, respectively.

[0178] The historical data collected for the multi-input multi-output control system model with added virtual control strategy are as follows:

[0179]

[0180] in, and The expression is as follows:

[0181]

[0182] in, and These are the first and second parameters from the collected historical data, respectively.

[0183] In step S4, the specific method for optimizing the initial regulator and initial controller based on the acquired historical data is as follows:

[0184] The initial regulator matrix set (X, U) is optimized based on the acquired historical data to obtain the optimized regulator matrix set (X). * U * The initial regulator is optimized using the following method:

[0185] Using the acquired historical data, the optimized regulator matrix set (X) is obtained by solving the following formula. * U * ):

[0186]

[0187] in,

[0188] Solve according to the following formula. and Specifically:

[0189]

[0190]

[0191] in,

[0192] The acquired historical data is used for data learning, and the initial controller is optimized. The optimized controller is as follows:

[0193]

[0194] Among them, u * (k) represents the optimized controller. and satisfy:

[0195]

[0196]

[0197] in, and The control gain matrix of the optimized controller;

[0198] In step S4, the optimal tracking controller obtained based on the optimized regulator and controller specifically refers to:

[0199] The optimal tracking controller is specifically:

[0200]

[0201] in, and This is the control gain matrix of the optimized controller.

[0202] In the specific implementation process, the first step is to establish a multi-input multi-output control system model under linear discrete time, specifically as follows:

[0203] x(k+1)=Ax(k)+Bu(k)+Dv(k)

[0204] v(k+1)=Ev(k)

[0205] y(k)=Cx(k)

[0206] y d (k)=Fv(k)

[0207] e(k)=y(k)-y d (k)

[0208] Where x(k) is the state variable of the multi-input multi-output control system at time k, u(k) is the input data of the multi-input multi-output control system at time k, v(k) is the external state variable of the multi-input multi-output control system at time k, and y(k) and y d (k) represents the output data and reference signal of the multi-input multi-output control system at time k, and e(k) represents the output error variable of the multi-input multi-output control system at time k; A, B, C, D, E, F are the first, second, third, fourth, fifth and sixth constant coefficient matrices of the multi-input multi-output control system, respectively.

[0209] The established linear discrete-time multi-input multi-output control system model satisfies the following constraints:

[0210] The constant coefficient matrix pair (A, B) is stable; the constant coefficient matrix pair (A, C) is observable; the eigenvalues ​​of the fifth constant coefficient matrix E are all outside the unit circle; the matrix Full rank, where λ i Let I be the i-th eigenvalue of the first constant coefficient matrix A, and let I be the identity matrix.

[0211] The state variable x(k), external state variable v(k), output error variable e(k), input data u(k), and output data y(k) of the multi-input multi-output control system at time k can be detected and recorded;

[0212] Then, the initial regulator and initial controller of the multi-input multi-output control system model are set, and the Sylvester mapping of the initial regulator is obtained;

[0213] The initial regulator set is as follows:

[0214]

[0215]

[0216] Where (X,U) is the initial regulator matrix set calculated based on the constant coefficient matrix, satisfying the following equation:

[0217] XE-AX=BU+D

[0218] CX = F

[0219] Where (X,U) is the initial regulator matrix group;

[0220] The initial controller set is as follows:

[0221] u(k) = -K x x(k)-K v v(k)

[0222] Among them, Kx and K v These are the first and second optimal control gain matrices of the initial controller, respectively;

[0223] The Sylvester mapping of the initial regulator in step S2 is specifically as follows:

[0224]

[0225] Among them, X i The definition is as follows:

[0226] Define X0 as a matrix with all zero elements and the same dimension as matrix X;

[0227] Define X1 as the seventh constant coefficient matrix satisfying CX1 = F, I q ={X2,X3,…X d+1 To satisfy CX i The set of constant coefficient matrices with constant coefficients equal to 0, where X, X1, X2, ..., X d+1 Satisfy the following expression:

[0228]

[0229] According to the Sylvester mapping of the initial regulator, we have:

[0230]

[0231] Where d is the number of elements in the set of constant coefficient matrices;

[0232] Add a virtual control strategy with a specified convergence rate to the multi-input multi-output control system model, specifically:

[0233]

[0234] in, K 0 Given a known initially stable control strategy. A =γA, B =γB, where γ is the convergence rate parameter;

[0235] The virtual control strategy with a specified convergence rate satisfies the following Bellman equation:

[0236]

[0237] in,

[0238]

[0239] in, The parameters of the Bellman equations are given by a virtual control strategy with a specified convergence rate.

[0240] matrix The expressions for each part of the matrix block are as follows:

[0241]

[0242]

[0243]

[0244]

[0245]

[0246]

[0247] The virtual control strategy with a specified convergence rate satisfies the following Bellman equation and can be expressed in the following form:

[0248]

[0249] in, Specifically:

[0250]

[0251] Specifically:

[0252]

[0253] Where γ is the convergence rate parameter;

[0254] Then, historical data of the multi-input multi-output control system model with added virtual control strategy were collected. The purpose of collecting historical data was to optimize the initial regulator and initial controller using reinforcement learning.

[0255] To generate a unique, optimal tracking controller that meets the system's dynamic performance requirements during reinforcement learning, a necessary amount of historical data needs to be collected. The specific amount of historical data collected is as follows:

[0256]

[0257] Where n, m, and q are the dimensions of the state variables, input variables, and external state variables of the multi-input multi-output control system model with added virtual control strategy, respectively.

[0258] The historical data collected for the multi-input multi-output control system model with added virtual control strategy are as follows:

[0259]

[0260] in, and The expression is as follows:

[0261]

[0262] in, and These are the first and second parameters from the collected historical data, respectively.

[0263] The initial regulator and initial controller are optimized based on the acquired historical data, specifically as follows:

[0264] The initial regulator matrix set (X, U) is optimized based on the acquired historical data to obtain the optimized regulator matrix set (X). * U * The initial regulator is optimized using the following method:

[0265] Using the acquired historical data, the optimized regulator matrix set (X) is obtained by solving the following formula. * U * ):

[0266]

[0267] in,

[0268] Solve according to the following formula. and Specifically:

[0269]

[0270]

[0271] in,

[0272] The acquired historical data is used for data learning, and the initial controller is optimized. The optimized controller is as follows:

[0273]

[0274] Among them, u * (k) represents the optimized controller. and satisfy:

[0275]

[0276]

[0277] in, and The control gain matrix of the optimized controller;

[0278] Finally, the optimal tracking controller is obtained based on the optimized regulator and controller. Specifically, the optimal tracking controller is:

[0279]

[0280] And the optimal tracking controller is used to perform tracking control on the multi-input multi-output control system model;

[0281] The application of this method in the control of LCL-coupled inverter-type distributed generation systems is explained in detail below:

[0282] like Figure 2 As shown, the LCL-coupled inverter type distributed generation system includes a three-phase inverter, a coupled LC inverter, and a load network;

[0283] The LCL-coupled inverter-type distributed generation system is used as the multi-input multi-output control system in this embodiment. The specific mathematical model is as follows:

[0284]

[0285]

[0286]

[0287] Transforming the above mathematical model into a state-space model yields the following formula:

[0288] x(k)=[I L V C I o ] T

[0289] v(k)=[V D I R ] T

[0290] y(k)=I O

[0291] u(k)=V I

[0292] Among them, V C V is the capacitor voltage. I For the input voltage, I L I is the inductor current. O V is the output current of the inverter. D For system voltage disturbance signal, IR This serves as a reference signal for the system output current.

[0293] The state-space model of the discrete system described above can be expressed as:

[0294] x(k+1)=Ax(k)+Bu(k)+Dv(k)

[0295] v(k+1)=Ev(k)

[0296] y(k)=Cx(k)

[0297] y d (k)=Fv(k)

[0298] e(k)=y(k)-y d (k)

[0299] The constant coefficient matrix of this system is unknown;

[0300] Collect historical operating data of the system, including the system's state variable x(k), the system's control input u(k), and the system's external state variable v(k), and treat the historical data at each time point as a set of data;

[0301] Calculate the constant coefficient matrix X0,X1,…,X that conforms to the system output equation. d+1 For the LCL-coupled inverter distributed generation system described above, d = 5;

[0302] Calculate ξ i (k), where

[0303]

[0304]

[0305] To ensure the uniqueness of the optimal control strategy, the number of data points s collected must be at least 1.

[0306]

[0307] The collected data is represented in the following form:

[0308]

[0309] in, φ in i (k+1) is:

[0310]

[0311] in, A constant coefficient matrix that matches any given dimension;

[0312] The collected data is calculated using the following formula to obtain...

[0313]

[0314] according to Calculate the control gain matrix of the optimized controller. And calculate the optimized regulator matrix group (X) according to the following formula. * U * ):

[0315]

[0316] The optimal tracking controller obtained at the end can be represented as:

[0317]

[0318] like Figure 3 The diagram illustrates the process of an LCL system output tracking a reference signal. It represents the process of the system output current tracking the reference current signal. The solid sine line represents the actual system output, the dashed line represents the reference signal, and the smooth solid line represents the system output error. The time interval 0-100 represents the system's learning phase. During this phase, the system operates using an initially stable control strategy, but is affected by noise signals, causing the system output to not accurately track the reference signal. After collecting historical data from this part of the system's operation, the method proposed in this embodiment is used to learn from this historical data, obtain the optimal control strategy with a specified convergence rate, and apply it to the system, resulting in the system output curve for time intervals 100-200. At this point, the system can stably track the reference signal.

[0319] like Figure 4 The image shows a comparison between the system error and the specified convergence speed. The smooth solid line is the reference line for the specified convergence speed, and the other solid line is the actual error output of the system. It can be seen that the system error cannot converge in the first half. After applying the optimal control strategy, the system can converge to 0 faster than the specified convergence speed.

[0320] like Figure 5 The figure shows the number of iterations required for the calculation by this method. It can be seen that after several iterations, the optimal control strategy can be obtained. Therefore, the method in this embodiment does not require a lot of computing resources.

[0321] This method utilizes reinforcement learning and data-driven design of optimal controllers, making it suitable for systems with unknown dynamics and a wider range of applications. Furthermore, this method ensures the convergence and uniqueness of the designed optimal controller, independent of specific system internal dynamics. When the optimal controller, regulator, and their parameters are applied to a multiple-input multiple-output (MIMO) control system, it can stably track a specified reference signal and ensure that the convergence speed of the system output meets the requirements during transient response.

[0322] Example 3

[0323] like Figure 6 As shown, this embodiment provides an optimal tracking control system with a specified convergence rate, applying the optimal tracking control method with a specified convergence rate described in Embodiment 1 or 2, including:

[0324] Model building unit 301: Used to build a multi-input multi-output control system model in linear discrete time;

[0325] Initialization unit 302: used to set the initial regulator and initial controller of the multi-input multi-output control system model, and to obtain the Sylvester mapping of the initial regulator;

[0326] Virtual control unit 303: used to add a virtual control strategy with a specified convergence rate to the multi-input multi-output control system model;

[0327] Data acquisition and optimization unit 304: used to acquire historical data of the multi-input multi-output control system model with added virtual control strategy, optimize the initial regulator and initial controller based on the acquired historical data, obtain the optimal tracking controller based on the optimized regulator and controller, and perform tracking control on the multi-input multi-output control system model.

[0328] In the specific implementation process, firstly, the model building unit 301 establishes a multi-input multi-output control system model under linear discrete time; then, the initialization unit 302 sets the initial regulator and initial controller of the multi-input multi-output control system model and obtains the Sylvester mapping of the initial regulator; the virtual control unit 303 adds a virtual control strategy with a specified convergence rate to the multi-input multi-output control system model; finally, the data acquisition and optimization unit 304 acquires historical data of the multi-input multi-output control system model with added virtual control strategy, optimizes the initial regulator and initial controller based on the acquired historical data, obtains the optimal tracking controller based on the optimized regulator and controller, and performs tracking control on the multi-input multi-output control system model;

[0329] This system utilizes reinforcement learning and data-driven design for optimal controllers, making it suitable for systems with unknown dynamics and thus having a wider range of applications. Furthermore, the system guarantees the convergence and uniqueness of the designed optimal controller, independent of specific system-level dynamic knowledge. When the optimal controller, regulator, and their parameters are applied to a multiple-input multiple-output (MIMO) control system, it can stably track a specified reference signal and ensure that the convergence speed of the system output meets requirements during transient response.

[0330] The same or similar labels correspond to the same or similar parts;

[0331] The terms used to describe positional relationships in the accompanying drawings are for illustrative purposes only and should not be construed as limiting this patent.

[0332] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. An optimal tracking control method with a specified convergence rate, characterized in that, Includes the following steps: S1: Establish a multi-input multi-output control system model under linear discrete time; S2: Set the initial regulator and initial controller of the multi-input multi-output control system model, and obtain the Sylvester mapping of the initial regulator; S3: Add a virtual control strategy with a specified convergence rate to the multi-input multi-output control system model; wherein, the virtual control strategy with a specified convergence rate is specifically: in, , , Given a known initially stable control strategy. , , ,in For convergence speed parameters; The virtual control strategy with a specified convergence rate satisfies the following Bellman equation: in, in, , The parameters of the Bellman equations are given by a virtual control strategy with a specified convergence rate. matrix The expressions for each part of the matrix block are as follows: The virtual control strategy with a specified convergence rate satisfies the following Bellman equation and can be expressed in the following form: in, Specifically: Specifically: in, For convergence speed parameters; S4: Collect historical data of the multi-input multi-output control system model with virtual control strategy added, optimize the initial regulator and initial controller based on the acquired historical data, obtain the optimal tracking controller based on the optimized regulator and controller, and perform tracking control on the multi-input multi-output control system model.

2. The optimal tracking control method with a specified convergence rate according to claim 1, characterized in that, The linear discrete-time multi-input multi-output control system model established in step S1 is specifically as follows: in, For multi-input multi-output control systems in The state variable at time t. For multi-input multi-output control systems in Input data at any given time, For multi-input multi-output control systems in The external state variables at time t, and For multi-input multi-output control systems Output data and reference signal at time t. For multi-input multi-output control systems in Output error variable at time t; These are the first, second, third, fourth, fifth, and sixth constant coefficient matrices of a multi-input multi-output control system; The established linear discrete-time multi-input multi-output control system model satisfies the following constraints: constant coefficient matrix pair Stable; constant coefficient matrix pair Observable; Fifth constant coefficient matrix The eigenvalues ​​of the matrix are all outside the unit circle; Full rank, among which, The first constant coefficient matrix The i-th eigenvalue, It is the identity matrix; Multiple-input multiple-output control systems in State variables at time 1 External state variables Output error variables Input data and output data It can be detected and recorded.

3. The optimal tracking control method with a specified convergence rate according to claim 2, characterized in that, The initial regulator set in step S2 is specifically: in, The initial regulator matrix set, calculated based on the constant coefficient matrix, satisfies the following equation: in, This is the initial regulator matrix group.

4. The optimal tracking control method with a specified convergence rate according to claim 2, characterized in that, The initial controller set in step S2 is specifically as follows: in, and These are the first and second optimal control gain matrices of the initial controller, respectively.

5. The optimal tracking control method with a specified convergence rate according to claim 3, characterized in that, The Sylvester mapping of the initial regulator in step S2 is specifically as follows: in, The definition is as follows: definition For all elements to be zero, dimension and matrix Consistent matrix; definition To meet The seventh constant coefficient matrix, To meet The set of constant coefficient matrices, where, Satisfy the following expression: According to the Sylvester mapping of the initial regulator, we have: Where d is the number of elements in the set of constant coefficient matrices.

6. The optimal tracking control method with a specified convergence rate according to claim 5, characterized in that, In step S4, historical data of a multi-input multi-output control system model with added virtual control strategy is collected. The specific method is as follows: The specific amount of historical data collected is as follows: in, These represent the dimensions of the state variables, input variables, and external state variables in a multi-input multi-output control system model with added virtual control strategies. The historical data collected for the multi-input multi-output control system model with added virtual control strategy are as follows: in, and The expression is as follows: in, and These are the first and second parameters from the collected historical data, respectively.

7. The optimal tracking control method with a specified convergence rate according to claim 6, characterized in that, In step S4, the specific method for optimizing the initial regulator and initial controller based on the acquired historical data is as follows: Based on the acquired historical data, the initial regulator matrix group was adjusted. Optimize and obtain the optimized regulator matrix set. The initial regulator was optimized using the following method: Using the acquired historical data, the optimized regulator matrix set is obtained by solving the following formula. : in, , ; Solve according to the following formula. and Specifically: in, ; The acquired historical data is used for data learning, and the initial controller is optimized. The optimized controller is as follows: in, For the optimized controller, and satisfy: in, and This is the control gain matrix of the optimized controller.

8. The optimal tracking control method with a specified convergence rate according to claim 7, characterized in that, In step S4, the optimal tracking controller obtained based on the optimized regulator and controller specifically refers to: The optimal tracking controller is specifically: in, and This is the control gain matrix of the optimized controller.

9. An optimal tracking control system with a specified convergence rate, employing the optimal tracking control method with a specified convergence rate according to any one of claims 1 to 8, characterized in that, include: Model building unit: used to build a model of a multi-input multi-output control system in linear discrete time; Initialization unit: used to set the initial regulator and initial controller of the multi-input multi-output control system model, and to obtain the Sylvester mapping of the initial regulator; Virtual control unit: used to add a virtual control strategy with a specified convergence rate to the multi-input multi-output control system model; Data acquisition and optimization unit: used to acquire historical data of the multi-input multi-output control system model with virtual control strategy added, optimize the initial regulator and initial controller based on the acquired historical data, obtain the optimal tracking controller based on the optimized regulator and controller, and perform tracking control on the multi-input multi-output control system model.

Citation Information

Patent Citations

  • Self-adaptive incremental optimization fault-tolerant control method for nonlinear system actuator faults

    CN113093536A

  • Batch process model-free deorbit strategy optimal tracking control method in packet loss environment

    CN114200834A