Optimal output tracking control method and system based on fixed convergence rate
By designing a dynamic controller and a dynamic error linear system for the linear continuous-time system, incorporating a fixed convergence rate, reconstructing the state using input and output data, and iteratively calculating the optimal control gain, the applicability and stability issues of the optimal output tracking control method for continuous-time systems in the existing technology are solved, and fast and stable output tracking control is achieved.
Patent Information
- Application Number
- CN202511259226.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-04
AI Technical Summary
In existing industrial production processes, the optimal output tracking control method is only applicable to discrete systems and is easily disturbed by external factors. It is difficult to apply to continuous-time systems, especially in complex systems such as chemical reactors, large generator sets or aircraft engines, where internal state variables cannot be obtained, resulting in poor stability.
An optimal output tracking control method based on a fixed convergence rate is designed. By designing a dynamic controller for a linear continuous-time system, a dynamic error linear system is constructed, and a fixed convergence rate is incorporated. The input and output data are used to reconstruct the state of the dynamic system, and the optimal control gain is obtained through iterative calculation to achieve output tracking of the target trajectory.
The optimal output tracking control method is extended to linear systems to ensure that the system converges stably and quickly under external disturbances. It is applicable to continuous-time systems and improves the stability and smooth progress of industrial production processes.
Smart Images

Figure CN120802636A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of industrial process control, and more particularly, to an optimal output tracking control method and system based on fixed convergence speed. BACKGROUND
[0002] When designing a controller for modern industrial production, it is usually desired that the output of the controlled system can accurately track the desired trajectory. In actual industrial production processes, this problem is usually simplified as an output regulation problem to analyze and solve. The regulation goal usually includes: regulating the stable controller, making the output signal asymptotically stable under the error of the given reference trajectory and being able to overcome the influence of external system disturbances on the system. However, solving the output regulation problem usually requires accurate system state parameters. Due to the complexity of modern production processes, the state parameters of the system are often difficult to measure. Therefore, it is more desirable to be able to find the optimal control strategy of the system through the measurement of the output information of the system.
[0003] The development of Reinforce Learning (RL) and Adaptive Dynamic Programming (ADP) theories is of great significance to solve the problem of model-free optimal output tracking. RL is a machine learning strategy that adjusts actions according to specified rewards through the interaction of agents and the environment, thereby gradually learning the optimal action or control strategy. In the absence of an accurate environment or system model, ADP approximates the performance index function and control strategy in the dynamic programming equation through a function approximation structure. The emergence of these control algorithms makes it possible to solve the optimal control problem without relying on accurate system dynamics, thereby expanding the application range of these technologies.
[0004] Many existing algorithms mainly rely on state feedback learning methods. These methods usually assume that complete measurement values of system states can be accessed, which is not always feasible in some engineering scenarios, and most learning designs do not consider the convergence speed, resulting in that the control system is more susceptible to disturbance factors and poor stability. To overcome the above problems, the prior art discloses an optimal tracking control method and system with a specified convergence speed, first establishes a multi-input multi-output control system model under linear discrete time, sets an initial regulator and an initial controller, and obtains a Sylvester mapping of the initial regulator; a virtual control strategy with a specified convergence speed is added to the established model; historical data of the multi-input multi-output control system model to which the virtual control strategy is added are collected, the initial regulator and the initial controller are optimized according to the obtained historical data, an optimal tracking controller is obtained according to the optimized regulator and the controller, and the multi-input multi-output control system model is tracked. The prior art is based on data driving to design an optimal controller, is suitable for a dynamically unknown system, has a wider application range, can stably track a specified signal, and ensures that the convergence speed of the system output in the transient response process meets the requirements. However, the prior art only proposes a method for solving the tracking control problem of a discrete system, and in the method, the controller needs to be designed based on the state variables that can be recorded by the system. In a large number of actual physical control systems, the internal dynamic characteristics are continuously changed in nature, and it is difficult to measure the intermediate state variables. For example, in a complex chemical reaction kettle, a large generator set or an aircraft engine, an engineer can usually only measure part of the output quantities (such as temperature, speed and pressure) and input quantities of the system, and cannot obtain all the internal state variables that describe the complete dynamics of the system. SUMMARY
[0005] To solve the problems that the optimal output tracking control method in the existing industrial production process is only suitable for a discrete system and is susceptible to external factor disturbance, the present application proposes an optimal output tracking control method based on a fixed convergence speed, applies the optimal output tracking control method to a linear system, improves the system stability, and ensures the smooth progress of the industrial production process.
[0006] To achieve the above technical effects, the technical solutions of the present application are as follows: In a first aspect, the present application proposes an optimal output tracking control method and system based on a fixed convergence speed, including the following steps: S1. Design a dynamic tracking controller based on an established linear continuous-time system model and a target trajectory model, and use the dynamic controller to generate a control signal to make the output of the linear continuous-time system model track the target trajectory; S2. Construct a dynamic error linear system model of the tracking target trajectory, and integrate a fixed convergence speed into the dynamic error linear system model. S3. Designing a behavior strategy of a dynamic error linear system model, designing a dynamic system model based on the behavior strategy and the linear continuous-time system model, running the dynamic system model by using input data generated by the behavior strategy, and obtaining output data; S4. Reconstructing a state of the dynamic system model by using the input data-output data; S5. Running the dynamic system model after the state is reconstructed by using the behavior strategy, obtaining output data of the dynamic system model, inputting the output data of the dynamic system model into the calculation equation based on the optimal control gain, and iteratively calculating to obtain the optimal control gain; S6. Generating a control signal based on the optimal control gain, and using the control signal to make the output of the linear continuous-time system model track the target trajectory.
[0007] In the technical solution, a dynamic controller is first designed for a linear continuous-time system to generate a control signal, and output tracking target trajectory control is constructed. A dynamic error linear system is designed to run a dynamic system by integrating a fixed convergence speed, and the state of the dynamic system is reconstructed by using input data-output data. The behavior strategy is applied to the reconstructed dynamic system to collect data again, the optimal control gain is solved by iterative calculation, and the control signal based on the optimal control gain is used to make the output of the linear continuous-time system model track the target trajectory. The method expands the optimal output tracking control method to a linear system, and the setting of the fixed convergence speed accelerates the convergence of the system, avoids disturbance from external factors, and is conducive to the smooth progress of industrial production processes.
[0008] Preferably, the expression of the linear continuous-time system model is:
[0009]
[0010] wherein, denotes a state matrix of the linear continuous-time system model with an initial state of , the dimension of , denotes a derivative of the state matrix of the linear continuous-time system model with respect to time, denotes a control input matrix, the dimension of , denotes an output matrix of the linear continuous-time system model, and the dimension of the matrix , denotes a system coefficient matrix with a dimension of , denotes a system coefficient matrix with a dimension of denotes a system coefficient matrix with a dimension of an input coefficient matrix of the target trajectory model, an output coefficient matrix of the target trajectory model, an output coefficient matrix of the target trajectory model; the target trajectory model is expressed as:
[0011]
[0012] wherein, denotes a state matrix of the target trajectory system model, denotes a derivative of the state matrix of the target trajectory system model with respect to time, denotes an output matrix of the target trajectory model, denotes a target trajectory coefficient matrix of dimension denotes a target trajectory coefficient matrix of dimension denotes a target trajectory output coefficient matrix of dimension denotes a target trajectory output coefficient matrix of dimension
[0013] Preferably, the dynamic tracking controller is expressed as:
[0014]
[0015] wherein, denotes a state coefficient matrix of the dynamic tracking controller, denotes a tracking error coefficient matrix, and a matrix pair comprises a minimal p-copy internal model of the matrix S, denotes an intermediate state matrix of the dynamic tracking controller, denotes a derivative of the intermediate state matrix of the dynamic tracking controller with respect to time, denotes a feedforward gain matrix that satisfies an observable, and denotes a feedback gain matrix, denotes a tracking error matrix of the linear continuous-time system model, the control signal generated by the dynamic controller controls the output of the linear continuous-time system model to track the target trajectory , .
[0016] Preferably, the dynamic error linear system model is expressed as:
[0017]
[0018] wherein, represents the state of a linear continuous-time system model error matrix of the target trajectory model state , represents derivative with respect to time , , , represents a set fixed convergence speed , , , represents an identity matrix of dimension .
[0019] Preferably, the expression of the behavior policy is:
[0020]
[0021] wherein, represents the internal state of the dynamic controller in the behavior policy, represents the derivative with respect to time of the internal state, represents the exploration noise matrix, , represents that the free variable of the exploration noise is a combination of sine and cosine signals of different amplitudes and frequencies; the expression of the dynamic system model is:
[0022]
[0023] wherein, represents the state under the linear continuous-time system running with the behavior policy, , represents derivative with respect to time , .
[0024] Preferably, the process of reconstructing the state of the dynamic system model with the input data-output data is: inputting the input data into the dynamic system model to obtain the output data and the exploration noise matrix ; based on the input data , the output data and the exploration noise matrix , calculating the state , , respectively represent the state driven by , and ; Substitute the state into the reconstruction formula, the re-expressed state is calculated, and the expression of the reconstruction formula is:
[0025] wherein, is a parameterized matrix.
[0026] Preferably, the output data of the acquired power system model comprises:
[0027]
[0028]
[0029]
[0030]
[0031]
[0032] wherein, denotes a Kronecker product, denotes a starting time of setting data collection, denotes a termination time of data collection, is a positive integer and needs to satisfy the expression:
[0033] wherein, ; The process of iterative calculation of the optimal control gain is: S51. Set an initial iteration number and a convergence error , substitute the output data of the power system model into the expression, and calculate the control gain , and the expression is:
[0034]
[0035]
[0036] wherein, , , ; S52. judge whether the error matrix norm between the control gain of the first iteration and the control gain of the second iteration satisfies the expression: the control gain of the second iteration the control gain of the second iteration
[0037] If not, return to execute step S51; if yes, execute S53; S53. take the control gain of the second iteration as the optimal control gain based on the optimal control gain, generate a control signal, and use the control signal to make the output of the linear continuous-time system model track the target trajectory; The expression of the control strategy for generating the control signal based on the optimal control gain is:
[0038]
[0039] In a second aspect, the application further provides an optimal output tracking control system based on a fixed convergence speed, which comprises: a dynamic tracking controller design unit, configured to design a dynamic tracking controller based on the constructed linear continuous-time system model and the target trajectory model, and use the dynamic controller to generate a control signal to make the output of the linear continuous-time system model track the target trajectory; a dynamic error linear system model construction unit, configured to construct a dynamic error linear system model of the target trajectory, and integrate the fixed convergence speed into the dynamic error linear system model; a dynamic system model construction and operation unit, configured to design a behavior strategy of the dynamic error linear system model, design a dynamic system model based on the behavior strategy and the linear continuous-time system model, and use the input data generated by the behavior strategy to operate the dynamic system model to obtain output data; a dynamic system model state reconstruction unit, configured to reconstruct the state of the dynamic system model using the input data and the output data; an optimal control gain calculation unit, configured to operate the dynamic system model with the reconstructed state using the behavior strategy, obtain the output data of the dynamic system model, input the output data of the dynamic system model into the constructed calculation equation based on the optimal control gain, and iteratively calculate to obtain the optimal control gain; a target trajectory tracking output unit, configured to generate a control signal based on the optimal control gain, and use the control signal to make the output of the linear continuous-time system model track the target trajectory.
[0040] In a third aspect, the present application further provides a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor executes the computer program to implement the optimal output tracking control method based on the fixed convergence speed.
[0041] In a fourth aspect, the present application further provides a computer storage medium, which stores a computer program, wherein the computer program comprises program instructions, and when the program instructions are executed by a computer, the computer is caused to execute the optimal output tracking control method based on the fixed convergence speed.
[0042] Compared with the prior art, the present application has the following beneficial effects: The present application provides an optimal output tracking control method and system based on a fixed convergence speed. Firstly, a dynamic controller is designed for a linear continuous-time system to generate a control signal, and an output tracking target trajectory control is constructed to build a dynamic error linear system with a fixed convergence speed. Then, a behavior strategy is designed to run a dynamic system, and the state of the dynamic system is reconstructed by using input data and output data. The data collected by the behavior strategy are applied to the reconstructed dynamic system again, and the optimal control gain is solved by iterative calculation. Finally, the output of the linear continuous-time system model tracks the target trajectory by using the control signal based on the optimal control gain. The present application expands the optimal output tracking control method to linear systems, and the setting of the fixed convergence speed accelerates the convergence of the system, avoids the disturbance of external factors, and is conducive to the smooth progress of industrial production processes. BRIEF DESCRIPTION OF DRAWINGS
[0043] Figure 1 FIG. 1 shows a flowchart of an optimal output tracking control method based on a fixed convergence speed according to an embodiment of the present application; Figure 2 FIG. 2 shows a flowchart of a method for obtaining an optimal control gain by iterative calculation according to an embodiment of the present application; Figure 3 FIG. 3 shows a structural diagram of a comparison chart of convergence errors of an optimal control gain and non-optimal control gains according to an embodiment of the present application; Figure 4 FIG. 4 shows a diagram of a relationship between an output tracking target trajectory control of a linear continuous-time system model and time according to an embodiment of the present application; Figure 5 FIG. 5 shows a diagram of an implementation of a fixed convergence speed according to an embodiment of the present application; Figure 6 FIG. 6 shows a structural diagram of an optimal output tracking control system based on a fixed convergence speed according to an embodiment of the present application; Figure 7Fig. 1 is a structural schematic diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION
[0044] The accompanying drawings are only used for illustrative purposes and should not be construed as limiting the patent; In order to better illustrate the present embodiment, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the actual size; For those skilled in the art, it is understandable that some well-known descriptions in the drawings may be omitted.
[0045] The technical solutions of the present application will be further described below in combination with the drawings and embodiments.
[0046] The positional relationship described in the drawings is only used for illustrative purposes and should not be construed as limiting the patent; Embodiment 1 The present embodiment proposes an optimal output tracking control method based on fixed convergence speed, and the flowchart of the method is shown in Figure 1 , comprising the following steps: S1. Based on the constructed linear continuous-time system model and the target trajectory model, a dynamic tracking controller is designed, and a control signal is generated by using the dynamic controller to make the output of the linear continuous-time system model track the target trajectory; S2. A dynamic error linear system model tracking the target trajectory is constructed, and a fixed convergence speed is integrated into the dynamic error linear system model; S3. The behavior strategy of the dynamic error linear system model is designed, and based on the behavior strategy and the linear continuous-time system model, a dynamic system model is designed, the input data generated by the behavior strategy is used to run the dynamic system model, and the output data is obtained; S4. The state of the dynamic system model is reconstructed by using the input data-output data; S5. The dynamic system model after the state reconstruction is run by using the behavior strategy, the output data of the dynamic system model is obtained, the output data of the dynamic system model is input into the constructed calculation equation based on the optimal control gain, and the optimal control gain is obtained by iterative calculation; S6. A control signal is generated based on the optimal control gain, and the output of the linear continuous-time system model tracks the target trajectory by using the control signal.
[0047] In this embodiment, first, a dynamic controller is designed for a linear continuous-time system to generate a control signal, and an output trajectory control is constructed to track a target trajectory. A dynamic error linear system with a fixed convergence speed is built. A behavior strategy is designed to run the dynamic system, and the state of the dynamic system is reconstructed using input data and output data. The behavior strategy is applied to the reconstructed dynamic system again to collect data, and the optimal control gain is calculated by iteration. The output of the linear continuous-time system model is tracked by the target trajectory using the control signal based on the optimal control gain. This method expands the optimal output tracking control method to linear systems, and the fixed convergence speed setting accelerates the convergence of the system, avoids disturbance from external factors, and is conducive to the smooth progress of industrial production processes.
[0048] In this embodiment, the expression of the linear continuous-time system model is:
[0049]
[0050] wherein, represents a state matrix of the linear continuous-time system model with an initial state of , has a dimension of , represents a derivative of the state matrix of the linear continuous-time system model with respect to time, represents a control input matrix, has a dimension of , represents an output matrix of the linear continuous-time system model, and the matrix has a dimension of , represents a system coefficient matrix with a dimension of , represents an input coefficient matrix with a dimension of , represents an output coefficient matrix with a dimension of ; The expression of the target trajectory model is:
[0051]
[0052] wherein, represents a state matrix of the target trajectory system model, represents a derivative of the state matrix of the target trajectory system model with respect to time, represents an output matrix of the target trajectory model, represents a target trajectory coefficient matrix with a dimension of , represents a target trajectory output coefficient matrix of dimension .
[0053] In particular, the following assumptions are made for the linear continuous-time system model and the target trajectory model parameters: 1、 is controllable, which means that all states inside the linear continuous-time system model can be changed by applying control inputs; is observable, which means that all states inside the linear continuous-time system model can be derived by analyzing 2、matrix does not have eigenvalues with negative real parts; 3、for any ; 4、matrix has dimension ; In this embodiment, the expression of the dynamic tracking controller is:
[0054]
[0055] wherein, represents a state coefficient matrix of the dynamic tracking controller, represents a tracking error coefficient matrix, and the matrix pair contains a minimum p-copy internal model of matrix S, represents an intermediate state matrix of the dynamic tracking controller, represents a derivative of the intermediate state matrix of the dynamic tracking controller with respect to time, represents a feedforward gain matrix satisfying is observable, and represents a feedback gain matrix, represents a tracking error matrix of the linear continuous-time system model, and a control signal generated by the dynamic controller controls the output of the linear continuous-time system model to track the target trajectory , .
[0056] In this embodiment, the expression of the dynamic error linear system model is:
[0057]
[0058] wherein, represents a linear continuous-time system model state error matrix of a target trajectory model state , represents derivative with respect to time , , , represents a set fixed convergence speed , , , represents a unit matrix with dimension .
[0059] Specifically, the expression of the optimal performance index is:
[0060] wherein , is a semi-positive definite matrix .
[0061] Specifically, the expression of the error matrix is:
[0062] wherein , has a dimension of , , is the unique solution of the following expression:
[0063] wherein ,
[0064] In this embodiment, the expression of the behavior policy is:
[0065]
[0066] wherein represents the internal state of the dynamic controller in the behavior policy, represents the derivative with respect to time of the internal state, represents an exploration noise matrix, , represents that the free variable of the exploration noise is a combination of sine and cosine signals with different amplitudes and frequencies; The expression of the dynamic system model is:
[0067]
[0068] wherein, denotes the state of the linear continuous-time system running under the behavioral policy, , denotes derivative with respect to time, , .
[0069] In the embodiment, the process of reconstructing the state of the dynamic system model by using the input data-output data is as follows: inputting the input data into the dynamic system model to obtain the output data and the exploration noise matrix ; based on the input data , the output data and the exploration noise matrix , calculating the state , , respectively denoting the states driven by , and ; substituting the state into the reconstruction formula to obtain the reconstructed state , and the expression of the reconstruction formula is as follows:
[0070] wherein, is a parameterized matrix.
[0071] Specifically, , denotes derivative with respect to time; , denotes derivative with respect to time; , denotes derivative with respect to time, , and the initial values of the states of ;
[0072] wherein, the eigenvalues of are negative real numbers, and is a positive coefficient.
[0073] In the embodiment, the output data of the acquired power system model comprises:
[0074]
[0075]
[0076]
[0077]
[0078]
[0079] wherein, denotes a Kronecker product, denotes a set data collection start time, denotes a data collection end time, is a positive integer and needs to satisfy the expression:
[0080] wherein, ; Specifically, the expression of the behavior strategy adopted when acquiring the output data of the power system model is:
[0081] wherein, , is stable.
[0082] The flowchart of the method for iteratively calculating the optimal control gain is shown in Figure 2 , and the process is as follows: S51. Set the initial iteration number and the convergence error , substitute the output data of the power system model into the expression to calculate the control gain , and the expression is:
[0083]
[0084]
[0085] wherein, , , ; S52. Determine whether the control gain of the first iteration is equal to the control gain of the first iteration. whether the error matrix norm between the two satisfies the expression:
[0086] If not, then , return to step S51; if yes, then execute S53; S53. Take as the optimal control gain , generate a control signal based on the optimal control gain, and use the control signal to make the linear continuous-time system model output track the target trajectory; The control strategy expression for generating a control signal based on the optimal control gain is:
[0087]
[0088] Specifically, before running the state-reconstructed dynamic error linear system model again using the behavior strategy, there is a model-free pre-collection phase, in which the dynamic error linear system runs for a long enough time before data collection begins, so that .
[0089] Specifically, for the state , define
[0090] For a matrix of dimension , is defined as:
[0091] For a matrix of dimension , is defined as:
[0092] Specifically, a comparison chart of the convergence error of the optimal control gain and the convergence error of the non-optimal control gain is shown in Figure 3 The convergence error of the optimal control gain and the convergence error of the non-optimal control gain have similar rates of change, but after obtaining the optimal control gain, the convergence error remains unchanged.
[0093] Specifically, the control relationship between the linear continuous-time system model output tracking the target trajectory using the control signal based on the optimal control gain and time is shown in Figure 4 Figure 4 In the figure, the ordinate represents the tracking control of the output of the linear continuous-time system model, and the abscissa represents time. After the control strategy of the optimal control gain is applied, the output of the tracking control is completely the same as the target trajectory control.
[0094] Specifically, the implementation of the fixed convergence speed is as shown in the figure, and the ordinate represents the tracking control of the output of the linear continuous-time system Figure 5 is completely the same as the target trajectory control. is completely the same as the target trajectory control. is completely the same as the target trajectory control.
[0095] Embodiment 2 The embodiment proposes an optimal output tracking control method based on a fixed convergence speed, comprising the following steps: S1. Based on the constructed linear continuous-time system model and the target trajectory model, a dynamic tracking controller is designed, and a control signal is generated by using the dynamic controller to make the output of the linear continuous-time system model track the target trajectory. In this embodiment, the linear continuous-time system is the F-16 aircraft dynamics system.
[0096] S2. A dynamic error linear system model tracking the target trajectory is constructed, and a fixed convergence speed is integrated into the dynamic error linear system model. S3. A behavior strategy of the dynamic error linear system model is designed, and based on the behavior strategy and the linear continuous-time system model, a dynamic system model is designed, the input data generated by the behavior strategy is used to run the dynamic system model, and output data is obtained. S4. The state of the dynamic system model is reconstructed by using the input data-output data. S5. The dynamic system model after the state is reconstructed is run by using the behavior strategy, the output data of the dynamic system model is obtained, the output data of the dynamic system model is input into the constructed calculation equation based on the optimal control gain, and the optimal control gain is obtained by iterative calculation. S6. A control signal is generated based on the optimal control gain, and the output of the linear continuous-time system model is made to track the target trajectory by using the control signal.
[0097] In this embodiment, the expression of the linear continuous-time system model is:
[0098]
[0099] wherein, denotes the state matrix of the linear continuous-time system model in initial state has dimension , denotes the derivative of the state matrix of the linear continuous-time system model with respect to time, denotes the control input matrix, has dimension , denotes the output matrix of the linear continuous-time system model, the matrix has dimension , denotes the system coefficient matrix of dimension , denotes the input coefficient matrix of dimension , denotes the output coefficient matrix of dimension ; the expression of the target trajectory model is:
[0100]
[0101] wherein, denotes the state matrix of the target trajectory system model, denotes the derivative of the state matrix of the target trajectory system model with respect to time, denotes the output matrix of the target trajectory model, denotes the target trajectory coefficient matrix of dimension , denotes the target trajectory output coefficient matrix of dimension .
[0102] In the present embodiment, specifically, the expression of the F-16 aircraft dynamics model is:
[0103]
[0104] the expression of the target trajectory model is:
[0105]
[0106] Specifically, for the linear continuous-time system model and the target trajectory model parameters have the following assumptions: 1、 controllable, meaning that all states inside the linear continuous-time system model can be changed by applying control inputs; observable, meaning that all states inside the linear continuous-time system model can be derived by analyzing ; 2. The matrix has no eigenvalue with negative real part; 3. For any ; 4. The matrix has dimension ; In this embodiment, the expression of the dynamic tracking controller is:
[0107]
[0108] wherein, denotes the state coefficient matrix of the dynamic tracking controller, denotes the tracking error coefficient matrix, and the matrix pair contains the minimal p-copy internal model of the matrix S, denotes the intermediate state matrix of the dynamic tracking controller, denotes the derivative of the intermediate state matrix of the dynamic tracking controller with respect to time, denotes the feedforward gain matrix satisfying is observable, and denotes the feedback gain matrix, denotes the tracking error matrix of the linear continuous-time system model, and the control signal generated by the dynamic controller controls the output of the linear continuous-time system model to track the target trajectory , .
[0109] In this embodiment, the expression of the dynamic error linear system model is:
[0110]
[0111] wherein, denotes the error matrix of the state of the linear continuous-time system model and the state of the target trajectory model, , denotes the derivative with respect to time, , , , represents a fixed convergence rate set, , , , represents a unit matrix of dimension .
[0112] Specifically, the expression of the optimal performance index is:
[0113] wherein, , is a positive semi-definite matrix, .
[0114] In this embodiment, the expression of the error matrix is:
[0115] wherein, , is of dimension , , is the unique solution of the following expression:
[0116] wherein, ,
[0117] In this embodiment, the expression of the behavior policy is:
[0118]
[0119] wherein, represents the internal state of the dynamic controller in the behavior policy, represents the derivative of the internal state with respect to time, represents the exploration noise matrix, , represents that the free variable of the exploration noise is a combination of sine and cosine signals of different amplitudes and frequencies; The expression of the dynamic system model is:
[0120]
[0121] wherein, represents the state of the linear continuous-time system running under the behavior policy, , denotes derivative with respect to time, , .
[0122] In the embodiment, the process of reconstructing the state of the power system model using the input-output data is as follows: inputting the input data into the power system model to obtain the output data and the exploration noise matrix ; based on the input data , the output data and the exploration noise matrix , calculating the state , , respectively represent the states driven by , and ; substituting the state into the reconstruction formula to obtain the reconstructed state , and the expression of the reconstruction formula is as follows:
[0123] wherein is a parameterized matrix.
[0124] Specifically, , denotes derivative with respect to time; , denotes derivative with respect to time; , denotes derivative with respect to time, , and the initial values of the states are all 0; ;
[0125] wherein the eigenvalues of are all negative real numbers, and is a positive coefficient.
[0126] In the embodiment, the output data of the power system model obtained includes:
[0127]
[0128]
[0129]
[0130]
[0131]
[0132] in, represents the Kronecker product, Indicates the start time for setting data collection. Indicates the end time of data collection. Is a positive integer and needs to satisfy the expression:
[0133] in, ; Specifically, the expression of the behavior strategy used when obtaining the output data of the power system model is:
[0134] in, , is stable.
[0135] The process of the method for iteratively calculating the optimal control gain is as follows: S51. Set the initial number of iterations and convergence error , substitute the output data of the dynamic system model into the expression to calculate the control gain , the expression is:
[0136]
[0137]
[0138] in, , , ; S52. Determine The control gain of the iteration With the The control gain of the iteration Does the error matrix norm between satisfy the expression:
[0139] If not satisfied, , return to step S51; if satisfied, execute S53; S53. The method of S52, wherein the optimal control gain is determined based on the optimal control gain expression a control signal is generated based on the optimal control gain, and the linear continuous-time system model output is caused to track the target trajectory using the control signal The control strategy expression for generating the control signal based on the optimal control gain is
[0140]
[0141] Specifically, before running the state-reconstructed dynamic error linear system model using the behavior policy again, a model-free pre-collection phase is further included, in which the dynamic error linear system runs for a long enough time before data collection starts, so that .
[0142] Specifically, for the state , the definition is
[0143] For the matrix with dimension , the definition is
[0144] For the matrix with dimension , the definition is
[0145] Specifically, as shown in the comparison chart of the convergence error of the optimal control gain and the convergence error of the non-optimal control gain, the convergence error of the optimal control gain is represented by triangular points, and the convergence error of the non-optimal control gain is represented by circular points. The two convergence errors have similar variation rates, but after obtaining the optimal control gain, the convergence error is constant.
[0146] Specifically, as shown in Figure 4 , the vertical coordinate is the tracking target trajectory control of the linear continuous-time system model output, and the horizontal coordinate is time. After applying the control strategy of the optimal control gain, the output tracking control is exactly the same as the target trajectory control . After applying the control strategy of the optimal control gain, the output tracking control of the linear continuous-time system is exactly the same as the target trajectory control .
[0147] Specifically, as shown in Figure 5The ordinate represents the tracking control of the linear continuous-time system output fully converges to the target trajectory control The ordinate represents the tracking control of the linear continuous-time system output fully converges to the target trajectory control The ordinate represents the tracking control of the linear continuous-time system output fully converges to the target trajectory control The ordinate represents the tracking control of the linear continuous-time system output
[0148] Embodiment 3 The embodiment proposes an optimal output tracking control system based on a fixed convergence speed, in which the system is used to implement an optimal output tracking control method based on a fixed convergence speed, and a structural schematic diagram is as shown in Figure 6 The embodiment comprises: A dynamic tracking controller design unit is configured to design a dynamic tracking controller based on a constructed linear continuous-time system model and a target trajectory model, so that the linear continuous-time system model output tracks the target trajectory control. An optimal performance index construction unit is configured to construct a dynamic error linear system model of tracking the target trajectory based on the dynamic tracking controller and the linear continuous-time system model, integrate a fixed convergence speed into the dynamic error linear system model, and set an optimal performance index of the tracking control based on the dynamic error linear system model. A dynamic system model construction unit is configured to design a behavior strategy of the dynamic error linear system model, design a dynamic system model based on the behavior strategy and the linear continuous-time system model, run the dynamic system with input data generated by the behavior strategy, and obtain output data. A dynamic system model reconstruction unit is configured to reconstruct a state of the dynamic system model with input-output data, and obtain a reconstructed dynamic system model state. An optimal control gain calculation unit is configured to run the reconstructed state-represented dynamic system model again with the behavior strategy, collect data and input the data into an optimal control gain calculation equation, iteratively calculate an optimal control gain, and use a control signal based on the optimal control gain to make the linear continuous-time system model output track the target trajectory control.
[0149] Embodiment 4 In this embodiment, a computer device is proposed, which comprises a memory 101, a processor 102, and a computer program stored on the memory 101 and executable by the processor 102, wherein the processor 102 executes the computer program to implement an optimal output tracking control method based on a fixed convergence speed, and a structural schematic diagram of the device is as shown in Figure 7
[0150] In the embodiment, a computer storage medium is also provided, and the computer storage medium stores a computer program including program instructions, and the program instructions, when executed by a computer, cause the computer to execute the optimal output tracking control method based on the fixed convergence speed.
[0151] Obviously, the above embodiments of the present application are only examples for clearly illustrating the present application, and are not intended to limit the implementation modes of the present application. Based on the above description, other different forms of changes or variations can be made by those skilled in the art. Here, all the implementation modes are not required or can not be exhausted. Any modification, equivalent replacement and improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the claims of the present application.
Claims
1. An optimal output tracking control method based on a fixed convergence rate, characterized in that: The following steps are involved: S1. Based on the constructed linear continuous-time system model and target trajectory model, design a dynamic tracking controller. Use the dynamic controller to generate control signals so that the output of the linear continuous-time system model tracks the target trajectory. S2. Construct a dynamic error linear system model for tracking the target trajectory and incorporate a fixed convergence rate into the dynamic error linear system model; S3. Design a behavioral strategy for a dynamic error linear system model. Based on the behavioral strategy and the linear continuous-time system model, design a dynamic system model. Run the dynamic system model using the input data generated by the behavioral strategy to obtain output data. S4. Reconstruct the state of the power system model using input data-output data; S5. Using the behavioral strategy to run the reconstructed power system model, obtain output data of the power system model, input the output data of the power system model into the constructed calculation equation based on the optimal control gain, and iteratively calculate the optimal control gain; S6. Generate a control signal based on the optimal control gain, and use the control signal to make the output of the linear continuous-time system model track the target trajectory.
2. The optimal output tracking control method based on a fixed convergence rate according to claim 1, characterized in that: The expression of the linear continuous-time system model is: in, Indicates the initial state is The state matrix of the linear continuous-time system model is The dimension is , represents the time derivative of the state matrix of a linear continuous-time system model, represents the control input matrix, The dimension is , Represents the output matrix of the linear continuous-time system model, the matrix The dimension is , Indicates the dimension The system coefficient matrix of Indicates the dimension The input coefficient matrix of Indicates the dimension The output coefficient matrix of The expression of the target trajectory model is: in, represents the target trajectory system model state matrix, represents the time derivative of the target trajectory system model state matrix, represents the output matrix of the target trajectory model, Indicates the dimension The target trajectory coefficient matrix, Indicates the dimension The target trajectory output coefficient matrix.
3. The optimal output tracking control method based on a fixed convergence rate according to claim 2, characterized in that: The expression of the dynamic tracking controller is: in, represents the state coefficient matrix of the dynamic tracking controller, represents the tracking error coefficient matrix, and The matrix pair The minimal p-copy internal model containing the matrix S, represents the intermediate state matrix of the dynamic tracking controller, represents the time derivative of the intermediate state matrix of the dynamic tracking controller, Express satisfaction The observable feedforward gain matrix, and represents the feedback gain matrix, Represents the tracking error matrix of the linear continuous-time system model. The control signal generated by the dynamic controller controls the output of the linear continuous-time system model. Tracking target trajectory , .
4. The optimal output tracking control method based on a fixed convergence rate according to claim 3, characterized in that: The expression of the dynamic error linear system model is: in, Represents the state of a linear continuous-time system model and target trajectory model state The error matrix, , express The time derivative, , , , Indicates the fixed convergence speed set. , , , Indicates the dimension The identity matrix of .
5. The optimal output tracking control method based on a fixed convergence rate according to claim 4, characterized in that: The expression of the behavior strategy is: in, represents the internal state of the dynamic controller in the behavior strategy, represents the time derivative of the internal state, represents the exploration noise matrix, , The free variables representing the exploratory noise are combinations of sine and cosine signals of different amplitudes and frequencies; The expression of the power system model is: in, represents the state of a linear continuous-time system running with a behavioral strategy, , express The time derivative, , .
6. The optimal output tracking control method based on a fixed convergence rate according to claim 5, characterized in that: The process of reconstructing the state of the power system model using input data-output data is as follows: Input data Input into the power system model to obtain output data and explore the noise matrix ; Based on input data , output data and explore the noise matrix , calculation status , , Respectively represented by , and The state of the drive; Will state Substitute into the reconstruction formula and calculate the re-expression state , the expression of the reconstruction formula is: in, is a parameterized matrix.
7. The optimal output tracking control method based on a fixed convergence rate according to claim 6, characterized in that: The output data of the power system model obtained includes: in, represents the Kronecker product, Indicates the start time for setting data collection. Indicates the end time of data collection. Is a positive integer and needs to satisfy the expression: in, ; The process of iterative calculation to obtain the optimal control gain is: S51. Set the initial number of iterations and convergence error , substitute the output data of the dynamic system model into the expression to calculate the control gain , the expression is: in, , , ; S52. Determine The control gain of the iteration With the The control gain of the iteration Does the error matrix norm between satisfy the expression: If not satisfied, , return to step S51; if satisfied, execute S53; S53. As the optimal control gain , generating a control signal based on the optimal control gain, and using the control signal to make the linear continuous-time system model output track the target trajectory; The control strategy expression for generating the control signal based on the optimal control gain is: 。 8. An optimal output tracking control system based on a fixed convergence rate, characterized in that: The system comprises: A dynamic tracking controller design unit is used to design a dynamic tracking controller based on the constructed linear continuous-time system model and the target trajectory model, and to generate a control signal using the dynamic controller so that the output of the linear continuous-time system model tracks the target trajectory; A dynamic error linear system model building unit is used to build a dynamic error linear system model for tracking the target trajectory and integrate a fixed convergence rate into the dynamic error linear system model; The dynamic system model construction and operation unit is used to design the behavioral strategy of the dynamic error linear system model, design the dynamic system model based on the behavioral strategy and the linear continuous time system model, and run the dynamic system model using the input data generated by the behavioral strategy to obtain the output data; A power system model state reconstruction unit, used to reconstruct the state of the power system model using input data-output data; An optimal control gain calculation unit is used to operate the power system model after the reconstructed state using the behavioral strategy, obtain output data of the power system model, input the output data of the power system model into the constructed calculation equation based on the optimal control gain, and iteratively calculate the optimal control gain; The target trajectory tracking output unit is used to generate a control signal based on the optimal control gain, and use the control signal to make the output of the linear continuous-time system model track the target trajectory.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable by the processor. The processor executes the computer program to implement the method according to any one of claims 1 to 7.
10. A computer storage medium, characterized in that A computer program is stored thereon, the computer program including program instructions, which, when executed by a computer, enable the computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Tracking control method based on measurement data under denial of service attack
CN115826414A
Optimal tracking control method and system with specified convergence speed
CN116382080A
Model-free online reinforcement learning method and system for solving H infinity control problem
CN117590742A
High-dynamic sliding mode prediction double-layer control method based on error optimization strategy
CN119846976A
Electromagnetic micromirror output feedback tracking control method and device based on adaptive Q learning
CN120491471A