A data-driven optimal switching control method for switching systems
By employing a data-driven optimal switching control method, HJB equations are constructed using state data to derive the optimal control strategy for the switching system. This solves the problem of unknown subsystem models in the switching system and achieves optimal switching control with minimum cost over an infinite time region.
Patent Information
- Application Number
- CN202211279534.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-19
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2042-10-19
AI Technical Summary
In switching system control, uncertainties make it impossible to obtain subsystem models, and traditional model-based methods cannot guarantee good performance. Existing technologies cannot effectively solve this problem.
Using a data-driven approach, the optimal control strategy is designed, and the state data of the switching system is used to construct infinite and finite time domain HJB equations. The expression of the optimal control strategy is derived, and the unknowns are replaced by approximate functions. The weights are estimated using the state data matrix to achieve optimal switching control.
It achieves optimal switching control of the switching system without the need for a subsystem model, simplifies the derivation process of the finite-time HJB equation, is applicable to a variety of practical application scenarios, and has better control performance than traditional methods.
Smart Images

Figure CN115755595B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of switching control, in particular to a data-driven optimal switching control method of a switching system. BACKGROUND
[0002] In switching system control, due to the existence of various uncertainties, it may be difficult to obtain a subsystem model or an accurate subsystem model, at which time the traditional model-based method cannot solve the problem or is difficult to guarantee good performance. Therefore, if the system model cannot be accurately obtained during control, a switching control method independent of the system model needs to be studied.
[0003] A large amount of process data is generated in industrial processes, including valuable state information. Using these online and offline data, controllers can be directly designed, performance can be evaluated, and decisions can be made. The data-driven switching control method of the present application uses these data to replace the switching subsystem model to design the controller. SUMMARY
[0004] The present application provides a data-driven optimal switching control method of a switching system for the case where the subsystem model of the switching system is unknown. The optimal switching control of the switching system can be realized without the subsystem model, only with state data.
[0005] The data-driven optimal switching control method of the switching system of the present application first designs an optimal control strategy to minimize the switching system in the infinite time region;
[0006] Secondly, the expression of the infinite time optimal control strategy (v, μ) is constructed using the infinite time HJB equation of the switching system; wherein v of the HJB equation represents the subsystem of the switching system, and μ represents the control amount of the subsystem;
[0007] Thirdly, the expression of the finite time optimal control strategy is constructed using the finite time HJB equation of the switching system; based on the expression of the optimal solution of the finite time HJB equation, the approximation of the value function is obtained according to the partial derivative from a positive definite function; some unknown quantities in the approximation, including the control amount μ and the value function, are replaced by an approximate function, which is the product of a basis function and a weight; the weight in the approximation is estimated using a state data matrix to obtain an approximate optimal weight; and the optimal value function is obtained based on the approximate optimal weight.
[0008] Finally, the optimal value function is brought into the expression of the infinite time optimal control strategy to obtain the optimal control strategy.
[0009] Preferably, the switching system is:
[0010]
[0011] where x(t) ∈ R n is the system state, which is measurable; v ∈ V represents the index of the current active subsystem; V = {1, 2,..., N} is the index set of all subsystems, and N is the number of subsystems; f v : R n × R m → R n is the unknown model of subsystem v; u(t) is the system control input; x(0) = x0 ∈ Ω is the initial state, is the state region to be studied, including the origin;
[0012] The infinite-time region cost function is:
[0013]
[0014] where Q: R n × R m → R is a positive definite function; the control input is defined as a state feedback form: μ(x(t)) = u(t);
[0015] In the optimization process, the cost from the state x(t) ∈ Ω at time t to the infinite-time region is defined as:
[0016]
[0017] satisfies:
[0018]
[0019] The corresponding infinite-time HJB equation is:
[0020]
[0021] where “ * ” represents the optimal value, x = x(t), and when (·) is a scalar, The corresponding optimal switching control strategy (v * , μ * ) is:
[0022]
[0023] The finite-time region HJB equation is:
[0024]
[0025] where the finite-time region cost function f from time t and state x(t) to time t The corresponding optimal switching control strategy is:
[0026]
[0027] More preferably, the control method specifically comprises the following steps:
[0028] Step 1: According to the optimal solution relationship based on HJB equation, wherein the unknown quantity is represented by the base function weight and the approximation error is ignored, a data matrix is introduced:
[0029]
[0030]
[0031] Wherein t1 < t2 < … < t l is a selected time point, l is a positive integer, r = 1, 2, … l; Φ(x) is a vector composed of a set of linearly independent base functions; Ψ(x, μ) is a vector composed of another set of linearly independent base functions.
[0032] The weight estimation value is as follows:
[0033]
[0034]
[0035] Wherein is the weight vector estimation value corresponding to the subsystem v and Ψ(x, μ); is the weight vector estimation value corresponding to Φ(x); Time t r satisfies the condition: v 0 (x(t r )) = v; Θ(x) is a vector composed of a set of linearly independent base functions, L v (s), v ∈ V is the weight vector corresponding to the subsystem v and Θ(x); Q is a positive definite function.
[0036] Further, the estimation value of the weight derivative is as follows:
[0037]
[0038] Wherein Time t r may be any one of the selected time points t1, t2, …, t l .
[0039] Step 2: Set the initial switching control strategy (v 0 (x), μ0 (x)) and the initial weight vector W(0), at time s=0.
[0040] Step 3: Apply the switching control strategy (v 0 (x),μ 0 (x)) Switch systems and then obtain status data. (Settings) By definition, using state data, Φ(t) is calculated for r = 1, 2, ..., l. r ), and g(t) r (for all v∈V), then use Φ(t) r ), and g(t) r )Calculate H for all v∈V v And K.
[0041] Step 4: Calculate for all v∈V using equations (I) and (II) respectively. and
[0042] Step 5: Calculate using equation (III) And calculate
[0043] Step 6: If Then let Then proceed to step 7. Otherwise, let s = s + δs, and return to step 4.
[0044] Step 7: Calculate the approximate optimal cost Near-optimal switching control strategy:
[0045]
[0046] Beneficial effects:
[0047] (1) This invention first determines the optimal control strategy for switching control that minimizes the cost of the switching system in the infinite time domain; then, it constructs the optimal solution of the infinite time domain HJB equation based on the finite-time domain HJB equation. Specifically, it derives the expression for the optimal solution based on the finite-time domain HJB equation, and starting from a certain positive number, obtains the approximation of the value function based on the partial derivative; simultaneously, it introduces an approximation function in the form of a basis function multiplied by the weights to replace the unknowns in the approximation; the weights of the approximation function in the approximation can then be estimated using the state data matrix; finally, it continuously updates the weight estimates until the approximate optimal weights are obtained, and then substitutes them into the infinite time domain HJB equation to calculate the optimal cost and the optimal switching control strategy. This method only requires state data and does not require a subsystem model to achieve optimal switching control of the switching system. It is independent of the system model and is suitable for situations where the subsystem model of the switching system is unknown.
[0048] (2) The application adopts an infinite time region cost function expression mode, and can be applied to more actual application scenarios.
[0049] (3) The finite time region HJB equation is directly derived from the infinite time region HJB equation, the derivation process of the finite time region HJB equation is simplified, and convenience and speed are achieved. BRIEF DESCRIPTION OF DRAWINGS
[0050] Figure 1 It is a value comparison graph of MATLAB simulation example 1.
[0051] Figure 2 It is a state trajectory comparison graph of MATLAB simulation example 1.
[0052] Figure 3 It is a switching control strategy comparison graph of MATLAB simulation example 1.
[0053] Figure 4 It is a state trajectory comparison graph of MATLAB simulation example 2 online application algorithm.
[0054] Figure 5 It is a state trajectory comparison graph of MATLAB simulation example 2 from the initial time application obtained switching control strategy.
[0055] Figure 6 It is a flow chart of the application. DETAILED DESCRIPTION
[0056] The application will be described in detail below with reference to the drawings and examples.
[0057] The application provides a data-driven optimal switching control method of a switching system.
[0058] Consider the following switching system with continuous time subsystems:
[0059]
[0060] Wherein, x(t)∈R n is a system state, the state quantity is measurable; v∈V represents an index of a currently active subsystem; V={1,2,...,N} is an index set of all subsystems, and N is a number of subsystems; f v :R n ×R m →R n is an unknown model of the subsystem v. x(0)=x0∈Ω is an initial state, is a state region to be studied of the application, including an origin.
[0061] The present application designs optimal switching control strategy by using data of each subsystem, so that the switching system (1) has minimum cost in infinite time region.
[0062] Therefore, we need to define the infinite time region cost function as follows:
[0063]
[0064] Where Q: R n ×R m →R is a positive definite function. The control input can be defined as state feedback in the form of: μ (x (t) ) = u (t).
[0065] In the optimization process, the cost from the state x (t) ∈Ω at time t to infinite time region is defined as:
[0066]
[0067] Satisfies:
[0068]
[0069] The present application is aimed at the case where the subsystem model is unknown. The whole optimal switching control method includes relationship derivation based on HJB equation, estimation of approximate function weight, and data-driven optimal switching control algorithm. Firstly, the optimal solution based on finite time domain HJB equation is derived, which starts from a positive definite function, and continuously updates the value function according to the partial derivative; and the approximate function is introduced to replace the unknown quantity, and the approximate function weight is estimated by using the relationship and data matrix; finally, the weight estimation value is continuously updated until the approximate optimal value is obtained, and the optimal cost and optimal switching control strategy are calculated. This method only needs state data, without the need of subsystem model, and can realize the optimal switching control of the switching system. The following will be described in detail.
[0070] (1) Relationship derivation based on HJB equation.
[0071] The infinite time domain HJB equation is:
[0072]
[0073] Where x = x (t), when (·) is a scalar, The corresponding optimal switching control strategy (v * , μ * ) is:
[0074]
[0075] The finite time domain HJB equation is:
[0076]
[0077] its corresponding optimal switching control strategy (v * ,μ * ) is:
[0078]
[0079] On this basis, define function VF1(·,s) as follows:
[0080]
[0081] Its initial condition is VF1(·,0) = V 0 (·), where V 0 (·) is a positive definite function, and has
[0082]
[0083] The optimal switching control strategy can be transformed as:
[0084]
[0085] where The optimal solution is based on the finite time domain HJB equation, from a positive definite function, according to the partial derivative to update the value function constantly.
[0086] In view of the function VF1(x,s) is unknown, for auxiliary calculation, the function VF1(x,s) is replaced by the following form:
[0087] VF1(x,s) = W(s)Φ(x) + e Φ (x,s) (12)
[0088] where Φ(x) = [Φ1(x),Φ2(x),...,Φ Nw (x)] T is a vector composed of a set of linearly independent basis functions, and the basis functions are Φ j (x):R n →R(j = 1,2,...,N w ), is the weight vector, and e Φ (x,s) is the approximation error. N w is the number of basis functions, the greater the value, the smaller the approximation error, and when the value is large enough, it can guarantee that the error approaches zero. A set of basis functions can constitute a basis of a function space and can almost approximate any function in the function space. When the subsystem model f v (x,μ), v∈V is unknown, for all (x,μ) ∈Ω xu , introduce another approximate function to represent the unknown variable
[0089]
[0090] where is a vector composed of a set of linearly independent basis functions, and the basis functions are j (x, μ): R n × R m → R (j = 1, 2,..., N c ), C v (s), v e V is the weight vector corresponding to the subsystem v, e Ψ,v (x, s, μ) is the approximation error. N c is the number of basis functions Ψ j (x, μ), the greater the value, the smaller the approximation error, and when the value is large enough to ensure that the error is approximated to zero.
[0091] Substituting the approximation functions (12) and (13) into equation (9) gives:
[0092]
[0093] Substituting equation (12) into (13) and integrating gives
[0094]
[0095] where When (·) = [(·)1, (·)2,..., (·) m ] T e R m ,
[0096] (2) Estimation of the weight of the approximation function
[0097] In order to solve each unknown weight vector from (14) and (15), a large amount of data of the obtained state trajectory is required, and some state-related data matrices are defined according to equation (14) as follows:
[0098]
[0099]
[0100] where t1 < t2 <... < t l are selected time points, and l is a positive integer, r = 1, 2,..., l, and
[0101] Since δt is very small, the value of v 0 (x) in the time interval [t r , t rwhich can be regarded as constant in the interval (t, t + δt). Using the data matrix, the estimate of C
[0102]
[0103] where t r satisfies the condition: v 0 (x(t r )) = v.
[0104] Using the data matrix, the estimate of C v (s) is obtained by equation (16) as follows:
[0105]
[0106] where t r satisfies the condition: v0(x(t r )) = v.
[0107] In equation (14), the control input is unknown, and an approximation function is introduced to replace it. Let μ 2,v (x, s) be the control input of subsystem v, which can be expressed as:
[0108] μ 2,v (x, s) = L v (s) Θ(x) + e Θ,v (x, s) (18)
[0109] where Θ(x) = [Θ1(x), Θ2(x),..., Θ Nl (x)] T is a vector composed of linearly independent basis functions, and Θ j (x): R n → R m (j = 1, 2,..., N l ) is a basis function, is a weight vector, and e Θ,v (x, s) is an approximation error, and N l is the number of basis functions Θ j .
[0110] Using the approximation equation (18) and according to equation (14), the corresponding weight L 1,v (s) is obtained when the right side takes the minimum value:
[0111]
[0112] Further, the estimate of C and is obtained by equation (14) as follows:
[0113]
[0114] where time t r may be any one of the selected time t1, t2,..., t l starting from the initial value is calculated by equation (20) where The value of satisfies: is positive definite. Based on the optimal solution of the finite time domain HJB equation, the approximate function is introduced to replace the unknown quantity, and the approximate function weight is estimated by using the relationship between them and the data matrix.
[0115] (3) Data-driven optimal switching control algorithm.
[0116] Based on the above derivation and estimation, the weight estimate value is continuously updated until the approximate optimal value is obtained, and the optimal cost and optimal switching control strategy are calculated. The data-driven optimal switching control algorithm can be obtained as follows:
[0117] a) Set the initial switching control strategy (v 0 (x), μ 0 (x)) and the initial weight vector W(0), and set the time s = 0.
[0118] b) Apply the switching control strategy (v 0 (x), μ 0 (x)) to the switching system, and then obtain the state data. Set According to the definition, using the state data, calculate Φ(t r ), g(t r ) for r = 1, 2,..., l, and then use Φ(t r ), g(t r ) to calculate H v and K for all v ∈ V. c) For all v ∈ V, calculate and
[0119]
[0120] d) Calculate by equation (20), and calculate
[0121] e) If , then and exit. Otherwise, let s = s + δs, and return to c).
[0122] The For the approximate optimal cost, the approximate optimal switching control policy is The algorithm uses the data generated by the initial switching control policy, starts with a certain positive definite function, and gradually approaches the optimal solution by constantly updating the weight and its derivative, without prior knowledge of the subsystem model.
[0123] The following uses matlab to simulate the effectiveness of the above method, the following two examples are used to illustrate.
[0124] Example 1: First consider a relatively simple switching system with two continuous autonomous subsystems:
[0125]
[0126] The system parameters are set as: x(0) = 2 and Q(x(t)) = x 2 (t). The goal is to find the optimal switching control policy to minimize the cost function. According to the literature, the optimal switching policy is
[0127]
[0128] The initial switching control policy is chosen as v 0 (x) = (-1) floor(x / 0.03) / 2 + 1.5, and the sampling period is δt = 0.002s. Apply the data-driven optimal switching control method, use the state data from t = 0 to t = 0.5s, then after 0.5s, complete 277 iterations, and obtain the approximate optimal cost and its corresponding approximate optimal switching control policy.
[0129] Figure 1 The cost corresponding to the initial switching control policy is shown The approximate optimal cost is shown And the optimal cost V * , obviously the approximate optimal cost is very close to the optimal cost V * . Figure 2 The state trajectory of the initial switching control policy v 0 , the approximate optimal switching control policy and the optimal switching control policy v * is shown, where the state trajectory of almost coincides with the state trajectory of v * . Figure 3 The initial switching control policy v 0 , the approximate optimal switching control policy and the optimal switching control policy v * are shown in the interval x ∈ [-2, 2], where the approximate optimal switching control policy with the optimal switching control policy v * are almost the same. The simulation verifies the effectiveness of the proposed method.
[0130] Example 2: Consider a switching system with two continuous subsystems:
[0131]
[0132] The system parameters are set as: x(0)=[1,-1] and The goal is to find the optimal switching control policy so as to minimize the cost function.
[0133] The initial switching control policy is chosen as: when v 0 (x)=2; when v 0 (x)=1, and μ 0 (x)=-x1+x2. The sampling period is δt=0.02s. The data-driven optimal switching control method is applied, using the state data from t=0 to t=10s, and then after 5s, 44 iterations are completed to obtain the approximate optimal cost of the system and the corresponding approximate optimal switching control policy, and then the approximate optimal switching control policy is applied to the system.
[0134] Figure 4 The state trajectory of always applying the initial switching control policy (v 0 (·), μ 0 (·)) and applying the approximate optimal switching control policy after t=15s is shown, wherein the state trajectory of applying converges to the origin quickly after t=15s, while the state trajectory of applying (v 0 (·), μ 0 (·)) oscillates around the origin continuously. Figure 5 The state trajectory graph of applying (v 0 (·), μ 0 (·)) and from t=0 in the system is shown, and the effect is similar to Figure 4 . The simulation verifies the effectiveness of the method.
[0135] In summary, the above is only a preferred embodiment of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A data-driven optimal switching control method for a switching system, characterized by, Design the optimal control strategy to minimize the cost of switching the system over an infinite time region; An expression for the infinite time-domain optimal control strategy (v, μ) is constructed using the infinite time-domain HJB equations of the switching system; where v in the HJB equations represents the subsystem of the switching system, and μ represents the control quantity of the subsystem. A finite-time domain optimal control strategy expression is constructed using the finite-time domain HJB equations of the switching system. Based on the optimal solution expression of the finite-time domain HJB equations, starting from the positive definite function, an approximation of the value function is obtained by using partial derivatives. Some unknowns in the approximation are replaced by approximation functions, where the unknowns include the control variable μ and the value function. The value function is the region cost, and the approximation function is the product of the basis function and the weights. The weights in the approximation are estimated using the state data matrix to obtain approximately optimal weights. The optimal value function is then obtained based on the approximately optimal weights. Substituting the optimal value function into the expression for the optimal control strategy in the infinite time domain yields the optimal control strategy.
2. The data-driven optimal switching control method of a switching system according to claim 1, wherein, Switching to the following system: where x(t) ∈ R n is the system state, which is measurable; u(t) is the system control input; v ∈ V represents the index of the current active subsystem, V = {1, 2,..., N} is the index set of all subsystems, and N is the number of subsystems; f v : R n × R m → R n is the unknown model of subsystem v; x(0) = x0 ∈ Ω is the initial state, is the state region to be studied, including the origin; The cost function for an infinite time region is: where Q:R n xR m R is a positive definite function; the control input is defined in the form of state feedback, i.e.: μ(x(t)) = u(t); During the optimization process, the cost from the state x(t)∈Ω at time t to the infinite time region is defined as: satisfy:
3. The data-driven optimal switching control method for a switching system as described in claim 2, characterized in that, The infinite-time domain HJB equation is: in The superscript * represents the optimal value, x = x(t), when (·) is a scalar. The corresponding optimal switching control strategy (v * ,μ * )for:
4. The data-driven optimal switching control method for the switching system as described in claim 3, characterized in that, The finite-time HJB equation is: Wherein, from time t and state x(t) to time t f The finite-time cost function VF1(x,s) is: VF1(x,s)=W(s)Φ(x)+e Φ (x,s), where Let be a vector consisting of a set of linearly independent basis functions, where the basis functions are Φ. j (x):R n →R(j=1,2,...,N w ), It is a weight vector, e Φ (x,s) is the approximation error, N w It is the number of basis functions; The control input μ of subsystem v is denoted as μ 2,v (x,s): μ 2,v (x,s)=L v (s)Θ(x)+e Θ,v (x,s), where, v∈V is a weight vector; It is a vector composed of linearly independent basis functions, Θ j (x):R n →R m (j = 1, 2, ..., N) l ) are basis functions; e Θ,v (x,s) is the approximation error; N l It is the basis function Θ j Quantity; Its corresponding optimal switching control strategy for:
5. The data-driven optimal switching control method for a switching system as described in claim 4, characterized in that, Includes the following steps: Step 1: Based on the optimal solution relation of the HJB equation, where the unknowns are represented by the sum of basis function weights and the approximation error is ignored, a data matrix is introduced: where t1 < t2 <... < tl l are selected time points, / is a positive integer, r = 1, 2,... / ; Φ(x) is a vector composed of a set of linearly independent basis functions; Ψ(x, μ) is a vector composed of another set of linearly independent basis functions; The formula for obtaining the weight estimate is as follows: in It is the estimated value of the weight vector corresponding to subsystem v and Ψ(x,μ); This is an estimate of W(s); Time t r Conditions met: v 0 (x(t r ))=v;Q is a positive definite function; Then use the estimated value and Obtain the weighted derivative The estimated values are as follows: in Time t r For selected times t1, t2, ..., t l Any one of them; Step 2: Set initial handover control policy (v 0 (x), μ 0 (x)) and initial weight vector W(0), time s = 0; Step 3: Apply the switching control strategy (v 0 (x),μ 0 (x)) Switching systems and then acquiring status data; setting Calculate Φ(t) using state data r ), and g(t) r Then use Φ(t) r ), and g(t) r )Calculate H for all v∈V v and K; Step 4: Calculate for all v∈V using equations (I) and (II) respectively. and Step 5: Calculate using equation (III) And calculate Step 6: If Then let And proceed to step 7; otherwise, let s = s + δs, and return to step 4; Step 7: Calculate the approximate optimal cost Near-optimal switching control strategy:
Citation Information
Patent Citations
Model-free optimal switching method of switching system
CN110262235A