Approximate dynamic programming control method based on support vector regression

By applying support vector regression technology and approximate dynamic programming methods in energy systems, the problem that traditional optimal control methods are difficult to cope with complex energy systems is solved, and more efficient and stable control strategy approximation is achieved.

CN120161728AActive Publication Date: 2025-06-17NORTHEASTERN UNIV CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510637337.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-17
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

Traditional optimal control methods are difficult to effectively deal with complexity and dynamic changes in large-scale or high-dimensional energy systems, especially in the face of nonlinearity and uncertainty.

Method used

The approximate dynamic programming control method based on support vector regression is adopted to build the state space of a nonlinear dynamic system, and a continuous time iteration algorithm and hard ε-bonded support vector regressor (HESVR) are used to approximate the value function and control strategy.

Benefits of technology

It improves the approximation accuracy of the optimal control strategy and the stability of the algorithm, can handle complex and nonlinear systems more effectively, and is suitable for real-time control of large-scale energy systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120161728A_ABST
    Figure CN120161728A_ABST
Patent Text Reader

Abstract

The invention provides an approximate dynamic programming control method based on support vector regression, relates to the technical field of approximate dynamic programming, and provides a method for solving the optimal control problem of a continuous nonlinear system. The method integrates a continuous time IRL Bellman equation and a support vector regression SVR technology; the use of IRL allows the algorithm to estimate a value function through an integral term under the condition that part of system dynamics is unknown, a powerful tool is provided for online learning, SVR is applied to ADP for the first time, a complex function approximation problem is converted into a convex optimization problem, and the existence of an optimal solution and the stability of the algorithm are ensured. By adopting the HESVR, the algorithm can more accurately approach the cost function and the control strategy, and meanwhile, the weight of the HESVR is adjusted by utilizing the least square method, so that the efficiency and the accuracy of the algorithm are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of approximate dynamic programming, and specifically to an approximate dynamic programming control method based on support vector regression. Background Art

[0002] The optimal control problem, as a key branch in control theory, aims to find a control strategy that can maximize or minimize the system performance index under specific constraints. Since the mid-20th century, with the continuous progress of space technology, automation, and various engineering applications, it has greatly promoted the development of optimal control theory and played an important role in many fields. The application of optimal control strategies can be found everywhere, from the trajectory optimization of spacecraft to the precise regulation of industrial processes.

[0003] In the energy field, especially for large-scale or high-dimensional energy systems such as smart grids, wind farms, or solar power generation systems, traditional control strategies often struggle to handle their complexity and dynamic changes. For example, a smart grid needs to respond in real-time to supply and demand changes, power quality control, and the volatility of renewable energy. These systems are usually highly nonlinear and uncertain, with strict requirements for the response time of control algorithms. With the growth of modern system complexity, especially the emergence of nonlinear systems, it has become difficult to use traditional linear optimal control methods to handle problems. Due to the inherent complex dynamic characteristics of nonlinear systems, the design of control strategies has become extremely difficult. The Hamilton-Jacobi-Bellman (HJB) equation is a nonlinear partial differential equation, and its solution can provide an optimal feedback control law for the system. However, since the analytical solution of the HJB equation is often difficult to obtain in high-dimensional spaces or complex dynamic systems, this has posed a major obstacle to solving the optimal control problem.

[0004] To address this challenge, researchers have developed a series of approximate analytical methods: fuzzy logic control, which uses fuzzy set theory to handle the uncertainty and ambiguity of the system and approximates the optimal control strategy through fuzzy rules. Neural network control utilizes its powerful nonlinear mapping ability to learn the control strategy through training data. ADP is a more direct numerical method that iteratively approximates the solution of the HJB equation to solve for the optimal control strategy.

[0005] The ADP method is based on the principles of reinforcement learning and dynamic programming, and realizes the gradual optimization of the control strategy through policy iteration or value iteration algorithms. The policy iteration algorithm starts from a good initial policy and gradually approaches the optimal policy through successive policy evaluation and improvement steps. The value iteration algorithm, on the other hand, starts from an initial value function and gradually improves the estimate of the value function through iterative updates until it converges to the optimal value function. These methods have achieved certain success in discrete-time or some specific types of continuous-time systems, but when dealing with high-dimensional or extremely complex dynamic systems, they still face problems such as slow convergence speed and high computational costs. In addition, existing approximate analytical methods also face some other challenges in practical applications. For example, although fuzzy logic control can handle uncertainties, its design process often relies on experience and is difficult to adapt to rapidly changing or unforeseen system behaviors. Although neural network control has good generalization ability, its training process requires a large amount of data, and the choice of network structure has a significant impact on performance. Although the ADP method is theoretically attractive, in practical applications, issues such as how to select appropriate iterative algorithms, how to design approximators, and how to ensure the convergence and stability of the algorithms still need further study.

[0006] As a core topic in control theory, the exploration of solutions to the optimal control problem has always been the focus of engineering and scientific research. In this field, methods such as dynamic programming, fuzzy logic control, neural network control, and ADP have been successively proposed and applied to various control scenarios. However, in the energy context, especially in systems containing renewable energy, problems of model uncertainty are often faced. The unpredictability of wind speed and solar radiation can lead to fluctuations in energy output, while traditional control methods are often based on deterministic models and are difficult to adapt to this uncertainty.

[0007] Dynamic programming is a classical method for solving optimal control strategies. It solves complex problems by decomposing them into a series of simple problems. As the scale of the system increases, such as the expansion of the power grid or the integration of multiple energy types, the dimension of the state space increases sharply, resulting in the traditional dynamic programming-based methods facing the "curse of dimensionality", and the computational resources and storage space requirements increase exponentially. Dynamic programming algorithms usually require prior knowledge of all the information of the system, which is often difficult to meet in practical applications.

[0008] As an effective means of dealing with uncertainty problems, fuzzy logic control describes the state and control rules of the system through fuzzy set theory, thus realizing the handling of uncertainty and fuzziness. Although fuzzy logic control performs well in some nonlinear or difficult-to-model systems, its generalization ability is limited and it is often difficult to adapt to system behaviors that exceed the designer's expectations. At the same time, the design of fuzzy logic controllers usually relies on expert experience, which limits its application in a wider range of fields.

[0009] Neural network control utilizes the powerful non - linear mapping ability of neural networks, which can approximate complex functional relationships to achieve effective control of the system. However, the robustness and model transparency of neural network control are insufficient, which limits its application in many application scenarios with extremely high requirements for safety and reliability. In addition, the training of neural networks usually requires a large amount of data, and the choice of network structure has a significant impact on control performance, which increases the difficulty of design and application.

[0010] As a numerical method that combines the ideas of reinforcement learning and dynamic programming, ADP solves the optimal control problem by iteratively approximating the HJB equation. The policy iteration (PI) algorithm in the ADP method needs to start from a good initial policy and approximate the optimal policy through successive policy evaluation and improvement steps. However, the PI algorithm may be difficult to converge to the global optimal solution, especially in non - convex cases or when there are multiple local optimal solutions. And the convergence speed of the PI algorithm may be slow and sensitive to the choice of the initial policy.

[0011] The value iteration algorithm, as another form of the ADP method, can start from any initial value and approximate the optimal value function by iteratively updating the value function. However, the value iteration algorithm may face problems of stability and convergence speed, especially in continuous - time or high - dimensional space systems. In addition, the value iteration algorithm may require a large number of iterations in practical applications, which undoubtedly increases the computational cost.

[0012] It should be noted that the existing methods also have limitations in terms of computational resources and real - time performance. As the system scale increases, the required computational resources may make the methods infeasible, especially in real - time systems that require fast response. Real - time systems have strict requirements for the response time of control algorithms, and the existing methods may be difficult to meet these requirements due to high computational complexity. The existing methods also face challenges of model uncertainty and external disturbances in practical applications. In practical engineering applications, system models often have uncertainties and are vulnerable to external environmental disturbances. The existing methods may perform poorly in dealing with these uncertainties and disturbances and need to be further improved to enhance their robustness. Summary of the Invention

[0013] Aiming at the deficiencies of the existing technology, the object of the present invention is to propose an approximate dynamic programming control method based on support vector regression, including: Step 1: Construct the state space of a non - linear dynamic system with a time - invariant affine structure, and the state space of the non - linear dynamic system is expressed as: ; where t represents the time, Denote the state space of the nonlinear dynamic system, Denote the state vector of the nonlinear dynamic system at t time instant, Denote the nonlinear drift term, Denote the control gain matrix, Denote t the control input vector at time instant; Step 2: Based on the state space of the nonlinear dynamic system, construct the formulas of the value function and the control strategy on the continuous-time iteration algorithm; Among them, the formula of the value function is expressed as: ; Among them, Denote the state value of the nonlinear dynamic system at the state vector i at the +1-th iteration, that is, the terminal state value, T Denote the transpose of the matrix, Denote the state vector of the nonlinear dynamic system at time instant, Denote the optimal control strategy of the nonlinear dynamic system at the state at the i-th iteration, R is a symmetric positive definite control weight matrix, and Q is a state positive definite function; Among them, the formula of the control strategy is expressed as: ; Among them, Denote the optimal control strategy of the nonlinear dynamic system at the state vector i at the -th iteration, Denote the set of admissible control strategies on the state space , is a preset time interval, T Denote the transpose, Denote the state vector of the nonlinear dynamic system at time instant, Denote the control input vector at time instant, Denote the state value of the nonlinear dynamic system at the state vector i at the -th iteration; Step 3: According to the formulas of the value function and the control strategy, determine the final optimal control strategy of the nonlinear dynamic system and the final state value of the nonlinear dynamic system.

[0014] Optionally, step 3 specifically includes iterating based on the formulas of the value function and the control strategy to obtain the final optimal control strategy of the nonlinear dynamic system and the final state value of the nonlinear dynamic system, which specifically includes the following steps: Step A1: Set the initial state value , set the initial iteration number i = 0, take the initial iteration number as the current iteration number, and take the initial state value as the state value of the current iteration number ; Step A2: Substitute the state value of the current iteration number into the formula of the control strategy, and calculate the optimal control strategy of the current iteration number ; Step A3: Substitute the optimal control strategy of the current iteration number into the formula of the value function, and calculate the state value of the next iteration of the current iteration number ; Step A4: Calculate the first difference quantity , which is specifically implemented through the following formula: ; Step A5: Determine whether the first difference quantity is less than the preset threshold ε. In the case where the first difference quantity is less than the preset threshold ε, take as the final state value, and take the optimal control strategy of the current iteration number as the final optimal control strategy. In the case where the first difference quantity is not less than the preset threshold ε, determine whether the current iteration number is less than or equal to the preset number. In the case where the current iteration number is less than or equal to the preset number, take as the final state value, and take the optimal control strategy of the current iteration number as the final optimal control strategy. In the case where the current iteration number is greater than the preset number, add one to the current iteration number as the new current iteration number, and return to execute step 3.2.

[0015] Optionally, step 3 specifically includes improving the formulas of the value function and the control strategy, and then iterating based on the improved formulas to obtain the final optimal control strategy of the nonlinear dynamic system and the final state value of the nonlinear dynamic system, which specifically includes the following steps: Step B1: Improve the formulas of the value function and the control strategy based on the hard ε-key support vector regression machine HESVR to obtain the value function formula based on HESVR and the control strategy formula based on HESVR; Step B2: Process the value function formula based on HESVR using the least squares method to obtain the weight formula of HESVR; Step B3: According to the weight formula of HESVR, the value function formula based on HESVR, and the control strategy formula based on HESVR, calculate the optimal control strategy of the nonlinear dynamic system and the final state value of the nonlinear dynamic system through an iterative algorithm.

[0016] Optionally, the value function formula based on HESVR in Step B1 is expressed as: ; where is the state value based on HESVR, is the approximate state value of the nonlinear dynamic system at the state vector i at the -th iteration, is the nonlinear basis function vector at the state .

[0017] Optionally, the control strategy formula based on HESVR in Step B1 is expressed as: ; where represents the optimal control strategy based on HESVR, is the inverse matrix of the control weight, is the system dynamic gradient, is the value function gradient at the i -th iteration, represents the transpose of the gradient matrix of the nonlinear basis function vector at the state x ( t ), is the weight vector obtained through HESVR training in the i -th iteration.

[0018] Optionally, the weight formula of HESVR in Step B2 is expressed as: ; where ; ; where is the nonlinear basis function vector at the state , is the nonlinear basis function vector at the state , and M is the number of sampling points in the least squares method. For the non - linear basis function vector at state , For the total trajectory cost from state x ( t ) when executing the control policy within time. For the total trajectory cost from state when executing the control policy within time.

[0019] Optionally, step B3 specifically includes: Step B3.1: Set the initial HESVR weight , set the initial state value , set the initial iteration number i = 0, take the initial iteration number as the current iteration number, take the initial HESVR weight as the current HESVR weight , take the initial state value as the state value of the current iteration number ; Step B3.2: Determine whether the current iteration number is greater than the preset number . If the current iteration number is greater than the preset number , execute step 3.3.9; if the current iteration number is not greater than the preset number , execute step 3.3.3; Step B3.3: Substitute the HESVR weight into the control policy formula based on HESVR, and calculate the optimal control policy based on HESVR for the current iteration number ; Step B3.4: Obtain the state vector through the Runge - Kutta algorithm; Step B3.5: Substitute the state vector into the weight formula of HESVR, and calculate the updated weight of HESVR; Step B3.6: Substitute the updated weight of HESVR into the value function formula based on HESVR, and calculate the state value based on HESVR for the next iteration number of the current iteration number ; Step B3.7: Calculate the second difference quantity , which is specifically implemented through the following formula: ; Step B3.8: Determine whether the second difference quantity is greater than the preset threshold ε. If the second difference quantity When it is greater than the preset threshold ε, step 3.3.9 is executed, and in the second difference When it is not greater than the preset threshold ε, the current iteration count is incremented by one as the new current iteration count, and step 3.3.2 is returned for execution; Step B3.9: Take as the final state value, and take the optimal control strategy based on HESVR of the current iteration count as the final optimal control strategy.

[0020] The beneficial effects produced by adopting the above technical solution are as follows: The present invention proposes a method for the optimal control problem of continuous nonlinear systems. This method combines the continuous-time IRL Bellman equation and the support vector regression SVR technology to form a novel numerical iteration method. The use of IRL allows the algorithm to estimate the value function through the integral term when part of the system dynamics is unknown, providing a powerful tool for online learning. And the present invention first applies SVR to ADP, transforming the complex function approximation problem into a convex optimization problem, ensuring the existence of the optimal solution and the stability of the algorithm. By adopting HESVR, the algorithm can approximate the cost function and control strategy more accurately, and at the same time adjusts the weights of HESVR using the least squares method, improving the efficiency and accuracy of the algorithm. Description of the Drawings

[0021] Figure 1 Schematic flowchart of an approximate dynamic programming control method based on support vector regression in an embodiment of the present invention. Detailed Embodiments

[0022] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0023] Aiming at the problems existing in the prior art, the purpose of the present invention is to develop an innovative control strategy to effectively solve the optimal control problem of continuous-time nonlinear systems. In this problem, especially for large-scale smart grids or complex energy networks integrating multiple energy sources, traditional control methods often have difficulty coping with the system complexity and dynamic characteristics. To solve this problem, the present invention adopts a hard ε-kernel support vector regression machine (HESVR) to approximate the cost function and control strategy of continuous-time systems.

[0024] As an advanced regression technique, HESVR demonstrates significant advantages in handling the cost-benefit analysis and control decisions of energy systems due to its superior generalization ability and robustness. By introducing HESVR, the present invention can accurately capture the cost function and control strategy of the system, and maintain a high-precision approximation even in the face of highly nonlinear system behaviors. Furthermore, the present invention combines the SVR technique with the ADP algorithm to propose a new numerical iteration method. As an approximate method of dynamic programming, the iterative characteristics of the ADP algorithm combined with the approximation ability of HESVR form a powerful tool for solving the optimal control problem of continuous-time nonlinear systems. This method not only improves the approximation accuracy of the optimal control strategy, but also gradually optimizes the control strategy through the iterative process until the optimum is reached.

[0025] The numerical iteration method proposed by the present invention pays particular attention to convergence and stability. Convergence ensures that the algorithm can find a satisfactory control strategy within a finite number of iterations, while stability means that the obtained control strategy remains effective even in the case of dynamic changes in the energy system or the presence of external disturbances. This dual focus on convergence and stability makes the present invention more reliable and practical in practical applications, and can be widely applied to different types of energy systems, including but not limited to power systems, energy production, distribution, and storage systems. Whether it is in improving energy efficiency, optimizing energy distribution, or promoting the integration of renewable energy, the present invention can provide effective optimal control strategies to promote the intelligent and sustainable development of energy systems.

[0026] Specifically, the present invention provides an approximate dynamic programming control method based on support vector regression, combined with Figure 1 , which may include the following steps: Step 1: Construct the state space of a nonlinear dynamic system with a time-invariant affine structure, where the nonlinear dynamic system can be a mechanical system, and the state space of the nonlinear dynamic system is represented as: ; where t represents the time instant, represents the state space of the nonlinear dynamic system, , represents the system state vector (such as the position / speed of a mechanical system, the voltage / frequency of a power grid, etc.), represents the state vector of the nonlinear dynamic system at t the time instant, represents the nonlinear drift term, where , represents the control gain matrix, where , represents tThe control input vector at a moment, , where m is the dimension of the control output, representing the control input vector (such as motor torque, output of reactive power compensation device, etc.); Among them, satisfies the Lipschitz continuity condition on the compact set , where n is the dimension of the state variable, and represent the nonlinear drift term and the control gain matrix respectively, describing the natural dynamics of the system and characterizing the regulation ability of the actuator, and there exists an equilibrium point (that is ). It is assumed that the system satisfies the controllability condition, that is, there exists at least one continuous control law , which can achieve the asymptotic stability of the closed-loop system within .

[0027] Step 2: Based on the state space of the nonlinear dynamic system, construct the formulas of the value function and the control strategy on the continuous-time iterative algorithm; Among them, the formula of the value function is expressed as: ; Among them, represents the state value of the nonlinear dynamic system at the state vector i at the +1th iteration, that is, the terminal state value, T represents the transpose of the matrix, represents the state vector of the nonlinear dynamic system at moment, represents the optimal control strategy of the nonlinear dynamic system at the state at the i-th iteration, R is a symmetric positive definite control weight matrix, R is used to represent the control weight, Q is a state positive definite function, and Q is used to represent the state weight; Among them, the formula of the control strategy is expressed as: ; Among them, represents the optimal control strategy of the nonlinear dynamic system at the state vector i at the th iteration, represents the set of admissible control strategies on the state space , is a preset time interval, T represents the transpose, represents the state vector of the nonlinear dynamic system at moment, represents the control input vector at moment, Denote the state value of the non - linear dynamic system at the state vector i at the -th iteration; Among them, the formulas of the value function and the control strategy are derived and constructed through the following steps: Define the infinite - horizon performance index of the system as: ; Among them, is the initial state of the system at the initial time \(t = 0\), is a positive - definite state function, representing the state weight, is a symmetric positive - definite control - weight matrix, representing the control weight. The core of this optimal control problem lies in solving the control strategy that satisfies the following conditions: 1. Ensure the asymptotic stability of the closed - loop system; 2. Make the performance index reach the global minimum.

[0028] For the admissible control strategy , its mathematical characteristics should satisfy: 1. The control law is continuously differentiable on ; 2. Satisfy at the equilibrium point; 3. The value function is finite.

[0029] Based on the value function, given an admissible strategy , for any initial time \(t\) and finite time interval \(T\), the value function (i.e., the performance index of the system) can be constructed in the form of rolling - horizon optimization in the time domain: ; Among them, represents the state - tracking cost, represents the control - energy consumption cost, represents the terminal cost, represents the zero - initial condition, and in the integral term is an arbitrary time (dynamic time variable) within the rolling horizon, \(T\) represents the vector / matrix transpose (e.g., is the transpose of the state vector).

[0030] This formula essentially constructs the Bellman equation under the framework of continuous - time inverse reinforcement learning, and its core feature is to combine the finite - horizon cost and the terminal cost. According to the Bellman optimality principle, the optimal value function needs to satisfy the recursive minimization condition: ; Among them, represents the state - tracking cost, represents the control - energy consumption cost, represents the optimal terminal cost.

[0031] The explicitly derived optimal control strategy can be expressed in a dual optimization form: ; ; To obtain the optimal solutions of the above three formulas, the continuous-time value iteration algorithm will iterate in the following two steps. Given the initial value function , for the iteration number the algorithm performs: (Policy improvement) For convenience of distinction, the present invention formulates the control law as follows: ; where, represents the optimal control strategy of the system at state at the i-th iteration, which is consistent with the definition of .

[0032] represents the finite-horizon cost, represents the terminal cost at the i-th iteration represents the inverse matrix of the control weight, represents the system dynamic gradient, represents the value function gradient.

[0033] (Value update) The value function estimation method is: ; where, represents the receding horizon optimization cost represents the value function of the previous iteration, and thus it is deduced that: ; where, represents the actual trajectory cost, represents the terminal state value.

[0034] Changing the integration range of the above formula from [t, t + to [t, , there is no terminal state value at this time, and finally the formula of the value function is obtained: .

[0035] Step 3: Determine the final optimal control strategy of the nonlinear dynamic system and the final state value of the nonlinear dynamic system according to the formula of the value function and the formula of the control strategy.

[0036] In the first implementation manner, step 3 specifically includes iterating based on the formula of the value function and the formula of the control strategy to obtain the final optimal control strategy of the nonlinear dynamic system and the final state value of the nonlinear dynamic system, which specifically includes the following steps: Step A1: Set the initial state value , set the initial iteration number i = 0, take the initial iteration number as the current iteration number, and take the initial state value as the state value of the current iteration number ; Step A2: Substitute the state value of the current iteration number into the formula of the control strategy, and calculate the optimal control strategy of the current iteration number ; Step A3: Substitute the optimal control strategy of the current iteration number into the formula of the value function, and calculate the state value of the next iteration of the current iteration number ; Step A4: Calculate the first difference quantity , which is specifically implemented through the following formula: ; Step A5: Determine whether the first difference quantity is less than the preset threshold ε. In the case where the first difference quantity is less than the preset threshold ε, take as the final state value, and take the optimal control strategy of the current iteration number as the final optimal control strategy. In the case where the first difference quantity is not less than the preset threshold ε, determine whether the current iteration number is less than or equal to the preset number. In the case where the current iteration number is less than or equal to the preset number, take as the final state value, and take the optimal control strategy of the current iteration number as the final optimal control strategy. In the case where the current iteration number is greater than the preset number, add 1 to the current iteration number as the new current iteration number, and return to execute step 3.2.

[0037] Among them, the selection conditions of the initial value function: , where represents the set of Lyapunov functions. Convergence guarantee conditions: When the system satisfies , .

[0038] In the first implementation manner, the idea of the IRL method is utilized. In each iteration, the newly updated function can be expressed as the time interval The cost integral term on and the function obtained in the previous iteration step The sum.

[0039] In the second implementation manner, step 3 specifically includes improving the formulas of the value function and the control strategy, and then performing iterations based on the improved formulas to obtain the final optimal control strategy of the nonlinear dynamic system and the final state value of the nonlinear dynamic system. The second implementation manner is an innovative numerical optimization method specifically designed to adjust the weights of HESVR. When facing a complex nonlinear system or requiring online learning, the present invention uses an algorithm that combines SVR to achieve efficient approximation. This algorithm is the core technology for implementing the key steps in continuous-time ADP, aiming to accurately approximate the value function and control strategy of complex systems. The second implementation manner can be regarded as an extension of the first implementation manner, and solves the problem of dependence on model accuracy of traditional methods by introducing SVR. Both are based on the value iteration framework, and the second implementation manner achieves a wider applicability through machine learning.

[0040] Specifically, the second implementation manner specifically includes the following steps: Step B1: Improve the formulas of the value function and the control strategy based on the hard ε-bond support vector regression machine (HESVR) to obtain the value function formula based on HESVR and the control strategy formula based on HESVR; Among them, the value function formula based on HESVR is expressed as: ; Among them, is the state value based on HESVR, is at the i th iteration, the approximate state value of the nonlinear dynamic system at the state vector , is at the state the nonlinear basis function vector at.

[0041] Among them, the control strategy formula based on HESVR is expressed as: ; Among them, represents the optimal control strategy based on HESVR, is the inverse matrix of the control weight, is the system dynamic gradient, is the value function gradient at the i-th iteration, represents at the state x ( tThe transpose of the gradient matrix of the non - linear basis function vector at is the weight vector obtained by training with HESVR in the \(i\) - th iteration.

[0042] Step B2: Based on the least - squares method, process the value function formula based on HESVR to obtain the weight formula of HESVR; Among them, the weight formula of HESVR is expressed as: ; Among them, ; ; Among them, is the non - linear basis function vector at state , is the non - linear basis function vector at state , \(M\) is the number of sampling points in the least - squares method, is the non - linear basis function vector at state , is the total trajectory cost from state x ( t ) starting, executing the control strategy within time, is the total trajectory cost from state starting, executing the control strategy within time.

[0043] Specifically, processing the value function formula based on HESVR to obtain the weight formula of HESVR includes the following steps: Since the value iteration algorithm based on HESVR starts from , so set the initial value to , the control strategy will be approximated by the control strategy formula based on HESVR, will be evaluated by the value function formula based on HESVR, which means that only one HESVR is used in the approximation process. In each iteration process, the HESVR weight is adjusted in the value function formula based on HESVR, and the training sample is known as .

[0044] Note 's rotation, actually, is to find the approximate optimal solution of HESVR with an error of \(\varepsilon\). In the tuning process of HESVR, the least - squares format is adopted, where the system dynamics and According to the value function formula based on HESVR, we have: ; where is the total trajectory cost from state x(t) when executing the control strategy within time.

[0045] In the i th iteration, keep the strategy unchanged, sample M points from the state space with the same sampling time interval t, where the parameter M should satisfy such that can be recognized. That is, in the least squares sense, the weight update of HESVR is , where: ; ; Define the system dynamics parameters f, g and the cost function matrices Q, R; the integration time interval ; the number of samples M for each iteration; the initial state vector ; the non - linear mapping function ; the iteration termination criterion ; the maximum number of iterations ; the approximate optimal cost function sequence ; the approximate optimal control strategy sequence .

[0046] Step B3: According to the weight formula of HESVR, the value function formula based on HESVR and the control strategy formula based on HESVR, calculate the optimal control strategy of the non - linear dynamic system and the final state value of the non - linear dynamic system through an iterative algorithm.

[0047] Step B3.1: Set the initial HESVR weight , set the initial state value , set the initial iteration number i = 0, take the initial iteration number as the current iteration number, take the initial HESVR weight as the current HESVR weight , and take the initial state value as the state value of the current iteration number ; Step B3.2: Judge whether the current iteration number is greater than the preset number . If the current iteration number is greater than the preset number , execute Step 3.3.9; if the current iteration number is not greater than the preset number , execute Step 3.3.3; Step B3.3: Substitute the HESVR weights into the control strategy formula based on HESVR to calculate the optimal control strategy based on HESVR for the current iteration number ; Step B3.4: Obtain the state vector through the Runge-Kutta algorithm; Step B3.5: Substitute the state vector into the weight formula of HESVR to calculate the updated weights of HESVR; Step B3.6: Substitute the updated weights of HESVR into the value function formula based on HESVR to calculate the state value based on HESVR for the next iteration number of the current iteration number ; Step B3.7: Calculate the second difference quantity , which is specifically implemented through the following formula: ; Step B3.8: Determine whether the second difference quantity is greater than the preset threshold ε. In the case where the second difference quantity is greater than the preset threshold ε, execute Step 3.3.9. In the case where the second difference quantity is not greater than the preset threshold ε, add one to the current iteration number as the new current iteration number, and return to execute Step 3.3.2; Step B3.9: Take as the final state value, and take the optimal control strategy based on HESVR for the current iteration number as the final optimal control strategy.

[0048] The solution proposed by the present invention includes two implementation methods. The first implementation method is responsible for the value iteration process based on the continuous-time system, while the second implementation method optimizes the HESVR weights through the least squares method. The second implementation method is an engineering extension of the first implementation method. The two work together, combining the theoretical framework with machine learning tools, solving the computational bottleneck in the actual system, and providing an efficient and stable numerical solution for solving the optimal control problem of continuous-time nonlinear systems.

[0049] By transforming the optimal control problem into a convex optimization problem, the present invention significantly improves the computational efficiency during the solution process, effectively reducing the dependence on computing resources when dealing with large-scale or complex systems. This transformation not only optimizes the computational process but also, through the application of convex optimization theory, ensures the existence and reachability of the global optimal solution, thereby improving the quality and reliability of the solution. Compared with traditional optimal control methods, the present invention achieves a significant simplification in the design process of control strategies. By introducing advanced mathematical programming techniques, the present invention reduces the dependence on complex mathematical modeling and control theory, making the design of control strategies more intuitive and convenient, and applicable to various applications in energy systems, such as load forecasting in smart grids, optimization of energy distribution, control strategies for integration of renewable energy, and management of energy storage systems. Guided by the simulation results, the present invention provides a clear path for the conversion between theoretical algorithms and actual energy control systems, greatly reducing the difficulty from theoretical research to engineering application and accelerating the speed of technology transfer.

[0050] Another advantage of the present invention is its wide applicability and flexibility. It is not limited to a specific system type or structure but can be flexibly applied to different systems, including linear systems, nonlinear systems, discrete-time systems, and continuous-time systems. This generality makes the present invention potentially applicable in multiple fields and industries, such as aerospace, robotics, intelligent manufacturing, economic system optimization, etc. This adaptability not only improves the practicality of the algorithm but also provides more options and flexibility for energy system operators when facing different control challenges, thus better coping with the fluctuations in the energy market and policy changes and achieving efficient, stable, and optimized operation of the energy system.

[0051] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with (but not limited to) the technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. An approximate dynamic programming control method based on support vector regression, characterized in that: include: Step 1: Construct a state space of a nonlinear dynamic system with a time-invariant affine structure. The state space of the nonlinear dynamic system is expressed as: ; in, t Indicates the time, represents the state space of a nonlinear dynamic system, Represents a nonlinear dynamic system in t The state vector at time, represents the nonlinear drift term, represents the control gain matrix, express t The control input vector at time t; Step 2: Based on the state space of the nonlinear dynamic system, the formula of the value function and the formula of the control strategy are constructed on the continuous time iterative algorithm; The formula of the value function is expressed as: ; in, Indicated in i At +1 iteration, the nonlinear dynamic system has a state vector The state value at , that is, the terminal state value, T represents the transpose of a matrix, Represents a nonlinear dynamic system in The state vector at time, Indicates that at the i-th iteration, the nonlinear dynamic system is in the state The optimal control strategy at , R is the symmetric positive definite control weight matrix, Q is the state positive definite function; Wherein, the formula of the control strategy is expressed as: ; in, Indicates i Iterations, the nonlinear dynamic system in the state vector The optimal control strategy at Represented in state space The set of admissible control strategies on , For the preset time interval, T represents transpose, Represents a nonlinear dynamic system in The state vector at time, express The control input vector at time , Indicated in i At the iteration, the nonlinear dynamic system has a state vector The state value of the location; Step 3: According to the formula of the value function and the formula of the control strategy, determine the final optimal control strategy of the nonlinear dynamic system and the final state value of the nonlinear dynamic system.

2. The approximate dynamic programming control method based on support vector regression according to claim 1 is characterized in that: Step 3 specifically includes iterating based on the formula of the value function and the formula of the control strategy to obtain the final optimal control strategy of the nonlinear dynamic system and the final state value of the nonlinear dynamic system, specifically including the following steps: Step A1: Set initial state value , set the initial iteration number i=0, use the initial iteration number as the current iteration number, and use the initial state value as the state value of the current iteration number ; Step A2: Set the current iteration state value Substitute into the control strategy formula and calculate the optimal control strategy for the current number of iterations ; Step A3: The optimal control strategy for the current number of iterations Substitute the formula of the value function to calculate the state value of the next iteration of the current iteration number ; Step A4: Calculate the first difference , which is specifically achieved through the following formula: ; Step A5: Determine the first difference Is it less than the preset threshold ε? When it is less than the preset threshold ε, As the final state value, the optimal control strategy for the current number of iterations is As the final optimal control strategy, in the first difference If the current number of iterations is less than or equal to the preset number of times, the current number of iterations is judged to be less than or equal to the preset number of times. If the current number of iterations is less than or equal to the preset number of times, As the final state value, the optimal control strategy for the current number of iterations is As the final optimal control strategy, when the current number of iterations is greater than the preset number, the current number of iterations is increased by one as the new current number of iterations, and the process returns to step 3.

2.

3. The approximate dynamic programming control method based on support vector regression according to claim 1 is characterized in that: Step 3 specifically includes improving the formula of the value function and the formula of the control strategy, and then iterating based on the improved formula to obtain the final optimal control strategy of the nonlinear dynamic system and the final state value of the nonlinear dynamic system, which specifically includes the following steps: Step B1: Based on the hard ε bond support vector regression machine HESVR, the formula of the value function and the formula of the control strategy are improved to obtain the value function formula based on HESVR and the control strategy formula based on HESVR; Step B2: Based on the least square method, the value function formula based on HESVR is processed to obtain the weight formula of HESVR; Step B3: According to the HESVR weight formula, the HESVR-based value function formula and the HESVR-based control strategy formula, the optimal control strategy of the nonlinear dynamic system and the final state value of the nonlinear dynamic system are calculated through an iterative algorithm.

4. The approximate dynamic programming control method based on support vector regression according to claim 3 is characterized in that: The value function formula based on HESVR in step B1 is expressed as: ; in, is the state value based on HESVR, For the i At the iteration, the nonlinear dynamic system has a state vector Approximate state value, For the status The nonlinear basis function vector at .

5. The approximate dynamic programming control method based on support vector regression according to claim 3 is characterized in that: The control strategy formula based on HESVR in step B1 is expressed as: ; in, represents the optimal control strategy based on HESVR, is the inverse matrix of the control weights, is the system dynamic gradient, For the i The gradient of the value function of the iteration, Indicates in status x ( t ), For the i The weight vector obtained by HESVR training in the iteration.

6. The approximate dynamic programming control method based on support vector regression according to claim 3 is characterized in that: The weight formula of HESVR in step B2 is expressed as: ; in, ; ; in, For the status The nonlinear basis function vector at , For the status The nonlinear basis function vector at , M is the number of sampling points in the least squares method, For the status The nonlinear basis function vector at , From the state x ( t ) to execute the control strategy exist The total trajectory cost in time, From the state Start and execute control strategy exist The total trajectory cost in time.

7. The approximate dynamic programming control method based on support vector regression according to claim 3 is characterized in that: Step B3 specifically includes: Step B3.1: Setting initial HESVR weights , set the initial state value , set the initial iteration number i=0, use the initial iteration number as the current iteration number, and use the initial HESVR weight as the current HESVR weight , take the initial state value as the state value of the current iteration ; Step B3.2: Determine whether the current number of iterations is greater than the preset number , when the current number of iterations is greater than the preset number In the case of, execute step 3.3.9, if the current number of iterations is not greater than the preset number If yes, proceed to step 3.3.3; Step B3.3: Set HESVR weights Substitute into the HESVR-based control strategy formula and calculate the optimal control strategy based on HESVR for the current number of iterations ; Step B3.4: Obtain the state vector using the Runge-Kutta algorithm ; Step B3.5: Transform the state vector Substitute it into the weight formula of HESVR to calculate the updated weight of HESVR; Step B3.6: Substitute the updated HESVR weight into the HESVR-based value function formula to calculate the HESVR-based state value for the next iteration of the current iteration. ; Step B3.7: Calculate the second difference , which is specifically achieved through the following formula: ; Step B3.8: Determine the second difference Is it greater than the preset threshold ε? If the difference is greater than the preset threshold ε, execute step 3.3.

9. If it is not greater than the preset threshold ε, the current number of iterations is increased by one as the new current number of iterations, and the process returns to step 3.3.2; Step B3.9: As the final state value, the optimal control strategy based on HESVR for the current iteration is As the final optimal control strategy.

Citation Information

Patent Citations

  • Electrical-power-system post-disturbance frequency dynamic-state prediction method based on support vector regression

    CN104333005A

  • Method for predicting net load of distributed power supply power distribution network

    CN105678415A

  • Self-adaptive dynamic programming control method for specified time

    CN118192224A

  • Emergency control method, device and equipment for power guarantee, medium and product

    CN118473075A