Adaptive dynamic programming method and system based on Chebyshev polynomials

Through the Chebyshev polynomial fitting adaptive dynamic programming method, the implementation difficulties of neural networks in adaptive dynamic programming are solved, and fast and accurate optimal control is achieved, which is suitable for digital control in industrial production, aerospace and power grid regulation.

CN119024686BActive Publication Date: 2025-09-26NR ENG CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310598691.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-24
Publication Date
2025-09-26
Estimated Expiration
2043-05-24

AI Technical Summary

Technical Problem

The use of neural networks in existing adaptive dynamic programming methods results in poor algorithm implementation, slow iteration and convergence speeds, and requires high debugging skills, making it difficult to meet the needs of industrial sites.

Method used

Chebyshev polynomials are used to replace neural networks, and optimal control is achieved by fitting performance index functions. The adaptive dynamic programming method of Chebyshev polynomials, including the fitting of models, actions and evaluation functions, is used in combination with the Newton iteration method to improve the convergence speed and accuracy of the algorithm.

Benefits of technology

The algorithm's execution efficiency and control accuracy are improved, the implementation difficulty is reduced, and fast optimal control is achieved. It is suitable for online implementation of digital control units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119024686B_ABST
    Figure CN119024686B_ABST
Patent Text Reader

Abstract

A Chebyshev polynomial-based adaptive dynamic programming method and system obtains Chebyshev polynomial fitting parameters for a model function, an action function, and an evaluation function. Within the range of state variables and control variables, the zero points of the Chebyshev polynomials for each function are used as the first, second, and third interpolation points, respectively. The model function is fitted at the first interpolation point. If initial iteration values ​​for the action and evaluation functions have been obtained, an inner iteration of the evaluation function is performed at the third interpolation point to obtain the latest evaluation function. During the outer iteration, the action function is updated with the control variable corresponding to the latest evaluation function that meets the convergence criterion. If no initial iteration value has been obtained, the second interpolation point is used as the initial iteration value at the initial iteration state. The control variable is obtained using the updated action function based on the state variable and output to the controlled object model. The present invention shortens the control time of a digital control unit and improves the control accuracy of the digital control unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of digital control technology, and in particular relates to an adaptive dynamic programming method and system based on Chebyshev polynomials. Background Art

[0002] In the field of digital control technology in industrial production, aerospace, and power grid regulation, it is often necessary to consider the optimal control problem, that is, to use the minimum control resources to make the controlled quantity reach the control target as quickly as possible.

[0003] In the prior art, dynamic programming is essentially a nonlinear programming method, centered on the Bellman optimality principle. While it is relatively easy to solve discrete-time optimal control problems using dynamic programming, its application is limited by the "curse of dimensionality" problem when dealing with complex problems. Subsequently, adaptive dynamic programming (ADP) was proposed, which utilizes function approximation structures to obtain optimal control solutions. There are three basic types of ADP methods: heuristic dynamic programming (HDP), dual heuristic programming (DHP), and globalized dual heuristic programming (GDHP). Each type of ADP method consists of three modules: an evaluation module, a model module, and an action module, each implemented using a neural network. However, neural network implementation requires a large number of parameters, such as network topology, weights, and initial threshold values. The learning process also exhibits a certain degree of randomness, making the output difficult to interpret. The algorithm also exhibits a "black box" nature, making debugging inconvenient and significantly impacting its effectiveness. Neural networks need to initialize weights and states. Unreasonable initialization parameters will directly affect the convergence of the ADP algorithm. The iteration and convergence speed of the neural network algorithm are slow. The iterative convergence depends on the selection of initial values. It requires high debugging skills for the algorithm user, which is not conducive to on-site use in engineering projects. Summary of the Invention

[0004] To address the deficiencies in the prior art, the present invention provides an adaptive dynamic programming method and system based on Chebyshev polynomials, which approximates the performance index function in dynamic programming based on Chebyshev polynomials, thereby obtaining the optimal performance index function and optimal control to satisfy the Bellman optimality principle, improve the convergence speed of the algorithm, effectively shorten the control time of the digital control unit, and improve the control accuracy of the digital control unit.

[0005] The present invention adopts the following technical solutions.

[0006] An adaptive dynamic programming method based on Chebyshev polynomials is used to fit and control a controlled object model in a digital control unit, wherein the controlled object model includes a model function, an action function, and an evaluation function; and includes:

[0007] Step 1: Obtain the fitting parameters of the Chebyshev polynomial of each function, including: fitting order, value range of state quantity, and value range of control quantity;

[0008] Step 2: within the range of the state variable and the control variable, obtain the zero points of the Chebyshev polynomials of the model function, the action function, and the evaluation function as the first interpolation point, the second interpolation point, and the third interpolation point, respectively; and fit the model function using the Chebyshev polynomial at the first interpolation point.

[0009] Step 3: Determine whether the initial iteration values ​​of the action function and the evaluation function have been obtained. If so, proceed to step 4. If not, when the iteration state is the initial state, select the second interpolation point within the range of the control variable as the initial iteration value.

[0010] Step 4: Based on the initial value of the iteration, perform inner iteration of the Chebyshev polynomial of the evaluation function at the third interpolation point. After the inner iteration, the latest evaluation function is obtained. In the outer iteration, the control amount corresponding to the latest evaluation function under the convergence criterion is used to update the action function to obtain the updated action function.

[0011] Step 5: Obtain the control quantity based on the state quantity using the updated action function, and output the control quantity to the controlled object model.

[0012] Preferably, in step 1, the Chebyshev polynomials include one-dimensional Chebyshev polynomials and multi-dimensional Chebyshev polynomials.

[0013] Preferably, in step 2, for the one-dimensional Chebyshev polynomial, the zero point of the Chebyshev polynomial is determined according to the fitting order;

[0014] For multidimensional Chebyshev polynomials, the multidimensional Chebyshev polynomials are expressed as the tensor product function of multiple one-dimensional Chebyshev polynomials, where the tensor product function is the product of trigonometric functions of multidimensional angular variables. The values ​​of the multidimensional angular variables are determined according to the zero values ​​of the trigonometric functions, and the values ​​of the multidimensional angular variables are converted into the values ​​of the multidimensional variables as the zero points of the multidimensional Chebyshev polynomials.

[0015] Preferably, step 2 also includes: sampling at the first interpolation point to obtain a test data set, using the test data set to calculate the error of the Chebyshev polynomial of the model function, and when the absolute value of the error is greater than a limit value, increasing the fitting order of the model function, wherein the limit value is not greater than 0.01.

[0016] Preferably, in step 3, if the initial value of the iteration is not obtained and the iteration state is the initial state, the following steps are included:

[0017] Step 3.1, within the value range of the state quantity and the value range of the control quantity, construct an artificial evaluation function using the continuous positive definite function of the state quantity and the control quantity and the model function;

[0018] Step 3.2, fitting the action function using Chebyshev polynomials at the second interpolation point;

[0019] Step 3.3, update the evaluation function using the control amount corresponding to the action function obtained in step 3.2;

[0020] Step 3.4, selecting multiple artificial iteration initial values ​​within the range of the control variable; based on different artificial iteration initial values, using the Newton iteration method to obtain the minimum value of the artificial evaluation function and the corresponding control variable;

[0021] Step 3.5, compare the minimum values ​​of different artificial evaluation functions, and take the control variable corresponding to the minimum value among the minimum values ​​of the artificial evaluation function as the minimum control variable;

[0022] Step 3.6, based on the minimum control variable, update the positive definite function. When the updated positive definite function is not greater than the positive definite function before the update, the minimum control variable and its corresponding evaluation function are used as the initial value of the iteration.

[0023] Preferably, the manual evaluation function g(x,u) is expressed as:

[0024]

[0025] Where, U(x k ,u k ) represents the k-th dimension state quantity x k and the control variable u k The corresponding evaluation function is,

[0026] F(x k ,u k ) represents the k-th dimension state quantity x k , control variable u k The corresponding model function is,

[0027] is the model function F(x k ,u k ) corresponds to a sufficiently large continuous positive definite function, and the positive definite function exists for the Chebyshev zeros of both the action function and the evaluation function, as follows:

[0028]

[0029] Where, is the k-th dimension state quantity x k is a continuous positive definite function of .

[0030] Preferably, based on the Newton iteration method, the control variable when the artificial evaluation function takes the minimum value is obtained by the following relationship:

[0031]

[0032] Where,

[0033] u l and u l+1 are the control variables at the lth and l+1th iterations respectively,

[0034] g′(x,u l ) and g″(x,u l ) are the first-order derivative and second-order derivative of the artificial evaluation function g(x,u), respectively.

[0035] Preferably, after obtaining the minimum control amount, the positive definite function is updated:

[0036]

[0037] Where, are positive definite functions before and after the update respectively.

[0038] Preferably, step 4 includes: performing inner layer iteration of the Chebyshev polynomial of the evaluation function at the zero point of the Chebyshev polynomial of the evaluation function based on the initial iteration value; when the number of inner layer iterations reaches a preset number, obtaining the latest internal evaluation function;

[0039] Based on the latest internal evaluation function, the Newton iteration method is used to perform outer iteration of the control quantity; before the number of outer iterations reaches the preset number, if for all zeros of the Chebyshev polynomials of the internal evaluation function, there are two times when the absolute value of the difference between the internal evaluation functions is less than the set value, the outer iteration ends;

[0040] When both inner and outer layer iterations are completed, the latest evaluation function and the optimal control function are obtained.

[0041] Preferably, the inner iteration of the evaluation function is performed according to the following relationship:

[0042]

[0043] Where,

[0044] i is the number of outer layer iterations,

[0045] j i is the number of inner layer iterations corresponding to the i-th outer layer iteration,

[0046] x k is the k-th dimension state quantity,

[0047] v i (x k ) is the k-th dimension state quantity x at the i-th outer iteration k The corresponding control quantity,

[0048] F(x k ,v i (x k )) is the k-th dimension state quantity x k , control quantity v i (x k ) corresponds to the model function, and serves as the k+1th dimension state quantity x k+1 ,

[0049] is the jth value of the i-th outer iteration i +1 inner iteration k-th dimension state x k The corresponding internal evaluation function is,

[0050] is the jth value of the i-th outer iteration i The k+1th dimension state x at the outer iteration k+1 The corresponding internal evaluation function is,

[0051] U(x k ,v i (x k )) is the state quantity x k , control quantity v i (x k ) is the corresponding evaluation function.

[0052] Preferably, based on the latest internal evaluation function, the Newton iteration method is used for outer iteration to fit the control quantity with the following relationship:

[0053]

[0054] Where,

[0055] U(x k ,u k ) is the state quantity x k , control variable u k The corresponding evaluation function is,

[0056] F(x k ,u k ) is the k-th dimension state quantity x k , control variable u k The corresponding model function is,

[0057] V i-1 (F(x k ,u k )) is the model function F(x) at the i-1th outer layer iteration k ,u k ) corresponds to the internal evaluation function.

[0058] Preferably, the preset number of times is not less than 1, and the set value ε is not greater than 0.01.

[0059] An adaptive dynamic programming system based on Chebyshev polynomials includes: an input module, an interpolation point module, a model function fitting module, an iterative initial value module, an action function and evaluation function fitting module, an output module,

[0060] An input module is used to obtain the fitting parameters of the Chebyshev polynomials of each function, including: fitting order, value range of state quantity, and value range of control quantity;

[0061] An interpolation point module is used to obtain the zero points of the Chebyshev polynomials of the model function, action function and evaluation function as the first interpolation point, the second interpolation point and the third interpolation point within the value range of the state quantity and the control quantity respectively;

[0062] A model function fitting module, used for fitting the model function using Chebyshev polynomials at the first interpolation point;

[0063] The iteration initial value module is used to determine whether the iteration initial values ​​of the action function and the evaluation function have been obtained. If the iteration initial values ​​have been obtained, the evaluation function module is called; if the iteration initial values ​​have not been obtained, when the iteration state is the initial state, the second interpolation point is selected within the value range of the control quantity as the iteration initial value;

[0064] The action function and evaluation function fitting module is used to perform inner layer iteration of the Chebyshev polynomial of the evaluation function at the third interpolation point based on the initial value of the iteration, obtain the latest evaluation function through the inner layer iteration, and update the action function with the control amount corresponding to the latest evaluation function under the convergence criterion in the outer layer iteration to obtain the updated action function;

[0065] The output module obtains the control quantity according to the state quantity using the updated action function, and outputs the control quantity to the controlled object model.

[0066] A terminal includes a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method.

[0067] A computer-readable storage medium stores a computer program thereon, which implements the steps of the method when executed by a processor.

[0068] The beneficial effect of the present invention is that, compared with the prior art, the adaptive dynamic programming method based on Chebyshev polynomials proposed in the present invention realizes the optimal control of the controlled object. The use of Chebyshev polynomials to replace the original neural network to realize the fitting of the model object, evaluation function, and action function improves the execution efficiency of the algorithm, avoids multiple forward and back iteration processes when training the neural network, reduces the difficulty of implementing the algorithm, and makes the evaluation function and action function have a clear polynomial structure, which is convenient for implementation using a digital control unit. The Newton iteration method is used to obtain the extreme value of the algorithm during the calculation process of the algorithm, further improving the implementation accuracy and calculation efficiency of the algorithm, and avoiding obtaining the extreme value by searching for a large number of sample values. And the convergence speed of the algorithm is improved by iterating the outer control strategy and iterating the inner evaluation value, making the online implementation of the adaptive dynamic programming algorithm possible. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 This is a flow chart of an adaptive dynamic programming method based on Chebyshev polynomials proposed by the present invention;

[0070] Figure 2 is a Chebyshev polynomial fitting flow chart in an embodiment of the present invention;

[0071] Figure 3 1 is an algorithm architecture diagram of the heuristic adaptive dynamic programming (HDP) in an embodiment of the present invention;

[0072] Figure 4 This is a diagram showing the execution effect of the adaptive dynamic programming method based on the Chebyshev function in an embodiment of the present invention. DETAILED DESCRIPTION

[0073] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. The embodiments described in this application are only part of the embodiments of the present invention, not all of them. Based on the spirit of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0074] The present invention proposes to use Chebyshev polynomials to replace function approximation to approximate the Hamilton-Jacobi-Bellman (HJB) equation to solve the dynamic programming problem based on the Bellman optimality principle, and realize time-forward dynamic programming to overcome the "dimensionality curse" problem.

[0075] Adaptive dynamic programming (ADP) based on Chebyshev polynomials involves fitting execution, evaluation, and model functions, and is an optimal control solution. ADP does not require a precise model of the nonlinear system; it only requires adaptive learning based on external evaluation information.

[0076] The present invention proposes an adaptive dynamic programming method based on Chebyshev polynomials for fitting and controlling a controlled object model, wherein the controlled object model includes a model function, an action function and an evaluation function; Figure 1 As shown, the specific steps of the method are as follows:

[0077] Step 1: Obtain the fitting parameters of the Chebyshev polynomial of each function, including: fitting order, value range of state quantity, and value range of control quantity.

[0078] Specifically, the model function, action function and evaluation function set corresponding parameters such as fitting order, value range of state quantity, value range of control quantity and convergence criterion, which are used to provide necessary parameter information for the Chebyshev polynomial fitting function and provide a parameter adjustment interface to avoid affecting the versatility of the dynamic programming method due to parameter solidification.

[0079] The fitting order is used to determine the order of the Chebyshev polynomial. The Chebyshev polynomials of each function may be different, so the fitting order determined according to the fitting conditions of each function is also different. A larger fitting order can improve the accuracy of the fitting function, but a larger fitting order increases the calculation time and causes the algorithm to overfit.

[0080] State quantities are physical quantities that describe the state of a material system. Examples include position, velocity, momentum, kinetic energy, angular velocity, angular momentum, pressure, temperature, volume, and potential energy. Under external influences, the state of a material system changes over time, and the state quantities describing the system also change over time. The integral of any closed-loop state quantity is zero. In other words, the change in a state quantity after any process is related only to the initial and final states, and not to the path of the process.

[0081] Furthermore, the value range of the state quantity and the value range of the control quantity are set according to the actual on-site use requirements of the digital control unit, and the state quantity and the control quantity are normalized to facilitate the fitting of the Chebyshev polynomial.

[0082] Furthermore, the convergence criterion is used to determine whether the adaptive dynamic programming method has converged, and the value is a given reasonable small value.

[0083] Step 2: Within the range of state variables and control variables, obtain the zero points of the Chebyshev polynomials of the model function, action function, and evaluation function as the first interpolation point, the second interpolation point, and the third interpolation point, respectively; and fit the model function using the Chebyshev polynomial at the first interpolation point.

[0084] When fitting each function, it is necessary to obtain the function's characteristics. Interpolation points are points selected on the function to reflect the function's characteristic information. After obtaining the interpolation points, fitting can only be performed when the characteristic information of the function being fitted is known. The present invention proposes using Chebyshev polynomials to characterize the functions involved in the adaptive dynamic programming algorithm. Therefore, using Chebyshev polynomials to fit the function can achieve better fitting results through a specific method of selecting interpolation points. Selecting the zero points of the Chebyshev polynomials as interpolation points can reduce interpolation errors and effectively avoid the Runge phenomenon.

[0085] Specifically, the system state equation is regarded as a black box, and the zero point of the Chebyshev polynomial is selected as the interpolation point according to the fitting order obtained in step 1. Sampling is performed at the interpolation point, and the model function is initially fitted using the Chebyshev polynomial. The three functions are fitted in the subsequent iterative process.

[0086] Specifically, Chebyshev polynomials include one-dimensional Chebyshev polynomials and multidimensional Chebyshev polynomials. Among them, multidimensional Chebyshev polynomials are more widely used. Using multidimensional Chebyshev polynomials to fit model functions, action functions, and evaluation functions has a wider range of applications.

[0087] Specifically, for a one-dimensional Chebyshev polynomial, the zero points of the Chebyshev polynomial are determined according to the fitting order.

[0088] Specifically, for multidimensional Chebyshev polynomials, the multidimensional Chebyshev polynomials are expressed as the tensor product function of multiple one-dimensional Chebyshev polynomials, where the tensor product function is the product of trigonometric functions of multidimensional angle variables. The value of the multidimensional angle variable is determined according to the zero value of the trigonometric function, and the value of the multidimensional angle variable is converted into the value of the multidimensional variable as the zero point of the multidimensional Chebyshev polynomial.

[0089] The multidimensional Chebyshev polynomial f(ξ) is expressed as:

[0090]

[0091] Where,

[0092] ξ is a k-dimensional variable, ξ1,ξ2,…,ξ k are the values ​​of ξ in each dimension,

[0093] i1,i2,…,i k are ξ1,ξ2,…,ξk The coefficient of i k =0,1,2,…,n,

[0094] n is a natural number, which is the highest order of the Chebyshev fitting function selected in step 1.

[0095] s is the order,

[0096] is the tensor product function, expressed as:

[0097]

[0098] Among them, θ1, θ2,…, θ k are respectively related to ξ1,ξ2,…,ξ k The associated angle variable.

[0099] is a fixed coefficient, expressed as:

[0100]

[0101]

[0102] in:

[0103]

[0104] Where,

[0105] j1,j2,…,j k are the coefficients of the variables corresponding to each dimension, j k =1,2,…,m,

[0106] m is the number of Chebyshev zeros, and the value of m is not less than n+1;

[0107] is the k-th dimension variable ξ k The upper limit of the value of

[0108] χ k is the k-th dimension variable ξ k The lower limit of the value of

[0109] When i1,…,i k When there are m 0s in , the corresponding s=m.

[0110] Use multidimensional Chebyshev polynomials to fit multidimensional system objects. After fitting, use the test data set to calculate the accuracy of the model function, action function and evaluation function. If the accuracy is lower than the set fitting accuracy, the accuracy of the Chebyshev polynomial fitting is improved by increasing the fitting order until it exceeds the set fitting accuracy.

[0111] The specific process of the multidimensional Chebyshev polynomial fitting algorithm is as follows Figure 2 As shown, it is expressed as:

[0112] 1) Before function fitting, set the parameters of the Chebyshev polynomial, including the Chebyshev polynomial order n and the value range of the variable x = [a, b];

[0113] 2) Calculate the Chebyshev interpolation zero point. First, calculate the angle at the Chebyshev interpolation zero point, and then obtain the Chebyshev interpolation zero point based on the value range of the variable.

[0114] 3) Get the function value of the fitting object at the interpolation zero point;

[0115] 4) Calculate the fixed coefficients of Chebyshev polynomials according to the formula;

[0116] 5) Use the obtained coefficients and basis functions to construct Chebyshev polynomials to fit the function object.

[0117] Step 3: Determine whether the initial iteration values ​​of the action function and the evaluation function have been obtained. If so, proceed to step 4. If not, when the iteration state is the initial state, select the second interpolation point within the range of the control quantity as the initial iteration value.

[0118] Determine whether it is necessary to obtain the initial value of the iteration based on the iteration state. When the iteration state is the initial state or the iteration state is non-convergent, select the Chebyshev zero point of the action function as the initial value of the iteration. Use the initial value of the iteration to fit the initial fit obtained in step 2, and use the Chebyshev polynomial to fit and update the controlled object model at the interpolation point until the iteration converges, and obtain the Chebyshev polynomials of the action function and the evaluation function.

[0119] In a non-restrictive preferred embodiment, the algorithm process is called when it is necessary to obtain the initial value of the iteration, the Chebyshev zero point of the action function is selected as the initial value of the iteration within the value range of the control quantity, and the initial control function and evaluation function are obtained using the initial value of the iteration to provide feasible numerical values ​​for the evaluation value and the control strategy iteration process.

[0120] Specifically, if the iteration initial value is not obtained and the iteration state is the initial state, step 3 includes:

[0121] Step 3.1: Within the value range of the state quantity and the value range of the control quantity, use the continuous positive definite function of the state quantity and the control quantity and the model function to construct the artificial evaluation function as follows:

[0122]

[0123] Where,

[0124] U(x k ,u k ) represents the k-th dimension state quantity x k and the control variable u k The corresponding evaluation function is,

[0125] F(x k ,u k ) represents the k-th dimension state quantity x k , control variable u k The corresponding model function is,

[0126] is the model function F(x k ,u k ) corresponds to a sufficiently large continuous positive definite function, and the positive definite function exists for the Chebyshev zeros of both the action function and the evaluation function, as follows:

[0127]

[0128] Where, is the k-th dimension state quantity x k is a continuous positive definite function of .

[0129] Step 3.2: Fit the action function using Chebyshev polynomials at the second interpolation point.

[0130] Step 3.3, update the evaluation function using the control amount corresponding to the action function obtained in step 3.2.

[0131] In step 3.4, multiple artificial iteration initial values ​​are selected within the range of the control variable; based on different artificial iteration initial values, the Newton iteration method is used to obtain the minimum value of the artificial evaluation function and the corresponding control variable.

[0132] Based on the Newton iteration method, the control variable when the artificial evaluation function takes the minimum value is obtained by the following relationship:

[0133]

[0134] Where,

[0135] u l and u l+1 are the control variables at the lth and l+1th iterations respectively,

[0136] g′(x,u l ) and g″(x,u l ) are the first-order derivative and second-order derivative of the artificial evaluation function g(x,u), respectively.

[0137] When the iteration converges, the minimum value of the evaluation function g(x,u l) state of the control quantity u k .

[0138] Step 3.5, compare the minimum values ​​of different artificial evaluation functions, and take the control variable corresponding to the minimum value among the minimum values ​​of the artificial evaluation function as the minimum control variable.

[0139] From the minimum value of the evaluation function g(x,u k ) to extract the control quantity corresponding to the minimum value As the minimum control amount.

[0140] Step 3.6, based on the minimum control variable, update the positive definite function. When the updated positive definite function is not greater than the positive definite function before the update, the minimum control variable and its corresponding evaluation function are used as the initial value of the iteration.

[0141] After obtaining the minimum control amount, the positive definite function is updated:

[0142]

[0143] Where, are the evaluation functions before and after updating respectively.

[0144] Both represent evaluation functions.

[0145] The present invention realizes the fitting update of the evaluation function and the action function based on the Chebyshev polynomials, and the fitting update process is consistent with the fitting process of the controlled object model based on the Chebyshev polynomials in step 2.

[0146] Step 4: Based on the initial value of the iteration, perform inner iteration of the Chebyshev polynomial of the evaluation function at the third interpolation point. After the inner iteration, the latest evaluation function is obtained. In the outer iteration, the action function is updated with the control quantity corresponding to the latest evaluation function under the convergence criterion to obtain the updated action function.

[0147] The evaluation function and action function can be understood as having a symbiotic relationship: finding the optimal evaluation function also leads to finding the optimal action function, and vice versa. The ultimate goal of the adaptive dynamic programming algorithm is to obtain the optimal action function, which can also be understood as obtaining the optimal evaluation function. This process is achieved through an iterative approach: iterating the evaluation function within an inner loop and the action function within an outer loop. To initiate the iteration process, an initial value must be input, but this initial value is unknown and must be determined algorithmically.

[0148] Therefore, in step 4, based on the initial value of the iteration, the inner layer iteration of the Chebyshev polynomial of the evaluation function is performed at the zero point of the Chebyshev polynomial of the evaluation function; when the number of inner layer iterations reaches the preset number, the latest internal evaluation function is obtained;

[0149] Based on the latest internal evaluation function, the Newton iteration method is used to perform outer iteration of the control quantity; before the number of outer iterations reaches the preset number, if for all zeros of the Chebyshev polynomials of the internal evaluation function, there are two times when the absolute value of the difference between the internal evaluation functions is less than the set value, the outer iteration ends;

[0150] When both inner and outer layer iterations are completed, the latest evaluation function and the optimal control function are obtained.

[0151] Among them, the preset number of times is not less than 1 time, and the set value ε is not greater than 0.01.

[0152] Specifically, step 4 includes:

[0153] Step 4.1: Using the Chebyshev polynomials of the action function and evaluation function obtained in step 3 as the initial function for fitting iteration, perform inner iteration of the evaluation function and outer iteration of the action function at the interpolation point.

[0154] The inner iteration of the evaluation function is performed according to the following relationship:

[0155]

[0156] Where,

[0157] i is the number of outer layer iterations,

[0158] j i is the number of inner layer iterations corresponding to the i-th outer layer iteration,

[0159] x k is the k-th dimension state quantity,

[0160] v i (x k ) is the k-th dimension state quantity x at the i-th outer iteration k The corresponding control quantity,

[0161] F(x k ,v i (x k )) is the k-th dimension state quantity x k , control quantity v i (x k ) corresponds to the model function, and serves as the k+1th dimension state quantity x k+1 ,

[0162] is the jth value of the i-th outer iteration i +1 inner iteration k-th dimension state x k The corresponding internal evaluation function is,

[0163] is the jth value of the i-th outer iteration i The k+1th dimension state x at the outer iteration k+1 The corresponding internal evaluation function is,

[0164] U(x k ,v i (x k )) is the state quantity x k , control quantity v i (x k ) is the corresponding evaluation function.

[0165] In step 4.2, based on the latest internal evaluation function, the Newton iteration method is used to perform outer iteration and fit the control quantity with the following relationship:

[0166]

[0167] Where,

[0168] U(x k ,u k ) is the state quantity x k , control variable u k The corresponding evaluation function is,

[0169] F(x k ,u k ) is the k-th dimension state quantity x k , control variable u k The corresponding model function is,

[0170] V i-1 (F(x k ,u k )) is the model function F(x) at the i-1th outer layer iteration k ,u k ) corresponds to the internal evaluation function.

[0171] Step 4.3, perform less than N i The outer iteration is repeated until the absolute value of the difference between two evaluation functions is less than the preset value ε for all Chebyshev zeros, that is, |V i (x k )-V i-1 (x k )|≤ε, the iteration ends.

[0172] The control function obtained at this time is considered to be the optimal control function, and the evaluation function obtained is considered to be the optimal evaluation function. The optimal control function and the optimal evaluation function are fitted in the form of Chebyshev polynomials.

[0173] Step 5: Obtain the control quantity based on the state quantity using the updated action function, and output the control quantity to the controlled object model.

[0174] Specifically, given the initial value of the state quantity, after substituting it into the action function, the control quantity of the corresponding state quantity is obtained, and the control quantity is applied to the controlled object. Then, a new state quantity can be obtained in the next sampling period, and the new state quantity is brought into the action function to obtain a new control quantity, and so on, until the state quantity finally tends to a stable value and the algorithm ends.

[0175] The present invention also proposes an adaptive dynamic programming system based on Chebyshev polynomials, comprising: an input module, an interpolation point module, a model function fitting module, an iterative initial value module, an action function and evaluation function fitting module, and an output module;

[0176] An input module is used to obtain the fitting parameters of the Chebyshev polynomials of each function, including: fitting order, value range of state quantity, and value range of control quantity;

[0177] An interpolation point module is used to obtain the zero points of the Chebyshev polynomials of the model function, action function and evaluation function as the first interpolation point, the second interpolation point and the third interpolation point within the value range of the state quantity and the control quantity respectively;

[0178] A model function fitting module, used for fitting the model function using Chebyshev polynomials at the first interpolation point;

[0179] The iteration initial value module is used to determine whether the iteration initial values ​​of the action function and the evaluation function have been obtained. If the iteration initial values ​​have been obtained, the evaluation function module is called; if the iteration initial values ​​have not been obtained, when the iteration state is the initial state, the second interpolation point is selected within the value range of the control quantity as the iteration initial value;

[0180] The action function and evaluation function fitting module is used to perform inner layer iteration of the Chebyshev polynomial of the evaluation function at the third interpolation point based on the initial value of the iteration, obtain the latest evaluation function through the inner layer iteration, and update the action function with the control amount corresponding to the latest evaluation function under the convergence criterion in the outer layer iteration to obtain the updated action function;

[0181] The output module obtains the control quantity according to the state quantity using the updated action function, and outputs the control quantity to the controlled object model.

[0182] The controlled object is a real physical entity. The model function simply fits the object's behavioral characteristic function using a polynomial, facilitating a mathematical description of the problem. This allows the use of an algorithm to derive the action function used for control. Once the action function is obtained, a control signal is generated based on the state information and applied to the controlled object, achieving control of the object.

[0183] The output module ultimately outputs a control signal. Figure 3 The algorithm architecture diagram, after the algorithm iteration is completed, for each state vector x k , the corresponding control vector u can be obtained through the action network k , so that u k Act on the physical system to achieve optimal control of the physical system.

[0184] Further, if Figure 3 The heuristic adaptive dynamic programming (HDP) shown is a specific implementation structure of ADP. It uses Chebyshev polynomials to fit the evaluation function and action function, reducing the size of preset parameters, improving the algorithm convergence speed and storage resource consumption, and making the obtained optimal evaluation network and action network have a clear polynomial form, which is convenient for the implementation of digital processing chips.

[0185] Dynamic programming is essentially a nonlinear programming method proposed by American mathematician Bellman. The core of this programming algorithm is the Bellman Optimality Principle. While dynamic programming is relatively easy to use to solve discrete-time optimal control problems, its application is limited by the "curse of dimensionality" when dealing with complex problems. Subsequently, Werbos proposed the theory of adaptive dynamic programming, which utilizes function approximation structures to obtain optimal control solutions. Based on this, he used Chebyshev polynomials to improve the adaptive dynamic programming algorithm and proposed an adaptive dynamic programming algorithm framework based on Chebyshev polynomials.

[0186] The principle of adaptive dynamic programming is to utilize a function approximation structure and Chebyshev polynomials to approximate the performance indicator function in dynamic programming, thereby obtaining the optimal performance indicator function and optimal control to satisfy the optimality principle. The adaptive dynamic programming structure consists of three modules: an evaluation module, a model module, and an execution module. Each module can be approximated using Chebyshev polynomials. The execution network is used to approximate the optimal control strategy, while the evaluation network approximates the optimal performance indicator function. After the control / execution is applied to the dynamic system (or controlled object), the rewards / penalties generated by the controlled object at different stages influence the judgment function. Chebyshev polynomials are used to approximate the execution and judgment functions. This allows the optimal control function to be obtained simultaneously with the optimal evaluation function. The parameter update of the judgment function is based on the Bellman optimality principle. This not only reduces forward computation time but also allows online response to dynamic changes in the system, automatically adjusting the fitting parameters in the network structure.

[0187] Reference Figure 4, is the execution effect of the adaptive dynamic programming method based on the Chebyshev function. Assuming that the controlled object is a torsion pendulum system, its discretized system equation can be expressed as:

[0188]

[0189] where [x 1k ,x 2k ] T is the system state equation, u k is the system control variable, f d is the friction coefficient, and f is taken during the simulation process d =0.2=0.2, the initial state of the simulation is taken as x0=[2,-1], and the adaptive dynamic programming method based on neural network and the method of the present invention are used for iteration to obtain the optimal control rate.

[0190] In the traditional iterative method based on BP neural network, the sampling value of the state quantity is selected as 10000, and the iteration of the value function and the action function takes 5254.84 seconds and occupies 200MB of memory to obtain the optimal action function; while using the method described in the present invention, it only takes 39.617 seconds and occupies 4MB of memory to obtain the optimal action function. After using the method described in the present invention, the computing time and memory occupied are greatly reduced, and the construction of complex neural network equations is avoided. Finally, the action function in polynomial form can be obtained, which is convenient for application in digital control units. And by Figure 4 It can be seen that under the control rate described in the present invention, the state quantity of the system will eventually converge to [0, 0], realizing the control of the object.

[0191] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0192] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0193] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0194] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. An adaptive dynamic programming method based on Chebyshev polynomials for fitting a controlled object model in a digital control unit, wherein: The controlled object model includes a model function, an action function, and an evaluation function; and is characterized by including: Step 1: Obtain the fitting parameters of the Chebyshev polynomial of each function, including: fitting order, value range of state quantity, and value range of control quantity; Step 2: within the range of the state variable and the control variable, obtain the zero points of the Chebyshev polynomials of the model function, the action function, and the evaluation function as the first interpolation point, the second interpolation point, and the third interpolation point, respectively; and fit the model function using the Chebyshev polynomial at the first interpolation point. Step 3, determine whether the iterative initial values ​​of the action function and the evaluation function have been obtained. If the iterative initial values ​​have been obtained, proceed to step 4; if the iterative initial values ​​have not been obtained, when the iteration state is the initial state, it includes: step 3.1, within the value range of the state quantity and the value range of the control quantity, use the continuous positive definite function of the state quantity and the control quantity and the model function to construct an artificial evaluation function; step 3.2, use the Chebyshev polynomial to fit the action function at the second interpolation point; step 3.3, use the control quantity corresponding to the action function obtained in step 3.2 to update the evaluation function number; Step 3.4, select multiple artificial iteration initial values ​​within the range of the control variable; based on different artificial iteration initial values, use the Newton iteration method to obtain the minimum value of the artificial evaluation function and the corresponding control variable; Step 3.5, compare the minimum values ​​of different artificial evaluation functions, and use the control variable corresponding to the minimum value among the minimum values ​​of the artificial evaluation function as the minimum control variable; Step 3.6, based on the minimum control variable, update the positive definite function, and when the updated positive definite function is not greater than the positive definite function before the update, use the minimum control variable and its corresponding evaluation function as the iteration initial value; Step 4, based on the initial value of the iteration, perform inner iteration of the Chebyshev polynomial of the evaluation function at the third interpolation point, obtain the latest evaluation function through the inner iteration, and update the action function of the control quantity corresponding to the latest evaluation function under the convergence criterion in the outer iteration to obtain the updated action function; including: based on the initial value of the iteration, perform inner iteration of the Chebyshev polynomial of the evaluation function at the zero point of the Chebyshev polynomial of the evaluation function; when the number of inner iterations reaches a preset number, obtain the latest internal evaluation function; based on the latest internal evaluation function, use the Newton iteration method to perform outer iteration of the control quantity; before the number of outer iterations reaches the preset number, for all zero points of the Chebyshev polynomial of the internal evaluation function, there are two times the absolute value of the difference of the internal evaluation function is less than the set value, then the outer iteration ends; when both the inner and outer iterations end, the latest evaluation function and the optimal control function are obtained; Step 5: Obtain the control quantity using the updated action function according to the state quantity, and output the control quantity to the controlled object model.

2. The adaptive dynamic programming method based on Chebyshev polynomials according to claim 1, characterized in that: In step 1, the Chebyshev polynomials include one-dimensional Chebyshev polynomials and multi-dimensional Chebyshev polynomials.

3. The adaptive dynamic programming method based on Chebyshev polynomials according to claim 2, characterized in that: In step 2, for the one-dimensional Chebyshev polynomial, the zero point of the Chebyshev polynomial is determined according to the fitting order; For multidimensional Chebyshev polynomials, the multidimensional Chebyshev polynomials are expressed as the tensor product function of multiple one-dimensional Chebyshev polynomials, where the tensor product function is the product of trigonometric functions of multidimensional angular variables. The values ​​of the multidimensional angular variables are determined according to the zero values ​​of the trigonometric functions, and the values ​​of the multidimensional angular variables are converted into the values ​​of the multidimensional variables as the zero points of the multidimensional Chebyshev polynomials.

4. The adaptive dynamic programming method based on Chebyshev polynomials according to claim 1, characterized in that: Step 2 also includes: sampling at the first interpolation point to obtain a test data set, using the test data set to calculate the error of the Chebyshev polynomial of the model function, and when the absolute value of the error is greater than a limit value, increasing the fitting order of the model function, where the limit value is not greater than 0.

01.

5. The adaptive dynamic programming method based on Chebyshev polynomials according to claim 1, characterized in that: The manual evaluation function g(x,u) is expressed as: Where, U(x k ,u k ) represents the k-th dimension state quantity x k and the control variable u k The corresponding evaluation function, F(x k ,u k ) represents the k-th dimension state x k , control variable u k The corresponding model function is, is the model function F(x k ,u k ) corresponds to a sufficiently large continuous positive definite function, and the positive definite function exists for the Chebyshev zeros of both the action function and the evaluation function, as follows: Where, is the k-th dimension state quantity x k is a continuous positive definite function of .

6. The adaptive dynamic programming method based on Chebyshev polynomials according to claim 5, characterized in that: Based on the Newton iteration method, the control variable when the artificial evaluation function takes the minimum value is obtained by the following relationship: Where u l and u l+1 are the control variables at the lth and l+1th iterations, g′(x,u l ) and g″(x,u l ) are the first-order derivative and second-order derivative of the artificial evaluation function g(x,u), respectively.

7. The adaptive dynamic programming method based on Chebyshev polynomials according to claim 6, characterized in that: After obtaining the minimum control amount, the positive definite function is updated: Where, are positive definite functions before and after the update respectively.

8. The adaptive dynamic programming method based on Chebyshev polynomials according to claim 1, characterized in that: The inner iteration of the evaluation function is performed according to the following relationship: Where i is the number of outer layer iterations, j is i is the number of inner layer iterations corresponding to the i-th outer layer iteration, x k is the k-th dimension state quantity, v i (x k ) is the k-th dimension state quantity x at the i-th outer iteration k The corresponding control quantity, F(x k ,v i (x k )) is the k-th dimension state quantity x k , control quantity v i (x k ) corresponds to the model function, and serves as the k+1th dimension state quantity x k+1 , (x k ) is the jth outer iteration of the i-th i +1 inner iteration k-th dimension state x k The corresponding internal evaluation function is, (F(x k ,v i (x k ))) is the jth value of the i-th outer iteration i The k+1th dimension state x at the outer iteration k+1 The corresponding internal evaluation function, U(x k ,v i (x k )) is the state quantity x k , control quantity v i (x k ) is the corresponding evaluation function.

9. The adaptive dynamic programming method based on Chebyshev polynomials according to claim 8, characterized in that: Based on the latest internal evaluation function, the Newton iteration method is used for outer iteration to fit the control quantity with the following relationship: Where, U(x k ,u k ) is the state quantity x k , control variable u k The corresponding evaluation function, F(x k ,u k ) is the k-th dimension state quantity x k , control variable u k The corresponding model function, V i-1 (F(x k ,u k )) is the model function F(x) at the i-1th outer layer iteration k ,u k ) corresponds to the internal evaluation function.

10. The adaptive dynamic programming method based on Chebyshev polynomials according to claim 1, characterized in that: The preset number of times is not less than 1, and the set value ε is not greater than 0.

01.

11. An adaptive dynamic programming system based on Chebyshev polynomials using the method according to any one of claims 1 to 10, characterized in that: The system includes: input module, interpolation point module, model function fitting module, iterative initial value module, action function and evaluation function fitting module, output module, An input module is used to obtain the fitting parameters of the Chebyshev polynomials of each function, including: fitting order, value range of state quantity, and value range of control quantity; An interpolation point module is used to obtain the zero points of the Chebyshev polynomials of the model function, action function and evaluation function as the first interpolation point, the second interpolation point and the third interpolation point within the value range of the state quantity and the control quantity respectively; A model function fitting module, used for fitting the model function using Chebyshev polynomials at the first interpolation point; The iterative initial value module is used to determine whether the iterative initial values ​​of the action function and the evaluation function have been obtained. If the iterative initial values ​​have been obtained, the evaluation function module is called; if the iterative initial value has not been obtained, when the iterative state is the initial state, it includes: within the value range of the state quantity and the value range of the control quantity, using the continuous positive definite function of the state quantity and the control quantity and the model function to construct an artificial evaluation function; using Chebyshev polynomials to fit the action function at the second interpolation point; using the control quantity corresponding to the obtained action function to update the evaluation function; selecting multiple artificial iterative initial values ​​within the value range of the control quantity; based on different artificial iterative initial values, using the Newton iteration method to obtain the minimum value of the artificial evaluation function and the corresponding control variable; comparing the minimum values ​​of different artificial evaluation functions, and taking the control variable corresponding to the minimum value of the minimum value of the artificial evaluation function as the minimum control variable; based on the minimum control variable, updating the positive definite function, and when the updated positive definite function is not greater than the positive definite function before the update, taking the minimum control variable and its corresponding evaluation function as the iterative initial value; The action function and evaluation function fitting module is used to perform inner iteration of the Chebyshev polynomial of the evaluation function at the third interpolation point based on the initial value of the iteration, obtain the latest evaluation function through the inner iteration, and update the action function with the control quantity corresponding to the latest evaluation function under the convergence criterion in the outer iteration to obtain the updated action function; including: performing inner iteration of the Chebyshev polynomial of the evaluation function at the zero point of the Chebyshev polynomial of the evaluation function based on the initial value of the iteration; obtaining the latest internal evaluation function when the number of inner iterations reaches a preset number; performing outer iteration of the control quantity using the Newton iteration method based on the latest internal evaluation function; before the number of outer iterations reaches the preset number, for all zero points of the Chebyshev polynomial of the internal evaluation function, there are two times when the absolute value of the difference between the internal evaluation functions is less than the set value, then the outer iteration ends; when both the inner and outer iterations end, the latest evaluation function and the optimal control function are obtained; The output module obtains the control quantity according to the state quantity using the updated action function, and outputs the control quantity to the controlled object model.

12. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Robot motion optimization method and device, computer equipment and storage medium

    CN109318230A

  • Inertial navigation resolving method and system based on matrix form function iteration

    CN114996652A