Neural network control law online learning method
By constructing a dynamic model of a hypersonic vehicle and designing an Actor-Critic neural network basis, the difficulty of solving the attitude control law of a hypersonic vehicle was solved, and robust and optimal adaptive control in complex environments was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING AEROSPACE AUTOMATIC CONTROL RES INST
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-21
AI Technical Summary
Traditional methods are difficult to effectively solve the problem of attitude control laws for hypersonic vehicles, especially when facing complex flight environments and system uncertainties, it is difficult to balance the robustness and optimality of the control system.
An online learning method for neural network control laws is adopted, combining the derivative rule of rotating coordinate system and momentum distance theorem to construct an aircraft dynamics model. Based on game theory and Weierstrass fundamental theorem, an Actor-Critic neural network basis is designed. The neural network weights are updated through discretization approximation method to realize intelligent control strategy.
This reduces the difficulty of solving attitude control laws, improves the robustness and optimality of the control system, realizes adaptive dynamic programming, and enhances the design efficiency and robustness of neural network control laws.
Smart Images

Figure CN121900162A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an online learning method for neural network control laws, belonging to the field of aircraft control system design. Background Technology
[0002] Intelligent control technology can make spacecraft more intelligent, significantly improve key control performance indicators, and continuously enhance the generalization ability of controllers to adapt to different objects through learning. This can compensate for the limitations of programmed control strategies, enhance the spacecraft's ability to adapt to complex flight environments and cope with sudden interference, improve the robustness of control strategies, and solve problems that are difficult to overcome by traditional methods.
[0003] As a newly emerging approximate optimal control method in recent years, adaptive dynamic programming technology is essentially based on the principle of reinforcement learning. It organically combines control theory with computational intelligence and effectively solves the optimal control / differential game problem of nonlinear systems by adopting a function approximation structure. Furthermore, with the continuous emergence of various new intelligent aircraft and the rapid development of aerospace technology, traditional control methods are no longer sufficient to meet the needs of future complex aerospace missions and the multi-objective decision-making requirements for adaptive trajectory adjustment in complex environments. Summary of the Invention
[0004] The technical problem solved by this invention is to overcome the shortcomings of the prior art and propose an online learning method for neural network control laws. This method solves the problem that the traditional method of obtaining the allowable attitude control law of the aircraft dynamics model by solving the HJI equation is difficult, and reduces the difficulty of solving the allowable attitude control law of the attitude control system with uncertain terms.
[0005] The technical solution of this invention is: an online learning method for neural network control laws, comprising: Based on the differentiation rule of rotating coordinate systems and the momentum distance theorem, a dynamic model of the aircraft is constructed; Based on the aircraft dynamics model and the game theory approach, an adaptive dynamic programming problem is constructed and the allowable control law is solved. In the admissible control law, based on the fundamental theorem of Weierstrass, a polynomial basis automation scheme is designed, and the Actor-Critic neural network basis is selected. Based on the discretization approximation method, the online weights of the Actor-Critic neural network are updated to obtain an intelligent control strategy based on a single-evaluation neural network.
[0006] The process of constructing a dynamic model of the aircraft based on the differentiation rules of rotating coordinate systems and the momentum distance theorem includes: Based on the angular momentum theorem and the axisymmetric shape of the aircraft, a system with torque input is established. , , A dynamic model of a hypersonic vehicle rotating about a fixed axis was constructed, and the angular velocity vector of the hypersonic vehicle in the body coordinate system was calculated. State equations as system state variables; By analyzing the forces in the airflow coordinate system, the dynamic equation of the aircraft's center of mass is obtained; By integrating the attitude angular dynamics and angular velocity dynamics of hypersonic vehicles, a set of equations is obtained, with the angle of attack, sideslip angle, roll angle, and angular velocity vector based on the body coordinate system as the system state variables.
[0007] The angular velocity vector of the hypersonic vehicle in the body coordinate system The state equation, which is the system state variable, is:
[0008] In the formula , , These are respectively the rolling moment, pitching moment, and yaw moment. , , These are the moments of inertia.
[0009] The dynamic equations of the spacecraft's center of mass are as follows:
[0010] In the formula, m For quality, g It is the acceleration due to gravity. L , Y These represent the lift and lateral forces acting on a hypersonic vehicle, respectively. For the trajectory inclination angle of a hypersonic vehicle, For the ballistic deflection angle of a hypersonic vehicle, For the angle of attack of hypersonic vehicles, For the sideslip angle of a hypersonic vehicle, The tilt angle of a hypersonic vehicle. Let be the angular velocity vector of the hypersonic vehicle in the body coordinate system. This refers to flight speed.
[0011] The system of equations using angle of attack, sideslip angle, roll angle, and angular velocity vector based on body coordinates as the system state variables is as follows: .
[0012] The process of constructing an adaptive dynamic programming problem and solving for an admissible control law based on game theory includes: Based on the controlled model with attitude angle and angular velocity as state variables, a dynamic model of the tracking error is established:
[0013] In the formula, For attitude tracking error, , , For the input matrix, Let be the partial derivative matrix of the torque with respect to the rudder deflection angle. For rudder deflection control, , It is a comprehensive nonlinear quantity in attitude angular dynamics, encompassing factors other than torque, factors other than rudder deflection in torque, model uncertainties, and external disturbances. For the desired attitude information, The superscript d indicates the expected value; The augmented system of the tracking error dynamics model is represented as:
[0014] in, To broaden the variables, , , , It is a generalized disturbance term, that is, the comprehensive uncertainty quantity that counteracts rudder deflection. It is the input matrix corresponding to the generalized interference term; Based on the augmented system, a performance index function for solving the admissible control law is constructed:
[0015] In the formula The integral index represents the integral from the initial moment to infinity. , , These are the weight matrices for the error vector, control input, and disturbance vector, respectively. At the initial moment of system evolution, For integration variables; The objective function of the bilateral optimal problem is:
[0016] In the formula, This represents the optimal integral index. , These represent the minimum control force and the maximum generalized disturbance, respectively. Based on the concept of game theory, the external disturbance term and the reinforcement learning control strategy are regarded as two sides in a game, and a differential game problem is constructed based on the objective function. Based on the classical variational method, the index-optimal solution under dynamic constraints is obtained, leading to the admissible control law: , , in, and These represent optimal control and optimal generalized disturbance, respectively.
[0017] The aforementioned scheme for automating polynomial basis design, based on Weierstrass's fundamental theorem and selecting an Actor-Critic neural network basis, includes: Based on the dynamics model and accuracy requirements, determine the expected order of magnitude of the error between the approximate Actor-Critic neural network and the true value; Based on the obtained error order of magnitude, determine the highest power of the polynomial basis of the polynomial neural network and initialize the parameters of the Actor-Critic neural network. Generate a polynomial basis by combining and arranging the highest power obtained; Based on the actual and approximate quantities of the dynamic model, the obtained Actor-Critic polynomial neural network parameters and the obtained polynomial basis are straightened to achieve the backpropagation condition of the neural network. The neural network parameters trained by the integral reinforcement learning algorithm are combined with information such as relative parameter error, zero-point deviation, and prior conditions to perform basis screening, resulting in a polynomial basis suitable for the dynamic model and accuracy requirements. Using the highest power of the obtained polynomial basis and the number of variables in the dynamic model as inputs, and the polynomial basis vectors and their derivative matrices as outputs, an automated polynomial basis scheme is constructed. The optimal value of the polynomial basis automation scheme is selected as the polynomial basis of the Actor-Critic neural network.
[0018] The step of updating the online weights of the Actor-Critic neural network according to the discretization approximation method to obtain an intelligent control strategy based on a single-evaluation neural network includes: Considering the addition of a discount factor to the infinite time-domain index, the recursive form is obtained from the Bellman equation:
[0019] in, Indicates policy-based pairs From the first Expected value from start to finish, time interval , It is designed as a reward function for optimal game control problems; definition A function, specifically a value function, uses both the controlled state and the policy as independent variables:
[0020] in, ; Based on the Bellman optimal equation, the optimal function and Conditions to be met:
[0021] In the formula, This represents a generalized control strategy. Indicates by strategy The minimum value obtained, Let f(h) represent the probability matrix for transitioning from step h to the next step. Based on reinforcement learning theory, the Bellman equation for the Q function is obtained:
[0022] Based on the Bellman optimal equation, the optimal control strategy is obtained. , in Indicates the first Step, based on the current state The optimal strategy that converges conditionally with probability 1. This indicates convergence with probability 1; Based on the discretization approximation method, a marker of the gradient direction with the same form as the traditional discrete time difference error is calculated; Based on the radial basis function (RBF), an approximate value for the Q function is calculated:
[0023] in, Let be a radial basis function vector, satisfying:
[0024] in, and It is the first The center point and width of a Gaussian function , ; It is the number of RBF neural networks; Based on the approximation of the Q function, the squared error is obtained. E Relative to parameter vector gradient
[0025] in, , Further, the gradient direction is identified.
[0026]
[0027] Update using gradient descent. Optimal parameters ,Will As parameters of the Actor-Critic neural network; Based on the obtained Actor-Critic neural network parameters, the optimal strategy is... Parameterized, thereby obtaining an intelligent control strategy based on a single-evaluation neural network:
[0028]
[0029] in, These are the basis function vectors used to evaluate the network.
[0030] The calculation of the gradient direction indicator, which has the same form as the traditional discrete time difference error, includes: Define the reward function Used to reflect the quality of the previously chosen actions.
[0031] in, , They represent , The learning rate; Define internal return as TD error for:
[0032] in, As a discount factor, ; Approaching Euler from behind. for:
[0033] in, , ; Define the flag for the gradient direction :
[0034] in, , .
[0035] The advantages of this invention compared to the prior art are: (1) In view of the fact that the trajectory and attitude of the aircraft change drastically during the flight mission, and the aircraft is affected by the uncertainty of the system model parameters, and the system's disturbance rejection and output tracking capabilities are not necessarily optimal at the same time, a neural network differential game strategy is introduced. The matching disturbance term caused by the uncertainty of the model parameters and the attitude reinforcement learning control strategy are used as the two sides of the game. A control method based on Actor-Critic neural network differential game is proposed. This reduces the tediousness of manually designing the weights of the disturbance rejection index and the error index, overcomes the conflict between the robustness and optimality of the control system, realizes the adaptive dynamic programming controller design, and solves the problem of efficient design of neural network control law.
[0036] (2) A polynomial uniform approximation method based on Weierstrass's fundamental theorem is used to select the Actor-Critic network basis. The expected error magnitude between the approximate network and the true value is determined based on the specific dynamic model and accuracy requirements. The polynomial basis is then raised to a higher power, and combinations of the highest powers are used to generate the polynomial basis. The polynomial neural network parameters and basis are straightened according to the actual model and approximations to achieve the backpropagation condition of the neural network. The parameters are trained using an integral reinforcement learning algorithm to screen the basis, obtaining a polynomial basis suitable for the dynamic model and accuracy requirements. This approach significantly improves the efficiency of dynamic polynomial basis selection and reduces the trial-and-error cost of dynamic polynomial basis selection.
[0037] (3) An online weight update law for the Actor-Critic neural network is designed for updating the entire neural network structure. This ensures that the neural network weights converge to the ideal weights while also guaranteeing the stability of the entire closed-loop control. This mechanism avoids the problem of neural network weight non-convergence and greatly improves the robustness of the neural network architecture.
[0038] (4) The continuous-time reinforcement learning given by the Bellman equation requires an integrator and the analytical solution is difficult to process. The discretization approximation method proposed in this paper has good operability and can provide a useful reference for the design of control laws based on the reinforcement learning framework. It has good application value.
[0039] (5) In addition to being used for aircraft attitude control laws, this method can also be used as a generalized anti-disturbance control framework. It transforms the problem of balancing the disturbance terms caused by the uncertainty of nonlinear systems and the system control strategy into a game problem. By minimizing the quadratic performance function containing tracking error, control quantity and disturbance, the solution of the HJI equation is obtained to obtain the robust optimal solution. The index function can be flexibly set according to the complexity of the problem, and the approximate reinforcement learning analytical solution can be discretized according to the index requirements to complete the design of control strategies with different index requirements, thereby improving the intelligence level of the neural network differential game control algorithm. Attached Figure Description
[0040] Figure 1 This is a block diagram of the control principle.
[0041] Figure 2 A flowchart for dynamic selection of polynomial bases.
[0042] Figure 3 This is the basis selection process for multinomial neural networks.
[0043] Figure 4 This is a flowchart of an automated solution for polynomial bases.
[0044] Figure 5 This refers to the tracking status of the angle of attack command under nominal conditions.
[0045] Figure 6 This is a graph showing the neural network weight update for the nominal case.
[0046] Figure 7 Tracking the angle of attack command with parameters adjusted downwards.
[0047] Figure 8 The graph shows the weight update curve of the neural network under parameter bias. Detailed Implementation
[0048] The specific embodiments of the present invention, as well as specific examples under nominal and parameter bias scenarios, will be described in further detail below with reference to the accompanying drawings.
[0049] (1) Constructing an uncertain nonlinear system model Considering the dynamic equations of rotational motion of a rigid body aircraft are
[0050] In the formula This is the torque vector acting on the aircraft. The missile has a symmetrical shape, therefore...
[0051] Therefore, the dynamic equation of the aircraft around its center is:
[0052] In the formula , , These are the rolling moment, pitching moment, and yaw moment, respectively.
[0053] By analyzing the forces in the airflow coordinate system, the dynamic equations of the aircraft's center of mass can be obtained.
[0054] In the formula, m For quality, g It is the acceleration due to gravity. L , Y The lift and lateral forces acting on a hypersonic vehicle. D The drag experienced by hypersonic vehicles For the ballistic tilt angle of hypersonic vehicles, For the ballistic deflection angle of a hypersonic vehicle, For the angle of attack of hypersonic vehicles, For the sideslip angle of a hypersonic vehicle, The tilt angle of a hypersonic vehicle. Let be the angular velocity vector of the hypersonic vehicle in the body coordinate system. This refers to flight speed.
[0055] In summary, the model for the uncertain nonlinear system is as follows:
[0056] (2) Design of a differential game controller based on adaptive dynamic programming For a model with external wind field disturbances, the vector form of its attitude control system is expressed as follows:
[0057] in, , , For the input matrix, Let be the partial derivative matrix of the torque with respect to the rudder deflection angle. For rudder deflection control, , It is a comprehensive nonlinear quantity combining factors other than torque in attitude angle dynamics, factors other than rudder deflection in torque, model uncertainties, and external disturbances.
[0058] Figure 1 A block diagram of the control principle is given. Assume the desired attitude information is represented as follows: Its derivative The superscript and subscript 'd' denote the expected value. Therefore, the tracking error is defined as... Its attitude error system can be expressed as:
[0059] For the attitude tracking error system, define auxiliary variables. The augmented system representation of the attitude tracking control system is further expressed as follows:
[0060] in, To broaden the variables, , , , It is a generalized disturbance term, that is, the comprehensive uncertainty quantity that counteracts rudder deflection. It is the input matrix corresponding to the generalized interference term.
[0061] For adaptive dynamic programming control problems with matching uncertainties, the goal is to find an admissible control law. This minimizes the following performance function.
[0062] In the formula The integral index represents the integral from the initial moment to infinity. , , These are the weight matrices for the error vector, control input, and disturbance vector, respectively. At the initial moment of system evolution, Let be the integral variable. The objective function of the bilateral optimal problem is:
[0063] In the formula, Optimal integral index , These represent the minimum control force and the maximum generalized disturbance, respectively.
[0064] The variational method can be used to solve the optimization problem under nonlinear dynamic constraints. , , and These represent optimal control and optimal generalized disturbance, respectively.
[0065] The optimal index function mentioned above needs to be obtained by solving the following HJI equation, which is as follows:
[0066] Considering the direct algebraic solution of the HJI equation to obtain The solution is difficult, so this patent uses a single Critic network to approximate the optimal game strategy formula, thereby reducing the difficulty of solving the problem.
[0067] (3) Selection of Critic network basis Dynamic polynomial base selection primarily aims to reduce redundant polynomial bases provided by the front end to obtain the minimum effective polynomial base. The dynamic polynomial base selection process is divided into importance-based base selection and brute-force base selection. The process is shown in the attached diagram. Figure 2 As shown.
[0068] The method for selecting the basis of the multinomial neural network used in this project is shown in the attached figure. Figure 3 As shown, firstly, the expected error magnitude between the approximate network and the true value is determined based on the specific dynamic model and accuracy requirements. Then, based on the error magnitude and mathematical methods, the highest power of the polynomial basis of the polynomial neural network is determined. The polynomial basis is generated by combining and arranging the highest powers. The parameters and basis of the polynomial neural network are straightened according to the actual model and approximations to achieve the backpropagation condition of the neural network. Finally, the parameters are trained based on the integral reinforcement learning algorithm, and the basis is selected by fusing information such as relative parameter error, zero-point deviation, and prior conditions to obtain a polynomial basis suitable for the dynamic model and accuracy requirements.
[0069] The process of the polynomial basis automation solution is shown in the appendix. Figure 4 As shown, the method takes the highest power of the polynomial basis and the number of variables in the dynamic model as input, and outputs the polynomial basis vectors and their derivative matrices. Its main function is to automatically generate polynomial basis vectors and their derivative matrices. This scheme significantly improves the efficiency of dynamic polynomial basis selection and reduces the trial-and-error cost. The automated polynomial basis scheme is mainly based on symbolic computation, and its implementation relies heavily on library function calls, making the process relatively easy to implement. After integral reinforcement learning training, the parameters corresponding to each polynomial neural network basis can be obtained. From the training results, the relative error of the variation amplitude of each basis and the zero-point deviation can be calculated, providing a basis for judging the approximation effect of each polynomial basis. Then, based on the descent power and point-value representation, n+1 polynomials with powers from 0 to n are selected as a basis multiple times. Finally, the optimal value among them is selected as the basis of the polynomial neural network through the approximation results.
[0070] (4) Online weight update of Critic neural network Define the following value function for an admissible control strategy.
[0071] The recurrence relation can be obtained from the Bellman equation.
[0072] Among them, among them, Indicates policy-based pairs From the first The expected value from the start to the final moment. , It is designed as the reward function for the optimal game control problem.
[0073] definition The function, that is, the value function with the controlled state and policy as independent variables, is defined as follows:
[0074] in, Based on the Bellman optimal equation, the optimal function... and satisfy
[0075]
[0076] According to reinforcement learning theory, for any fixed policy The following formula holds true:
[0077] in, The Bellman equation for the Q function
[0078] Based on the Bellman optimality equation, the optimal control strategy is obtained.
[0079] in Indicates the first Step, based on the current state The optimal strategy that converges conditionally with probability 1. This indicates convergence with probability 1.
[0080] reward function Used to reflect the quality of previously chosen actions, defined as:
[0081] in, , They represent , The learning rate. The internal reward is the TD error. Defined as:
[0082] in, As a discount factor, This patent takes Afterwards, it approaches Euler. for
[0083] in, , The above formula is
[0084] The above formula has the same form as the traditional discrete time difference error.
[0085] in, , . It serves as a marker of gradient direction to improve motion performance. However, The Q-function still needs to be calculated, and it is approximated using radial basis functions (RBFs). The approximation of the Q-function is given by... Given, among which
[0086] in, and It is the first The center point and width of a Gaussian function , ; This represents the number of RBF neural networks. Next, we will derive the update formulas for each parameter.
[0087] Update using gradient descent. Parameters in The squared error relative to the parameter vector The gradient is
[0088] in, , Further obtained
[0089] Optimal differential strategy It is parameterized, thereby obtaining an intelligent control strategy based on a single-evaluation neural network.
[0090]
[0091]
[0092] in, These are the basis function vectors used to evaluate the network.
[0093] Example 1: Attitude tracking under nominal conditions Under nominal conditions, the aerodynamic coefficients (axial force coefficient, normal force coefficient, lateral force coefficient, pitching moment coefficient, yaw moment coefficient, roll moment coefficient) and moments of inertia (directional moment of inertia, directional moment of inertia, directional moment of inertia) are all ideal values. The designed controller is applied to the warhead to achieve trajectory tracking. The tracking results at three attitude angles—angle of attack, sideslip angle, and roll angle—are shown in the figure. It can be seen that when the trajectory changes at a certain attitude angle, the attitude angle can quickly track the current trajectory. (See attached figure.) Figure 5 and attached Figure 6 .
[0094] Example 2: Attitude tracking under parameter bias Considering deviations in rotational inertia caused by fuel sloshing and consumption, and atmospheric density deviations due to factors such as temperature and season, parameter pull tests were conducted on the attitude system. The parameter deviations and their ranges are listed in Table 1 below, and the control effects are shown in the attached figure. Figure 7 and attached Figure 8 As shown.
[0095] Table 1. Error band of parameters used in simulation
[0096] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications to the technical solutions of the present invention based on the above-disclosed technical content without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall fall within the protection scope of the technical solutions of the present invention.
Claims
1. An online learning method for neural network control laws, characterized in that, include: Based on the differentiation rule of rotating coordinate systems and the momentum distance theorem, a dynamic model of the aircraft is constructed; Based on the aircraft dynamics model and the game theory approach, an adaptive dynamic programming problem is constructed and the allowable control law is solved. In the admissible control law, based on the fundamental theorem of Weierstrass, a polynomial basis automation scheme is designed, and the Actor-Critic neural network basis is selected. Based on the discretization approximation method, the online weights of the Actor-Critic neural network are updated to obtain an intelligent control strategy based on a single-evaluation neural network.
2. The online learning method for neural network control laws according to claim 1, characterized in that, The process of constructing a dynamic model of the aircraft based on the differentiation rules of rotating coordinate systems and the momentum distance theorem includes: Based on the angular momentum theorem and the axisymmetric shape of the aircraft, a system with torque input is established. , , A dynamic model of a hypersonic vehicle rotating about a fixed axis was constructed, and the angular velocity vector of the hypersonic vehicle in the body coordinate system was calculated. State equations as system state variables; By analyzing the forces in the airflow coordinate system, the dynamic equation of the aircraft's center of mass is obtained; By integrating the attitude angular dynamics and angular velocity dynamics of hypersonic vehicles, a set of equations is obtained, with the angle of attack, sideslip angle, roll angle, and angular velocity vector based on the body coordinate system as the system state variables.
3. The online learning method for neural network control laws according to claim 2, characterized in that, The angular velocity vector of the hypersonic vehicle in the body coordinate system The state equation, which is the system state variable, is: In the formula , , These are respectively the rolling moment, pitching moment, and yaw moment. , , These are the moments of inertia.
4. The online learning method for neural network control laws according to claim 3, characterized in that, The dynamic equations of the spacecraft's center of mass are as follows: In the formula, m For quality, g It is the acceleration due to gravity. L , Y These represent the lift and lateral forces acting on a hypersonic vehicle, respectively. For the trajectory inclination angle of a hypersonic vehicle, For the ballistic deflection angle of a hypersonic vehicle, For the angle of attack of hypersonic vehicles, For the sideslip angle of a hypersonic vehicle, The tilt angle of a hypersonic vehicle. Let be the angular velocity vector of the hypersonic vehicle in the body coordinate system. This refers to flight speed.
5. The online learning method for neural network control laws according to claim 4, characterized in that, The system of equations using angle of attack, sideslip angle, roll angle, and angular velocity vector based on body coordinates as the system state variables is as follows: 。 6. The online learning method for neural network control laws according to claim 5, characterized in that, The process of constructing an adaptive dynamic programming problem and solving for an admissible control law based on game theory includes: Based on the controlled model with attitude angle and angular velocity as state variables, a dynamic model of the tracking error is established: In the formula, For attitude tracking error, , , For the input matrix, Let be the partial derivative matrix of the torque with respect to the rudder deflection angle. For rudder deflection control, , It is a comprehensive nonlinear quantity in attitude angular dynamics, encompassing factors other than torque, factors other than rudder deflection in torque, model uncertainties, and external disturbances. For the desired attitude information, The superscript d indicates the expected value; The augmented system of the tracking error dynamics model is represented as: in, To broaden the variables, , , , It is a generalized disturbance term, that is, the comprehensive uncertainty quantity that counteracts rudder deflection. It is the input matrix corresponding to the generalized interference term; Based on the augmented system, a performance index function for solving the admissible control law is constructed: In the formula The integral index represents the integral from the initial moment to infinity. , , These are the weight matrices for the error vector, control input, and disturbance vector, respectively. At the initial moment of system evolution, For integration variables; The objective function of the bilateral optimal problem is: In the formula, This represents the optimal integral index. , These represent the minimum control force and the maximum generalized disturbance, respectively. Based on the concept of game theory, the external disturbance term and the reinforcement learning control strategy are regarded as two sides in a game, and a differential game problem is constructed based on the objective function. Based on the classical variational method, the index-optimal solution under dynamic constraints is obtained, leading to the admissible control law: , , in, and These represent optimal control and optimal generalized disturbance, respectively.
7. The online learning method for neural network control laws according to claim 6, characterized in that, The aforementioned scheme for automating polynomial basis design, based on Weierstrass's fundamental theorem and selecting an Actor-Critic neural network basis, includes: Based on the dynamics model and accuracy requirements, determine the expected order of magnitude of the error between the approximate Actor-Critic neural network and the true value; Based on the obtained error order of magnitude, determine the highest power of the polynomial basis of the polynomial neural network and initialize the parameters of the Actor-Critic neural network. Generate a polynomial basis by combining and arranging the highest power obtained; Based on the actual and approximate quantities of the dynamic model, the obtained Actor-Critic polynomial neural network parameters and the obtained polynomial basis are straightened to achieve the backpropagation condition of the neural network. The neural network parameters trained by the integral reinforcement learning algorithm are combined with information such as parameter relative error, zero-point deviation, and prior conditions to perform basis screening, resulting in a polynomial basis suitable for the dynamic model and accuracy requirements. Using the highest power of the obtained polynomial basis and the number of variables in the dynamic model as inputs, and the polynomial basis vectors and their derivative matrices as outputs, an automated polynomial basis scheme is constructed. The optimal value of the polynomial basis automation scheme is selected as the polynomial basis of the Actor-Critic neural network.
8. The online learning method for neural network control laws according to claim 7, characterized in that, The step of updating the online weights of the Actor-Critic neural network according to the discretization approximation method to obtain an intelligent control strategy based on a single-evaluation neural network includes: Considering the addition of a discount factor to the infinite time-domain index, the recursive form is obtained from the Bellman equation: in, Indicates policy-based pairs From the first Expected value from start to finish, time interval , It is designed as a reward function for optimal game control problems; definition A function, specifically a value function, uses both the controlled state and the policy as independent variables: in, ; Based on the Bellman optimal equation, the optimal function and Conditions to be met: In the formula, This represents a generalized control strategy. Indicates by strategy The minimum value obtained, Let f(h) represent the probability matrix for transitioning from step h to the next step. Based on reinforcement learning theory, the Bellman equation for the Q function is obtained: Based on the Bellman optimal equation, the optimal control strategy is obtained. , in Indicates the first Step, based on the current state The optimal strategy that converges conditionally with probability 1. This indicates convergence with probability 1; Based on the discretization approximation method, a marker of the gradient direction with the same form as the traditional discrete time difference error is calculated; Based on the radial basis function (RBF), an approximate value for the Q function is calculated: in, Let be a radial basis function vector, satisfying: in, and It is the first The center point and width of a Gaussian function , ; It is the number of RBF neural networks; Based on the approximation of the Q function, the squared error is obtained. E Relative to parameter vector gradient in, , Further, the gradient direction is identified. Update using gradient descent. Optimal parameters ,Will As parameters of the Actor-Critic neural network; Based on the obtained Actor-Critic neural network parameters, the optimal strategy is... The parameters are then used to obtain an intelligent control strategy based on a single-evaluation neural network. in, These are the basis function vectors used to evaluate the network.
9. The online learning method for neural network control laws according to claim 8, characterized in that, The calculation of the gradient direction indicator, which has the same form as the traditional discrete time difference error, includes: Define the reward function Used to reflect the quality of the previously chosen actions. in, , They represent , The learning rate; Define internal return as TD error for: in, As a discount factor, ; Approaching Euler from behind. for: in, , ; Define the flag for the gradient direction : in, , .