An autonomous driving control method based on BLF-SRL
Through the safety reinforcement learning algorithm based on the obstacle Liyapunov function, a safety reinforcement learning method for the autonomous driving control system is constructed, and the safety and efficiency problems in the learning process in the autonomous driving control system are solved, and safety assurance in changing scenarios is achieved.
Patent Information
- Application Number
- CN202210712700.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-22
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-06-22
AI Technical Summary
The existing reinforcement learning methods in the autonomous driving control system have problems such as strong data dependence, low online learning efficiency, easy learning in non-stationary environments, and difficulty in ensuring safety during the learning process.
Using the secure reinforcement learning algorithm (BLF-SRL) based on the obstacle Liyapunov function, a nonlinear system with strict feedback form is designed, combined with the reverse-step optimization method and the Actor-Critic framework, a secure reinforcement learning algorithm is designed to achieve effective updates of system state constraints and error signals to ensure the safety in the learning process.
It improves the safety and learning efficiency of the autonomous driving control system in changing scenarios, ensures the safety performance in the learning process, and solves the safety and efficiency problems of reinforcement learning methods in autonomous driving.
Smart Images

Figure CN115016278B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automatic driving control systems, and in particular to an automatic driving control method based on BLF-SRL. Background Art
[0002] The field of reinforcement learning has gone through a period of research. Initially, it was mostly based on table learning of discrete states and actions. However, when it comes to learning methods involving continuous state and action spaces, the high-dimensional space formed will cause the curse of dimensionality. It is usually necessary to use function approximation methods to represent state-value functions and state-action-value functions. With the development of deep learning technology, based on the powerful function approximation ability of deep neural networks, deep reinforcement learning has been applied and developed in strategy games and control. Algorithms such as DQN and DDPG have been proposed and effectively verified. Since autonomous vehicles need to face complex dynamic environments and multi-scenario generalization and interactive characteristics, existing research widely uses reinforcement learning with interactive feedback for decision-making and control.
[0003] However, autonomous driving control systems are a type of system with safety-critical (SC) characteristics. Existing reinforcement learning methods have difficulties in making adaptive interactive behavior decisions, and the safety and adaptive performance of motion control systems under changing working conditions are also difficult to guarantee. Therefore, it is necessary to propose a method to solve the problems of reinforcement learning based on trial and error, such as strong data dependence, low online learning efficiency, easy failure of learning based on non-stationary environments, and difficulty in ensuring safety during the learning process. Summary of the Invention
[0004] The purpose of the present invention is to provide an automatic driving control method based on BLF-SRL in order to overcome the defects of the above-mentioned prior art.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] A BLF-SRL-based automatic driving control method comprises the following steps:
[0007] Step 1: Construct a secure reinforcement learning algorithm based on the barrier Lyapunov function;
[0008] Step 2: Model the autonomous driving control system as a nonlinear system in strict feedback form;
[0009] Step 3: Use the barrier Lyapunov function-based safety reinforcement learning algorithm in step 1 to achieve the safety of some system state constraints during the learning and updating process of the autonomous driving control system and the effectiveness of the error signal in each backstepping subsystem.
[0010] In step 1, the process of the secure reinforcement learning algorithm based on the barrier Lyapunov function specifically includes the following steps:
[0011] Step 101: Reconstruct the nonlinear system in strict feedback form into an error system;
[0012] Step 102: Use backstepping optimization method and BLF to design the optimal control law for each subsystem;
[0013] Step 103: defining the Bellman optimality condition for each subsystem according to the Bellman optimality principle;
[0014] Step 104: Lyapunov analysis is used to design the error update signal for each subsystem. During the learning process, the virtual control of the subsystem is optimized by iteratively updating the unknown function terms in each subsystem to achieve optimization of the overall system control.
[0015] The subsystem includes z1 subsystem, z i (i=2,...,n-1) subsystem and z n subsystem.
[0016] In step 101, the nonlinear system in strict feedback form is:
[0017]
[0018] Among them, f j (j=1,2,...,n) and g j (j=1,2,...,n) are the models required to define the second-order strict feedback form of nonlinear system, n is the number of subsystems, is the state variable, is the state vector, is the control input, is the system output;
[0019] In order to optimize the system control to achieve the desired system output y d , introduce the virtual control α to be optimized i (i=1,...,n-1), define the error state z1=x1-y d and z i =x i -α i-1 (i=2,...,n), the nonlinear system to be optimized is re-established as an error system:
[0020]
[0021] Among them, z j(j=1,2,...,n) is the error state of the jth subsystem, f j (j=1,2,...,n) and g j (j=1,2,...,n) are the models required to define the second-order strict feedback form of nonlinear system, n is the number of subsystems, y d is the expected output of the system;
[0022] The error system presents a cascade structure, and each virtual control α introduced by optimization i (i=1,...,n-1) Finally, the overall control of the optimized system is achieved, and all state variables z=[z1,...,z n ] T Divided into state variables to be constrained and free state variables Among them, n s In order to ensure the continuity of segmentation points, the learning problem is described as:
[0023] Throughout the learning process, the optimization system controls the tracking system's expected output y d At the same time, some state variables z i ,(i=1,...,n s ) Always stay in the safe area of the design Inside, among them, Is a positive number.
[0024] In step 102, the process of using the backstepping optimization method and BLF to design the optimal control law for each subsystem is specifically as follows:
[0025] Based on the backstepping optimization method, the Actor-Critic framework of reinforcement learning is adopted in each subsystem, which is defined as Sub-Actor and Sub-Critic respectively. The backstepping subsystem in which the virtual control quantity is located is designed based on the barrier Lyapunov function; for the free state variables The backstepping subsystem is designed for virtual control or system control input based on the quadratic Lyapunov function.
[0026] In step 103, the process of defining the Bellman optimality condition of each subsystem according to the Bellman optimality principle is specifically as follows:
[0027] Sub-Actor and Sub-Critic are decomposed into BLF / QLF terms and unknown function terms approximated by independent neural networks respectively, and the Bellman optimality conditions of the subsystems are defined according to the Bellman optimality principle.
[0028] In steps 102 to 104, for the z1 subsystem, the backstepping optimization method and BLF are used to design the optimal control law of the z1 subsystem, and the Bellman optimal condition of the z1 subsystem is defined. Then, the process of designing the error update signal is specifically as follows:
[0029] The virtual control to be optimized is introduced into the z1 subsystem, and the optimal performance index function of the z1 subsystem is defined as:
[0030]
[0031] in, is the optimal performance index function of the z1 subsystem, is the cost function, is the optimal virtual control, κ 1s and κ 1c are weight coefficients respectively, and the corresponding HJB equation The expression is:
[0032]
[0033] in, represents the partial derivative of the optimal performance index function with respect to z1, and f1 and g1 are the models required to establish the nonlinear system to be optimized;
[0034] because Established and has a unique solution, by solving Get optimal virtual control for:
[0035]
[0036] Optimal Virtual Control The design is broken down into:
[0037]
[0038] in, is the unknown continuous function to be learned, κ1 is a positive constant, and the optimal virtual control after decomposition design is The partial derivative of the optimal performance index function can be obtained The expression is:
[0039]
[0040] In the z1 subsystem, the partial derivative of the optimal performance index function and optimal virtual control They are all unknown functions, and the uncertain terms are approximated by independent neural networks. According to the optimal virtual control after decomposition design and the partial derivative of the optimal performance index function Get its estimated value and Then, under the Actor-Critic framework, strategy evaluation and strategy improvement are carried out. is the estimated value of the optimal virtual control, defined as Sub-Actor a1, is the partial derivative of the optimal performance index function The estimated value of is defined as Sub-Criticc1;
[0041] Due to the nonlinear characteristics of the HJB equation, the optimal solution of the analytical form cannot be directly obtained. In order to iteratively obtain its numerical solution, two independent neural networks are first used to approximate the partial derivatives of the optimal performance index function. and optimal virtual control The unknown term in breaks the partial derivative of the optimal performance index function Optimal Virtual Control The correlation between them; then the neural network is iteratively updated through strategy evaluation and strategy improvement under the Actor-Critic framework to update the estimated value and Eventually, the two gradually meet the correlation relationship Then the system can be optimized and controlled;
[0042] Optimal virtual control Estimated value of The expression is:
[0043]
[0044] in, is the expected output of Sub-Actor NN;
[0045] Partial derivatives of the optimal performance index function Estimated value of The expression is:
[0046]
[0047] in, is the expected output of the Sub-Critic NN;
[0048] Optimal Virtual Control Estimated value of The expression of and the partial derivative of the optimal performance index function Estimated value of Substitute the expression into the HJB equation Then we get the estimated value of HJB equation The expression is:
[0049]
[0050] Obtain the Bellman optimality condition in the z1 subsystem. The expression of the Bellman optimality condition in the z1 subsystem is:
[0051]
[0052] In Sub-Criticc1, perform current virtual control The ultimate goal of strategy evaluation is to make the estimated value of the HJB equation To reach the optimal value, In Sub-Actora1, Sub-Criticc1 strategy evaluation is used to improve the strategy, and finally Bellman optimality conditions are achieved through iterative learning;
[0053] Define Bellman residual The expression is:
[0054]
[0055] The expressions of the update equations of Sub-Critic NN and Sub-Actor NN are:
[0056]
[0057]
[0058] in, The error variable required for the Sub-Critic NN update equation, The error variable required for the Sub-Actor NN update equation, and are the learning rates of Sub-Critic NN and Sub-Actor NN respectively;
[0059] Finally, in the z1 subsystem, the optimal virtual control and the partial derivative of the optimal performance index function The estimation is performed, and the Sub-Critic NN and Sub-Actor NN are further iteratively learned through their update equations to meet the Bellman optimality condition.
[0060] In the steps 102 to 104, for z i Subsystem, the backstepping optimization method and BLF are used to design the optimal control law of the z1 subsystem, and the z iThe Bellman optimal condition of the subsystem and the process of designing the error update signal are as follows:
[0061] In z i The virtual control α to be optimized is introduced into the subsystem i , its optimal value is The optimal performance index function is defined as:
[0062]
[0063] in, is the cost function, κ is and κ ic are weight coefficients respectively, and the corresponding HJB equation The expression is:
[0064]
[0065] in, Denotes the optimal performance index function for z i Find the partial derivative by solving Get optimal virtual control The expression is:
[0066]
[0067] Optimal Virtual Control The design is broken down into:
[0068]
[0069] Among them, κ i is a positive constant, is the unknown continuous function to be learned, For dummy control variables Estimated value of Derivative, α i,aux is an auxiliary dummy control variable, and its expression is:
[0070]
[0071] in, For z i-1 The expected output of the Sub-Actor NN corresponding to the subsystem, and Corresponding to z i and z i-1 The coefficient of the subsystem to be calibrated, n s To ensure the continuity of segmentation points, represents the optimal virtual control The cost function of
[0072] The optimal virtual control after decomposition Get the partial derivative of the optimal performance index function The expression is:
[0073]
[0074] z i The subsystem is similar to the z1 subsystem. and The uncertainties in the equation are approximated by independent neural networks, and the optimal virtual control and the partial derivative of the optimal performance index function Get its estimated value and Then, under the Actor-Critic framework, strategy evaluation and strategy improvement are carried out. is a dummy control variable The estimated value of Sub-Actor i , is the partial derivative of the optimal performance index function The estimated value of Sub-Criticc i ;
[0075] Dummy control variables Estimated value of The expression is:
[0076]
[0077] in, is the expected output of Sub-Actor NN;
[0078] Partial derivatives of the optimal performance index function Estimated value of The expression is:
[0079]
[0080] in, is the expected output of Sub-Critic NN;
[0081] The dummy control variable Estimated value of The expression of and the partial derivative of the optimal performance index function Estimated value of Substitute the expression of into the partial derivative of the optimal performance index function Estimated value of The expression of , and then get the estimated value of HJB equation The expression is:
[0082]
[0083] Get in z i Bellman optimality conditions in the subsystem, at z i The expression of Bellman optimality condition in the subsystem is:
[0084]
[0085] z i The Bellman optimality condition of the subsystem is achieved through iterative calculation of policy evaluation and policy improvement in the Actor-Critic framework. i In the current virtual control The ultimate goal of strategy evaluation is to make the estimated value of the HJB equation To reach the optimal value, In Sub-Actora i In the algorithm, Critic strategy evaluation is used to improve the strategy, and finally Bellman optimality condition is achieved through iterative learning;
[0086] Define Bellman residual The expression is:
[0087]
[0088] The update equations of Sub-Critic NN and Sub-Actor NN are:
[0089]
[0090]
[0091] in, The error variable required for the Sub-Critic NN update equation, The error variable required for the Sub-Actor NN update equation, and are the learning rates of Sub-Critic NN and Sub-Actor NN respectively;
[0092] In z i In the subsystem, the optimal virtual control and the partial derivative of the optimal performance index function The estimation is performed, and the Sub-Critic NN and Sub-Actor NN are further iteratively learned through their update equations, and finally the Bellman optimality condition is satisfied.
[0093] In the steps 102 to 104, for z n Subsystem, the backstepping optimization method and BLF are used to design the optimal control law of the z1 subsystem, and the z n The Bellman optimal condition of the subsystem and the process of designing the error update signal are as follows:
[0094] In z n In the subsystem, the optimal system control input u is u * , define z n The expression of the optimal performance index function of the subsystem is:
[0095]
[0096] in, z n The optimal performance index function of the subsystem, is the cost function, κ ns and κ nc are all weight coefficients, and the corresponding HJB equation The expression is:
[0097]
[0098] in, Denotes the optimal performance index function for z n Find the partial derivative, optimal system control input u * By solving get:
[0099]
[0100] The optimal system control input u * Breaks down to:
[0101]
[0102] in, is the unknown continuous function to be learned, κ n is a positive constant, α n,aux is an auxiliary dummy control variable, and its expression is:
[0103]
[0104] in, zn-1 The expected output of the Sub-Actor NN corresponding to the subsystem, and Corresponding to z n and z n-1 The coefficient of the subsystem to be calibrated, n s To ensure the continuity of the segmentation points;
[0105] The optimal system control input u after decomposition * Get the partial derivative of the optimal performance index function The expression is:
[0106]
[0107] In z n subsystem, with subsystem z1 and z i Similar to the subsystem, the partial derivative of the optimal performance index function and the optimal system control input u * The uncertain terms in are approximated by independent neural networks, and the optimal system control input u * and the partial derivative of the optimal performance index function Get the optimal system control input u * and the partial derivative of the optimal performance index function Estimated value of and Then, under the Actor-Critic framework, strategy evaluation and strategy improvement are carried out. is the estimated value of the optimal system control input, defined as Sub-Actora i , is the estimated value of the partial derivative of the optimal performance indicator function, defined as Sub-Criticc i ;
[0108] Optimal system control input u * Estimated value of The expression is:
[0109]
[0110] in, is the expected output of Sub-actor NN;
[0111] Partial derivatives of the optimal performance index function Estimated value of The expression is:
[0112]
[0113] in, is the expected output of Sub-critic NN;
[0114] The optimal system control input u * Estimated value of The expression of and the partial derivative of the optimal performance index function Estimated value of Substitute the expression into the HJB equation Then we get the estimated value of the HJB equation The expression is:
[0115]
[0116] Get z n Bellman optimality conditions in the subsystem, z n The expression of Bellman optimality condition in the subsystem is:
[0117]
[0118] z n The Bellman optimality condition in the subsystem is achieved through iterative calculation of policy evaluation and policy improvement in the Actor-Critic framework. n In the current system control input The ultimate goal of strategy evaluation is to make the estimated value of the HJB equation To reach the optimal value, In Sub-Actora n In the algorithm, Critic strategy evaluation is used to improve the strategy, and finally Bellman optimality condition is achieved through iterative learning;
[0119] Define Bellman residual
[0120]
[0121] The update equations of Sub-Critic NN and Sub-Actor NN are:
[0122]
[0123]
[0124] in, The error variable required for the Sub-Critic NN update equation, The error variable required for the Sub-Actor NN update equation, and are the learning rates of Sub-Critic NN and Sub-Actor NN respectively;
[0125] In z n In the subsystem, the optimal system control input and the partial derivative of the optimal performance index function The estimation is performed, and the Sub-Critic NN and Sub-Actor NN are further iteratively learned through their update equations, and finally the Bellman optimality condition is satisfied.
[0126] In step 2, a kinematic model and a dynamic model of the four-wheel drive autonomous vehicle are established. Assuming that the longitudinal speed of the autonomous vehicle is constant, the autonomous driving control system is modeled as a nonlinear system in a strict feedback form:
[0127]
[0128]
[0129] Among them, f1, g1, f2 and g2 are the models required to establish a second-order strict feedback motion control system. represents the lateral position and heading angle of the vehicle, Indicates the lateral velocity and yaw rate of the vehicle, u=[δ f ,M z ] T Indicates that the control input is the front wheel angle and the additional yaw moment. For a four-wheel drive vehicle, the longitudinal driving force of the left and right wheels can be independently controlled by the in-wheel motor to generate an additional yaw moment. z :=(F x,fr -F x,fl )d / 2+(F x,rr -F x,rl )d / 2 is the additional yaw moment, d is the distance between the two wheels, F x,fl 、F x,fr 、F x,rl and F x,rr are the longitudinal tire forces of the left front wheel, right front wheel, left rear wheel, and right rear wheel respectively;
[0130] When establishing a strict feedback controller model (motion control system), a linear tire force model is used. However, the tires in actual vehicles have nonlinear characteristics and are affected by different working conditions, which causes the model f i and g i The dynamic model of the real system f i p and There is a model mismatch between the two systems. The expression of the tire force of the real system is:
[0131]
[0132] in, is the tire force of the real system, F y ,· is the tire force in the controller model, and β is the relationship coefficient.
[0133] Compared with the prior art, the present invention has the following beneficial effects:
[0134] In response to the demand for adaptive learning of model parameter changes in reinforcement learning under changing scenarios and working conditions, the present invention constructs a safe reinforcement learning algorithm based on the obstacle Lyapunov function. That is, on the basis of improving the backstepping optimization control method, a hierarchical learning architecture is established based on the model. By introducing the obstacle Lyapunov function to consider the constraints, the analytical form and auxiliary function of the adaptively learnable safety control law are designed, the learning part update equation is derived, and the safe reinforcement learning algorithm based on the obstacle Lyapunov function is applied to the autonomous driving control system. By continuously affecting the safety performance of the entire learning and control process, the safety of the autonomous driving control system in the learning process is guaranteed, so as to solve the problems of strong data dependence, low online learning efficiency, easy failure of learning based on non-stationary environment, and difficulty in ensuring safety in the learning process in the reinforcement learning based on trial and error. BRIEF DESCRIPTION OF THE DRAWINGS
[0135] Figure 1 It is a structural schematic diagram of the present invention. DETAILED DESCRIPTION
[0136] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0137] To address the problems of reinforcement learning based on trial and error, such as strong data dependence, low online learning efficiency, easy failure of learning in non-stationary environments, and difficulty in ensuring safety during the learning process, the present invention addresses the difficulty of consistently meeting the state constraint requirements of typical SC systems throughout the entire learning process under uncertainty. The present invention proposes a safe reinforcement learning algorithm based on a barrier Lyapunov function (BLF-SRL). Its characteristic is that it can constrain some state variables within a designed constraint region during the learning process. This is because the barrier Lyapunov function method is a constrained control method. Its basic principle is that when the variable approaches the region boundary, the value of the Lyapunov function tends to infinity, thereby ensuring the constraint of the variable. By combining the barrier Lyapunov function with backstepping control, which has been widely used in nonlinear systems, the response speed of SC systems such as autonomous driving control can be accelerated and the robustness to system uncertainty and external interference can be improved.
[0138] like Figure 1 As shown, the present invention proposes a safety reinforcement learning autonomous driving control method. Through the theoretical methods of backstepping optimization, adaptive dynamic programming and obstacle Lyapunov function, a hierarchical safety control law with an analytical form and an adaptive learning equation are established to solve the comprehensive problem of the lack of safety assurance in the learning process of existing reinforcement learning methods. The method includes the following steps:
[0139] Step 1: Obtain a barrier-based Lyapunov function secure reinforcement learning algorithm (BLF-SRL);
[0140] Step 2: Model the autonomous driving control system as a nonlinear system in strict feedback form;
[0141] Step 3: Use the barrier Lyapunov function-based safety reinforcement learning algorithm in step 1 to achieve the safety of some system state constraints during the learning and updating process of the autonomous driving control system and the effectiveness of the error signal in each backstepping subsystem.
[0142] like Figure 1 As shown, Figure 1 Where OC is the Bellman optimality condition, PE is the policy evaluation, and PI is the policy improvement. In step 1, the process of the secure reinforcement learning algorithm based on the barrier Lyapunov function specifically includes the following steps:
[0143] Step 101: Reconstruct the nonlinear system in strict feedback form into an error system;
[0144] Step 102: Use backstepping optimization method and BLF to design the optimal control law for each subsystem;
[0145] Step 103: defining the Bellman optimality condition for each subsystem according to the Bellman optimality principle;
[0146] Step 104: Lyapunov analysis is used to design the error update signal for each subsystem. During the learning process, the virtual control of the subsystem is optimized by iteratively updating the unknown function terms in each subsystem to achieve optimization of the overall system control.
[0147] The nonlinear system in strict feedback form is reconstructed into an error system, and the backstepping optimization method and BLF are used to design the z1 subsystem, z i (i=2,...,n-1) subsystem and z n The optimal control law of the subsystem is defined, and the z1 subsystem, z i (i=2,...,n-1) subsystem and z n The Bellman optimal conditions of the subsystem are then used to design the error update signal.
[0148] In step 101, the nonlinear system in strict feedback form is:
[0149]
[0150] in, is the state variable, is the state vector, is the control input, is the system output;
[0151] In order to optimize the system control to achieve the desired system output y d , introduce the virtual control α to be optimized i (i=1,...,n-1), define the error state z1=x1-y d and z i =x i -α i-1 (i=2,...,n), the nonlinear system to be optimized is re-established as an error system:
[0152]
[0153] The entire nonlinear system to be optimized presents a cascade structure. Each virtual control α introduced by optimization i (i=1,...,n-1) Finally, the overall control of the optimized system is achieved, and all state variables z=[z1,...,z n ] T Divided into state variables to be constrained and free state variables Therefore, the learning problem is described as: During the entire learning process, the optimal system control tracking system expected output y d At the same time, some state variables z i,(i=1,...,n s ) Always stay in the safe area of the design within, among them Is a positive number.
[0154] In step 102, the backstepping optimization method and BLF are used to design the optimal control law. The backstepping optimization method is used to adopt the Actor-Critic framework of reinforcement learning in each backstepping subsystem, which are defined as Sub-Actor and Sub-Critic respectively. For the state variables to be constrained The backstepping subsystem in which the virtual control quantity is located is designed based on the barrier Lyapunov function (BLF); for the free state variables The backstepping subsystem is designed based on the quadratic Lyapunov function (QLF) for virtual control or system control input);
[0155] In step 103, the sub-actor and sub-critic are decomposed into BLF / QLF terms and unknown function terms approximated by independent neural networks (NNs), and the Bellman optimality conditions of the subsystems are defined according to the Bellman optimality principle.
[0156] In steps 102 to 104, the backstepping optimization method and BLF are used to design the z1 subsystem, z i (i=2,...,n-1) subsystem and z n The optimal control law of the subsystem is defined, and the z1 subsystem, z i (i=2,...,n-1) subsystem and z n The Bellman optimal condition of the subsystem and the specific process of designing the error update signal are as follows:
[0157] Introduce the virtual control to be optimized in the z1 subsystem and define the optimal performance index function as:
[0158]
[0159] in, is the optimal performance indicator function, is the cost function, is the optimal virtual control, κ 1s and κ 1c are weight coefficients respectively, and the corresponding HJB equation The expression is:
[0160]
[0161] in, represents the partial derivative of the optimal performance index function with respect to z1, and f1 and g1 are the models required to establish the nonlinear system to be optimized;
[0162] because Established and has a unique solution, by solving Get optimal virtual control for:
[0163]
[0164] Optimal Virtual Control The design is broken down into:
[0165]
[0166] in, is the unknown continuous function to be learned, κ1 is a positive constant, and the optimal virtual control after decomposition design is The partial derivative of the optimal performance index function can be obtained The expression is:
[0167]
[0168] In the z1 subsystem, the partial derivative of the optimal performance index function and optimal virtual control They are all unknown functions, and the uncertain terms are approximated by independent neural networks. According to the optimal virtual control after decomposition design and the partial derivative of the optimal performance index function Get its estimated value and Then, under the Actor-Critic framework, strategy evaluation and strategy improvement are carried out. is the estimated value of the optimal virtual control, defined as Sub-Actor a1, is the partial derivative of the optimal performance index function The estimated value of is defined as Sub-Criticc1;
[0169] Due to the nonlinear characteristics of the HJB equation, the optimal solution of the analytical form cannot be directly obtained. In order to iteratively obtain its numerical solution, two independent neural networks are first used to approximate the partial derivatives of the optimal performance index function. and optimal virtual control The unknown term in breaks the partial derivative of the optimal performance index function Optimal Virtual Control The correlation between them; then the neural network is iteratively updated through strategy evaluation and strategy improvement under the Actor-Critic framework to update the estimated value and Eventually, the two gradually meet the correlation relationship Then the system can be optimized and controlled;
[0170] Optimal virtual control Estimated value of The expression is:
[0171]
[0172] in, is the expected output of Sub-Actor NN;
[0173] Partial derivatives of the optimal performance index function Estimated value of
[0174]
[0175] in, is the expected output of Sub-Critic NN;
[0176] Optimal Virtual Control Estimated value of The expression of and the partial derivative of the optimal performance index function Estimated value of Substitute the expression into the HJB equation Then we get the estimated value of HJB equation The expression is:
[0177]
[0178] Obtain the Bellman optimality condition in the z1 subsystem. The expression of the Bellman optimality condition in the z1 subsystem is:
[0179]
[0180] In Sub-Criticc1, perform current virtual control The ultimate goal of strategy evaluation is to make the estimated value of the HJB equation To reach the optimal value, In Sub-Actora1, Sub-Criticc1 strategy evaluation is used to improve the strategy, and finally Bellman optimality conditions are achieved through iterative learning;
[0181] Define Bellman residual The expression is:
[0182]
[0183] The expressions of the update equations of Sub-Critic NN and Sub-Actor NN are:
[0184]
[0185]
[0186] in, The error variable required for the Sub-Critic NN update equation, The error variable required for the Sub-Actor NN update equation, and are the learning rates of Sub-Critic NN and Sub-Actor NN respectively;
[0187] Finally, in the z1 subsystem, the optimal virtual control and the partial derivative of the optimal performance index function The estimation is performed, and the Sub-Critic NN and Sub-Actor NN are further iteratively learned through their update equations to meet the Bellman optimality condition.
[0188] Similarly, in z i The virtual control α to be optimized is introduced into the subsystem i , its optimal value is The optimal performance index function is defined as:
[0189]
[0190] in, is the cost function, κ is and κ ic are weight coefficients respectively, and the corresponding HJB equation Expressed as:
[0191]
[0192] in, Denotes the optimal performance index function for z i Find the partial derivative by solving Get optimal virtual control
[0193]
[0194] Optimal Virtual Control Decomposed into:
[0195]
[0196] Among them, κ i is a positive constant, is the unknown continuous function to be learned, α i,aux is an auxiliary dummy control variable, and its expression is:
[0197]
[0198] in, and Corresponding to z i and z i-1 The coefficient of the subsystem to be calibrated, n s To ensure the continuity of the segmentation points;
[0199] The optimal virtual control after decomposition Get the partial derivative of the optimal performance index function The expression is:
[0200]
[0201] In z i In the subsystem, similar to the z1 subsystem, and The uncertainties in the equation are approximated by independent neural networks, and the optimal virtual control and the partial derivative of the optimal performance index function Get its estimated value and Then, under the Actor-Critic framework, strategy evaluation and strategy improvement are carried out. is a dummy control variable The estimated value of Sub-Actor i , is the partial derivative of the optimal performance index function The estimated value of Sub-Criticc i ;
[0202] Dummy control variables Estimated value of The expression is:
[0203]
[0204] in, is the expected output of Sub-Actor NN;
[0205] Partial derivatives of the optimal performance index function Estimated value of The expression is:
[0206]
[0207] in, is the expected output of Sub-Critic NN.
[0208] The dummy control variable Estimated value of The expression of and the partial derivative of the optimal performance index function Estimated value of Substitute the expression of into the partial derivative of the optimal performance index function Estimated value of The expression of , and then get the estimated value of HJB equation The expression is:
[0209]
[0210] Get in z i Bellman optimality conditions in the subsystem, at z i The expression of Bellman optimality condition in the subsystem is:
[0211]
[0212] z i The Bellman optimality condition of the subsystem is achieved through iterative calculation of policy evaluation and policy improvement in the Actor-Critic framework. i In the current virtual control The ultimate goal of strategy evaluation is to make the estimated value of the HJB equation To reach the optimal value, In Sub-Actora i In the algorithm, Critic strategy evaluation is used to improve the strategy, and finally Bellman optimality condition is achieved through iterative learning;
[0213] Define Bellman residual
[0214]
[0215] The update equations of Sub-Critic NN and Sub-Actor NN are:
[0216]
[0217]
[0218] in, and are the learning rates of Sub-Critic NN and Sub-Actor NN respectively.
[0219] Based on the above analysis, we can draw the following conclusions:
[0220] In z i In the subsystem, the optimal virtual control and the partial derivative of the optimal performance index function The estimation is performed, and the Sub-Critic NN and Sub-Actor NN are further iteratively learned through their update equations, and finally the Bellman optimality condition is satisfied.
[0221] Similarly, in z n In the subsystem, the optimal system control input u is u * , define z n The expression of the partial derivative of the optimal performance index function of the subsystem is:
[0222]
[0223] in, is the cost function, κ ns and κ nc are all weight coefficients, and the corresponding HJB equation The expression is:
[0224]
[0225] in, Represents the partial derivative of the optimal performance index function with respect to z n Find the partial derivative, optimal system control input u * By solving get:
[0226]
[0227] The optimal system control input u * Breaks down to:
[0228]
[0229] in, is the unknown continuous function to be learned, κ n is a positive constant, α n,aux is an auxiliary dummy control variable, and its expression is:
[0230]
[0231] The optimal system control input u after decomposition * Get the partial derivative of the optimal performance index function The expression is:
[0232]
[0233] In z n In the subsystem, z1 and z i Similar to the subsystem, the partial derivative of the optimal performance index function and the optimal system control input u * The uncertain terms in are approximated by independent neural networks, and the optimal system control input u * and the partial derivative of the optimal performance index function Get the optimal system control input u * and the partial derivative of the optimal performance index function Estimated value of and Then, under the Actor-Critic framework, strategy evaluation and strategy improvement are carried out. is the estimated value of the optimal system control input, defined as Sub-Actora i , is the estimated value of the partial derivative of the optimal performance indicator function, defined as Sub-Criticc i ;
[0234] Optimal system control input u * Estimated value of The expression is:
[0235]
[0236] in, is the expected output of Sub-actor NN;
[0237] Partial derivatives of the optimal performance index function Estimated value of The expression is:
[0238]
[0239] in, is the expected output of Sub-critic NN;
[0240] The optimal system control input u * Estimated value of The expression of and the partial derivative of the optimal performance index function Estimated value of Substitute the expression into the HJB equation Then we get the estimated value of the HJB equation The expression is:
[0241]
[0242] Get z n Bellman optimality conditions in the subsystem, z n The expression of Bellman optimality condition in the subsystem is:
[0243]
[0244] Similar to the above, z n The Bellman optimality condition in the subsystem is achieved through iterative calculation of policy evaluation and policy improvement in the Actor-Critic framework. n In the current system control input The ultimate goal of strategy evaluation is to make the estimated value of the HJB equation To reach the optimal value, In Sub-Actora n In the algorithm, Critic strategy evaluation is used to improve the strategy, and finally Bellman optimality condition is achieved through iterative learning;
[0245] Define Bellman residual
[0246]
[0247] The update equations of Sub-Critic NN and Sub-Actor NN are:
[0248]
[0249]
[0250] in, and are the learning rates of Sub-Critic NN and Sub-Actor NN respectively.
[0251] Based on the above analysis, we can draw the conclusion that: n In the subsystem, the optimal system control input and the partial derivative of the optimal performance index function The estimation is performed, and the Sub-Critic NN and Sub-Actor NN are further iteratively learned through their update equations, and finally the Bellman optimality condition is satisfied.
[0252] In summary, the nonlinear system in strict feedback form is reconstructed into an error system. After the above derivation, the virtual control and system control inputs are iteratively optimized through the Actor-Critic framework to meet the Bellman optimality conditions in each subsystem, thereby ensuring the system's safety and optimal control performance requirements.
[0253] In step 2, a kinematic model and a dynamic model of the four-wheel drive autonomous vehicle are established. Assuming that the longitudinal speed of the autonomous vehicle is constant, the autonomous driving control system (a type of system with safety-critical characteristics) is modeled as a nonlinear system in a strict feedback form. The expression of the nonlinear system in a strict feedback form is:
[0254]
[0255]
[0256] Among them, f1, g1, f2 and g2 are the models required to establish a second-order strict feedback motion control system. represents the lateral position and heading angle of the vehicle, v = [v y ,ω r ] T Indicates the lateral velocity and yaw rate of the vehicle, u=[δ f ,M z ] T Indicates that the control input is the front wheel angle and the additional yaw moment. For a four-wheel drive vehicle, the longitudinal driving force of the left and right wheels can be independently controlled by the in-wheel motor to generate an additional yaw moment. z :=(F x,fr -F x,fl )d / 2+(F x,rr -F x,rl )d / 2 is the additional yaw moment, d is the distance between the two wheels, F x,fl 、F x,fr 、F x,rl and F x,rr are the longitudinal tire forces of the left front wheel, right front wheel, left rear wheel, and right rear wheel respectively;
[0257] When establishing a strict feedback controller model (motion control system), a linear tire force model is used. However, the tires in actual vehicles have nonlinear characteristics and are affected by different working conditions, which causes the model f i and g i The dynamic model of the real system f i p and There is a model mismatch between the two systems. The expression of the tire force of the real system is:
[0258]
[0259] in, is the tire force of the real system, F y,· is the tire force in the controller model, and β is the relationship coefficient.
[0260] Design ablation experiments based on the established autonomous driving control system:
[0261] The Barrier Lyapunov Function-based Secure Reinforcement Learning algorithm (BLF-SRL) improves performance through two main parts: decomposing the optimal control using a BLF-based backstepping optimization method to ensure the safety of some state constraints of the system during the learning update process (denoted as Ablation A) and deriving the error signal according to the Bellman optimality condition in each backstepping subsystem (denoted as Ablation B). Ablation A specifically refers to the z i α in the subsystem i,aux Set to 0, ablation B specifically refers to not using the updated error signal, as shown in Tables 1 and 2 for the experimental settings and experimental results of ablation experiments under various experimental conditions:
[0262] Table 1 Experimental settings and experimental results of ablation experiment under experimental condition #D
[0263]
[0264] #D The settings of the experimental conditions are:
[0265] #D1:β=1,δ=0
[0266] #D2:β~N(1,0.8),δ=0.4
[0267] #D3:β~N(1,0.4),δ=0.4
[0268] #D4:β~N(1.2,0.6),δ=0.4
[0269] #D5:
[0270] #D6:
[0271] Where β is the proportional coefficient between the tire force of the real system and the tire force of the controller model. The system uncertainty considered is the model mismatch caused by the parameter mismatch between the controlled object and the model. In this embodiment, the boundary of the parameter β set in the simulation is [1-δ,1+δ], where δ is the boundary parameter. is the tire force defined by Fiala's formula.
[0272] Table 2 Experimental settings and experimental results of ablation experiment under experimental condition #E
[0273]
[0274]
[0275] #E The settings of the experimental conditions are:
[0276] #E1:β=1,δ=0
[0277] #E2:β~N(1,0.4),δ=0.4
[0278] #E3:β~N(1,0.8),δ=0.4;
[0279] #E4:
[0280] Table 1 records the state variables under 20 repeated simulations. and The probability of exceeding the safe area, where ablation part A is the main part to ensure safety during learning, therefore, the BLF-SRL method and ablation part B can ensure that the state variables and Constrained in the safe area of the design, at the same time, it can be seen from Table 1 that the state variables and Exceeding the designed safety zone only occurs in the ablation A or ablation AB portion.
[0281] Table 2 records the estimated values of the HJB equation during the entire simulation process under 5 repeated experiments. The average of the maximum, minimum, mean and standard deviation of The displacement y in the y-axis direction is G The estimated value of the HJB equation and the heading angle The corresponding HJB equation estimate, the velocity v in the y-axis direction y The estimated value of the HJB equation and the yaw rate ω r The corresponding HJB equation estimates are shown in Table 2. The minimum values of all HJB function estimates are kept in a small range. When part B is eliminated, the maximum and average values of the HJB function estimates increase significantly. This shows that when there is no part B, the estimated value of the HJB function can only converge to 0 under control. When part B is used for learning and updating, the estimated value of the HJB function can gradually converge to 0 at each moment with the learning update.
[0282] like Figure 1The BLF-SRL algorithm control block diagram shown in the figure reconstructs the nonlinear system in strict feedback form into an error system, then uses the backstepping optimization method and BLF to design the optimal control law, further defines the Bellman optimality condition of the subsystem, and finally designs the update signal, where OC is the Bellman optimality condition, PE is the policy evaluation, and PI is the policy improvement.
[0283] The present invention establishes a hierarchical learning architecture based on the model, introduces the obstacle Lyapunov function to consider the constraints, designs the analytical form of the safety control law and auxiliary functions that can be adaptively learned, derives the learning partial update equation to achieve the optimization of the overall system control, and establishes an autonomous driving control system based on this method. Through ablation experiments, it is verified that this method can ensure the safety of some state constraints of the system during the learning update process and the validity of the error signal in each backstepping subsystem.
[0284] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.
Claims
1. A BLF-SRL based automatic driving control method, characterized in that: The method comprises the following steps: Step 1: Construct a secure reinforcement learning algorithm based on the barrier Lyapunov function; Step 2: Model the autonomous driving control system as a nonlinear system in strict feedback form; Step 3: Use the barrier Lyapunov function-based safety reinforcement learning algorithm in step 1 to ensure the safety of some state constraints of the autonomous driving control system during the learning and updating process, as well as the validity of the error signal in each backstepping subsystem. In step 1, the process of the secure reinforcement learning algorithm based on the barrier Lyapunov function specifically includes the following steps: Step 101: Reconstruct the nonlinear system in strict feedback form into an error system; Step 102: Use backstepping optimization method and BLF to design the optimal control law for each subsystem; Step 103: defining the Bellman optimality condition for each subsystem according to the Bellman optimality principle; Step 104: Using Lyapunov analysis, the error update signal of each subsystem is designed. During the learning process, the virtual control of each subsystem is optimized by iteratively updating the unknown function terms in each subsystem to achieve optimization of the overall system control. The subsystem includes z1 subsystem, z i (i=2,...,n-1) subsystems and z n subsystem; In step 101, the nonlinear system in strict feedback form is: Among them, f j (j=1,2,…,n) and g j (j=1,2,...,n) are the models required to define the second-order strict feedback form of nonlinear system, n is the number of subsystems, is the state variable, is the state vector, is the control input, is the system output; In order to optimize the system control to achieve the desired system output y d , introduce the virtual control α to be optimized i (i=1,...,n-1), define the error state z1=x1-y d and z i =x i -α i-1 (i=2,...,n), the nonlinear system to be optimized is re-established as an error system: Among them, z j (j=1,2,...,n) is the error state of the jth subsystem, f j (j=1,2,...,n) and g j (j=1,2,...,n) are the models required to define the second-order strict feedback form of nonlinear system, n is the number of subsystems, y d is the expected output of the system; The error system presents a cascade structure, and each virtual control α introduced by optimization i (i=1,...,n-1) Finally, the overall control of the optimized system is achieved, and all state variables z=[z1,...,z n ] T Divided into state variables to be constrained and free state variables Among them, n s In order to ensure the continuity of segmentation points, the learning problem is described as: During the entire learning process, the optimization system controls the tracking system's expected output y d At the same time, some state variables z i ,(i=1,...,n s ) Always stay in the safe area of the design Inside, among them, is a positive constant; In step 103, the process of defining the Bellman optimality condition of each subsystem according to the Bellman optimality principle is specifically as follows: The sub-actor and sub-critic are decomposed into BLF / QLF terms and unknown function terms approximated by independent neural networks, and the Bellman optimality conditions of the subsystems are defined according to the Bellman optimality principle. In steps 102 to 104, for the z1 subsystem, the backstepping optimization method and BLF are used to design the optimal control law of the z1 subsystem, and the Bellman optimal condition of the z1 subsystem is defined. Then, the process of designing the error update signal is specifically as follows: The virtual control to be optimized is introduced into the z1 subsystem, and the optimal performance index function of the z1 subsystem is defined as: in, is the optimal performance index function of the z1 subsystem, is the cost function, is the optimal virtual control, κ 1s and κ 1c are weight coefficients respectively, and the corresponding HJB equation The expression is: in, represents the partial derivative of the optimal performance index function with respect to z1, and f1 and g1 are the models required to establish the nonlinear system to be optimized; because Established and has a unique solution, by solving Get optimal virtual control for: Optimal Virtual Control The design is broken down into: in, is the unknown continuous function to be learned, κ1 is a positive constant, and the optimal virtual control after decomposition design is The partial derivative of the optimal performance index function can be obtained The expression is: In the z1 subsystem, the partial derivative of the optimal performance index function and optimal virtual control They are all unknown functions, and the uncertain terms are approximated by independent neural networks. According to the optimal virtual control after decomposition design and the partial derivative of the optimal performance index function Get its estimated value and Then, under the Actor-Critic framework, strategy evaluation and strategy improvement are carried out. is the estimated value of the optimal virtual control, defined as Sub-Actor a1, is the partial derivative of the optimal performance index function The estimated value of is defined as Sub-Criticc1; Due to the nonlinear characteristics of the HJB equation, the optimal solution of the analytical form cannot be directly obtained. In order to iteratively obtain its numerical solution, two independent neural networks are first used to approximate the partial derivatives of the optimal performance index function. and optimal virtual control The unknown term in breaks the partial derivative of the optimal performance index function Optimal Virtual Control The correlation between them; then the neural network is iteratively updated through strategy evaluation and strategy improvement under the Actor-Critic framework to update the estimated value and Eventually, the two gradually meet the correlation relationship Then the system can be optimized and controlled; Optimal virtual control Estimated value of The expression is: in, is the expected output of Sub-Actor NN; Partial derivatives of the optimal performance index function Estimated value of The expression is: in, is the expected output of the Sub-Critic NN; Optimal Virtual Control Estimated value of The expression of and the partial derivative of the optimal performance index function Estimated value of Substitute the expression into the HJB equation Then we get the estimated value of HJB equation The expression is: Obtain the Bellman optimality condition in the z1 subsystem. The expression of the Bellman optimality condition in the z1 subsystem is: In Sub-Criticc1, perform current virtual control The ultimate goal of strategy evaluation is to make the estimated value of the HJB equation To reach the optimal value, In Sub-Actora1, Sub-Criticc1 strategy evaluation is used to improve the strategy, and finally Bellman optimality conditions are achieved through iterative learning; Define Bellman residual The expression is: The expressions of the update equations of Sub-Critic NN and Sub-Actor NN are: in, The error variable required for the Sub-Critic NN update equation, The error variable required for the Sub-Actor NN update equation, and are the learning rates of Sub-Critic NN and Sub-Actor NN respectively; Finally, in the z1 subsystem, the optimal virtual control and the partial derivative of the optimal performance index function Estimation is performed, and the Sub-Critic NN and Sub-Actor NN are further iteratively learned through their update equations to satisfy the Bellman optimality condition; The method further includes designing an ablation test based on the established autonomous driving control system. In the ablation test, the security of some state constraints of the system during the learning update process is recorded as ablation A, and the error signal derived according to the Bellman optimality condition in each backstepping subsystem is recorded as ablation B. Ablation A specifically refers to the z i α in the subsystem i,aux Set to 0, Ablation B specifically means not using the updated error signal. Set multiple experimental conditions for ablation experiments. The settings of each experimental condition are as follows: #D1:β=1,δ=0 #D2:β~N(1,0.8),δ=0.4 #D3:β~N(1,0.4),δ=0.4 #D4:β~N(1.2,0.6),δ=0.4 Among them, β is the proportional coefficient between the real system tire force and the tire force of the actuator model. The boundary of parameter β is [1-δ,1+δ], δ is the boundary parameter, is the tire force defined by Fiala's formula.
2. The BLF-SRL-based automatic driving control method according to claim 1, characterized in that: In step 102, the process of using the backstepping optimization method and BLF to design the optimal control law for each subsystem is specifically as follows: Based on the backstepping optimization method, the Actor-Critic framework of reinforcement learning is adopted in each subsystem, which is defined as Sub-Actor and Sub-Critic respectively. The backstepping subsystem in which the virtual control quantity is located is designed based on the barrier Lyapunov function; for the free state variables The backstepping subsystem is designed for virtual control or system control input based on the quadratic Lyapunov function.
3. The automatic driving control method based on BLF-SRL according to claim 1, characterized in that: In the steps 102 to 104, for z i Subsystem, the backstepping optimization method and BLF are used to design the optimal control law of the z1 subsystem, and the z i The Bellman optimal condition of the subsystem and the process of designing the error update signal are as follows: In z i The virtual control α to be optimized is introduced into the subsystem i , its optimal value is The optimal performance index function is defined as: in, is the cost function, κ is and κ ic are weight coefficients respectively, and the corresponding HJB equation The expression is: in, Denotes the optimal performance index function for z i Find the partial derivative by solving Get optimal virtual control The expression is: Optimal Virtual Control The design is broken down into: Among them, κ i is a positive constant, is the unknown continuous function to be learned, For dummy control variables Estimated value of Derivative, α i,aux is an auxiliary dummy control variable, and its expression is: in, z i-1 The expected output of the Sub-Actor NN corresponding to the subsystem, and Corresponding to z i and z i-1 The coefficient of the subsystem to be calibrated, n s To ensure the continuity of segmentation points, represents the optimal virtual control The cost function of The optimal virtual control after decomposition Get the partial derivative of the optimal performance index function The expression is: z i The subsystem is similar to the z1 subsystem. and The uncertainties in the equation are approximated by independent neural networks, and the optimal virtual control and the partial derivative of the optimal performance index function Get its estimated value and Then, under the Actor-Critic framework, strategy evaluation and strategy improvement are carried out. is a dummy control variable The estimated value of Sub-Actor i , is the partial derivative of the optimal performance index function The estimated value of Sub-Criticc i ; Dummy control variables Estimated value of The expression is: in, is the expected output of Sub-Actor NN; Partial derivatives of the optimal performance index function Estimated value of The expression is: in, is the expected output of Sub-Critic NN; The dummy control variable Estimated value of The expression of and the partial derivative of the optimal performance index function Estimated value of Substitute the expression of into the partial derivative of the optimal performance index function Estimated value of The expression of , and then get the estimated value of HJB equation The expression is: Get in z i Bellman optimality conditions in the subsystem, at z i The expression of Bellman optimality condition in the subsystem is: z i The Bellman optimality condition of the subsystem is achieved through iterative calculation of policy evaluation and policy improvement in the Actor-Critic framework. i In the current virtual control The ultimate goal of strategy evaluation is to make the estimated value of the HJB equation To reach the optimal value, In Sub-Actora i In the algorithm, Critic strategy evaluation is used to improve the strategy, and finally Bellman optimality condition is achieved through iterative learning; Define Bellman residual The expression is: The update equations of Sub-Critic NN and Sub-Actor NN are: in, The error variable required for the Sub-Critic NN update equation, The error variable required for the Sub-Actor NN update equation, and are the learning rates of Sub-Critic NN and Sub-Actor NN respectively; In z i In the subsystem, the optimal virtual control and the partial derivative of the optimal performance index function The estimation is performed, and the Sub-Critic NN and Sub-Actor NN are further iteratively learned through their update equations, and finally the Bellman optimality condition is satisfied.
4. The BLF-SRL-based automatic driving control method according to claim 3, characterized in that: In the steps 102 to 104, for z n Subsystem, the backstepping optimization method and BLF are used to design the optimal control law of the z1 subsystem, and the z n The Bellman optimal condition of the subsystem and the process of designing the error update signal are as follows: In z n In the subsystem, the optimal system control input u is u * , define z n The expression of the optimal performance index function of the subsystem is: in, z n The optimal performance index function of the subsystem, is the cost function, κ ns and κ nc are all weight coefficients, and the corresponding HJB equation The expression is: in, Denotes the optimal performance index function for z n Find the partial derivative, optimal system control input u * By solving get: The optimal system control input u * Breaks down to: in, is the unknown continuous function to be learned, κ n is a positive constant, α n,aux is an auxiliary dummy control variable, and its expression is: in, z n-1 The expected output of the Sub-Actor NN corresponding to the subsystem, and Corresponding to z n and z n-1 The coefficient of the subsystem to be calibrated, n s To ensure the continuity of the segmentation points; The optimal system control input u after decomposition * Get the partial derivative of the optimal performance index function The expression is: In z n subsystem, with subsystem z1 and z i Similar to the subsystem, the partial derivative of the optimal performance index function and the optimal system control input u * The uncertain terms in are approximated by independent neural networks, and the optimal system control input u * and the partial derivative of the optimal performance index function Get the optimal system control input u * and the partial derivative of the optimal performance index function Estimated value of and Then, under the Actor-Critic framework, strategy evaluation and strategy improvement are carried out. is the estimated value of the optimal system control input, defined as Sub-Actora i , is the estimated value of the partial derivative of the optimal performance indicator function, defined as Sub-Criticc i ; Optimal system control input u * Estimated value of The expression is: in, is the expected output of Sub-actor NN; Partial derivatives of the optimal performance index function Estimated value of The expression is: in, is the expected output of Sub-critic NN; The optimal system control input u * Estimated value of The expression of and the partial derivative of the optimal performance index function Estimated value of Substitute the expression into the HJB equation Then we get the estimated value of the HJB equation The expression is: Get z n Bellman optimality conditions in the subsystem, z n The expression of Bellman optimality condition in the subsystem is: z n The Bellman optimality condition in the subsystem is achieved through iterative calculation of policy evaluation and policy improvement in the Actor-Critic framework. n In the current system control input The ultimate goal of strategy evaluation is to make the estimated value of the HJB equation To reach the optimal value, In Sub-Actora n In the algorithm, Critic strategy evaluation is used to improve the strategy, and finally Bellman optimality condition is achieved through iterative learning; Define Bellman residual The update equations of Sub-Critic NN and Sub-Actor NN are: in, The error variable required for the Sub-Critic NN update equation, The error variable required for the Sub-Actor NN update equation, and are the learning rates of Sub-Critic NN and Sub-Actor NN respectively; In z n In the subsystem, the optimal system control input and the partial derivative of the optimal performance index function The estimation is performed, and the Sub-Critic NN and Sub-Actor NN are further iteratively learned through their update equations, and finally the Bellman optimality condition is satisfied.
5. The automatic driving control method based on BLF-SRL according to claim 1, characterized in that: In step 2, a kinematic model and a dynamic model of the four-wheel drive autonomous vehicle are established. Assuming that the longitudinal speed of the autonomous vehicle is constant, the autonomous driving control system is modeled as a nonlinear system in a strict feedback form: Among them, f1, g1, f2 and g2 are the models required to establish a second-order strict feedback motion control system. represents the lateral position and heading angle of the vehicle, v = [v y ,ω r ] T Indicates the lateral velocity and yaw rate of the vehicle, u=[δ f ,M z ] T Indicates that the control input is the front wheel angle and the additional yaw moment. For a four-wheel drive vehicle, the longitudinal driving force of the left and right wheels can be independently controlled by the in-wheel motor to generate an additional yaw moment. z :=(F x,fr -F x,fl )d / 2+(F x,rr -F x,rl )d / 2 is the additional yaw moment, d is the distance between the two wheels, F x,fl 、F x,fr 、F x,rl and F x,rr are the longitudinal tire forces of the left front wheel, right front wheel, left rear wheel, and right rear wheel respectively; When establishing a strict feedback controller model (motion control system), a linear tire force model is used. However, the tires in actual vehicles have nonlinear characteristics and are affected by different working conditions, which causes the model f i and g i The dynamic model of the real system f i p and There is a model mismatch between the two systems. The expression of the tire force of the real system is: in, is the tire force of the real system, F y,· is the tire force in the controller model, and β is the relationship coefficient.