Autonomous vehicle hyper-stable sliding mode control method based on fusion of eso and deep reinforcement learning

By integrating ESO and deep reinforcement learning into a super-twisted sliding mode control method, the problem of lateral trajectory tracking of autonomous vehicles under unknown disturbances and parameter uncertainties is solved through real-time observation and online parameter tuning, thereby improving control accuracy and stability.

CN119575819BActive Publication Date: 2025-11-25FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411710876.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-11-25
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

Existing autonomous vehicles exhibit poor robustness in lateral trajectory tracking control, low tracking accuracy, and difficulty in parameter tuning when facing unknown external disturbances and uncertainties in tire lateral stiffness parameters.

Method used

A super-twisted sliding mode control method that integrates extended state observer (ESO) and deep reinforcement learning improves anti-interference and adaptive capabilities by observing uncertainties in real time and tuning parameters online.

Benefits of technology

It effectively solves the impact of unknown external disturbances and the uncertainty of tire lateral stiffness parameters on lateral trajectory tracking, and improves the control accuracy and stability of autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119575819B_ABST
    Figure CN119575819B_ABST
Patent Text Reader

Abstract

The application proposes an automatic driving vehicle super-torsion sliding mode control method fusing ESO and deep reinforcement learning, comprising the following steps: step S1: for an automatic driving vehicle, a single-track dynamic equation under the influence of external disturbance is established; step S2: a vehicle single-point preview deviation model is established in combination with a preview distance adaptive adjustment strategy; step S3: an extended state observer ESO is introduced to realize real-time observation of the uncertainty in the vehicle preview deviation model; step S4: in combination with the ESO observation result, a super-torsion sliding mode control method for automatic driving vehicle lateral tracking is proposed; step S5: a deep reinforcement learning problem is modeled and offline training is performed; step S6: the trained MLP neural network model is used for online self-tuning of the related key control parameters in the super-torsion sliding mode control method; the application can better solve the automatic driving vehicle lateral tracking control problem under the influence of unknown external disturbance and tire cornering stiffness parameter uncertainty.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control system technology for autonomous vehicles, and in particular to a super-twisted sliding mode control method for autonomous vehicles that integrates ESO and deep reinforcement learning. Background Technology

[0002] As a crucial component of future intelligent transportation systems, autonomous vehicles with controllable driving behavior offer advantages such as improved road safety, increased traffic efficiency, and reduced toll costs, and have garnered widespread attention in recent years. Lateral trajectory tracking control, a key technology for autonomous vehicles, aims to ensure vehicle stability and comfort while actively controlling tire steering to guide the vehicle along a predetermined path. Generally, pure tracking methods and Stanley control methods, based on kinematic models, are primarily suitable for lateral motion control of autonomous vehicles operating at low speeds with known geometric models. To better consider the actual dynamic characteristics of the tracking vehicle, dynamic model-based control methods are often necessary to achieve lateral trajectory tracking control under different operating conditions. Among existing vehicle dynamics control methods, PID control is the most commonly used and simplest model-free control algorithm. Lateral motion control can be achieved by appropriately setting the proportional, integral, and derivative gain parameters of the controller. However, when vehicle model parameters change significantly, such algorithms often require readjustment of the gain parameters to ensure system stability and control accuracy. In contrast, Model Predictive Control (MPC) based on rolling optimization is a control method that relies on vehicle dynamics models and can handle various constraints on the vehicle system during the control process. However, this method has high requirements for the system model and its control effect is easily affected by unmodeled factors. In fact, due to tire manufacturing technology, usage time, and other reasons, the tire lateral stiffness of autonomous vehicles is often unpredictable, and complex and variable road conditions often expose vehicles to various unknown external disturbances. To better solve the lateral trajectory tracking control problem of autonomous vehicles under the influence of the above uncertainties, people have turned their attention to sliding mode control methods with stronger anti-interference capabilities. However, it is worth mentioning that the control effect of traditional sliding mode control methods mainly depends on the selection of its key control parameters (such as reaching law parameters), and the selection of such control parameters is highly dependent on the uncertainty information of the system. If such key control parameters are not selected properly, it is easy to affect the overall control accuracy of the system or cause chattering problems, which is not conducive to the lateral tracking control of autonomous vehicles.

[0003] To address this, this invention utilizes an Extended State Observer (ESO) to observe uncertainties in the vehicle model in real time, thereby reducing their impact on system control performance. Combining the ESO observation results, a super-twisted sliding mode control method based on deep reinforcement learning is designed to effectively ensure the vehicle can better achieve the expected lateral tracking control objective. In the aforementioned control method, super-twisted sliding mode effectively solves the inherent chattering problem of sliding mode control. Furthermore, integrating deep reinforcement learning online parameter tuning technology into the super-twisted sliding mode controller further enhances the anti-interference and adaptive capabilities of the autonomous vehicle control system. Summary of the Invention

[0004] This invention proposes an ultra-torsional sliding mode control method for autonomous vehicles that integrates ESO and deep reinforcement learning, which can better solve the lateral tracking control problem of autonomous vehicles under the influence of unknown external disturbances and uncertain tire lateral stiffness parameters.

[0005] The present invention adopts the following technical solution.

[0006] A super-twisting sliding mode control method for autonomous vehicles integrating ESO and deep reinforcement learning is characterized by:

[0007] Includes the following steps;

[0008] Step S1: For autonomous vehicles, establish their single-track dynamic equations under the influence of external disturbances;

[0009] Step S2: Combine the pre-aiming distance adaptive adjustment strategy to establish a single-point pre-aiming deviation model for the vehicle;

[0010] Step S3: Introduce the Extended State Observer (ESO) to observe the uncertainties in the vehicle aiming deviation model in real time;

[0011] Step S4: Based on ESO observations, a super-torsion sliding mode control method for lateral tracking of autonomous vehicles is proposed.

[0012] Step S5: Model the deep reinforcement learning problem, build a multilayer perceptron (MLP) neural network model, and use the dual-delay deep deterministic policy gradient (TD3) algorithm to train the neural network model offline.

[0013] Step S6: Use the trained MLP neural network model for online self-tuning of relevant key control parameters in the super-twisted sliding mode control method, so that the vehicle can perform the expected lateral tracking task better.

[0014] The single-track dynamic equation established in step S1 is used to introduce the uncertainty of the vehicle tire lateral stiffness parameter in the subsequent controller design, and to re-describe the vehicle's dynamic equation.

[0015] In step S1, by setting up the geodetic coordinate system and the vehicle coordinate system, and using classical mechanics methods, the differential equation of single-track dynamics for the autonomous vehicle with active front-wheel steering is established; expressed as a formula:

[0016]

[0017] Where m is the total vehicle mass, I z Let z be the moment of inertia of the vehicle about z. l is the heading angle of the vehicle. i (i = f, r) represent the distances from the vehicle's center of gravity to the front and rear axles, respectively, and v x and v y C represents the vehicle's longitudinal velocity and lateral velocity, respectively. i (i = f, r) represent the lateral stiffness of the front and rear tires, respectively, and δ f f is the front wheel steering angle. j (j=1,2) represents the unknown external disturbances experienced by the vehicle while it is in motion. The unknown external disturbances include air resistance and the unmodeled parts of the system.

[0018] Let the vehicle tire lateral stiffness parameter C i (i=f,r) has uncertainty, and C is defined r0 C f0 These are the estimated values ​​of the lateral stiffness of the front and rear tires, ΔC. i (i = f, r) represents the deviation between the actual and estimated values ​​of the lateral stiffness of each tire, then C r =C r0 +ΔC r C f =C f0 +ΔC f The differential equation of the vehicle's single-track dynamics is then rewritten as follows:

[0019]

[0020] in,

[0021]

[0022]

[0023] The specific method of step S2 is as follows: use a pre-aiming distance selection strategy that is adaptively adjusted according to the vehicle's longitudinal speed and road curvature, expressed by the formula:

[0024]

[0025] Where L0 is the basic aiming distance, V x Longitudinal vehicle speed v x The magnitude of the target distance is given by ρ0, where ρ0 is the magnitude of the road curvature at the current vehicle position, η>0 is the correction coefficient related to the magnitude of the longitudinal vehicle speed, and μ>0 and ΔL>0 are the correction coefficient and correction distance related to the magnitude of the road curvature at the current vehicle position, respectively. The proposed target distance selection strategy introduces the influence of changes in the vehicle's longitudinal speed and road curvature on the target distance L, so as to achieve adaptive adjustment of the target distance L, which is beneficial to the subsequent lateral tracking control of autonomous vehicles.

[0026] Let (X) C ,Y C Let be the position coordinates of the vehicle's center of mass C in the geodetic coordinate system. Then, let the position coordinates of point P0, which is a distance L from the vehicle's center of mass C along the x-axis of the vehicle coordinate system (X). P0 ,Y P0 Expressed as a formula

[0027] Traverse vehicle reference trajectory T ref For all points on the map, find the point with the smallest Euclidean distance to point P0, and take that point as the vehicle's actual aiming point P; if we assume (X... P ,Y P ( ) represents the position coordinates of the aiming point P in the geodetic coordinate system. Let the angle between the tangent at the reference point P on the trajectory and the X-axis of the geodetic coordinate system be the position coordinates (x, y) of the reference point P in the vehicle coordinate system. e ,y e The angle between the tangent at the aiming point P and the x-axis of the vehicle coordinate system. They are respectively represented as

[0028]

[0029] x e y e , These represent the longitudinal deviation, lateral deviation, and heading angle deviation of the vehicle at the aiming point P in the vehicle coordinate system, respectively.

[0030] According to kinematic relationships, the first derivatives of the vehicle's lateral position deviation and heading angle deviation with respect to time are:

[0031] Where ρ is the road curvature at the aiming point;

[0032] right Taking the derivative again, we get Substitute the rewritten vehicle dynamics equations from step S1 into... The single-point aiming deviation model of the vehicle is obtained as follows

[0033] in,

[0034]

[0035]

[0036] It is considered an unknown integrated disturbance term in the system.

[0037] The implementation process of step S3 is as follows: using the lateral position deviation y of the vehicle at the pre-aiming point P... e To control the objective, an extended state observer (ESO) is designed for online estimation of the unknown integrated disturbance term ξ; ξ is assumed to be differentiable, i.e., it exists... And expand ξ as a system state variable, while letting If r3 = ξ, then the expanded vehicle dynamics equations can be rewritten as the following state equations.

[0038]

[0039] Where q is the control output of the system;

[0040] The following ESO is designed to observe the disturbance term ξ online:

[0041]

[0042]

[0043] Where z1, z2, and z3 are the estimated values ​​of the state variables r1 and r2 and the disturbance term ξ, respectively; σ j >0, κ j >0 (j=1,2) are all parameters of the function fal; τ γ >0 (γ=1,2,3) is the gain control parameter of ESO. By selecting an appropriate τ γ (γ=1,2,3) is used to ensure that ESO provides efficient estimates of r1, r2 and ξ; sgn(*) represents the sign function.

[0044] The implementation process of step S4 includes the following steps:

[0045] Step S4.1: To enable the autonomous vehicle to quickly and stably follow the reference trajectory, based on the ESO observation results from step S3, a super-torsion sliding mode control method for lateral tracking of the vehicle is adopted. Specifically, a sliding surface is defined, expressed by the formula:

[0046]

[0047] Where λ>0 is the slope parameter of the sliding surface.

[0048] The super-torsion sliding mode control method integrating ESO is adopted, which is expressed by the formula as follows:

[0049]

[0050] Among them, u sta The Super Twisting Algorithm (STA) is a sliding mode control term, and its specific expression is as follows:

[0051]

[0052] in, And α > 0, β > 0; u eso The compensation control term based on ESO is used to eliminate the influence of the disturbance term ξ, and its specific expression is as follows:

[0053]

[0054] Step S4.2: The stability of the control method is assessed using Lyapunov stability theory; specifically, by differentiating with respect to s and combining this with the vehicle dynamics equations, we obtain...

[0055]

[0056] The front wheel steering angle control law δ f Substitution Closed-loop system

[0057]

[0058] Under the compensation effect of ESO, ξ-z3→0, then the above equation is approximately simplified to

[0059]

[0060] Define the Lyapunov function V = χ T Pχ

[0061] in, and Here, P is the Hurwitz matrix; P is the equation A T A symmetric, positive definite solution to P+PA=-Q, where Q is a symmetric, positive definite matrix;

[0062] For V=χ T Differentiating Pχ, we have

[0063]

[0064] In summary, when ||χ||≠0, then The system will satisfy the asymptotic stability condition; when ||χ||=0, then s=0, y e =0.

[0065] The implementation process of step S5 includes the following steps:

[0066] Step S5.1: Model the optimization problem of relevant key control parameters in the super-twisted sliding mode control method as a deep reinforcement learning problem, and then define the corresponding state space, action space, reward function, and complete the setting of the training environment;

[0067] Define the state space as Considering that the switching gain parameter α and slope parameter λ of the sliding surface directly affect the lateral tracking accuracy, response speed, and stability of the super-twisted sliding mode control method, the motion space is defined as A = [α, λ].

[0068] In designing the reward function, the primary consideration is to minimize lateral and heading deviations while ensuring stable vehicle operation. Therefore, the lateral error is rewarded with R. y and heading deviation reward Designed as

[0069] Where h1>0 and h2>0 are reward weight coefficients; here the reward value is set to a negative value so that deep reinforcement learning can encourage the autonomous vehicle to move in the direction with smaller lateral error and heading deviation.

[0070] To avoid the potential adverse effects of an excessively large rate of change of parameter α in STA, the following reward is set for parameter α.

[0071] Where h3 > 0 is the reward weight coefficient. To control the rate of change of parameter α, The preset threshold for the rate of change of parameter α;

[0072] Considering ride comfort and the need to limit the frequency of changes in front wheel steering angle, the front wheel steering angle δ... f The rate of change is set as follows: Reward

[0073]

[0074] in, The front wheel steering angle δ f rate of change, The preset threshold for the rate of change of the front wheel steering angle;

[0075] The total reward function is then expressed as:

[0076] Based on the designed state space, action space, and reward function, a Markov decision process <S,A,P,R,γ> is established.

[0077] Where S represents the state, A represents the action, P represents the state transition process, R represents the reward, and γ represents the discount factor;

[0078] Step S5.2: Build a multilayer perceptron (MLP) neural network model and train the neural network model offline using the dual-delay deep deterministic policy gradient TD3 algorithm based on the actor-critic framework.

[0079] MLP neural network is a feedforward neural network model that includes an input layer, multiple hidden layers, and an output layer;

[0080] The TD3 algorithm structure includes an Actor original network and its target network Actor_t, a Critic original network Critic1, a Critic original network Critic2 and its target networks Critic1_t and Critic2_t. The Actor, Actor_t, Critic1, Critic2, Critic1_t, and Critic2_t networks are all built using MLP neural networks. The input of the Actor and Actor_t networks is the system state, and the output is the system action. The input of the Critic1, Critic2, Critic1_t, and Critic2_t networks is the system state and action, and the output is the score of that action.

[0081] The offline training process of the TD3 algorithm includes the following steps:

[0082] Step A1: Initialize the three original networks Actor, Critic1, and Critic2, and their corresponding target networks Actor_t, Critic1_t, and Critic2_t. The initialization parameters of the target networks are the same as those of the original networks. At the same time, initialize the optimizer, the experience replay pool, and their batch sizes in the algorithm. Explore the parameters of noise.

[0083] Step A2: At each time step, exploratory noise is applied to the Actor network, and the current state S is input into the Actor network to obtain action A; the vehicle is controlled using the super-twisted sliding mode control method based on action A to obtain reward R and new state S', and the obtained experience (S, A, R, S') is stored in the experience pool; the number of experience records N in the experience pool is compared with the preset batch size. ?like Repeat step A2; if Proceed to step A3.

[0084] Step A3: Randomly sample a batch of empirical data (S, A, R, S') from the experience pool; input the state S' into the Actor_t network to predict the next action A'; then simultaneously input the action A' and the state S' into the Critic1_t and Critic2_t networks to obtain scores Q1′ and Q2′ respectively; select the smaller of these two scores as the target value Q′; simultaneously input the state S and action A into the Critic1 and Critic2 networks to calculate the current scores Q1 and Q2; using the Bellman equation and mean squared error as the loss function, calculate the loss of the Critic1 network based on the target value Q′ and the current score Q1, and calculate the loss of the Critic2 network based on the target value Q′ and the current score Q2. Minimize the loss function using the backpropagation algorithm to update the parameters of the Critic1 and Critic2 networks;

[0085] Step A4: Determine if the network meets the delayed update condition. If the update condition is not met, proceed to step A3; if it is met, update the Actor network. Input the state S into the Actor network to predict the new action A. * State S and action A * The input is fed into the Critic1 network to calculate the score for the "state-action" relationship, and this score is used to update the parameters of the Actor network to optimize the strategy.

[0086] Step A5: Soft update the Actor network parameters to its target network (Actor_t), and soft update the Critic1 and Critic2 network parameters to their target networks (Critic1_t, Critic2_t).

[0087] The implementation process of step S6 is as follows: the trained MLP neural network model is deployed online into the super-twisted sliding mode controller designed in step S4. That is, according to the current operating state of the vehicle, the key parameters α and λ in the super-twisted sliding mode controller are adaptively adjusted online using deep reinforcement learning technology to better ensure the lateral tracking control performance of the autonomous vehicle with uncertainty.

[0088] In step S2, the vehicle's longitudinal speed is obtained through the on-board speed sensor; the road curvature information of the reference path is obtained by capturing the road markings ahead through the on-board camera and then calculating it through the upper computing unit. By combining the vehicle's longitudinal speed and road curvature information, the pre-aiming distance L is adaptively calculated using the pre-aiming distance selection strategy.

[0089] In deriving the vehicle anti-aiming deviation model, the vehicle's current position information and the position information of the anti-aiming point P are used. The vehicle's current position information is obtained through the on-board inertial measurement unit. The coordinates of point P0 are calculated from the anti-aiming distance L, and then the anti-aiming point P is obtained by traversing the reference trajectory information through P0. Subsequently, the longitudinal deviation, lateral deviation, and heading angle deviation of the vehicle at the anti-aiming point P in the vehicle coordinate system are calculated, and the anti-aiming deviation model is obtained by combining it with the dynamic model derived in step S1.

[0090] Compared to existing technologies, this invention offers the following advantages: Addressing the problems of poor robustness, low tracking accuracy, and difficulty in parameter tuning that may arise from traditional trajectory tracking control technologies when facing unknown external disturbances and system parameter perturbations, this invention provides a super-twisted sliding mode control method for autonomous vehicles that integrates ESO and deep reinforcement learning. This invention utilizes a preview distance selection strategy to achieve adaptive adjustment of the preview distance to a certain extent; furthermore, it introduces an extended state observer to compensate for the negative impact of system uncertainties, and integrates deep reinforcement learning technology into the super-twisted sliding mode controller, enabling the controller to dynamically adjust parameters according to the current vehicle state, effectively improving the lateral tracking capability of autonomous vehicles.

[0091] This invention utilizes an Extended State Observer (ESO) to observe uncertainties in the vehicle model in real time, thereby reducing their impact on system control performance. Combining the ESO observation results, a super-twisted sliding mode control method based on deep reinforcement learning is designed to effectively ensure that the vehicle can better achieve the expected lateral tracking control objective. In the aforementioned control method, super-twisted sliding mode effectively solves the inherent chattering problem of sliding mode control. Furthermore, integrating deep reinforcement learning online parameter tuning technology into the super-twisted sliding mode controller further improves the anti-interference and adaptive capabilities of the autonomous vehicle control system. Attached Figure Description

[0092] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0093] Appendix Figure 1 This is a schematic diagram of the design process of the super-twisted sliding mode control method for autonomous vehicles that integrates ESO and deep reinforcement learning in an embodiment of the present invention.

[0094] Appendix Figure 2 This is a schematic diagram of a single-point pre-aiming model for an autonomous vehicle in an embodiment of the present invention;

[0095] Appendix Figure 3 These are schematic diagrams of the MLP neural network structures in the TD3 algorithm of this invention; wherein... Figure 3(a) shows the schematic diagram of the structure used by the Actor and Actor_t networks. Figure 3 (b) is a schematic diagram of the structure used by the Critic1, Critic2, Critic1_t, and Critic2_t networks;

[0096] Appendix Figure 4 This is a schematic diagram of the network parameter update of the TD3 algorithm in an embodiment of the present invention;

[0097] Appendix Figure 5 This is a schematic diagram of the overall control block of the super-twisting sliding mode control method proposed in the embodiments of the present invention;

[0098] Appendix Figure 6 This is a schematic diagram of a simulation result obtained from an embodiment of the present invention; wherein... Figure 6 (a) shows the comparison of the lateral trajectory tracking performance of the autonomous vehicle obtained by the proposed control method and the comparative control method 1 (i.e., directly using a fixed aiming distance in the proposed control method); Figure 6 (b) shows the comparison of the lateral trajectory tracking performance of the autonomous vehicle obtained by the proposed control method and the comparative control method 2 (in the proposed control method, the gain parameter α and the slope parameter λ are directly changed to be fixed). Detailed Implementation

[0099] As shown in the figure, the super-twisting sliding mode control method for autonomous vehicles that integrates ESO and deep reinforcement learning is characterized by the following steps;

[0100] Step S1: For autonomous vehicles, establish their single-track dynamic equations under the influence of external disturbances;

[0101] Step S2: Combine the pre-aiming distance adaptive adjustment strategy to establish a single-point pre-aiming deviation model for the vehicle;

[0102] Step S3: Introduce the Extended State Observer (ESO) to observe the uncertainties in the vehicle aiming deviation model in real time;

[0103] Step S4: Based on ESO observations, a super-torsion sliding mode control method for lateral tracking of autonomous vehicles is proposed.

[0104] Step S5: Model the deep reinforcement learning problem, build a multilayer perceptron (MLP) neural network model, and use the dual-delay deep deterministic policy gradient (TD3) algorithm to train the neural network model offline.

[0105] Step S6: Use the trained MLP neural network model for online self-tuning of relevant key control parameters in the super-twisted sliding mode control method, so that the vehicle can perform the expected lateral tracking task better.

[0106] The single-track dynamic equation established in step S1 is used to introduce the uncertainty of the vehicle tire lateral stiffness parameter in the subsequent controller design, and to re-describe the vehicle's dynamic equation.

[0107] In step S1, by setting up the geodetic coordinate system and the vehicle coordinate system, and using classical mechanics methods, the differential equation of single-track dynamics for the autonomous vehicle with active front-wheel steering is established; expressed as a formula:

[0108]

[0109] Where m is the total vehicle mass, I z Let z be the moment of inertia of the vehicle about z. l is the heading angle of the vehicle. i (i = f, r) represent the distances from the vehicle's center of gravity to the front and rear axles, respectively, and v x and v y C represents the vehicle's longitudinal velocity and lateral velocity, respectively. i (i = f, r) represent the lateral stiffness of the front and rear tires, respectively, and δ f f is the front wheel steering angle. j (j=1,2) represents the unknown external disturbances experienced by the vehicle while it is in motion. The unknown external disturbances include air resistance and the unmodeled parts of the system.

[0110] Let the vehicle tire lateral stiffness parameter C i (i=f,r) has uncertainty, and C is defined r0 C f0 These are the estimated values ​​of the lateral stiffness of the front and rear tires, ΔC. i (i = f, r) represents the deviation between the actual and estimated values ​​of the lateral stiffness of each tire, then C r =C r0 +ΔC r C f =C f0 +ΔC f The differential equation of the vehicle's single-track dynamics is then rewritten as follows:

[0111]

[0112] in,

[0113]

[0114]

[0115] The specific method of step S2 is as follows: use a pre-aiming distance selection strategy that is adaptively adjusted according to the vehicle's longitudinal speed and road curvature, expressed by the formula:

[0116]

[0117] Where L0 is the basic aiming distance, V x Longitudinal vehicle speed v x The magnitude of the target distance is given by ρ0, where ρ0 is the magnitude of the road curvature at the current vehicle position, η>0 is the correction coefficient related to the magnitude of the longitudinal vehicle speed, and μ>0 and ΔL>0 are the correction coefficient and correction distance related to the magnitude of the road curvature at the current vehicle position, respectively. The proposed target distance selection strategy introduces the influence of changes in the vehicle's longitudinal speed and road curvature on the target distance L, so as to achieve adaptive adjustment of the target distance L, which is beneficial to the subsequent lateral tracking control of autonomous vehicles.

[0118] Let (X) C ,Y C Let be the position coordinates of the vehicle's center of mass C in the geodetic coordinate system. Then, let the position coordinates of point P0, which is a distance L from the vehicle's center of mass C along the x-axis of the vehicle coordinate system (X). P0 ,Y P0 Expressed as a formula

[0119] Traverse vehicle reference trajectory T ref For all points on the map, find the point with the smallest Euclidean distance to point P0, and take that point as the vehicle's actual aiming point P; if we assume (X... P ,Y P ( ) represents the position coordinates of the aiming point P in the geodetic coordinate system. Let the angle between the tangent at the reference point P on the trajectory and the X-axis of the geodetic coordinate system be the position coordinates (x, y) of the reference point P in the vehicle coordinate system. e ,y e The angle between the tangent at the aiming point P and the x-axis of the vehicle coordinate system. They are respectively represented as

[0120]

[0121] x e y e , These represent the longitudinal deviation, lateral deviation, and heading angle deviation of the vehicle at the aiming point P in the vehicle coordinate system, respectively.

[0122] According to kinematic relationships, the first derivatives of the vehicle's lateral position deviation and heading angle deviation with respect to time are:

[0123] Where ρ is the road curvature at the aiming point;

[0124] right Taking the derivative again, we get Substitute the rewritten vehicle dynamics equations from step S1 into... The single-point aiming deviation model of the vehicle is obtained as follows

[0125] in,

[0126]

[0127]

[0128] It is considered an unknown integrated disturbance term in the system.

[0129] The implementation process of step S3 is as follows: using the lateral position deviation y of the vehicle at the pre-aiming point P... e To control the objective, an extended state observer (ESO) is designed for online estimation of the unknown integrated disturbance term ξ; ξ is assumed to be differentiable, i.e., it exists... And expand ξ as a system state variable, while letting If r3 = ξ, then the expanded vehicle dynamics equations can be rewritten as the following state equations.

[0130]

[0131] Where q is the control output of the system;

[0132] The following ESO is designed to observe the disturbance term ξ online:

[0133]

[0134]

[0135] Where z1, z2, and z3 are the estimated values ​​of the state variables r1 and r2 and the disturbance term ξ, respectively; σ j >0, κ j >0 (j=1,2) are all parameters of the function fal; τ γ >0 (γ=1,2,3) is the gain control parameter of ESO. By selecting an appropriate τ γ (γ=1,2,3) is used to ensure that ESO provides efficient estimates of r1, r2 and ξ; sgn(*) represents the sign function.

[0136] The implementation process of step S4 includes the following steps:

[0137] Step S4.1: To enable the autonomous vehicle to quickly and stably follow the reference trajectory, based on the ESO observation results from step S3, a super-torsion sliding mode control method for lateral tracking of the vehicle is adopted. Specifically, a sliding surface is defined, expressed by the formula:

[0138]

[0139] Where λ>0 is the slope parameter of the sliding surface.

[0140] The super-torsion sliding mode control method integrating ESO is adopted, which is expressed by the formula as follows:

[0141]

[0142] Among them, u sta The Super Twisting Algorithm (STA) is a sliding mode control term, and its specific expression is as follows:

[0143]

[0144] in, And α > 0, β > 0; u eso The compensation control term based on ESO is used to eliminate the influence of the disturbance term ξ, and its specific expression is as follows:

[0145]

[0146] Step S4.2: The stability of the control method is assessed using Lyapunov stability theory; specifically, by differentiating with respect to s and combining this with the vehicle dynamics equations, we obtain...

[0147]

[0148] The front wheel steering angle control law δ f Substitution Closed-loop system

[0149]

[0150] Under the compensation effect of ESO, ξ-z3→0, then the above equation is approximately simplified to

[0151]

[0152] Define the Lyapunov function V = χ T Pχ

[0153] in, and Here, P is the Hurwitz matrix; P is the equation A T A symmetric, positive definite solution to P+PA=-Q, where Q is a symmetric, positive definite matrix;

[0154] For V=χ T Differentiating Pχ, we have

[0155]

[0156] In summary, when ||χ||≠0, then The system will satisfy the asymptotic stability condition; when ||χ||=0, then s=0, y e =0.

[0157] The implementation process of step S5 includes the following steps:

[0158] Step S5.1: Model the optimization problem of relevant key control parameters in the super-twisted sliding mode control method as a deep reinforcement learning problem, and then define the corresponding state space, action space, reward function, and complete the setting of the training environment;

[0159] Define the state space as Considering that the switching gain parameter α and slope parameter λ of the sliding surface directly affect the lateral tracking accuracy, response speed, and stability of the super-twisted sliding mode control method, the motion space is defined as A = [α, λ].

[0160] In designing the reward function, the primary consideration is to minimize lateral and heading deviations while ensuring stable vehicle operation. Therefore, the lateral error is rewarded with R. y and heading deviation reward Designed as

[0161] Where h1>0 and h2>0 are reward weight coefficients; here the reward value is set to a negative value so that deep reinforcement learning can encourage the autonomous vehicle to move in the direction with smaller lateral error and heading deviation.

[0162] To avoid the potential adverse effects of an excessively large rate of change of parameter α in STA, the following reward is set for parameter α.

[0163] Where h3 > 0 is the reward weight coefficient. To control the rate of change of parameter α, The preset threshold for the rate of change of parameter α;

[0164] Considering ride comfort and the need to limit the frequency of changes in front wheel steering angle, the front wheel steering angle δ... f The rate of change is set as follows: Reward

[0165]

[0166] in, The front wheel steering angle δ f rate of change, The preset threshold for the rate of change of the front wheel steering angle;

[0167] The total reward function is then expressed as: Based on the designed state space, action space, and reward function, a Markov decision process <S,A,P,R,γ> is established.

[0168] Where S represents the state, A represents the action, P represents the state transition process, R represents the reward, and γ represents the discount factor;

[0169] Step S5.2: Build a multilayer perceptron (MLP) neural network model and train the neural network model offline using the dual-delay deep deterministic policy gradient TD3 algorithm based on the actor-critic framework.

[0170] MLP neural network is a feedforward neural network model that includes an input layer, multiple hidden layers, and an output layer;

[0171] The TD3 algorithm structure includes an Actor original network and its target network Actor_t, a Critic original network Critic1, a Critic original network Critic2 and its target networks Critic1_t and Critic2_t. The Actor, Actor_t, Critic1, Critic2, Critic1_t, and Critic2_t networks are all built using MLP neural networks. The input of the Actor and Actor_t networks is the system state, and the output is the system action. The input of the Critic1, Critic2, Critic1_t, and Critic2_t networks is the system state and action, and the output is the score of that action.

[0172] The offline training process of the TD3 algorithm includes the following steps:

[0173] Step A1: Initialize the three original networks Actor, Critic1, and Critic2, and their corresponding target networks Actor_t, Critic1_t, and Critic2_t. The initialization parameters of the target networks are the same as those of the original networks. At the same time, initialize the optimizer, the experience replay pool, and their batch sizes in the algorithm. Explore the parameters of noise.

[0174] Step A2: At each time step, exploratory noise is applied to the Actor network, and the current state S is input into the Actor network to obtain action A; the vehicle is controlled using the super-twisted sliding mode control method based on action A to obtain reward R and new state S', and the obtained experience (S, A, R, S') is stored in the experience pool; the number of experience records N in the experience pool is compared with the preset batch size. ?like Repeat step A2; if Proceed to step A3.

[0175] Step A3: Randomly sample a batch of empirical data (S, A, R, S') from the experience pool; input the state S' into the Actor_t network to predict the next action A'; then simultaneously input the action A' and the state S' into the Critic1_t and Critic2_t networks to obtain scores Q1′ and Q2′ respectively; select the smaller of these two scores as the target value Q′; simultaneously input the state S and action A into the Critic1 and Critic2 networks to calculate the current scores Q1 and Q2; using the Bellman equation and mean squared error as the loss function, calculate the loss of the Critic1 network based on the target value Q′ and the current score Q1, and calculate the loss of the Critic2 network based on the target value Q′ and the current score Q2. Minimize the loss function using the backpropagation algorithm to update the parameters of the Critic1 and Critic2 networks;

[0176] Step A4: Determine if the network meets the delayed update condition. If the update condition is not met, proceed to step A3; if it is met, update the Actor network. Input the state S into the Actor network to predict the new action A. * State S and action A * The input is fed into the Critic1 network to calculate the score for the "state-action" relationship, and this score is used to update the parameters of the Actor network to optimize the strategy.

[0177] Step A5: Soft update the Actor network parameters to its target network (Actor_t), and soft update the Critic1 and Critic2 network parameters to their target networks (Critic1_t, Critic2_t).

[0178] The implementation process of step S6 is as follows: the trained MLP neural network model is deployed online into the super-twisted sliding mode controller designed in step S4. That is, according to the current operating state of the vehicle, the key parameters α and λ in the super-twisted sliding mode controller are adaptively adjusted online using deep reinforcement learning technology to better ensure the lateral tracking control performance of the autonomous vehicle with uncertainty.

[0179] In step S2, the vehicle's longitudinal speed is obtained through the on-board speed sensor; the road curvature information of the reference path is obtained by capturing the road markings ahead through the on-board camera and then calculating it through the upper computing unit. By combining the vehicle's longitudinal speed and road curvature information, the pre-aiming distance L is adaptively calculated using the pre-aiming distance selection strategy.

[0180] In deriving the vehicle anti-aiming deviation model, the vehicle's current position information and the position information of the anti-aiming point P are used. The vehicle's current position information is obtained through the on-board inertial measurement unit. The coordinates of point P0 are calculated from the anti-aiming distance L, and then the anti-aiming point P is obtained by traversing the reference trajectory information through P0. Subsequently, the longitudinal deviation, lateral deviation, and heading angle deviation of the vehicle at the anti-aiming point P in the vehicle coordinate system are calculated, and the anti-aiming deviation model is obtained by combining it with the dynamic model derived in step S1.

[0181] Example:

[0182] In this embodiment, the main parameters of the vehicle are shown in Table 2.

[0183] Table 2 Vehicle Parameters

[0184]

[0185] Vehicle reference trajectory T ref Set the double-track trajectory as follows:

[0186]

[0187] in,

[0188] Assume the vehicle's initial longitudinal velocity v x The velocity is 10 m / s², and its corresponding longitudinal acceleration is 0.8 m / s². 2 .

[0189] In the simulation, the estimated values ​​of the tire lateral stiffness are assumed to be as follows:

[0190] C f0 =ψC f C r0 =ψC r

[0191] Wherein, ψ is a random value among 0.5, 0.6, 0.7, and 0.8.

[0192] External disturbances to the vehicle f j Let (j = 1, 2) be denoted as

[0193]

[0194] The parameters in the aiming distance selection strategy are set as follows:

[0195] L0=4m, ΔL=1m, eta=8, μ=40.5

[0196] The parameters in the extended state observer are set as follows:

[0197] τ1=8, τ2=56, τ3=102, σ1=0.5, σ2=0.25, κ1=κ2=0.5

[0198] The parameters of the reward function in deep reinforcement learning are set as follows:

[0199] h1 = 60, h2 = 20, h3 = 1

[0200] When using the TD3 algorithm to train an MLP neural network offline, the range of the STA gain parameter α in the super-twisted sliding mode control method is set to (0,8] and the range of the slope parameter λ is set to [5,15]. At the beginning of training, the initial values ​​of the gain parameter α and the slope parameter λ are α(0) = 1 and λ(0) = 10, respectively; the other STA parameter β is set to β = 1.

[0201] Figure 6 This section compares the simulation results of lateral trajectory tracking for autonomous vehicles obtained in this invention. Figure 6 (a) is a schematic diagram showing the vehicle lateral trajectory tracking results obtained by the control method proposed in this paper and the control method 1 (i.e., directly using a fixed pre-aiming distance L = 4m in the proposed control method); Figure 6 As can be seen from (a), the adaptive anti-aiming distance selection strategy mentioned in this invention helps to improve the trajectory tracking performance of the vehicle at the double-movement turning point; Figure 6 (b) is a schematic diagram showing the vehicle lateral trajectory tracking results obtained by the proposed control method and the control method 2 (in the proposed control method, the fixed gain parameter α = 1 and slope parameter λ = 10 are directly used); Figure 6 As can be seen from (b), compared with control method 2, the control method proposed in this paper can better eliminate the influence of system uncertainty factors through online parameter tuning by deep reinforcement learning, and enable the vehicle to have better anti-interference and adaptive capabilities during driving.

[0202] The above are preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention shall fall within the protection scope of the present invention.

Claims

1. A super-twisting sliding mode control method for autonomous vehicles integrating ESO and deep reinforcement learning, characterized by: Includes the following steps; Step S1: For autonomous vehicles, establish their single-track dynamic equations under the influence of external disturbances; Step S2: Combine the pre-aiming distance adaptive adjustment strategy to establish a single-point pre-aiming deviation model for the vehicle; Step S3: Introduce the Extended State Observer (ESO) to observe the uncertainties in the vehicle aiming deviation model in real time; Step S4: Based on ESO observations, a super-torsion sliding mode control method for lateral tracking of autonomous vehicles is proposed. Step S5: Model the deep reinforcement learning problem, build a multilayer perceptron (MLP) neural network model, and use the dual-delay deep deterministic policy gradient (TD3) algorithm to train the neural network model offline. Step S6: Use the trained MLP neural network model for online self-tuning of relevant key control parameters in the super-twisted sliding mode control method, so that the vehicle can better complete the expected lateral tracking task. The implementation process of step S4 includes the following steps: Step S4.1: To enable the autonomous vehicle to quickly and stably follow the reference trajectory, based on the ESO observation results from step S3, a super-torsion sliding mode control method for lateral tracking of the vehicle is adopted. Specifically, a sliding surface is defined, expressed by the formula: Among them, y e λ is the ordinate of the aiming point in the vehicle coordinate system, and λ > 0 is the slope parameter of the sliding surface; The super-torsion sliding mode control method integrating ESO is adopted, which is expressed by the formula as follows: in, f y b are variables in the vehicle single-point anti-aiming deviation model. u sta The specific expression for the over-torque sliding mode control term STA is as follows: in, And α > 0, β > 0; u eso The compensation control term based on ESO is used to eliminate the influence of the disturbance term ξ, and its specific expression is as follows: Step S4.2: The stability of the control method is assessed using Lyapunov stability theory; specifically, by differentiating with respect to s and combining this with the vehicle dynamics equations, we obtain... The front wheel steering angle control law δ f Substitution Closed-loop system Under the compensation effect of ESO, ξ-z3→0, where z3 is an estimated value of ξ. Therefore, the above equation approximately simplifies to... Define the Lyapunov function V = χ T Pχ Where χ=[χ1,χ2] T , and Here, P is the Hurwitz matrix; P is the equation A T A symmetric, positive definite solution to P+PA=-Q, where Q is a symmetric, positive definite matrix; For V=χ T Differentiating Pχ, we have In summary, when ||χ||≠0, then The system will satisfy the asymptotic stability condition; when ||χ||=0, then s=0, y e =0.

2. The method for super-twisted sliding mode control of autonomous vehicles integrating ESO and deep reinforcement learning as described in claim 1, characterized in that: The single-track dynamic equation established in step S1 is used to introduce the uncertainty of the vehicle tire lateral stiffness parameter in the subsequent controller design, and to re-describe the vehicle's dynamic equation.

3. The method for super-twisted sliding mode control of autonomous vehicles integrating ESO and deep reinforcement learning as described in claim 2, characterized in that: In step S1, by setting up the geodetic coordinate system and the vehicle coordinate system, and using classical mechanics methods, the differential equation of single-track dynamics for the autonomous vehicle with active front-wheel steering is established; expressed as a formula: Where m is the total vehicle mass, I z Let z be the moment of inertia of the vehicle about z. l is the vehicle's heading angle. i Let i = f, r represent the distances from the vehicle's center of gravity to the front and rear axles, respectively, and v x and v y C represents the vehicle's longitudinal velocity and lateral velocity, respectively. i Let i = f, r be the lateral stiffness of the front and rear tires, respectively, and δ be the lateral stiffness of the rear tires. f f is the front wheel steering angle. j j = 1, 2 represents the unknown external disturbances experienced by the vehicle while it is in motion. These unknown external disturbances include air resistance and unmodeled parts of the system. Let the vehicle tire lateral stiffness parameter C i , i = f, r, there is uncertainty, and C is defined r0 C f0 These are the estimated values ​​of the lateral stiffness of the front and rear tires, ΔC. i Let i = f,r, be the deviation between the actual and estimated values ​​of the lateral stiffness of each tire, then C r =C r0 +ΔC r C f =C f0 +ΔC f The differential equation of the vehicle's single-track dynamics is then rewritten as follows: in, 4. The method for super-twisted sliding mode control of autonomous vehicles integrating ESO and deep reinforcement learning as described in claim 3, characterized in that: The specific method of step S2 is as follows: use a pre-aiming distance selection strategy that is adaptively adjusted according to the vehicle's longitudinal speed and road curvature, expressed by the formula: Where L0 is the basic aiming distance, V x Longitudinal vehicle speed v x The magnitude of the target distance is given by ρ0, where ρ0 is the magnitude of the road curvature at the current vehicle position, η>0 is the correction coefficient related to the magnitude of the longitudinal vehicle speed, and μ>0 and ΔL>0 are the correction coefficient and correction distance related to the magnitude of the road curvature at the current vehicle position, respectively. The proposed target distance selection strategy introduces the influence of changes in the vehicle's longitudinal speed and road curvature on the target distance L, so as to achieve adaptive adjustment of the target distance L, which is beneficial to the subsequent lateral tracking control of autonomous vehicles. Let (X) C ,Y C Let be the position coordinates of the vehicle's center of mass C in the geodetic coordinate system. Then, let the position coordinates of point P0, which is a distance L from the vehicle's center of mass C along the x-axis of the vehicle coordinate system (X). P0 ,Y P0 Expressed as a formula Traverse vehicle reference trajectory T ref For all points on the map, find the point with the smallest Euclidean distance to point P0, and take that point as the vehicle's actual aiming point P; if we assume (X... P ,Y P ( ) represents the position coordinates of the aiming point P in the geodetic coordinate system. Let the angle between the tangent at the reference point P on the trajectory and the X-axis of the geodetic coordinate system be the position coordinates (x, y) of the reference point P in the vehicle coordinate system. e ,y e The angle between the tangent at the aiming point P and the x-axis of the vehicle coordinate system. They are respectively represented as x e y e , These represent the longitudinal deviation, lateral deviation, and heading angle deviation of the vehicle at the aiming point P in the vehicle coordinate system, respectively. According to kinematic relationships, the first derivatives of the vehicle's lateral position deviation and heading angle deviation with respect to time are: Where ρ is the road curvature at the aiming point; right Taking the derivative again, we get Substitute the rewritten vehicle dynamics equations from step S1 into... The single-point aiming deviation model of the vehicle is obtained as follows in, It is considered an unknown integrated disturbance term in the system.

5. The method for super-twisted sliding mode control of autonomous vehicles integrating ESO and deep reinforcement learning as described in claim 4, characterized in that: The implementation process of step S3 is as follows: using the lateral position deviation y of the vehicle at the pre-aiming point P... e To control the objective, an extended state observer (ESO) is designed for online estimation of the unknown comprehensive disturbance term ξ; assuming... ξ is differentiable, that is, it exists. And expand ξ as the system state variable, while letting r1 = y e , If r3 = ξ, then the expanded vehicle dynamics equations can be rewritten as the following state equations. Where q is the control output of the system; The following ESO is designed to observe the disturbance term ξ online: Where z1, z2, and z3 are the estimated values ​​of the state variables r1 and r2 and the disturbance term ξ, respectively; σ j >0, κ j >0, j=1,2, are parameters of the function fal; τ γ >0, γ = 1, 2, 3, are the gain control parameters of the ESO. By selecting an appropriate τ γ γ = 1, 2, 3 to ensure that ESO provides efficient estimates of r1, r2 and ξ; sgn(*) represents the sign function.

6. The method for super-twisted sliding mode control of autonomous vehicles integrating ESO and deep reinforcement learning according to claim 5, characterized in that: The implementation process of step S5 includes the following steps: Step S5.1: Model the optimization problem of relevant key control parameters in the super-twisted sliding mode control method as a deep reinforcement learning problem, and then define the corresponding state space, action space, reward function, and complete the setting of the training environment; Define the state space as Considering that the switching gain parameter α and slope parameter λ of the sliding surface directly affect the lateral tracking accuracy, response speed, and stability of the super-twisted sliding mode control method, the motion space is defined as A = [α, λ]. In designing the reward function, the primary consideration is to minimize lateral and heading deviations while ensuring stable vehicle operation. Therefore, the lateral error is rewarded with R. y and heading deviation reward Designed as Where h1>0 and h2>0 are reward weight coefficients; here the reward value is set to a negative value so that deep reinforcement learning can encourage the autonomous vehicle to move in the direction with smaller lateral error and heading deviation. To avoid the potential adverse effects of an excessively large rate of change of parameter α in STA, the following reward is set for parameter α. Where h3 > 0 is the reward weight coefficient. To control the rate of change of parameter α, The preset threshold for the rate of change of parameter α; Considering ride comfort and the need to limit the frequency of changes in front wheel steering angle, the front wheel steering angle δ... f The rate of change is set as follows: Reward in, The front wheel steering angle δ f rate of change, The preset threshold for the rate of change of the front wheel steering angle; The total reward function is then expressed as: Based on the designed state space, action space, and reward function, a Markov decision process is established. <S,A,P,R,γ> Where S represents the state, A represents the action, P represents the state transition process, R represents the reward, and γ represents the discount factor; Step S5.2: Build a multilayer perceptron (MLP) neural network model and train the neural network model offline using the dual-delay deep deterministic policy gradient TD3 algorithm based on the actor-critic framework.

7. The method for super-twisted sliding mode control of autonomous vehicles integrating ESO and deep reinforcement learning as described in claim 6, characterized in that: MLP neural network is a feedforward neural network model that includes an input layer, multiple hidden layers, and an output layer; The TD3 algorithm structure includes an Actor original network and its target network Actor_t, a Critic original network Critic1, a Critic original network Critic2 and its target networks Critic1_t and Critic2_t. The Actor, Actor_t, Critic1, Critic2, Critic1_t, and Critic2_t networks are all built using MLP neural networks. The input of the Actor and Actor_t networks is the system state, and the output is the system action. The input of the Critic1, Critic2, Critic1_t, and Critic2_t networks is the system state and action, and the output is the score of that action. The offline training process of the TD3 algorithm includes the following steps: Step A1: Initialize the three original networks Actor, Critic1, and Critic2, and their corresponding target networks Actor_t, Critic1_t, and Critic2_t. The initialization parameters of the target networks are the same as those of the original networks. At the same time, initialize the optimizer, the experience replay pool, and their batch sizes in the algorithm. Explore the parameters of the noise; Step A2: At each time step, exploratory noise is applied to the Actor network, and the current state S is input into the Actor network to obtain action A; the vehicle is controlled using the super-twisted sliding mode control method based on action A to obtain reward R and new state S', and the obtained experience (S, A, R, S') is stored in the experience pool; the number of experience records N in the experience pool is compared with the preset batch size. ?like Repeat step A2; if Jump to step A3; Step A3: Randomly sample a batch of empirical data (S, A, R, S') from the experience pool; input the state S' into the Actor_t network to predict the next action A'; then simultaneously input the action A' and the state S' into the Critic1_t and Critic2_t networks to obtain scores Q1′ and Q2′ respectively; select the smaller of these two scores as the target value Q′; simultaneously input the state S and action A into the Critic1 and Critic2 networks to calculate the current scores Q1 and Q2; using the Bellman equation and the mean squared error as the loss function, calculate the loss of the Critic1 network based on the target value Q′ and the current score Q1, and calculate the loss of the Critic2 network based on the target value Q′ and the current score Q2; minimize the loss function through the backpropagation algorithm to update the parameters of the Critic1 and Critic2 networks; Step A4: Determine if the network meets the delayed update condition. If the update condition is not met, proceed to step A3; if it is met, update the Actor network by inputting the state S into the Actor network to predict the new action A. * State S and action A * The input is fed into the Critic1 network to calculate the score for the "state-action" relationship, and this score is used to update the parameters of the Actor network to optimize the strategy. Step A5: Soft update the Actor network parameters to its target network Actor_t, and soft update the Critic1 and Critic2 network parameters to their target networks Critic1_t and Critic2_t respectively.

8. The method for super-twisted sliding mode control of autonomous vehicles integrating ESO and deep reinforcement learning according to claim 7, characterized in that: The implementation process of step S6 is as follows: the trained MLP neural network model is deployed online into the super-twisted sliding mode controller designed in step S4. That is, according to the current operating state of the vehicle, the key parameters α and λ in the super-twisted sliding mode controller are adaptively adjusted online using deep reinforcement learning technology to better ensure the lateral tracking control performance of the autonomous vehicle with uncertainty.

9. The method for super-twisted sliding mode control of autonomous vehicles integrating ESO and deep reinforcement learning according to claim 8, characterized in that: In step S2, the vehicle's longitudinal speed is obtained through the on-board speed sensor; the road curvature information of the reference path is obtained by capturing the road markings ahead through the on-board camera and then calculating it through the upper computing unit. By combining the vehicle's longitudinal speed and road curvature information, the pre-aiming distance L is adaptively calculated using the pre-aiming distance selection strategy. In deriving the vehicle anti-aiming deviation model, the vehicle's current position information and the position information of the anti-aiming point P are used; The vehicle's current position information is obtained through the onboard inertial measurement unit; the position information of the aiming point P is obtained by calculating the coordinates of point P0 through the aiming distance L, and then the aiming point P is obtained by traversing the reference trajectory information through P0; subsequently, the longitudinal deviation, lateral deviation and heading angle deviation of the vehicle at the aiming point P in the vehicle coordinate system are calculated, and the aiming deviation model is obtained by combining the dynamic model derived in step S1.

Citation Information

Patent Citations

  • Finite time convergence second-order sliding mode control method

    CN111752157A

  • Underactuated unmanned ship track tracking control method based on adaptive sliding mode

    CN116339314A