An improved reinforcement learning adaptive fusion method for intelligent vehicle path control
By combining natural logarithmic sliding mode control and nonlinear model predictive control and using the attention mechanism to optimize the steering angle, the path tracking problem of autonomous vehicles in dynamic scenarios is solved, achieving high-precision and stable path tracking effects.
Patent Information
- Application Number
- CN202411682018.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-11-22
AI Technical Summary
Existing path tracking technology in autonomous vehicles has problems such as slow convergence, large data requirements, and insufficient stability. It is particularly difficult to achieve high-precision tracking in dynamic scenarios and conditions.
Combining natural logarithmic sliding mode control with nonlinear model predictive control, through the proximal strategy optimization algorithm, the attention mechanism is used to enhance the perception of error and curvature, optimize the steering angle control, and realize adaptive path tracking of intelligent vehicles.
It improves the path tracking accuracy and stability of intelligent vehicles in dynamic scenes and states, enhances the adaptability to complex environments, and improves real-time control performance.
Smart Images

Figure CN119536269B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of robotics technology, and in particular relates to an intelligent vehicle path control method based on improved reinforcement learning and adaptive fusion. Background Art
[0002] Autonomous vehicles (AVs) have garnered significant attention in recent years, offering a promising solution for improving road safety, traffic efficiency, and fuel economy. Consequently, AV technologies, which integrate positioning, sensing, path planning, and path tracking, have been widely adopted in AVs. As a key technology enabling autonomous vehicles to independently complete tasks without human intervention, path tracking plays an important role in AVs. This means that achieving high-precision path tracking is crucial for ensuring the safety of intelligent transportation systems. However, the real world is often complex, dynamic, and unpredictable, posing challenges for existing methods to achieve robust and accurate tracking.
[0003] As a classic nonlinear system, path following has been extensively explored to address the under-actuation problem associated with autonomous vehicles. Various control methods have been developed to improve path following performance. For example, polynomial control (PP) is popular for autonomous steering control due to its simplicity. However, PP struggles to find the optimal look-ahead distance. To address this challenge, other model-based controllers, such as linear quadrature controller (LQR), sequential multi-processor controller (SMC), and multi-processor controller (MPC), have been developed to achieve accurate tracking. However, LQR does not consider physical constraints, which limits its practical application. To address this issue, linear multi-processor controller (MPC) has been used for AV tracking through online optimization. However, linear models rely heavily on accurate modeling, which is difficult to achieve in practical applications. Fortunately, due to its input-output nonlinearity, nonlinear MPC has been widely used for autonomous vehicle path following in scenarios with large curvature. A comprehensive analysis of MPC and nonlinear multi-processor controllers (NMPC) shows that NMPC outperforms even when localization errors reach the centimeter level. Although NMPC exhibits greater adaptability to nonlinear systems, its real-time performance remains poor. To address this issue, researchers have employed genetic algorithms to optimize NMPC in real time, improving its solution time capability. Researchers used the ACADO toolkit to solve nonlinear optimal control problems, improving the convergence speed of NMPC. Due to solution rate limitations, NMPC suffers from insufficient stability and mediocre tracking performance, especially in cornering scenarios prone to understeer. Therefore, sliding mode control (SMC) offers an effective solution for achieving rapid convergence of path tracking. Due to its stable performance and ease of practical application, sliding mode control is widely used in autonomous vehicle systems. However, in linear SMC (LSM), the tracking error can only converge asymptotically to zero in infinite time. To improve convergence speed, terminal sliding mode (TSM) control was introduced, which uses a nonlinear sliding surface. However, a major limitation of TSM is the singularity problem. Furthermore, chatter remains a significant challenge in TSM / NTSM control. Fortunately, researchers proposed the global natural logarithm function (LnSM), which improves system chatter. However, due to the characteristics of the natural function, the convergence speed of LnSM is significantly slower than that of LSM. Due to limitations in convergence speed and chatter, reinforcement learning (RL) has been widely used in path-following control due to its ability to interact with the environment and adjust future behavior based on action feedback. While RL has shown great potential for adapting to dynamic environments, its application is limited by several challenges, such as its reliance on large amounts of training data and the instability of significant tracking errors. Based on the above analysis, different algorithms have different tracking performance characteristics. It is very challenging for a single controller to effectively track the various scenarios and states of autonomous driving.
[0004] To achieve accurate tracking in diverse scenarios and states, combined approaches are an effective solution. For example, researchers have combined PP and PID in a complementary manner. Researchers have proposed a controller that integrates MPC with a hybrid PID and LQR for multi-modal tracking control. However, this controller focuses solely on path curvature and fails to consider the impact of varying tracking errors. To address this issue, researchers have combined PP and MPC methods and tested them under dynamic error conditions. However, properly adjusting the weights of the different algorithms to adapt to dynamic scenarios remains a challenge. This means that tracking methods should be optimized based on the scenario and state for effective performance. Therefore, some researchers have investigated RL-based combined approaches to improve flexibility in dynamic scenarios and states while maintaining tracking performance. For example, researchers have proposed a method that combines baseline control with RL, but this method does not consider varying weights. Subsequently, researchers have proposed a human-robot cooperative control system that combines TD3 with an optimal preview drive model and MPC. However, this paper partially relies on steering angles obtained from humans, requiring a large empirical dataset. Furthermore, the TD3 algorithm used in this paper is highly sensitive to the choice of hyperparameters. Therefore, researchers proposed an adaptive weight adjustment method based on PPO to balance smoothness and accuracy by adjusting the weight between PID and PP. However, as a model-free algorithm, PPO makes decisions based on a single state information under model uncertainty, which may increase tracking risks in complex scenarios.
[0005] While these studies have achieved some application, challenges remain for proximal policy optimization control, including slow convergence and the large amount of data required. This is particularly true for autonomous intelligent vehicles, where path tracking performance crucially depends on real-time steering control. Traditional proximal policy optimization path tracking methods suffer from poor dynamics in practical applications, which in turn affects tracking accuracy. Therefore, improvements to proximal policy optimization algorithms are needed to enhance the real-time performance of path tracking. Summary of the Invention
[0006] To solve the above problems, the present invention discloses an intelligent vehicle path control method with improved reinforcement learning adaptive fusion. By integrating natural logarithm and nonlinear model predictive control, the path tracking accuracy is improved and the adaptability of intelligent vehicles in dynamic scenes is enhanced, the impact of tracking errors is reduced, the convergence speed of reinforcement learning is increased, and the real-time performance and stability of intelligent vehicle tracking are improved.
[0007] To achieve the above object, the technical solution of the present invention is as follows:
[0008] An intelligent vehicle path control method based on improved reinforcement learning adaptive fusion is proposed. The specific steps are as follows:
[0009] Step 1: Establish an error model for intelligent vehicles:
[0010]
[0011] where X e Expressing longitudinal error, Y e represents the lateral error, ψ e represents the vehicle heading angle error, L represents the vehicle wheelbase, v represents the vehicle speed, and δ represents the vehicle front wheel steering angle;
[0012] Step 2: Design a fast natural logarithmic sliding mode controller for the intelligent vehicle, including the following steps:
[0013] (2.1) Establish the fast natural logarithmic sliding mode surface of the vehicle:
[0014]
[0015] where ζ0, ζ1 and ζ2 are all positive numbers, and ln(·) represents the natural logarithm function.
[0016] (2.2) Establish a fast-converging sliding surface approach law:
[0017]
[0018] in C1 、 C2 are all positive numbers and satisfy 0≤ C1 ≤1, C2 ≥1.
[0019] (2.3) According to the sliding surface and the reaching law, the steering angle of the vehicle sliding mode controller is calculated:
[0020]
[0021]
[0022] Step 3: Design a nonlinear model predictive controller for the intelligent vehicle, including the following steps:
[0023] (3.1) Establish a prediction model for nonlinear model prediction:
[0024]
[0025] Where Yte represents the lateral error of the fitting curve, b is the quadratic coefficient of the polynomial curve, z(t)=k / (1+k 2 ), k is the slope of the reference point.
[0026] (3.2) Establish the steering constraint for nonlinear model prediction:
[0027] δ min ≤δ(k+n|k)≤δ max n=1,2…N (7)
[0028] Δδ min ≤Δδ(k+n|k)≤Δδ max (8)
[0029] (3.3) Establishing the objective function of nonlinear model prediction
[0030]
[0031] (3.4) Establish the steering angle predicted by the nonlinear model:
[0032] δ2(t)=δ2(t-1)+Δδ2(t) (10)
[0033] Step 4: Establish a proximal policy optimization fusion controller for the intelligent vehicle:
[0034]
[0035] Where δ1 represents the output steering angle of fast sliding mode control, and δ2 represents the output steering angle of nonlinear model predictive control.
[0036] Step 5: Use the proximal strategy optimization algorithm to quickly solve the problem in step 4. and The optimal solution is to predict the optimal state of the intelligent vehicle in the future and calculate the optimal parameters at the current moment.
[0037] Step 6: Set the current time The optimal gain parameters are substituted into the proximal strategy optimization fusion control model of the intelligent vehicle to obtain the optimal expected steering angle δ of the intelligent vehicle. opt , control the operation of the intelligent vehicle; determine whether to switch the reference point, jump to step 2 to continue the control at the next moment, until the intelligent vehicle reaches the end of the reference path.
[0038] As a further improvement of the present invention, the intelligent vehicle is an unmanned vehicle.
[0039] As a further improvement of the present invention, the step 5 comprises:
[0040] (5.1) Establish the cumulative return factor:
[0041]
[0042] Where γ represents the discount factor, and 0≤γ≤1, R t represents the reward at time t.
[0043] (5.2) Use the state value function and action value function to evaluate the expected returns of all possible decision sequences after the intelligent vehicle starts to perform an action:
[0044]
[0045]
[0046] Among them S t represents the state at time t, A t Represents the action at time t.
[0047] (5.3) Establish the advantage function under the selected action:
[0048] A π (s t ,a t )=Q π (s t ,a t )-V π (s t ) (15)
[0049] (5.4) Establish the estimation function of the advantage function:
[0050]
[0051] (5.5) Establish the policy gradient loss function:
[0052]
[0053]
[0054] Where r represents the probability ratio of the new and old strategies, and ε represents the clipping factor.
[0055] (5.5) Establish a state-decomposition attention mechanism to better focus on the impact of curvature and error on action selection:
[0056]
[0057] where Q Ni , K Ni represents Q and K in the attention mechanism, and d represents the corresponding dimension.
[0058] (5.6) Construct the loss function based on the physical information of the intelligent vehicle:
[0059]
[0060] in represents the error between the predicted model and the actual model, represents the data observation error, γpinn and γ data Indicates the corresponding coefficient.
[0061] Finally got and The optimal parameters
[0062] As a further improvement of the present invention, step 6 includes: the optimal parameters obtained by the near-end strategy optimization algorithm Bring in the fusion control algorithm to calculate the control quantity of the intelligent vehicle at the current moment:
[0063]
[0064] The optimal steering angle δ of the intelligent vehicle at the current moment can be obtained opt , to control the operation of smart vehicles.
[0065] The beneficial effects of the present invention are:
[0066] The intelligent vehicle path control method based on improved reinforcement learning adaptive fusion, described in this invention, utilizes an attention mechanism to incorporate error information and path curvature as state inputs for generating action outputs, significantly improving the adaptability of the proximal policy optimization algorithm to dynamic scenarios and conditions. The proposed improved proximal policy optimization algorithm combines a neural network with a vehicle model to extract intrinsic principles and physical constraints, enabling more reliable and stable tracking in dynamic scenarios and conditions. This capability enhances the generalization performance of the intelligent vehicle model and significantly improves the stability and adaptability of proximal policy optimization in path tracking systems under dynamic scenarios and conditions, making it effectively applicable to adaptive path tracking control of intelligent vehicles. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 This is a flow chart of the intelligent vehicle path control method based on improved reinforcement learning adaptive fusion disclosed by the present invention. DETAILED DESCRIPTION
[0068] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0069] The present invention discloses an intelligent vehicle path control method with improved reinforcement learning adaptive fusion, which can adaptively adjust the improved reinforcement learning control model according to the operating state of the intelligent vehicle, obtain the optimal intelligent vehicle steering angle, ensure the intelligent vehicle path tracking effect, and realize the precise operation of the intelligent vehicle.
[0070] As an embodiment of the present invention, the present invention discloses an intelligent vehicle path control method for improving reinforcement learning adaptive fusion, wherein the intelligent vehicle is an unmanned vehicle, and the flow chart is as follows: Figure 1 Shown, including:
[0071] Step 1: Establish an error model for intelligent vehicles:
[0072]
[0073] where X e Expressing longitudinal error, Y e represents the lateral error, ψ e represents the vehicle heading angle error, L represents the vehicle wheelbase, v represents the vehicle speed, and δ represents the vehicle front wheel steering angle;
[0074] Step 2: Design a fast natural logarithmic sliding mode controller for the intelligent vehicle, including the following steps:
[0075] (2.1) Establish the fast natural logarithmic sliding mode surface of the vehicle:
[0076]
[0077] where ζ0, ζ1 and ζ2 are all positive numbers, and ln(·) represents the natural logarithm function.
[0078] (2.2) Establish a fast-converging sliding surface approach law:
[0079]
[0080] in C1 、 C2 are all positive numbers and satisfy 0≤ C1 ≤1, C2 ≥1.
[0081] (2.3) According to the sliding surface and the reaching law, the steering angle of the vehicle sliding mode controller is calculated:
[0082]
[0083]
[0084] Step 3: Design a nonlinear model predictive controller for the intelligent vehicle, including the following steps:
[0085] (3.1) Establish a prediction model for nonlinear model prediction:
[0086]
[0087] Where Yte represents the lateral error of the fitting curve, b is the quadratic coefficient of the polynomial curve, z(t)=k / (1+k 2 ), k is the slope of the reference point.
[0088] (3.2) Establish the steering constraint for nonlinear model prediction:
[0089] δ min ≤δ(k+n|k)≤δ max n=1,2…N (7)
[0090] Δδ min ≤Δδ(k+n|k)≤Δδ max (8)
[0091] (3.3) Establishing the objective function of nonlinear model prediction
[0092]
[0093] (3.4) Establish the steering angle predicted by the nonlinear model:
[0094] δ2(t)=δ2(t-1)+Δδ2(t) (10)
[0095] Step 4: Establish a proximal policy optimization fusion controller for the intelligent vehicle:
[0096]
[0097] Where δ1 represents the output steering angle of fast sliding mode control, and δ2 represents the output steering angle of nonlinear model predictive control.
[0098] Step 5: Use the proximal strategy optimization algorithm to quickly solve the problem in step 4. and The optimal solution is to predict the optimal state of the intelligent vehicle in the future and calculate the optimal parameters at the current moment.
[0099] Step 6: Set the current time The optimal gain parameters are substituted into the proximal strategy optimization fusion control model of the intelligent vehicle to obtain the optimal expected steering angle δ of the intelligent vehicle. opt , control the operation of the intelligent vehicle; determine whether to switch the reference point, jump to step 2 to continue the control at the next moment, until the intelligent vehicle reaches the end of the reference path.
[0100] As a further improvement of the present invention, the intelligent vehicle is an unmanned vehicle.
[0101] As a further improvement of the present invention, the step 5 comprises:
[0102] (5.1) Establish the cumulative return factor:
[0103]
[0104] Where γ represents the discount factor, and 0≤γ≤1, R t represents the reward at time t.
[0105] (5.2) Use the state value function and action value function to evaluate the expected returns of all possible decision sequences after the intelligent vehicle starts to perform an action:
[0106]
[0107]
[0108] Among them S t represents the state at time t, A t Represents the action at time t.
[0109] (5.3) Establish the advantage function under the selected action:
[0110] A π (s t ,a t )=Q π (s t ,a t )-V π (s t ) (15)
[0111] (5.4) Establish the estimation function of the advantage function:
[0112]
[0113] (5.5) Establish the policy gradient loss function:
[0114]
[0115]
[0116] Where r represents the probability ratio of the new and old strategies, and ε represents the clipping factor.
[0117] (5.5) Establish a state-decomposition attention mechanism to better focus on the impact of curvature and error on action selection:
[0118]
[0119] where Q Ni , K Ni represents Q and K in the attention mechanism, and d represents the corresponding dimension.
[0120] (5.6) Construct the loss function based on the physical information of the intelligent vehicle:
[0121]
[0122] in represents the error between the predicted model and the actual model, represents the data observation error, γ pinn and γ data Indicates the corresponding coefficient.
[0123] Finally got and The optimal parameters
[0124] As a further improvement of the present invention, step 6 includes: the optimal parameters obtained by the near-end strategy optimization algorithm Bring in the fusion control algorithm to calculate the control quantity of the intelligent vehicle at the current moment:
[0125]
[0126] The optimal steering angle δ of the intelligent vehicle at the current moment can be obtained opt , to control the operation of smart vehicles.
[0127] It should be noted that the above content merely illustrates the technical idea of the present invention and cannot be used to limit the scope of protection of the present invention. For ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications all fall within the scope of protection of the claims of the present invention.
Claims
1. An intelligent vehicle path control method based on improved reinforcement learning adaptive fusion, characterized by: The specific steps are as follows: Step 1: Establish an error model for intelligent vehicles: where X e Expressing longitudinal error, Y e represents the lateral error, ψ e represents the vehicle heading angle error, L represents the vehicle wheelbase, v represents the vehicle speed, and δ represents the vehicle front wheel steering angle; Step 2: Design a fast natural logarithmic sliding mode controller for the intelligent vehicle, including the following steps: (2.1) Establish the fast natural logarithmic sliding mode surface of the vehicle: Where ζ0, ζ1 and ζ2 are all positive numbers, and ln(·) represents the natural logarithm function; (2.2) Establish a fast-converging sliding surface approach law: in are all positive numbers and satisfy (2.3) According to the sliding surface and the reaching law, the steering angle of the vehicle sliding mode controller is calculated: Step 3: Design a nonlinear model predictive controller for the intelligent vehicle, including the following steps: (3.1) Establish a prediction model for nonlinear model prediction: Where Yte represents the lateral error of the fitting curve, b is the quadratic coefficient of the polynomial curve, z(t)=k / (1+k 2 ), k is the slope of the reference point; (3.2) Establish the steering constraint for nonlinear model prediction: d min ≤δ(k+n∣k)≤δ max n=1,2…N (7) Dd min ≤Δδ(k+n∣k)≤Δδ max (8) (3.3) Establishing the objective function of nonlinear model prediction (3.4) Establish the steering angle predicted by the nonlinear model: δ2(t)=δ2(t-1)+Δδ2(t) (10) Step 4: Establish a proximal policy optimization fusion controller for the intelligent vehicle: Where δ1 represents the output steering angle of fast sliding mode control, and δ2 represents the output steering angle of nonlinear model predictive control; Step 5: Use the proximal strategy optimization algorithm to quickly solve the problem in step 4. and The optimal solution is to predict the optimal state of the intelligent vehicle in the future and calculate the optimal parameters at the current moment. Step 6: Set the current time The optimal gain parameters are substituted into the proximal strategy optimization fusion control model of the intelligent vehicle to obtain the optimal expected steering angle δ of the intelligent vehicle. opt , control the operation of the intelligent vehicle; determine whether to switch the reference point, jump to step 2 to continue the control at the next moment, until the intelligent vehicle reaches the end of the reference path.
2. The intelligent vehicle path control method based on improved reinforcement learning adaptive fusion according to claim 1 is characterized in that: The intelligent vehicle is an unmanned vehicle.
3. The intelligent vehicle path control method based on improved reinforcement learning adaptive fusion according to claim 1, characterized in that: The step 5 comprises: (5.1) Establish the cumulative return factor: Where γ represents the discount factor, and 0≤γ≤1, R t represents the reward at time t; (5.2) Use the state value function and action value function to evaluate the expected returns of all possible decision sequences after the intelligent vehicle starts to perform an action: Among them S t represents the state at time t, A t represents the action at time t; (5.3) Establish the advantage function under the selected action: A π (s t ,a t )=Q π (s t ,a t )-V π (s t ) (15) (5.4) Establish the estimation function of the advantage function: (5.5) Establish the policy gradient loss function: Where r represents the probability ratio of the new and old strategies, and ε represents the clipping factor; (5.5) Establish a state-decomposition attention mechanism to better focus on the impact of curvature and error on action selection: where Q Ni , K Ni Represents Q and K in the attention mechanism, and d represents the corresponding dimension; (5.6) Construct the loss function based on the physical information of the intelligent vehicle: in represents the error between the predicted model and the actual model, represents the data observation error, γ pinn and γ data represents the corresponding coefficient; Finally got and The optimal parameters 4. The intelligent vehicle path control method based on improved reinforcement learning adaptive fusion according to claim 1, characterized in that: The step 6 includes: obtaining the optimal parameters from the near-end strategy optimization algorithm Bring in the fusion control algorithm to calculate the control quantity of the intelligent vehicle at the current moment: Get the optimal steering angle δ of the intelligent vehicle at the current moment opt , to control the operation of smart vehicles.
Citation Information
Patent Citations
Finite time convergence second-order sliding mode control method
CN111752157A
Model predictive control trajectory tracking control system and method based on reinforcement learning
CN114967676A