A Reinforcement Learning-Based Method for Path Tracking and Obstacle Avoidance Control of Connected Vehicles

CN122569540APending Publication Date: 2026-08-14DALIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

对于正在执行路径跟踪的车辆来说,在路径跟踪控制与避障控制之间切换时确保车辆的稳定性,同时仍然保持安全行驶,是一个挑战

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122569540A_ABST
    Figure CN122569540A_ABST
Patent Text Reader

Abstract

This invention provides a method for path tracking and obstacle avoidance control of connected vehicles based on reinforcement learning strategies, relating to the technical field of vehicle path planning. The method includes the following steps: establishing a three-degree-of-freedom vehicle dynamics model and setting control objectives for the upper and lower level controllers; designing the upper level controller based on a Markov decision process model, employing a deep deterministic policy gradient algorithm to achieve longitudinal control of the vehicle to plan the ideal longitudinal speed, and employing a deep Q-network algorithm to achieve lateral control of the vehicle to plan the ideal yaw rate, and designing corresponding reward functions; designing a constant obstacle avoidance angle algorithm; and the lower level controller, through the design of a fuzzy rule table, using a fuzzy PID controller to accurately track the ideal longitudinal speed and yaw rate planned by the upper level controller.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of vehicle path planning, and more particularly to a method for connected vehicle path tracking and obstacle avoidance control based on a reinforcement learning strategy. Background Technology

[0002] In recent years, with the rapid development of vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communication technologies, connected vehicles (CVs) have been able to achieve seamless interconnection between vehicle-side, roadside and cloud systems [1]. Path tracking, as an important enabling technology in the field of intelligent transportation, refers to the trajectory tracking process of a vehicle traveling along a predetermined complex path in a multi-dimensional space without considering time constraints [2]. With the development of deep learning (DL), reinforcement learning (RL) based on deep neural networks can solve the optimal action of the system with the optimal strategy, which provides a new idea for solving the path tracking control problem. How to achieve high-precision tracking and control that maintains vehicle stability under different working conditions through RL has always been a research hotspot in the field of connected vehicle control [3].

[0003] At present, many progress has been made in the research of vehicle path tracking control. Reference [4] proposed a cooperative trajectory tracking control strategy and successfully completed the multi-vehicle cooperative tracking task. However, these control algorithms often ignore the following effect of vehicles, resulting in unreasonable vehicle position tracking error, or even negative position error, as shown in references [5] and [6]. For the vehicle trajectory tracking control problem, reference [7] innovatively adopted the Twin Delayed Deep Deterministic Policy Gradient (TD3) method to realize vehicle steering control. However, it should be noted that the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm has problems such as high computational load and relatively slow convergence rate, and its algorithm generalization performance has obvious limitations. Reference [8] adopted the Model Predictive Control (MPC) strategy to improve the vehicle spacing maintenance accuracy in platooning path tracking. However, it should be noted that the vehicle dynamics model constructed in this study only contains longitudinal dynamic components and does not consider the coupling effect of lateral dynamics. Reference [9] proposes a two-layer path tracking control strategy that takes into account vehicle energy efficiency. However, it should be noted that this strategy is only applicable to low-speed vehicle conditions and the lateral error increases significantly in high-speed scenarios. At the same time, it does not construct an obstacle avoidance mechanism in dynamic obstacle scenarios.

[0004] In addition, how connected vehicles avoid obstacles is a crucial issue that cannot be ignored. Reference

[10] proposes an autonomous driving obstacle avoidance decision control model based on deep reinforcement learning (DRL). This method innovatively uses deep deterministic policy gradient (DDPG) to construct an end-to-end obstacle avoidance model for autonomous driving. However, it should be noted that using a global model for control requires high computational resources and communication overhead, which significantly restricts the feasibility of actual vehicle deployment. Reference

[11] proposes an environmental perception and obstacle avoidance decision algorithm based on two-dimensional rotating pulse lidar (2D RPLiDAR), focusing on obstacle detection and obstacle avoidance technology for autonomous vehicles in unknown environments. However, it should be noted that the obstacle avoidance algorithm constructed in this paper has significant limitations in generalization performance for irregular obstacles, such as difficulty in effectively handling obstacles with irregular geometric contours.

[0005] Previous studies have rarely included a systematic research framework that integrates path tracking and obstacle avoidance strategies, and this technological gap has become a research direction that urgently needs to be addressed

[12] . For vehicles that are performing path tracking, ensuring vehicle stability while maintaining safe driving when switching between path tracking control and obstacle avoidance control is a challenge.

[0006] Reference documents:

[0007] [1] R. Yan, P. Li, D. Qin and S. Liu, "Trajectory Tracking Control for Autonomous Vehicle Based on Curvature Feedforward MPC," 2024 5thInternational Conference on Artificial Intelligence and ElectromechanicalAutomation (AIEA), Shenzhen, China, 2024, pp. 1018-1024. [2]Li YF, Tang CC, Li KZ, He XZ, Peeta S, Wang Y B. Consensus-based cooperative control for multiplatoon under the connected vehicles environment. IEEE Transactions on Intelligent Transportation Systems, 2019,20(6): 2220 2229. [3]Li YF, Tang CC, Peeta S, Wang Y B. Nonlinear consensus based connected vehicle platoon control incorporating car-following interactions and heterogeneous time delays. IEEE Transactions on IntelligentTransportation Systems, 2019, 20(6):2209 2219. [4]Z. Yan, X. Liu, J. Zhou and D. Wu, "Coordinated Target TrackingStrategy for Multiple Unmanned Underwater Vehicles With Time Delays," in IEEEAccess, vol. 6, pp. 10348-10357, 2018. [5]X. Yu and L. Liu, "Target Enclosing and Trajectory Tracking for aMobile Robot With Input Disturbances," in IEEE Control Systems Letters, vol.1, no. 2, pp. 221-226, Oct. 2017. [6]J. Wang, "Distributed Coordinated Tracking Control for a Class ofUncertain Multiagent Systems," in IEEE Transactions on Automatic Control,vol. 62, no. 7, pp. 3423-3429, July 2017. [7]Z. Li, "Optimal Control of Unmanned Surface Vehicle TrajectoryTracking Based on Reinforcement Learning," 2022 IEEE International Conferenceon Advances in Electrical Engineering and Computer Applications (AEECA),Dalian, China, 2022, pp. 1315-1323. [8]H. Liu, J. Sun and K. W. E. Cheng, "A Two-Layer Model PredictivePath-Tracking Control With Curvature Adaptive Method for High-SpeedAutonomous Driving," in IEEE Access, vol. 11, pp. 89228-89239, 2023. [9]Z. Zhou and Y. Bao, "Path tracking control of autonomous vehiclebased on MPC," 2024 3rd International Conference on Energy and PowerEngineering, Control Engineering (EPECE), Chengdu, China, 2024, pp. 211-216.

[10] B. Ben Elallid, A. Abouaomar, N. Benamar and A. Kobbane, "Vehicles Control: Collision Avoidance using Federated Deep ReinforcementLearning," GLOBECOM 2023 - 2023 IEEE Global Communications Conference, KualaLumpur, Malaysia, 2023, pp. 4369-4374.

[11] M. T.R. and A. M., "Obstacle Detection and Obstacle AvoidanceAlgorithm based on 2-D RPLiDAR," 2019 International Conference on ComputerCommunication and Informatics (ICCCI), Coimbatore, India, 2019, pp. 1-4.

[12] MS Wiig, KY Pettersen and TR Krogstad, "CollisionAvoidance for Underactuated Marine Vehicles Using the Constant AvoidanceAngle Algorithm," in IEEE Transactions on Control Systems Technology, vol.28, no. 3, pp. 951-966, May 2020. Summary of the Invention To address the technical problems mentioned in the background section, this invention provides a path tracking and obstacle avoidance control method for connected vehicles based on a reinforcement learning strategy. The invention proposes a two-layer control scheme integrating DDPG, DQN algorithms, and a fuzzy PID controller, and introduces a constant obstacle avoidance angle algorithm into the lateral controller. The designed controller can smoothly track the speed of the vehicle ahead in both path tracking and obstacle avoidance modes, maintaining a preset distance and ensuring safety.

[0008] The technical means employed in this invention are as follows: A method for path tracking and obstacle avoidance control of connected vehicles based on reinforcement learning strategy, employing a hierarchical control architecture, including an upper-layer controller and a lower-layer controller, the method comprising the following steps: Step 1: Establish a three-degree-of-freedom vehicle dynamics model and set the control objectives for the upper-level controller and the lower-level controller; Step 2: Design the upper-level controller based on the Markov decision process model, and use a deep deterministic policy gradient algorithm to achieve longitudinal control of the vehicle to plan the ideal longitudinal speed. Q A network algorithm is used to implement lateral control of the vehicle to plan the ideal yaw rate, and a corresponding reward function is designed. Step 3: Design a constant obstacle avoidance angle algorithm. When the vehicle sensors detect that the distance to the obstacle is less than the safety threshold, the upper longitudinal controller continues to work, and the lateral controller switches to execute the constant obstacle avoidance angle algorithm to control the front wheel steering angle to achieve obstacle avoidance. Step 4: The lower-level controller designs a fuzzy rule table and uses a fuzzy PID controller to accurately track the ideal longitudinal speed and yaw rate of the vehicle planned by the upper-level controller.

[0009] Furthermore, the vehicle dynamics model with three degrees of freedom includes: yaw rate, longitudinal velocity, and lateral velocity; The longitudinal control objective of the upper-level controller is to reduce the vehicle spacing error and the speed error between vehicles, and the lateral control objective is to reduce the lateral position deviation and heading deviation relative to the desired path; the control objective of the lower-level controller is to track the speed and yaw rate planned by the upper-level controller.

[0010] Furthermore, in the Markov decision process model, the initial state of the following vehicle is selected as the current state, and an action is selected according to the state-action mapping strategy. By executing the action, the vehicle reaches the next state and obtains a reward from the environment. The state space of the deep deterministic policy gradient algorithm includes the relative acceleration, relative longitudinal velocity, and relative position between vehicles, and the action space is the longitudinal velocity of the vehicle. The depth Q The state space of the network algorithm includes: the lateral displacement error of the vehicle, the lateral velocity error between vehicles, and the heading angle error of the vehicle relative to the reference path. The action space is the front wheel steering angle of the vehicle.

[0011] Furthermore, the reward function design of the deep deterministic policy gradient algorithm considers speed changes, driving comfort, and vehicle distance errors, including: a speed change reward term, a distance error reward term, a comfort reward term, and a termination penalty term.

[0012] Furthermore, the depth Q The reward function of the network algorithm considers lateral position error, action frequency penalty, and heading angle error, including: lateral position error reward term, heading angle error reward term, action frequency penalty term, and termination penalty term.

[0013] Furthermore, the constant obstacle avoidance angle algorithm includes the following steps: The obstacle is modeled as a circular region, the vehicle is modeled as a circular region, and the circular region is circumcircled to treat the vehicle as a point mass, thus obtaining the expanded obstacle radius; the two edges of the line-of-sight cone are rotated inward by a constant obstacle avoidance angle to form a new line-of-sight cone; Two velocity vectors are defined along the two edges of the new line-of-sight cone. When the obstacle is stationary, the smaller collision avoidance angle is selected as the desired obstacle avoidance angle. The vehicle's longitudinal velocity is still provided by the depth deterministic strategy gradient controller.

[0014] Furthermore, the initial parameters of the lower-level fuzzy PID controller are set to initial values ​​for the proportional coefficient, integral coefficient, and derivative coefficient; the output of the fuzzy PID controller is: ; in, , , Both represent the initial weight parameters; , , Both represent the weight increments output by the fuzzy logic controller based on the fuzzy rule table; This represents the error between the upper-level planned value and the actual output value of the vehicle.

[0015] Furthermore, the fuzzy rule table includes five fuzzy subsets, namely {negative large, negative small, zero, positive small, positive large}, with trigonometric functions serving as the membership functions of the fuzzy subsets.

[0016] Furthermore, when the vehicle switches to obstacle avoidance mode, the upper longitudinal controller continues to operate, and the lateral controller operates according to the constant obstacle avoidance angle algorithm. The control objective of the lateral controller is to track the desired obstacle avoidance angle, so that the vehicle's heading angle tracks the desired obstacle avoidance angle, and the maximum lateral position error of the vehicle during obstacle avoidance does not exceed the safety limit.

[0017] Furthermore, the connected vehicles exchange information through an onboard ad hoc network in a head-to-tail vehicle network topology. Each following vehicle communicates with adjacent vehicles through the onboard ad hoc network to obtain relative status information, and the vehicle-to-vehicle communication is stable.

[0018] Compared with the prior art, the present invention has the following advantages: This invention addresses the path tracking and obstacle avoidance problem in vehicle platooning on curved roads by proposing a two-layer control scheme that integrates DDPG, DQN algorithms, and a fuzzy PID controller. A constant obstacle avoidance angle algorithm is also introduced into the lateral controller. The designed controller can smoothly track the speed of the vehicle ahead in both path tracking and obstacle avoidance modes, maintaining a preset distance and ensuring safety. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This invention relates to a vehicle networking fleet control system.

[0021] Figure 2 This invention provides a two-layer control structure for formation.

[0022] Figure 3 This invention provides an obstacle avoidance-based path tracking controller based on reinforcement learning.

[0023] Figure 4 This is a schematic diagram of the extended line-of-sight cone of the present invention.

[0024] Figure 5 This is a schematic diagram of the Markov decision process of the present invention.

[0025] Figure 6 This is a schematic diagram of the lower-level fuzzy PID controller of the present invention.

[0026] Figure 7 This is a schematic diagram of DDPG rewards in an embodiment of the present invention.

[0027] Figure 8 This is a schematic diagram of DQN rewards in an embodiment of the present invention.

[0028] Figure 9 This is a schematic diagram of the vehicle's driving trajectory in an embodiment of the present invention.

[0029] Figure 10 This is a schematic diagram of the lateral position error of a vehicle in an embodiment of the present invention.

[0030] Figure 11 This is a schematic diagram of vehicle speed error in an embodiment of the present invention.

[0031] Figure 12 This is a schematic diagram of vehicle speed in an embodiment of the present invention.

[0032] Figure 13 This is a schematic diagram of vehicle spacing error in an embodiment of the present invention.

[0033] Figure 14 This is a schematic diagram of the vehicle travel distance in an embodiment of the present invention.

[0034] Figure 15 This is a schematic diagram of the overall process of the present invention. Detailed Implementation

[0035] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0036] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0037] like Figures 1-15 As shown, the present invention provides a method for path tracking and obstacle avoidance control of connected vehicles based on reinforcement learning strategy, which adopts a hierarchical control architecture, including an upper-level controller and a lower-level controller.

[0038] Consider a set of A convoy of vehicles is navigating a road with varying curvature, and the convoy must ensure it avoids obstacles as it travels. For example... Figure 1 As shown, vehicles exchange information through a vehicle-to-vehicle network (VANET) in a front-following (PF) vehicle network topology.

[0039] Assuming the curved road under study can be represented by the following... Polynomial representation of degree.

[0040] (1) in , For different horizontal coordinates There exists a unique In the curve Corresponding to it. It is a positive integer. It is a constant factor And the distance the following vehicle traveled along the curved route. Depend on It can be calculated using the following formula.

[0041] (2) in .

[0042] This invention includes the following steps: Step 1: Establish a three-degree-of-freedom vehicle dynamics model and set the control objectives for the upper and lower level controllers. To simplify the convoy control model, assume the convoy travels on a road with varying horizontal curvature, ignoring the influence of road slope. A two-wheeled, three-degree-of-freedom model (including yaw rate, longitudinal velocity, and lateral velocity) is used to track the desired trajectory, minimizing errors and outputting the desired curve. The vehicle dynamics model is as follows: ; in These are longitudinal velocity and lateral velocity. ; It is the vehicle's yaw rate. ; It is the steering angle of the vehicle's front wheels. ; The longitudinal and lateral forces exerted by the ground on the front wheels ; The longitudinal and lateral forces exerted by the ground on the rear wheels ; It is the moment of inertia about the z-axis. ; It's about the quality of the vehicle; It is the vehicle's heading angle; These are the horizontal and vertical coordinates of the geodetic coordinate system; It is the distance from the center of gravity to the front axle; It is the distance from the center of gravity to the rear axle. .

[0043] This invention employs a two-layer control structure to control the vehicle. The upper-layer controller plans the vehicle's longitudinal speed and yaw rate, while the lower-layer controller tracks the ideal vehicle target parameters planned by the upper-layer controller. The longitudinal control objective of the upper-layer controller is to reduce the vehicle spacing error and the speed error between vehicles, while the lateral control objective is to reduce the lateral position deviation and heading deviation relative to the desired path. The longitudinal control objective of the upper-layer controller is shown in formula (5), and the lateral control objective is shown in formula (6). (4) (5) (6) in Indicates the desired spacing between vehicles; Indicates the safe distance between vehicles when stationary; This indicates the expected time for using a fixed inter-vehicle spacing strategy; To represent the distance the vehicle travels along the desired path; These represent the permissible ranges for speed and spacing errors, respectively. Each represents a very small constant; This represents the angle between the geometric center of the vehicle projected onto a point on the desired path and the horizontal axis of the geodetic coordinate system. These represent the horizontal and vertical coordinates of the desired path, respectively.

[0044] In the lower-level controller, the control objective can be expressed as: (7) in These represent the speed and yaw rate planned by the upper-level controller, respectively. Each represents a very small constant.

[0045] When the vehicle encounters an obstacle and switches to obstacle avoidance mode, the upper longitudinal controller continues to operate, while the lateral controller switches to execute the CAA algorithm. The control objective of the longitudinal controller is shown in equation (5), and the control objective of the lateral controller is shown in equation (8). (8) in The table is a very small constant. For example... Figure 4 As shown, Indicates the desired obstacle avoidance angle.

[0046] Compared to direct control, hierarchical control has the advantages of centralized functions and clear control objectives at each level, which facilitates the adjustment of control parameters and functional testing of each module. Therefore, this application proposes such a hierarchical control method. Figure 2 The hierarchical control framework shown is designed to simultaneously plan the velocity and lateral angular velocity and track the ideal parameters planned in the upper layer.

[0047] Since path tracking control involves both lateral and longitudinal control of the vehicle, this application employs the DDPG algorithm for longitudinal vehicle control and the DQN algorithm for lateral vehicle control. The detailed control framework of this application is as follows: Figure 3 As shown: The upper-level structure includes longitudinal velocity and yaw rate planning controllers based on DDPG and DQN algorithms. These controllers utilize environmental observations to maximize reward values. It seeks the optimal control strategy. At each time step, it plans the reference vehicle longitudinal speed. and yaw rate The lower-level controller is designed as a fuzzy PID controller, while the longitudinal controller uses the vehicle's longitudinal speed planned in the upper level. With the actual speed of the vehicle The difference between the two, and the rate of change of that difference, serve as the input to the controller. The lower-level lateral controller uses the same design architecture. The output of the lower-level controller is then the vehicle's longitudinal acceleration. and the steering angle of the vehicle's front wheels ,like Figure 3 As shown.

[0048] Further, in step 2, the upper-level controller is designed based on the Markov decision process model, and a deep deterministic policy gradient algorithm is used to achieve longitudinal control of the vehicle to plan the ideal longitudinal speed. Q A network algorithm is used to implement lateral control of the vehicle to plan the ideal yaw rate, and a corresponding reward function is designed.

[0049] Markov decision processes (MDPs) are based on Markov properties and Markov processes

[13] . A Markov process is a stochastic process with Markov properties, which can be represented as a process consisting of tuples. The Markov property process is represented, where It is a finite set of states. This is the state transition probability matrix, meaning that a state can transition to another state based on the state transition probability matrix. Markov Decision Processes (MDPs) are as follows: Figure 5 As shown, the agent is in the initial state Next action This affects the environment. Based on the state transition probability... The state transitions to the new state. At the same time, the intelligent agent receives an instant reward. This cycle continues.

[0050] Within the Markov Decision Process (MDP) framework, this study proposes a reinforcement learning-based upper-level vehicle following policy to control each following vehicle. The initial state of the following vehicle is selected as follows: Then, according to the state-action mapping strategy Choose an action By performing this action, the vehicle reaches the next state. and get rewards from the environment. To obtain an optimal strategy Maximize the reward value The expected value of the cumulative discounted reward in each training round. The goal of reinforcement learning is to find such a... .

[0051] The motion of the vehicle is followed within the framework of a Markov Decision Process (MDP). Next, this application introduces the state space, action space, and reward function.

[0052] In order for the agent to observe enough information to determine the output action, the state space mainly contains environmental information and system state information. Under normal communication conditions, the state space of the DDPG algorithm for following the vehicle is shown in equation (12), and the state space of the DQN algorithm is shown in equation (13).

[0053] (12) (13) in It is a vehicle and The relative acceleration, relative longitudinal velocity, and relative position between them. Representing the lateral displacement error of the vehicle, respectively, the vehicle and vehicles The lateral velocity error between the two sides, and the heading angle error of the vehicle relative to the reference path. Indicates vehicle and vehicles In time and The front wheel steering angle. In this study, the appropriate range of speeds is defined as the action space. Furthermore, this application converts the vehicle's front wheel steering angle into radians as part of the motion space. .

[0054] In reinforcement learning, the reward function is a crucial design element; a good reward function determines the effectiveness of policy control. The reward function design of the longitudinal DDPG controller in this paper considers the following three aspects: speed variation, driving comfort, and vehicle distance error.

[0055] (14) (15) (16) (17) (18) in It is a hyperparameter of the reward function; Indicates vehicle With vehicles The spacing between them; This serves as the penalty condition and penalty value triggered during the termination of training. and The values ​​are as follows: (19) (20) The reward function of the lateral control DQN controller takes into account lateral position error. To prevent frequent movements or sharp turns from affecting driving comfort, a movement frequency penalty term is designed. Additionally, to ensure tracking accuracy, heading angle error is also considered.

[0056] (twenty one) (twenty two) (twenty three) (twenty four) (25) in It is a hyperparameter of the reward function; This serves as the penalty condition and penalty value triggered during the termination of training. As shown below: (26) When vehicle sensors detect that the distance to an obstacle is less than a safe threshold, longitudinal control remains managed by the DDPG and the lower-level fuzzy PID controller, while the upper-level lateral control switches to the CAA algorithm to control the steering angle of the front wheels. Meanwhile, to maintain traffic efficiency in vehicle platooning, the design of the longitudinal control reward function remains unchanged. The underlying design of the longitudinal controller is based on actual and planned speeds, accurately tracking the target speed by outputting acceleration. Similarly, the underlying design of the lateral controller receives the actual and planned yaw rates and tracks the target yaw rate by adjusting the front wheel angles.

[0057] Further, in step 3, a constant obstacle avoidance angle algorithm is designed. When the vehicle sensor detects that the distance to the obstacle is less than the safety threshold, the upper longitudinal controller continues to work, and the lateral controller switches to execute the constant obstacle avoidance angle algorithm to control the front wheel steering angle to achieve obstacle avoidance. like Figure 3 As shown in the upper-level controller, the CAA algorithm needs to be switched during obstacle avoidance. For the convenience of subsequent research, the obstacle is modeled as a circular region in this application. , where the radius is . This indicates the location information of the obstacle's center. Considering the vehicle's dimensions, the vehicle is modeled as an obstacle with a radius of... circular area , Related to the vehicle's length, its center is located at point. Before applying the collision cone theory, it is necessary to treat the circular region formed by the obstacle and vehicle model as a circumcircle. Therefore, the vehicle is treated as a point mass, and the circular region of the vehicle is... Circular area circumscribed by obstacles During the process, the radius of the obstacle after processing Satisfy the following equation: (9) Therefore, the distance from the vehicle to the expanding obstacle can be calculated using the following equation: (10) According to the CAA algorithm, in order to obtain an angle that can safely avoid obstacles, the line-of-sight cone is... The two edges rotate inward. This forms a new line of sight cone. .(See Figure 4 ) Along Two edges are defined with two velocity vectors. as follows: (11) When the obstacle is stationary This represents the longitudinal speed of the vehicle. It can be known that when… If the vehicle's heading angle can track any or Then the vehicle's trajectory will tend towards the circle determined by the obstacle model. Using the X-axis as a reference, a smaller avoidance angle is selected when avoiding stationary obstacles; therefore, this application selects... , It still refers to the speed provided to the longitudinal controller for vehicle path tracking control.

[0058] Furthermore, in step 4, the lower-level controller designs a fuzzy rule table and uses a fuzzy PID controller to accurately track the ideal longitudinal speed and yaw rate of the vehicle planned by the upper-level controller. The initial parameters of the fuzzy PID controller in the lower-level structure are set to...

[14] . The specific design is as follows: (27) in These are the weight parameters of the fuzzy PID controller; This refers to the weight increment used by the fuzzy logic controller (FLC) to update the PID. The basic idea in this application, when defining fuzzy rules, is to increase or decrease them based on system performance. and The value of . For a fuzzy PID controller with lateral control, This represents the error between the yaw rate planned by the upper layer and the actual yaw rate output by the vehicle. Similarly, for the longitudinal control of the vehicle, This represents the error between the vehicle's longitudinal speed as planned by the upper layer and the actual output longitudinal speed. For a fuzzy PID controller in lateral control, This represents the steering angle of the vehicle's front wheels. Similarly, in the longitudinal control of the vehicle... , Indicates the vehicle's acceleration .

[0059] To meet the real-time requirements of the controller and ensure the feasibility and robustness of the control system, this paper selects five variables "negative large, negative small, zero, positive small, and positive large" as fuzzy subsets, and uses trigonometric functions as the membership functions of the fuzzy subsets.

[0060] Table 1 Fuzzy rules

[0061] Table 2 Fuzzy rules

[0062] Table 3 Fuzzy rules

[0063] Example To verify the effectiveness of the controller, a simulation experiment was conducted in a vehicle platoon tracking control system consisting of one lead vehicle and four follower vehicles.

[0064] The expression for the vehicle's desired path is as follows: (28) in, The radius of obstacles on the desired path Set to 8, safe distance The value is set to 14. The vehicle parameters involved in the simulation experiment are shown in Table 4.

[0065] Table 4 Vehicle Parameters

[0066] Two agents are trained to control the vehicle's lateral and longitudinal movements respectively for path tracking. Figure 7 and Figure 8The changes in reward during training are shown. The DDPG longitudinal controller reached an average reward of approximately 5000 after 200 epochs (3000 iterations per epoch), while the DQN lateral controller reached an average reward of approximately 10000 under the same conditions. These results indicate that the algorithm has successfully converged.

[0067] To evaluate the obstacle avoidance and path tracking performance of vehicle platooning, the lead vehicle in the experiment was a virtual vehicle, and the desired path coordinates were used. An obstacle was placed at the location. In the simulation, there was a red obstacle on the path. When the distance between the vehicle and the obstacle fell below a safe threshold, the vehicle switched to obstacle avoidance mode. Figure 9 The results showed that all four test vehicles successfully avoided the obstacles and continued driving safely. Figure 10 This indicates that the maximum lateral position error of the vehicle in obstacle avoidance mode did not exceed the safety limit, verifying the effectiveness of the proposed control algorithm. From the simulation experiment... Figures 11 to 14 It is evident that the hierarchical control strategy algorithm designed in this paper, combined with the CAA obstacle avoidance algorithm, enables platooned vehicles to not only quickly maintain the preset distance and constant speed, but also accurately track the trajectory of the lead vehicle. Throughout the process, the vehicles can also effectively perform obstacle avoidance tasks.

[0068] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In the above embodiments of the present invention, the descriptions of each embodiment have their own emphasis; parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. It should be understood that the disclosed technical content in the several embodiments provided in this application can be implemented in other ways.

[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for path tracking and obstacle avoidance control of connected vehicles based on reinforcement learning strategy, characterized in that, The method employs a hierarchical control architecture, including an upper-level controller and a lower-level controller, and includes the following steps: Step 1: Establish a three-degree-of-freedom vehicle dynamics model and set the control objectives for the upper-level controller and the lower-level controller; Step 2: Design the upper-level controller based on the Markov decision process model. Use the deep deterministic policy gradient algorithm to realize the longitudinal control of the vehicle to plan the ideal longitudinal speed of the vehicle. Use the deep Q-network algorithm to realize the lateral control of the vehicle to plan the ideal yaw rate of the vehicle. Design the corresponding reward function. Step 3: Design a constant obstacle avoidance angle algorithm. When the vehicle sensors detect that the distance to the obstacle is less than the safety threshold, the upper longitudinal controller continues to work, and the lateral controller switches to execute the constant obstacle avoidance angle algorithm to control the front wheel steering angle to achieve obstacle avoidance. Step 4: The lower-level controller designs a fuzzy rule table and uses a fuzzy PID controller to accurately track the ideal longitudinal speed and yaw rate of the vehicle planned by the upper-level controller.

2. The method for connected vehicle path tracking and obstacle avoidance control based on reinforcement learning strategy according to claim 1, characterized in that, The vehicle dynamics model with three degrees of freedom includes: yaw rate, longitudinal velocity, and lateral velocity; The longitudinal control objective of the upper-level controller is to reduce the vehicle spacing error and the speed error between vehicles, and the lateral control objective is to reduce the lateral position deviation and heading deviation relative to the desired path; the control objective of the lower-level controller is to track the speed and yaw rate planned by the upper-level controller.

3. The method for connected vehicle path tracking and obstacle avoidance control based on reinforcement learning strategy according to claim 1, characterized in that, In the Markov decision process model, the initial state of the following vehicle is selected as the current state. An action is selected according to the state-action mapping strategy. By executing the action, the vehicle reaches the next state and obtains a reward from the environment. The state space of the deep deterministic policy gradient algorithm includes the relative acceleration, relative longitudinal velocity, and relative position between vehicles, and the action space is the longitudinal velocity of the vehicle. The depth Q The state space of the network algorithm includes: the lateral displacement error of the vehicle, the lateral velocity error between vehicles, and the heading angle error of the vehicle relative to the reference path. The action space is the front wheel steering angle of the vehicle.

4. The method for path tracking and obstacle avoidance control of connected vehicles based on reinforcement learning strategy according to claim 3, characterized in that, The reward function design of the deep deterministic policy gradient algorithm takes into account speed changes, driving comfort, and vehicle distance errors, and includes: a speed change reward term, a distance error reward term, a comfort reward term, and a termination penalty term.

5. The method for connected vehicle path tracking and obstacle avoidance control based on reinforcement learning strategy according to claim 3, characterized in that, The depth Q The reward function of the network algorithm considers lateral position error, action frequency penalty, and heading angle error, including: lateral position error reward term, heading angle error reward term, action frequency penalty term, and termination penalty term.

6. The method for connected vehicle path tracking and obstacle avoidance control based on reinforcement learning strategy according to claim 1, characterized in that, The constant obstacle avoidance angle algorithm includes the following steps: The obstacle is modeled as a circular region, the vehicle is modeled as a circular region, and the circular region is circumcircled to treat the vehicle as a point mass, thus obtaining the expanded obstacle radius; the two edges of the line-of-sight cone are rotated inward by a constant obstacle avoidance angle to form a new line-of-sight cone; Two velocity vectors are defined along the two edges of the new line-of-sight cone. When the obstacle is stationary, the smaller collision avoidance angle is selected as the desired obstacle avoidance angle. The vehicle's longitudinal velocity is still provided by the depth deterministic strategy gradient controller.

7. The method for connected vehicle path tracking and obstacle avoidance control based on reinforcement learning strategy according to claim 1, characterized in that, The initial parameters of the lower-level fuzzy PID controller are set to initial values ​​for the proportional coefficient, integral coefficient, and derivative coefficient; the output of the fuzzy PID controller is: ; in, , , Both represent the initial weight parameters; , , Both represent the weight increments output by the fuzzy logic controller based on the fuzzy rule table; This represents the error between the upper-level planned value and the actual output value of the vehicle.

8. The method for connected vehicle path tracking and obstacle avoidance control based on reinforcement learning strategy according to claim 1, characterized in that, The fuzzy rule table includes five fuzzy subsets: {negative large, negative small, zero, positive small, positive large}, with trigonometric functions serving as the membership functions of the fuzzy subsets.

9. The method for connected vehicle path tracking and obstacle avoidance control based on reinforcement learning strategy according to claim 1, characterized in that, When the vehicle switches to obstacle avoidance mode, the upper longitudinal controller continues to work, and the lateral controller operates according to the constant obstacle avoidance angle algorithm. The control objective of the lateral controller is to track the desired obstacle avoidance angle, so that the vehicle's heading angle tracks the desired obstacle avoidance angle, and the maximum lateral position error of the vehicle during obstacle avoidance does not exceed the safety limit.

10. The method for connected vehicle path tracking and obstacle avoidance control based on reinforcement learning strategy according to claim 1, characterized in that, The connected vehicles exchange information through an onboard ad hoc network in a head-to-tail vehicle network topology. Each following vehicle communicates with its neighboring vehicles through the onboard ad hoc network to obtain relative status information, and the communication between the vehicles is stable.