Distributed optimization control method for intelligent connected vehicles based on reinforcement learning strategy

Through a hierarchical reinforcement learning strategy, combined with distributed convex optimization and the Actor-Critic framework, the stability and safety issues in vehicle platoon control are solved, efficient vehicle platoon control under complex road conditions is achieved, and the comfort and safety of the fleet are improved.

CN119087808BActive Publication Date: 2025-09-23DALIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411210889.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-09-23
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

Existing vehicle platoon control algorithms have difficulty ensuring the stability and safety of vehicle platoons when dealing with complex road conditions and vehicle-to-vehicle communication anomalies. They also have high computational complexity and lack scalability and efficiency.

Method used

A distributed optimization control method for intelligent connected vehicles based on reinforcement learning strategy is adopted, with a hierarchical structure design. The upper layer uses a trajectory planning controller based on a distributed convex optimization algorithm combined with the internal model principle, and the lower layer adopts an optimal trajectory tracking controller based on the Actor-Critic framework. The Lyapunov stability criterion is used to ensure system stability and convergence.

Benefits of technology

It achieves stable tracking control of vehicle platoons under complex road conditions, reduces dependence on precise vehicle dynamics equations, improves the comfort, safety and stability of the fleet, and reduces computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119087808B_ABST
    Figure CN119087808B_ABST
Patent Text Reader

Abstract

The present invention provides a distributed optimization control method for intelligent connected vehicles based on a reinforcement learning strategy. The method has a hierarchical structure and includes the following steps: establishing a vehicle longitudinal dynamics model and setting upper and lower control objectives; designing the upper layer, namely the trajectory optimization control layer; a trajectory planning controller based on a distributed convex optimization algorithm combined with the internal model principle, wherein the controller designed in combination with the internal model principle removes external interference from the upper layer trajectory planning process; designing the lower layer, namely the tracking control layer; an optimal trajectory tracking controller based on the reinforcement learning actor-critic framework; analyzing the stability and convergence of the tracking control system using the Lyapunov stability criterion to ensure accurate tracking of the reference trajectory; and verifying the feasibility of the proposed algorithm through simulation experiments. The present invention studies the trajectory planning and trajectory tracking control problems in the control of connected vehicle platoons, breaking away from the reliance on precise vehicle dynamics equations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle optimization control methods, and more specifically, to a distributed optimization control method for intelligent connected vehicles based on a reinforcement learning strategy. Background Art

[0002] In recent years, the continuous optimization of road transportation systems and the continued growth in the number of motor vehicles have significantly improved travel convenience, but this has also been accompanied by problems such as traffic congestion, excessive energy consumption, and severe environmental pollution. Therefore, in the context of rapid urban development, the use of high-tech means to improve vehicle intelligence and road utilization has become a top priority.

[0003] With the rapid development of artificial intelligence and 5G communication technology, intelligent connected vehicles (ICV) [1] The development of has become a reality, opening a new chapter in the automotive industry. Vehicle platooning control is a micro-driving behavior that mainly refers to multiple vehicles driving in a platoon on the same lane while maintaining the desired vehicle spacing and the same driving speed. [2] , research shows that autonomous vehicle platoon control systems have the potential to significantly reduce traffic congestion, improve traffic efficiency, enhance driving safety, and improve fuel economy [3]-[4] . However, platoon control needs to consider issues such as vehicle stability, robustness, safety, and inter-vehicle communication failures at the same time, and it is difficult to take into account multiple objective controls. In the driving system of intelligent connected vehicles, especially in the micro-driving behavior decision-making process, the vehicle's perception of its own state and the external environment state has inherent deviations and disturbances. These factors will lead to inconsistencies in vehicle decision-making and collaborative control, thereby affecting the overall driving safety and efficiency. Vehicle trajectory planning and tracking control based on reinforcement learning can improve the generalization and flexibility of trajectory planning and tracking control and greatly reduce the complexity of calculations. Therefore, it is necessary to further study the behavioral decision-making and tracking control problems of vehicle platoon control in intelligent connected vehicles.

[0004] At present, scholars at home and abroad have conducted in-depth research on vehicle platoon control and have achieved certain research results. Peters A et al. [5] The communication delay problem was taken into consideration. The communication limitation problem was reduced by improving the transmission efficiency, but intelligent control was not achieved, and the fleet control scenario was not universal. Ge Guo team [6]A hierarchical control structure is proposed to achieve coordinated control of the convoy. The upper layer uses a quadratic spacing strategy and a speed optimization algorithm to plan the optimal speed, and the lower layer designs a PID sliding mode controller to achieve speed tracking. The aim is to effectively suppress the jitter of vehicle speed and inter-vehicle distance and ensure a fast convergence rate. However, this method does not consider the problem of external disturbances during vehicle driving.

[0005] Most existing vehicle platoon control algorithms require solving the optimal control problem online. Rule-based control methods have difficulty mapping all possible states and corresponding control strategies into mathematical models when dealing with variable working conditions. Therefore, they are limited in adaptability and generalization capabilities, and often involve complex calculation processes, resulting in insufficient scalability and efficiency. Reinforcement learning algorithms are effective in solving such problems and have therefore received widespread attention. Luo Ying and other scholars [7] An improved DDPG decision-making algorithm was proposed, focusing on the research of vehicle low-speed following behavior decision-making. The algorithm combines the DDPG algorithm with the CBF algorithm for safety compensation control, and successfully achieves low-speed close-range following of vehicles in the case of heavy traffic. However, this control strategy is only applicable to the case of heavy traffic and low speed. Liming Jiang et al. [8] A soft actor-critic (SAC) reinforcement learning algorithm was proposed. The algorithm designed a reward function that can reduce the frequent changes in speed and solve the safety and stability problems of the stop-and-go driving process of the convoy. They also pointed out the advantages of reinforcement learning in suppressing traffic oscillations. However, this control strategy does not consider the problem of abnormal inter-vehicle communication under complex road conditions. Ruidong Yan et al. [9] A switching control strategy combining CACC with Deep Deterministic Policy Gradient (DDPG) is proposed. This strategy ensures the basic stability of the vehicle's following performance and fully leverages DDPG's ability to explore complex environments. Simulation results show that this strategy achieves significant improvements in vehicle following performance compared to traditional DDPG and CACC. However, this control strategy increases the computational complexity of the agent's decision-making process, resulting in a slower convergence rate.

[0006] References:

[0007] [1]Sikai L,Yingfeng C,Long C,et al.A sharing deep reinforcementlearning method for efficient vehicle platooning control[J].IET IntelligentTransport Systems,2021,16(12):1697-1709.

[0008] [2]Qin Yanyan,Wang Hao,Wang Wei,and NIDai-heng.Summary of adaptivecruise control vehicle-following models[J].Journal of Traffic andTransportation Engineering,2017,17(3):121-130.

[0009] [3]Francisco J.Martinez,Chai-Keong Toh,Juan Carlos Cano,CarlosT.Calafate,Pietro Manzoni.Emergency Services in Future IntelligentTransporta-

[0010] tation Systems Based on Vehicular Communication Networks[J].Intelligent Transportation Systems Magazine,IEEE,2010,2(2):6-20.

[0011] [4]Alam AA,Gattami A,Johansson K H.An experimental study on the fuelreduction potential of

[15] Kia,Solmaz S,Cortés,Jorge,Martínez,Sonia.Distributed convex optimization via continuous-time coordination algorithms with discrete-time communication[J].Automatica,2015,55:254-264heavy duty vehicle platooning[C].Intelligent Transportation Systems(ITSC),2010,13th International IEEE Conference on.IEEE,2010:306-311.

[0012] [5]Andres A, Peters, Richard H. Middleton, Oliver Mason. Leader tracking in homogeneous vehicle platoons with broadcast delays [J]. AUTOMATICA, 2014, 50(1): 64-74.

[0013] [6]Ge Guo,Dongqi Yang,Renyongkang Zhang.Distributed TrajectoryOptimization and Platooning of Vehicles to Guarantee Smooth Traffic Flow[J].IEEE Transactions on Intelligent Vehicles,2023,8(1):684-695

[0014] [7] Luo Ying, Qin Wenhu, Zhai Jinfeng. Research on vehicle low-speed following behavior decision-making based on improved DDPG algorithm[J]. Measurement and Control Technology, 2019, 38(9): 19-23.

[0015] [8]Liming Jiang.Reinforcement Learning based cooperative longitudinalcontrol for reducing traffic oscillations and improving platoon stability[J].Transportation Research Part C:Emerging Technologies,2022,141:103744.

[0016] [9]Ruidong Yan,Rui Jiang,Bin Jia,Jin Huang,Diange Yang.Hybrid Car-Following Strategy Based on Deep Deterministic Policy Gradient andCooperative Adaptive Cruise Control[J].IEEE Transactions on AutomationScience and Engineering,2022,19(4):2816-2824.

[0017]

[10] Shixi Wen;Ge Guo.Distributed Trajectory Optimization and SlidingMode Control of Heterogenous Vehicular Platoons[J].IEEE Transactions onIntelligent Transportation Systems,2022,Vol.23(7):7096-7111

[0018]

[11] Daizhan Cheng.Nonlinear output regulation theory andapplications,Jie Huang,SIAM,Philadelphia,2004,318pp.ISBN 0-89871-562-8[J].International Journal of Robust and Nonlinear Control,2006,16(8):413-415

[0019]

[12] Jian Y D,Konglong W.Note on graph-based BCJ relation for Berends-Giele currents[J].Journal of High Energy Physics,2022,2022(12):34-27

[0020]

[13] Daizhan Cheng.Nonlinear output regulation theory andapplications,Jie Huang,SIAM,Philadelphia,2004,318pp.ISBN 0-89871-562-8[J].International Journal of Robust and Nonlinear Control,2006,16(8):413-415

[0021]

[14] De Persis,C,Jayawardhana,B.On the internal model principle in thecoordination of nonlinear systems(Article)[J].IEEE Transactions on Control ofNetwork Systems,2014,1(3):272-282

[0022]

[15] Kia,Solmaz S,Cortés,Jorge,Martínez,Sonia.Distributed convexoptimization via continuous-time coordination algorithms with discrete-timecommunication[J].Automatica,2015,55:254-264

[0023]

[16] Wang,

[0024]

[17] Lewis FL,Abu-Khalaf M.Nearly optimal control laws for nonlinearsystems with saturating actuators using a neural network HJB approach[J].Automatica,2005,41(5):779-791

[0025]

[18] Li Y, Tee PK, Yan R, et al. A Framework of Human-Robot CoordinationBased on Game Theory and Policy Iteration[J]. IEEE Trans.Robotics, 2016, 32(6): 1408-1418. Summary of the Invention

[0026] In response to the technical problems mentioned in the above background technology, a distributed optimization control method for intelligent connected vehicles based on reinforcement learning strategy is provided.

[0027] The technical means adopted in the present invention are as follows:

[0028] The distributed optimization control method for intelligent connected vehicles based on reinforcement learning strategy has a hierarchical structure and includes the following steps:

[0029] Step 1: Build a vehicle longitudinal dynamics model and set upper and lower control objectives

[0030] Step 2: Design the upper layer, the trajectory optimization control layer. A trajectory planning controller based on a distributed convex optimization algorithm combined with the internal model principle is designed to remove external interference from the upper trajectory planning process.

[0031] Step 3: Design the lower layer, the tracking control layer; an optimal trajectory tracking controller based on the reinforcement learning actor-critic framework, where the actor network is used to approximate the optimal tracking controller and the critic neural network is used to approximate the optimal cost function;

[0032] Step 4: Analyze the stability and convergence of the tracking control system using the Lyapunov stability criterion to ensure accurate tracking of the reference trajectory;

[0033] Step 5: Verify the feasibility of the proposed algorithm through simulation experiments.

[0034] Furthermore, in step 1, define x i , v i and a i Represents the position, speed and acceleration of the i-th vehicle respectively. The pilot vehicle is numbered 0 and the following vehicles are numbered from 1 to n. represents the communication topology of the convoy. Each following vehicle communicates with adjacent vehicles through the vehicle-mounted ad hoc network to obtain information, and the communication between vehicles is stable and reliable. Only the longitudinal dynamics model of the vehicle is considered. The longitudinal dynamics model is established to provide a basis for the design of the underlying controller. The longitudinal dynamics model of vehicle i is expressed as:

[0035]

[0036] in, θ i represents the uncertain time constant of the vehicle transmission system, u i2 (t) represents the control input of the lower-level system following vehicle i; the tracking error for vehicle i is given as follows:

[0037] δ i =x i-1 (t)-x i (t)-d i,i-1 ;

[0038] Among them, d i,i-1 >0 indicates the desired distance between vehicle i and vehicle i-1; a constant distance strategy is adopted, i.e., d i,i-1 =d;

[0039] The upper-level trajectory planning control objective is:

[0040]

[0041]

[0042] Among them, f(t) represents the cost function and is a convex function;

[0043] The system error equation of the lower layer is:

[0044]

[0045] in, Represent the upper layer, the reference position, reference velocity and reference acceleration respectively; the tracking error vector is defined as Therefore, the error dynamics equation of vehicle i is expressed as:

[0046]

[0047] in, B i =[0 0 ξ i ] T , define the control objective of the trajectory tracking control layer as:

[0048]

[0049] Furthermore, in step 2, the reference kinetic equation is defined as:

[0050]

[0051] in, represents the external interference on the upper layer of vehicle i, Represents about μ i polynomial, w represents the polynomial coefficient, λ1, λ2, λ3 represent the control gain; then μ i The disturbance is:

[0052]

[0053] in, Matrix S i All eigenvalues ​​of have non-negative real parts;

[0054] According to the reference-reference dynamic equation, the distributed output feedback controller is expressed as:

[0055]

[0056] Among them, ξ i (0)=0, represents f i The reference position error, reference velocity error and reference acceleration error are expressed as:

[0057]

[0058] According to the above formula, we can get:

[0059]

[0060] The derivative is:

[0061]

[0062] Then, the control objective is:

[0063]

[0064]

[0065] The distributed optimization problem has a feasible solution for any set vector Constant ρ>0, a distributed optimization controller u can be designed i1 , ensuring the solution of the closed-loop system i (t) = 1, ..., N converge to the same point a0∈X * , for any and

[0066] Furthermore, in step 2, in order to remove external interference, the distributed optimization problem is converted into a distributed stability problem with IM;

[0067] because so The expression is:

[0068]

[0069] The minimal polynomial of is expressed as:

[0070]

[0071] in, Represents real numbers, by defining The following expression is obtained:

[0072]

[0073] in,

[0074] Because (Ψ i , Φ i ) is observable, there exists a matrix G i , so that M i =Φ i +G i u i1 represents the Hurwitz matrix. At this time, the IM of vehicle i is:

[0075]

[0076] Define the coordinate transformation as:

[0077]

[0078] After importing:

[0079]

[0080] in

[0081]

[0082]

[0083] Defining a vector Expressed as

[0084]

[0085]

[0086]

[0087]

[0088] For any There exists a distributed output feedback controller:

[0089]

[0090] When i ′(0), for any Then the solution of the system is bounded [0,∞), and can converge to 0.

[0091] Furthermore, in step 3, the control input u i2 When is optimal, the system tracking error dynamics model is rewritten as:

[0092]

[0093] Then the optimal cost function is:

[0094]

[0095] Combined with the optimal control theory, the corresponding HJB equation is constructed by derivation:

[0096]

[0097] in, represents the optimal gradient value of the cost function; according to the optimal control theory, the HJB equation of the system should satisfy:

[0098]

[0099] according to Therefore, the control law of the ideal optimal controller is expressed as:

[0100]

[0101] Substituting into the HJB equation we get

[0102]

[0103] Furthermore, in the Critic neural network, the optimal cost function is approximately:

[0104]

[0105] in, represents the ideal weight vector of the optimal cost function, N represents the number of neurons, represents the neuron regression vector, and, represents the network approximation error;

[0106] The gradient corresponding to the optimal cost function is expressed as:

[0107]

[0108] in, Indicates about e i The gradient of ; the residual of the function approximation error is obtained, which is expressed as follows:

[0109]

[0110] As the number of hidden layers N increases, the residual gradually converges to zero; that is,

[0111]

[0112] Since the ideal weight vector is unknown, leading to an estimate of the cost function Approach To obtain the actual optimal cost function:

[0113]

[0114] in, represents an estimate of the ideal weight;

[0115] Then the HJB equation can be rewritten as:

[0116]

[0117] Given any control strategy u i2 , adjust the appropriate Make The square of is the smallest, and the objective function is defined as follows:

[0118]

[0119] According to the gradient descent algorithm, the weight update law of the Critic neural network is:

[0120]

[0121] in, Used for normalization, a1>0 represents the learning rate;

[0122] Define the weight estimation error The weight estimation error update law is obtained as:

[0123]

[0124] in,

[0125] Then there exist constants β1>0, β2>0, T>0, satisfying:

[0126]

[0127] Then when When , the weight estimation error converges to zero, or the bounded Bellman error can make the critic weight estimation error converge to the residual set.

[0128] Furthermore, the control law of the optimal control strategy approximated by the Actor network is:

[0129]

[0130] in, represents the ideal neural network weights The estimated value of , the weight update rate of the Actor network is:

[0131]

[0132] in, It is bounded. Right now π iis a set positive constant, F2>0, F1>0 are adjustment parameters, a2 represents the learning rate of the Actor network; the weight update law of the Critic neural network is:

[0133]

[0134] in,

[0135] Compared with the prior art, the present invention has the following advantages:

[0136] This paper addresses the trajectory planning and tracking control issues in connected vehicle platooning. First, distributed connected vehicle trajectory planning is solved by leveraging only information exchange between neighboring vehicles, minimizing vehicle spacing errors. Then, based on a reinforcement learning strategy, optimal tracking control of the planned reference trajectory is achieved, eliminating the reliance on precise vehicle dynamics equations. BRIEF DESCRIPTION OF THE DRAWINGS

[0137] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0138] Figure 1 Schematic diagram of vehicle spacing.

[0139] Figure 2 The present invention is a double-layer control structure for vehicle queues

[0140] Figure 3 Schematic diagram of the weight change curve of the Critic neural network of the present invention.

[0141] Figure 4 Schematic diagram of the weight change curve of the Actor neural network of the present invention.

[0142] Figure 5 Schematic diagram of the acceleration curve of the present invention.

[0143] Figure 6 Schematic diagram of the acceleration curve under interference of the present invention.

[0144] Figure 7 Schematic diagram of the vehicle acceleration curve when the controller in reference

[10] is applied.

[0145] Figure 8 Schematic diagram of the speed error curve of the present invention.

[0146] Figure 9Schematic diagram of the speed error curve under interference of the present invention.

[0147] Figure 10 Schematic diagram of the speed error curve when applying the controller in reference

[10] .

[0148] Figure 11 Schematic diagram of the position error curve of the present invention.

[0149] Figure 12 Schematic diagram of the position error curve under interference of the present invention.

[0150] Figure 13 Schematic diagram of the position error curve when applying the controller in reference

[10] . DETAILED DESCRIPTION

[0151] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0152] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0153] The distributed optimization control method for intelligent connected vehicles based on reinforcement learning strategy has a hierarchical structure and includes the following steps:

[0154] Step 1: Build a vehicle longitudinal dynamics model and set upper and lower control objectives

[0155] Step 2: Design the upper layer, the trajectory optimization control layer. A trajectory planning controller based on a distributed convex optimization algorithm combined with the internal model principle is designed to remove external interference from the upper trajectory planning process.

[0156] Step 3: Design the lower layer, the tracking control layer; an optimal trajectory tracking controller based on the reinforcement learning actor-critic framework, where the actor network is used to approximate the optimal tracking controller and the critic neural network is used to approximate the optimal cost function;

[0157] Step 4: Analyze the stability and convergence of the tracking control system using the Lyapunov stability criterion to ensure accurate tracking of the reference trajectory;

[0158] Step 5: Verify the feasibility of the proposed algorithm through simulation experiments.

[0159] Furthermore, based on the VANET, consider a vehicle platoon consisting of N+1 vehicles traveling on a horizontal road (see Figure 1 ). At present, vehicle control systems can be divided into direct and hierarchical types. Direct control directly calculates the throttle opening and braking force by analyzing state quantities such as vehicle spacing, vehicle acceleration, and vehicle speed. However, due to the nonlinear characteristics of the vehicle system, the robustness of direct control is poor, and under extreme conditions, it is easy to cause the overall failure of the control system, thereby causing safety problems. In contrast, the hierarchical control system is mainly divided into upper and lower control systems, and its structural framework is as follows: Figure 2 As shown. In step 1, define x i , v i and a i Represents the position, speed and acceleration of the i-th vehicle respectively. The pilot vehicle is numbered 0 and the following vehicles are numbered from 1 to n. represents the communication topology of the convoy. Each following vehicle communicates with adjacent vehicles through the vehicle-mounted ad hoc network to obtain information, and the communication between vehicles is stable and reliable. Only the longitudinal dynamics model of the vehicle is considered. The longitudinal dynamics model is established to provide a basis for the design of the underlying controller. The longitudinal dynamics model of vehicle i is expressed as:

[0160]

[0161] in, θ i represents the uncertain time constant of the vehicle transmission system, u i2 (t) represents the control input of the lower-level system following vehicle i; the tracking error for vehicle i is given as follows:

[0162] δ i =x i-1 (t)-x i (t)-d i,i-1 ;

[0163] Among them, d i,i-1 >0 indicates the desired distance between vehicle i and vehicle i-1; a constant distance strategy is adopted, i.e., di,i-1 =d;

[0164] The upper-level trajectory planning control objective is:

[0165]

[0166]

[0167] Among them, f(t) represents the cost function and is a convex function;

[0168] The system error equation of the lower layer is:

[0169]

[0170] in, Represent the upper layer, the reference position, reference velocity and reference acceleration respectively; the tracking error vector is defined as Therefore, the error dynamics equation of vehicle i is expressed as:

[0171]

[0172] in, B i =[0 0 ξ i ] T , define the control objective of the trajectory tracking control layer as:

[0173]

[0174] Furthermore, in step 2, the reference kinetic equation is defined as:

[0175]

[0176] in, represents the external interference on the upper layer of vehicle i, represents the polynomial about μi, w represents the polynomial coefficient, λ1, λ2, λ3 represent the control gains; then the disturbance of μi is:

[0177]

[0178] in, Matrix S i All eigenvalues ​​of have non-negative real parts;

[0179] According to the reference-reference dynamic equation, the distributed output feedback controller is expressed as:

[0180]

[0181] Among them, ξ i(0)=0, represents f i The reference position error, reference velocity error and reference acceleration error are expressed as:

[0182]

[0183] According to the above formula, we can get:

[0184]

[0185] The derivative is:

[0186]

[0187] Then, the control objective is:

[0188]

[0189]

[0190] The distributed optimization problem has a feasible solution for any set vector Constant ρ>0, a distributed optimization controller u can be designed i1 , ensuring the solution of the closed-loop system i (t) = 1, ..., N converge to the same point a0∈X * , for any and

[0191] Furthermore, in step 2, in order to remove external interference, the distributed optimization problem is converted into a distributed stability problem with IM;

[0192] because so The expression is:

[0193]

[0194] The minimal polynomial of is expressed as:

[0195]

[0196] in, Represents real numbers, by defining The following expression is obtained:

[0197]

[0198] in,

[0199] Because (Ψ i , Φ i ) is observable, there exists a matrix G i , so that M i =Φ i +G i u i1 represents the Hurwitz matrix. At this time, the IM of vehicle i is:

[0200]

[0201] Define the coordinate transformation as:

[0202]

[0203] After importing:

[0204]

[0205] in

[0206]

[0207]

[0208] Defining a vector Expressed as

[0209]

[0210]

[0211]

[0212]

[0213] For any There exists a distributed output feedback controller:

[0214]

[0215] When i ′(0), for any Then the solution of the system is bounded [0,∞), and can converge to 0.

[0216] Furthermore, in step 3, the control input u i2 When is optimal, the system tracking error dynamics model is rewritten as:

[0217]

[0218] Then the optimal cost function is:

[0219]

[0220] Combined with the optimal control theory, the corresponding HJB equation is constructed by derivation:

[0221]

[0222] in, represents the optimal gradient value of the cost function; according to the optimal control theory, the HJB equation of the system should satisfy:

[0223]

[0224] according to Therefore, the control law of the ideal optimal controller is expressed as:

[0225]

[0226] Substituting into the HJB equation we get

[0227]

[0228] Furthermore, in the Critic neural network, the optimal cost function is approximately:

[0229]

[0230] in, represents the ideal weight vector of the optimal cost function, N represents the number of neurons, represents the neuron regression vector, and, represents the network approximation error;

[0231] The gradient corresponding to the optimal cost function is expressed as:

[0232]

[0233] in, Indicates about e i The gradient of ; the residual of the function approximation error is obtained, which is expressed as follows:

[0234]

[0235] As the number of hidden layers N increases, the residual gradually converges to zero; that is,

[0236]

[0237] Since the ideal weight vector is unknown, leading to an estimate of the cost function Approach To obtain the actual optimal cost function:

[0238]

[0239] in, represents an estimate of the ideal weight;

[0240] Then the HJB equation can be rewritten as:

[0241]

[0242] Given any control strategy u i2 , adjust the appropriate Make The square of is the smallest, and the objective function is defined as follows:

[0243]

[0244] According to the gradient descent algorithm, the weight update law of the Critic neural network is:

[0245]

[0246] in, Used for normalization, a1>0 represents the learning rate;

[0247] Define the weight estimation error The weight estimation error update law is obtained as:

[0248]

[0249] in,

[0250] Then there exist constants β1>0, β2>0, T>0, satisfying:

[0251]

[0252] Then when When , the weight estimation error converges to zero, or the bounded Bellman error can make the critic weight estimation error converge to the residual set.

[0253] Furthermore, the control law of the optimal control strategy approximated by the Actor network is:

[0254]

[0255] in, represents the ideal neural network weights The estimated value of , the weight update rate of the Actor network is:

[0256]

[0257] in, It is bounded. Right now π i is a set positive constant, F2>0, F1>0 are adjustment parameters, a2 represents the learning rate of the Actor network; the weight update law of the Critic neural network is:

[0258]

[0259] in,

[0260] Preferably, in step 3, the following Lyapunov function of the trajectory tracking control system is constructed for the member vehicle i:

[0261]

[0262] in, Finding the time derivative, we can get:

[0263]

[0264] in, It is expressed as follows:

[0265]

[0266] Will α i2 Substituting the expression into the above formula, we can get

[0267]

[0268] in:

[0269]

[0270]

[0271] make Then, rewrite it into the following form

[0272]

[0273] make Will Substitute the expression of into the above formula, and we can get

[0274]

[0275] In summary, we can The expression can be further rewritten into the following form:

[0276]

[0277] in, The above formula can be further simplified into the following form

[0278]

[0279] in,

[0280]

[0281] Will The following expression can be obtained:

[0282]

[0283] Formula (47) and Substitute the above expression j and rewrite it into

[0284]

[0285] ε1(e i ) expression considering the Cauchy-Schwarz inequality:

[0286]

[0287] according to By definition, there exists a real number δ max ,Right now, The upper bound is Will The expression and formula ε1(e i ) into the following inequality:

[0288]

[0289] Substituting into the above formula we can get:

[0290]

[0291] in

[0292]

[0293]

[0294] In the above formula, For a normal number According to Young's inequality, we can get

[0295]

[0296]

[0297]

[0298]

[0299]

[0300] Simplified processing can be obtained as follows

[0301]

[0302] in:

[0303]

[0304]

[0305]

[0306] For p1, p2, p3, p4, p5, p6 exist

[0307] like There is the following inequality

[0308]

[0309]

[0310]

[0311] The following inequality can be obtained:

[0312]

[0313] in,

[0314]

[0315]

[0316]

[0317]

[0318] By adjusting the parameters, K is made into a positive definite matrix, and is a finite constant, then j is further expressed as

[0319]

[0320] Where υ=λ min .

[0321] We can get:

[0322]

[0323] Based on Lyapunov's theorem, the system is stable and the error e i It is convergent. It is also convergent.

[0324] Example 1

[0325] In order to verify the effectiveness of the two-layer controller designed in this chapter, this section will build a vehicle dynamics model and the corresponding control algorithm on the MATLAB / SIMULINK platform, and highlight the superiority of the algorithm designed in this paper by comparing it with the control algorithm proposed in the literature

[10] .

[0326] An experimental simulation is conducted on a convoy system consisting of one pilot vehicle and three follower vehicles. Assume that the pilot vehicle is numbered 0 and the remaining member vehicles are numbered 1-3. The experimental vehicle parameters are:

[0327] Reference kinetic model parameters (λ1, λ2, λ2) = (-0.8, 1, 1), interference:

[0328] q i (t) = d mi sin(ω i t)+s mi ,

[0329] (d m1 , d m2 , d m3 )=(5,8,4);

[0330] (ω1, ω2, ω3)=(π / 4, π / 5, π / 7);

[0331] (s m1 , s m2 , s m3 )=(3,2,1);

[0332] G i =[-8, -24, -32, -64] T ;

[0333] Safety vehicle distance d = 5m, controller parameters: Ψ i =[1,0,0,0],k=15,∧ e =I 3×3 , ∧ u =I3×3 , a1=0.85I 6×6 , a2=0.9I 6×6 , a2=0.9I 6×6 , F2=0.5I 6×6 ; Vehicle initial state: (a0, ...a3) = (0, 0, 0, 0), (v0, ...v3) = (8, 8, 8, 8), (x0, ...x3) = (10, 2, -5, -11).

[0334] Figure 3 and Figure 4 This is the update curve of the reinforcement learning Actor-Critic neural network weight designed by the present invention. It can be seen from the figure that the neural network has the characteristics of rapid convergence and can better meet actual needs.

[0335] Example 2

[0336] Figure 5 represents the acceleration curve of vehicles in the convoy, Figure 6 is the acceleration curve without interference, Figure 7 is the acceleration curve corresponding to the distributed speed optimization algorithm; it can be seen from the acceleration curve that using the algorithm in this paper, all cars can quickly adjust their own status to track the acceleration of the leader car. In addition, the algorithm can not only remove external interference and solve the problem of acceleration fluctuation caused by interference, but also minimize the acceleration change in the initial acceleration adjustment stage, thereby improving the comfort and safety of the ride.

[0337] By analyzing the speed error curve ( Figure 8 , Figure 9 , Figure 10 ) It is not difficult to find that the algorithm in this paper can quickly achieve the consistency of the driving speed of each member vehicle, minimize the speed error, and the vehicle driving speed is not affected by external interference. Although the convergence speed is slower, in the long run, the smaller the speed fluctuation of the vehicle queue, the more conducive it is to maintaining the stability and safety control of the queue.

[0338] Example 3

[0339] Position error curve is as follows Figure 11 , Figure 12 , Figure 13 As shown, from Figure 11 It can be seen that the error convergence of the algorithm in this paper gradually converges to zero at about 12s. In contrast, Figure 13 The error gradually converges to zero after 9 seconds. Although the convergence speed of the proposed algorithm is slightly reduced after the interference is removed, the error still converges when the acceleration changes during the period of 30s-60s, and the effect is better than Figure 13 .pass Figure 12It can be seen that when vehicles are subject to external interference, the inter-vehicle distance error is difficult to converge and may lead to safety hazards. Therefore, removing interference not only helps achieve the vehicle platoon control goal, but also reduces safety issues during platoon driving.

[0340] Based on the above analysis, the control algorithm proposed in this paper successfully eliminates external interference to vehicles and achieves the goal of vehicle platoon control. In addition, the algorithm further improves the comfort, safety, and stability of the platoon.

[0341] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0342] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0343] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0344] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0345] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0346] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.

[0347] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A distributed optimization control method for intelligent connected vehicles based on reinforcement learning strategy has a hierarchical structure, characterized by: The following steps are involved: Step 1: Establish a vehicle longitudinal dynamics model and set the upper and lower control targets; in step 1, define , and Respectively represent The position, speed and acceleration of the vehicle. The leader vehicle is numbered 0 and the following vehicles are numbered from 1 to ; Through a directed connected graph Represents the communication topology of the convoy. Each following vehicle communicates with the adjacent vehicle to obtain information through the vehicle-mounted ad hoc network, and the communication between the vehicles is stable and reliable. Only the longitudinal dynamics model of the vehicle is considered. The longitudinal dynamics model is established to provide a basis for the design of the underlying controller. The vehicle The longitudinal dynamic model of is expressed as: ; in, , represents the uncertain time constant of the vehicle transmission system, Indicates following car The lower system control input of the vehicle The tracking error is determined as follows: ; in, Indicates vehicle With vehicle The desired vehicle spacing between them; adopting a constant spacing strategy, that is, ; The upper-level trajectory planning control objective is: ; ; in, Represents the cost function, which is a convex function; The system error equation of the lower layer is: ; in, , , Represent the upper layer, the reference position, reference velocity and reference acceleration respectively; the tracking error vector is defined as , therefore, the vehicle The error dynamics equation is expressed as: ; in, , , define the control objective of the trajectory tracking control layer as: ; Step 2: Design the upper layer, i.e., the trajectory optimization control layer; a trajectory planning controller based on a distributed convex optimization algorithm combined with the internal model principle is designed. The controller designed in combination with the internal model principle removes external interference in the upper layer trajectory planning process; in step 2, the reference dynamic equation is defined as: ; in, Indicates vehicle External interference on the upper layer, Express about The polynomial of represents the polynomial coefficients, , , represents the control gain; then The disturbance is: ; in, ,matrix All eigenvalues ​​of have non-negative real parts; According to the reference-reference dynamic equation, the distributed output feedback controller is expressed as: ; in, , express The reference position error, reference velocity error and reference acceleration error are expressed as: ; According to the above formula, we can get: ; The derivative is: ; Then, the control objective is: ; The distributed optimization problem has a feasible solution for any set 、 、 ,vector , , ,constant , a distributed optimization controller can be designed , ensuring the solution of the closed-loop system Converge to the same point , for any and ; Step 3: Design the lower layer, the tracking control layer; the optimal trajectory tracking controller based on the reinforcement learning Actor-Critic framework, where the Actor network is used to approximate the optimal tracking controller and the Critic neural network is used to approximate the optimal cost function; in step 3, the control input When is optimal, the system tracking error dynamics model is rewritten as: ; Then the optimal cost function is: ; Combined with the optimal control theory, the corresponding HJB equation is constructed by derivation: ; in, represents the optimal gradient value of the cost function; according to the optimal control theory, the HJB equation of the system should satisfy: ; according to , so the control law of the ideal optimal controller is expressed as: ; Substituting into the HJB equation we get ; Step 4: Analyze the stability and convergence of the tracking control system using the Lyapunov stability criterion to ensure accurate tracking of the reference trajectory; Step 5: Verify the feasibility of the proposed algorithm through simulation experiments; In the Critic neural network, the optimal cost function is approximately: ; in, represents the ideal weight vector of the optimal cost function, , represents the number of neurons, represents the neuron regression vector, and, , , represents the network approximation error; The gradient corresponding to the optimal cost function is expressed as: ; in, , Express about The gradient of ; the residual of the function approximation error is obtained, which is expressed as follows: ; As the number of hidden layers N increases, the residual gradually converges to zero; that is, ; Since the ideal weight vector is unknown, leading to an estimate of the cost function Approach To obtain the actual optimal cost function: ; in, represents an estimate of the ideal weight; Then the HJB equation can be rewritten as: ; Given any control strategy , adjust the appropriate Make The square of is the smallest, and the objective function is defined as follows: ; According to the gradient descent algorithm, the weight update law of the Critic neural network is: ; in, , For normalization, represents the learning rate; Define the weight estimation error , the weight estimation error update law is obtained as: ; in, , ; Then there is a constant , , ,satisfy: ; Then when When , the weight estimation error converges to zero, or the bounded Bellman error can make the Critic weight estimation error converge to the residual set; the control law of the optimal control strategy approximated by the Actor network is: ; in, represents the ideal neural network weights The estimated value of , the weight update rate of the Actor network is: ; in, , is bounded, ,Right now , is a set positive constant, , To adjust the parameters, represents the learning rate of the Actor network; the weight update law of the Critic neural network is: ; in, , .

2. The distributed optimization control method for intelligent connected vehicles based on reinforcement learning strategy according to claim 1 is characterized in that: In step 2, in order to remove external interference, the distributed optimization problem is converted into a distributed stability problem with IM; because ,so The expression is: , ; The minimal polynomial of is expressed as: ; in, Represents a real number, by defining , we get the following expression: ; in, ; because It is observable, and there is a matrix , making Represents the Hurwitz matrix. At this time, the vehicle The IM is: ; Define the coordinate transformation as: ; After importing: ; in ; ; Defining a vector , , , Expressed as ; ; ; ; For any , , , There exists a distributed output feedback controller: ; when When any , then the solution of the system is bounded ,and can converge to 0.

Citation Information

Patent Citations

  • Robot trajectory tracking optimal control method based on an event trigger mechanism

    CN113093548A

  • Random multi-agent optimization control method and system based on reinforcement learning

    CN117130272A