Self-driving vehicle steering and suspension cooperative control method based on reinforcement learning
Through a reinforcement learning-based method, combined with the dynamic model of the steering system and the suspension system and PID control strategy, the steering and suspension coordinated control of autonomous vehicles is achieved, solving the problem that it is difficult for the vehicle to ensure comfort and stability at the same time in path tracking control, and achieving efficient, safe and comfortable path tracking effect.
Patent Information
- Application Number
- CN202411971777.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-30
AI Technical Summary
In the path tracking control, existing autonomous driving vehicles are difficult to ensure the comfort and stability of the vehicle during horizontal and vertical maneuvering, especially in complex driving environments.
The coordinated control method of steering and suspension of an autonomous vehicle based on reinforcement learning is adopted. By establishing a dynamic model of the vehicle steering system and suspension system, the control action and reward function is constructed, and the intelligent body is constructed using deep neural networks to realize the end-to-end learning process. In combination with PID control strategies, the coordinated control of the steering system and suspension system is realized.
It effectively realizes efficient, safe and comfortable path tracking control of the vehicle in complex environments, and improves the lateral roll stability and vertical smoothness of the vehicle.
Smart Images

Figure CN119975527A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of motion control of autonomous driving vehicles, and in particular to a method for coordinated steering and suspension control of autonomous driving vehicles based on reinforcement learning. Background Art
[0002] Intelligent car autonomous driving technology is an important way to alleviate traffic congestion, reduce traffic accidents, and improve road and vehicle utilization. Among them, high-level autonomous driving technology requires steering and other subsystems to achieve highly autonomous coordination. Designing a more comfortable and stable path tracking solution will be a necessary prerequisite for highly intelligent autonomous driving vehicles. At present, lateral control methods mainly focus on improving the stability of autonomous driving vehicles, while there are still some blind spots in the study of ride comfort of autonomous driving vehicles. As autonomous driving vehicles continue to develop towards a highly intelligent direction, considering comfort and stability will further enhance their safety and experience.
[0003] In the path tracking control of autonomous driving vehicles, the under-actuated process only executes the path tracking process of the vehicle by controlling the steering angle, which makes it difficult to ensure the comfort and stability of the vehicle in the lateral and longitudinal directions while tracking the reference path. In order to achieve safe, comfortable and intelligent autonomous driving control functions in complex driving environments, most studies use methods such as yaw torque and active steering control in a single path tracking task to improve the comfort and stability of autonomous driving vehicles in the lateral and longitudinal directions. However, there is a certain constraint relationship between the tracking ability and comfort stability in the plane model. Summary of the invention
[0004] The main purpose of the present invention is to overcome the above-mentioned defects of vehicle multi-system motion control and propose a method for coordinated control of steering and suspension of an autonomous driving vehicle based on reinforcement learning.
[0005] The present invention adopts the following technical solution:
[0006] A method for coordinated control of steering and suspension of an autonomous driving vehicle based on reinforcement learning comprises the following steps:
[0007] 1) Establish the dynamic model of the vehicle steering system and suspension system, and obtain the lateral and vertical vehicle motion state by controlling the steering angle and suspension force;
[0008] 2) Construct the control actions and reward functions of the vehicle steering system and suspension system to form a reinforcement learning exploration strategy for the coordinated control of vehicle steering and suspension;
[0009] 3) Use deep neural networks to build intelligent agents for the steering and suspension systems, and realize an end-to-end learning process from the motion state of the steering and suspension systems to the exploration strategy;
[0010] 4) A reinforcement learning environment is constructed by combining the dynamic model and the reward function module, and the coordinated control of the steering system and the suspension system is achieved through the reinforcement learning method of the intelligent agent of the steering system and the suspension system.
[0011] In step 1), the steering system dynamics model is constructed by combining the roll and yaw motion states in the two-degree-of-freedom steering system model:
[0012]
[0013] Among them, m and m s are the total vehicle mass and the sprung mass of the vehicle respectively; h is the distance from the center of mass of the vehicle to the roll center; is the road inclination; g is the acceleration due to gravity; l f With l r are the distances from the center of mass to the front and rear axles, respectively; I z is the yaw moment of inertia; r is the yaw angular velocity; is the yaw angular acceleration; v x With v y are the longitudinal and lateral velocities, respectively; is the lateral acceleration; is the roll acceleration; F yf With F yr are the lateral forces of the front and rear axle tires respectively; κ is the road curvature; e ψ is the heading deviation; and are the rates of change of lateral deviation and heading deviation respectively.
[0014] The magic formula model is used to describe the relationship between the tire lateral force and the sideslip angle. The tire model is expressed as:
[0015] F y =Dsin[Carctan(Bx-E(Bx-arctan(Bx)))]+S v
[0016] x=α+S h
[0017] D=A1F z 2 +A2F z
[0018] B=A3sin(A4arctan(A5F z )) / CD
[0019] E=A6F z 2 +A7F z +A8
[0020] Among them, F y is the tire lateral force; B is the curve stiffness factor; C is the curve shape factor; D is the curve peak factor; E is the curve curvature factor; α is the tire side slip angle; S h With S v are the lateral and longitudinal offsets of the curve respectively; F z is the vertical load of the tire; A1~A8 are the identification parameters of the tire model.
[0021] Furthermore, the front and rear tire slip angles can be expressed as:
[0022]
[0023] Among them, α f With α r are the front and rear tire side slip angles respectively; δ f is the front wheel steering angle.
[0024] So far, a vehicle steering system dynamics model considering suspension roll has been established, which can obtain the lateral vehicle motion state by controlling the steering angle.
[0025] The forces of the front and rear suspension systems are added together and concentrated on an equivalent roll axis, and the dynamic model of the suspension system is constructed as follows:
[0026]
[0027] Among them, z s is the displacement of the vehicle body; is the vehicle body acceleration; a y is the lateral acceleration; I x is the roll moment of inertia; is the vehicle roll angle; z s1 With z s2 are the displacement of the left and right sides of the vehicle body respectively; z w1 With z w2 are the left and right wheel displacements respectively; and are the left and right wheel accelerations respectively; k s1 With k s2 are the left and right suspension stiffness coefficients respectively; m w1 With m w2 are the unsprung masses on the left and right sides respectively; z r1 With z r2 are the road excitation displacements on the left and right sides respectively; k t1 With k t2 are the stiffness coefficients of the left and right tires respectively; F s1 With F s2 The left and right suspensions are powered respectively; T w is the wheel track.
[0028] The road surface excitation displacement generated by the filtered white noise method is:
[0029]
[0030] Among them, z r (t) is the road surface excitation displacement at time t; is the rate of change of the road surface excitation displacement at time t; η, σ are constants related to the road surface grade; G q (n0) is the road surface power spectrum density at the reference spatial frequency n0, that is, the road surface roughness coefficient, which can be obtained through the road surface roughness classification standard; w(t) is a unit white noise with a mean of 0 and a variance of 1.
[0031] At this point, a vehicle suspension system dynamics model associated with the steering system dynamics model has been established, which can obtain the vertical vehicle motion state by controlling the suspension force.
[0032] In step 2), the present invention combines the prior feedforward steering to improve the efficiency of reinforcement learning exploration, and applies feedforward steering in the steering control exploration process of reinforcement learning. Through the front and rear wheel side slip motion relationship, the relationship between the vehicle's feedforward steering angle and the front and rear tire side slip angle can be described as:
[0033] δ f ′=Lκ-α f ′+α r '
[0034] Among them, δ f ′ is the feedforward steering angle; L is the distance from the front axle to the rear axle of the vehicle; α f ′ and α r ′ are the feedforward tire slip angles of the front and rear wheels, respectively, which can be obtained by the feedforward tire force, which is expressed as:
[0035]
[0036] Among them, F yf ′ and F yr ′ are the feedforward tire forces of the front and rear wheels respectively.
[0037] The steering exploration command required for reinforcement learning is calculated by the sum of feedforward and learning feedback control. The control action of the steering system is expressed as:
[0038] a p =π p +δ f ′+N p
[0039] Among them, a p is the control action of the steering system; π pis the exploration strategy for reinforcement learning; N p Random exploration noise for reinforcement learning.
[0040] In order to further improve the path tracking stability and accuracy in the reinforcement learning process, the vehicle's envelope constraint performance is considered in the reward function. m Set to:
[0041] R m =-(k e e y 2 +k ψ e ψ 2 +k δ δ f 2 +k Δδ Δδ f 2 )
[0042] Among them, Δδ f is the control increment of the steering angle; k e , k ψ , k δ , k Δδ They are the penalty coefficients for lateral deviation, heading deviation, steering angle control amount and control increment respectively.
[0043] The constraints of the steering angle control amount and control increment are set as:
[0044]
[0045] Among them, R p1 With R p2 are the constraints for controlling the steering angle and its increment respectively; k δ′ With k Δδ′ are the penalty coefficients for controlling the steering angle and its incremental constraint respectively; δ fmax With Δδ fmax They are the maximum limit values for controlling the steering angle and its increment respectively.
[0046] Penalties are set for actions that exceed the vehicle's yaw and lateral motion constraints. The constraints are:
[0047]
[0048] Among them, R p3 With R p4 are the constraints of vehicle yaw and lateral motion respectively; k β With k r are the penalty coefficients for the sideslip angle and yaw rate constraints respectively; α s is the saturated tire slip angle; b1 and b2 are the constraint boundaries of the yaw angular velocity.
[0049] Combining the above objective function and constraint function, we get the reward function R of the vehicle steering system: p for:
[0050] R p =R m +R p1 +R p2 +R p3 +R p4
[0051] In order to better extend the reinforcement learning method to the complex suspension system environment, a fast-response PID control strategy is used to compensate for the uncertainty in the suspension exploration process. Two PID controllers are used to control the vertical and roll motions of the suspension system respectively. The control inputs are the body acceleration and roll acceleration error terms with the goals of smoothness and roll stability. The corresponding PID control design is:
[0052]
[0053] Among them, f z (t) are the control forces of vertical and rolling motion at time t respectively; e z (t) are the error inputs of vertical and rolling motion at time t respectively; K zp , K zi With K zd is the PID control coefficient of vertical motion; and is the PID control coefficient of the roll motion.
[0054] The control force relationship of the decoupled distribution can be expressed as:
[0055] f z (t) = f d1 (t)+f d2 (t),
[0056] Among them, f d1 (t) and f d2 (t) are the forces acting on the left and right suspensions at time t respectively.
[0057] The compensation force of the left and right suspensions can be calculated as:
[0058]
[0059] The required suspension system exploration command is calculated by the sum of the suspension compensation force and the learning feedback control force. The control action of the suspension system is expressed as:
[0060] a s1 =π s1 +f d1 +N s1 ,a s2 =π s2 +f d2 +N s2
[0061] Among them, a s1 with a s2 are the control actions of the left and right suspensions respectively; π s1 With π s2 are the exploration strategies of the left and right suspensions respectively; N s1 With N s2 are the random exploration noises of the left and right suspensions, respectively.
[0062] Since the dimensions of the body acceleration, suspension travel, body roll acceleration and suspension force of the suspension system are quite different, the various indicators in the reward process are normalized according to the dynamic response process of the passive suspension system to obtain the target reward function R n for:
[0063]
[0064] Among them, k a , k1, k2, k f1 , k f2 They are the penalty coefficients of vehicle speed acceleration, roll acceleration, left and right suspension travel, and suspension force; z 1nor 、z 2nor 、F 1nor 、F 2nor They are the normalized terms of vehicle speed acceleration, roll acceleration, left and right suspension travel and suspension force.
[0065] The constraints on the suspension travel are extended to the left and right suspensions. The constraints on the suspension travel need to meet the following requirements:
[0066]
[0067] Among them, R s1 With R s2 are the constraints of the left and right suspension travel respectively; k d1 With k d2 are the penalty coefficients of the left and right suspension travel constraints respectively; (z s1 -z w1 ) max ,(z s2 -z w2 ) maxThey are the maximum limit values of the left and right suspension dynamic travel respectively.
[0068] The control constraints of the left and right suspensions are set as:
[0069] |F s1 |≤F s1max ,|F s2 ≤F s2max
[0070] Among them, F s1max With F s2max They are the maximum limit values of the left and right suspension control quantities respectively.
[0071] Finally, the reward function R of the vehicle suspension system is obtained s for:
[0072] R s =R n +R s1 +R s2
[0073] In step 3), the present invention uses a deep deterministic policy gradient algorithm to construct an intelligent agent for the steering system and the suspension system, and the intelligent agent includes a behavior network and an evaluation network. The behavior network of the steering system consists of a state input layer, three hidden layers and a control output layer. Among them, the state of the state input layer is 6-dimensional, including lateral velocity, yaw angular velocity, lateral deviation, heading deviation, road curvature and control action at the previous moment, each hidden layer is composed of 100 neurons, and the control output layer is the exploration strategy of the steering system. The evaluation network of the steering system mainly includes a state input layer, a control input layer, three hidden layers and an output layer. Among them, the state input layer obtains the 6-dimensional state of the intelligent agent from the behavior network of the steering system, the control input layer is a 1-dimensional control action, each hidden layer is composed of 100 neurons, and the output layer is a 1-dimensional action value function for evaluating the path tracking control behavior of the steering system. The three hidden layers are connected in series in sequence, the state input layer is connected to the first hidden layer, the control input layer skips the first hidden layer, and is connected and transmitted with the second hidden layer.
[0074] The behavior network of the suspension system mainly includes a state input layer, three hidden layers and a control output layer. Among them, the state input layer is 6-dimensional, including vehicle body displacement, vehicle roll angle, left and right suspension dynamic travel and left and right suspension control actions at the previous moment. Each hidden layer is set with 100 neurons. The control output layer is the exploration strategy of the left and right suspensions. The control output layer scales the suspension force within the range of [-1,1] through the activation function. The evaluation network of the suspension system includes a state input layer, a control input layer, three hidden layers and an output layer. The state input layer of the evaluation network of the suspension system is consistent with the behavior network of the suspension system. The control input layer is a 2-dimensional left and right suspension control action. The three hidden layers are connected in series in sequence. The control input layer is directly connected to the second hidden layer, and the state input layer is connected to the first hidden layer. Each hidden layer consists of 100 neurons, and the output layer is the action value function for evaluating the behavior of the suspension system.
[0075] When the intelligent agent is being trained, if the vehicle exceeds a certain error range during the path tracking process, continuing to run will add too much invalid data. Therefore, when the vehicle tracking error exceeds 1m during the learning process, the intelligent agent training process will be terminated to start a new training cycle, effectively improving the convergence speed of the training learning process.
[0076] In step 4), the vehicle steering system and suspension system dynamics model, tracking error model, road model and steering and suspension system reward modules together constitute the reinforcement learning environment. In the path tracking control process of the autonomous driving vehicle, the lateral and vertical motion processes of the vehicle in the path tracking process are controlled by outputting the control actions of the steering and suspension systems of the controlled vehicle. The two agents respectively observe the motion states of the steering system and suspension system of the reinforcement learning environment, including the lateral velocity, yaw angular velocity, lateral deviation, heading deviation, and road curvature of the steering system, and the body displacement, vehicle roll angle, and suspension travel of the suspension system. The control actions of the steering system are respectively executed according to the strategies in the agents. p Control action of the suspension system s1 with a s2 After making a control decision, the motion state of the autonomous vehicle changes, and the control state enters the next control state. At the same time, the reward function module of the steering system is calculated by the lateral motion state s p (v y ,r,e y ,e ψ ,κ) calculates the reward value obtained after executing the steering control action strategy. The reward function module of the suspension system is based on the vertical motion state The reward value obtained after executing the suspension control action strategy is calculated, and the agent further updates the applied strategy based on the feedback reward value. This cycle continues until the agent obtains enough cumulative rewards to achieve coordinated control of the steering system and suspension system.
[0077] It can be seen from the above description of the present invention that, compared with the prior art, the present invention has the following beneficial effects:
[0078] (1) By constructing the lateral and vertical dynamic models of the steering and suspension systems, a multi-agent reinforcement learning method including the steering agent and the suspension agent is designed to reduce the data input dimension. At the same time, lateral motion constraints and vertical motion constraints are derived and applied to effectively achieve efficient, safe and comfortable multi-objective coordinated control of the two systems.
[0079] (2) In order to improve the exploration efficiency of the reinforcement learning method, the method of the present invention adds feedforward steering information to the steering system, reduces the workload of reinforcement learning feedback control, and reduces the invalid steering search process in the reinforcement learning tracking process. The PID compensation control process is added to the reinforcement learning of the complex suspension system, which makes it easier to search for the strategic actions of the coupling system, ensure the efficiency of strategy learning and reduce the learning cost.
[0080] (3) The method of the present invention comprehensively considers different vertical and lateral plane environments and different learning tasks, and takes into account the optimal motion control strategies for different working conditions. Under the premise of ensuring accurate path tracking, the lateral roll stability and vertical smoothness and comfort are improved by coordinating the steering system and the suspension system control. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 is a flow chart of the method of the present invention;
[0082] Figure 2 is a schematic diagram of a dynamic model of a steering system of the present invention;
[0083] Figure 3 is a schematic diagram of a dynamic model of a suspension system of the present invention;
[0084] Figure 4a is a behavioral network diagram of the steering system of the present invention;
[0085] Figure 4b is a schematic diagram of an evaluation network of a steering system of the present invention;
[0086] Figure 5a is a behavioral network diagram of the suspension system of the present invention;
[0087] Figure 5b is a schematic diagram of an evaluation network of a suspension system of the present invention;
[0088] Figure 6is a schematic diagram of a reinforcement learning method for coordinated control of a steering and suspension system of the present invention;
[0089] Figure 7 is a comparative schematic diagram of path tracking of the present invention;
[0090] Figure 8 2 is a schematic diagram for comparing suspension control of the present invention.
[0091] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. DETAILED DESCRIPTION
[0092] The present invention is further described below through specific implementation modes.
[0093] See also Figure 1 , the method of the present invention comprises the following steps:
[0094] 1) Establish the dynamic model of the vehicle steering system and suspension system, and obtain the lateral and vertical vehicle motion states by controlling the steering angle and suspension force.
[0095] Specifically, in order to consider the roll motion characteristics of the suspension in the steering system dynamics model, the vehicle roll degree of freedom is increased. Further combined with the yaw motion state in the two-degree-of-freedom steering system model, such as Figure 2 As shown in the figure, the steering system dynamics model is constructed under the small angle assumption:
[0096]
[0097] Among them, m and m s are the total vehicle mass and the sprung mass of the vehicle respectively; h is the distance from the center of mass of the vehicle to the roll center; is the road inclination; g is the acceleration due to gravity; l f With l r are the distances from the center of mass to the front and rear axles, respectively; I z is the yaw moment of inertia; r is the yaw angular velocity; is the yaw angular acceleration; v x With v y are the longitudinal and lateral velocities, respectively; is the lateral acceleration; is the roll acceleration; F yf With F yr are the lateral forces of the front and rear axle tires respectively; κ is the road curvature; e ψ is the heading deviation; and are the rates of change of lateral deviation and heading deviation respectively.
[0098] The complex nonlinear mechanical behavior in the vehicle steering system dynamics model mainly comes from the interaction between the tire and the ground. The present invention uses the magic formula model to describe the relationship between the tire lateral force and the sideslip angle. The tire model is expressed as:
[0099] F y =Dsin[Carctan(Bx-E(Bx-arctan(Bx)))]+S v
[0100] x=α+S h
[0101] D=A1F z 2 +A2F z
[0102] B=A3sin(A4arctan(A5F z )) / CD
[0103] E=A6F z 2 +A7F z +A8
[0104] Among them, F y is the tire lateral force; B is the curve stiffness factor; C is the curve shape factor; D is the curve peak factor; E is the curve curvature factor; α is the tire side slip angle; S h With S v are the lateral and longitudinal offsets of the curve respectively; F z is the vertical load of the tire; A1~A8 are the identification parameters of the tire model.
[0105] Furthermore, the front and rear tire slip angles can be expressed as:
[0106]
[0107] Among them, α f With α r are the front and rear tire side slip angles respectively; δ f is the front wheel steering angle.
[0108] So far, a vehicle steering system dynamics model considering suspension roll has been established, which can obtain the lateral vehicle motion state by controlling the steering angle.
[0109] In order to establish a suspension system model associated with the roll characteristics in the steering system dynamics model, the front and rear suspension system forces are added together and concentrated on an equivalent roll axis, such as Figure 3 As shown in the figure, the constructed suspension system dynamics model is:
[0110]
[0111] Among them, z s is the displacement of the vehicle body; is the vehicle body acceleration; a y is the lateral acceleration; I x is the roll moment of inertia; is the vehicle roll angle; z s1 With z s2 are the displacement of the left and right sides of the vehicle body respectively; z w1 With z w2 are the left and right wheel displacements respectively; and are the left and right wheel accelerations respectively; k s1 With k s2 are the left and right suspension stiffness coefficients respectively; m w1 With m w2 are the unsprung masses on the left and right sides respectively; z r1 With z r2 are the road excitation displacements on the left and right sides respectively; k t1 With k t2 are the stiffness coefficients of the left and right tires respectively; F s1 With F s2 The left and right suspensions are powered respectively; T w is the wheel track.
[0112] Road roughness is the main external interference factor that affects the dynamic characteristics of the suspension. Flat asphalt roads and washboard roads all show significant discrete random road roughness. The road excitation displacement generated by the filtered white noise method is:
[0113]
[0114] Among them, z r (t) is the road surface excitation displacement at time t; is the rate of change of the road surface excitation displacement at time t; η, σ are constants related to the road surface grade; G q (n0) is the road surface power spectrum density at the reference spatial frequency n0, that is, the road surface roughness coefficient, which can be obtained through the road surface roughness classification standard; w(t) is a unit white noise with a mean of 0 and a variance of 1.
[0115] At this point, a vehicle suspension system dynamics model associated with the steering system dynamics model has been established, which can obtain the vertical vehicle motion state by controlling the suspension force.
[0116] 2) Construct the control actions and reward functions of the vehicle steering system and suspension system to form a reinforcement learning exploration strategy for the coordinated control of vehicle steering and suspension.
[0117] Specifically, in the process of autonomous driving path tracking of the steering system, disordered learning exploration in high-dimensional continuous state space is prone to fall into local optimization. The present invention combines a priori feedforward steering to improve the efficiency of reinforcement learning exploration, applies feedforward steering in the process of reinforcement learning steering control exploration, reduces the workload of reinforcement learning feedback control, and reduces the ineffective steering search process. Through the relationship between the front and rear wheel side slip motion, the relationship between the vehicle's feedforward steering angle and the front and rear tire side slip angle can be described as:
[0118] δ f ′=Lκ-α f ′+α r '
[0119] Among them, δ f ′ is the feedforward steering angle; L is the distance from the front axle to the rear axle of the vehicle; α f ′ and α r ′ are the feedforward tire slip angles of the front and rear wheels, respectively, which can be obtained by the feedforward tire force, which is expressed as:
[0120]
[0121] Among them, F yf ′ and F yr ′ are the feedforward tire forces of the front and rear wheels respectively.
[0122] Feedforward steering predicts steering commands in advance based on road curvature and vehicle speed information, while reinforcement learning feedback control adjusts steering input based on tracking error. The required steering exploration command is calculated by the sum of feedforward and learning feedback control. The control action of the steering system is expressed as:
[0123] a p =π p +δ f ′+N p
[0124] Among them, a p is the control action of the steering system; π p is the exploration strategy for reinforcement learning; N p Random exploration noise for reinforcement learning.
[0125] In order to further improve the path tracking stability and accuracy in the reinforcement learning process, the vehicle envelope constraint performance is considered in the reward function, and the reward function of the vehicle steering system path tracking is constructed with reference to the model predictive control objective. In the process of designing the target value function, it is hoped that the reinforcement learning agent can minimize the tracking error and heading error, and complete the steering system path tracking behavior more smoothly. Therefore, the target function R m Set to:
[0126] R m=-(k e e y 2 +k ψ e ψ 2 +k δ δ f 2 +k Δδ Δδ f 2 )
[0127] Among them, Δδ f is the control increment of the steering angle; k e , k ψ , k δ , k Δδ They are the penalty coefficients for lateral deviation, heading deviation, steering angle control amount and control increment respectively.
[0128] The control state constraint is imposed in the reward function to actively constrain the steering angle and its increment within a reasonable range. The constraints on the steering angle control amount and the control increment are set as:
[0129]
[0130] Among them, R p1 With R p2 are the constraints for controlling the steering angle and its increment respectively; k δ′ With k Δδ′ are the penalty coefficients for controlling the steering angle and its incremental constraint respectively; δ fmax With Δδ fmax They are the maximum limit values for controlling the steering angle and its increment respectively.
[0131] At the same time, the envelope performance of the vehicle's motion state is further considered in the active constraint function, and a penalty is set for the action behavior that exceeds the vehicle's yaw and lateral motion constraints. The constraints are:
[0132]
[0133] Among them, R p3 With R p4 are the constraints of vehicle yaw and lateral motion respectively; k β With k r are the penalty coefficients for the sideslip angle and yaw rate constraints respectively; α s is the saturated tire slip angle; b1 and b2 are the constraint boundaries of the yaw angular velocity.
[0134] Combining the above objective function and constraint function, we get the reward function R of the vehicle steering system: p for:
[0135] Rp =R m +R p1 +R p2 +R p3 +R p4
[0136] Since the vehicle suspension system has a larger range of motion exploration, the difficulty of strategy exploration in reinforcement learning is increased. In order to better expand the reinforcement learning method to the complex suspension system environment, a fast-response PID control strategy is used to compensate for the uncertainty in the suspension exploration process. Two PID controllers are used to control the vertical and roll motions of the suspension system respectively. The control inputs are selected as the vehicle body acceleration and roll acceleration error terms with the goals of smoothness and roll stability. The corresponding PID control design is:
[0137]
[0138] Among them, f z (t) are the control forces of vertical and rolling motion at time t respectively; e z (t) are the error inputs of vertical and rolling motion at time t respectively; K zp , K zi With K zd is the PID control coefficient of vertical motion; and is the PID control coefficient of the roll motion.
[0139] In order to distribute the vertical and rolling motion control forces of the suspension system to the controllable actuators of the left and right suspensions, the control force relationship of the decoupled distribution can be expressed as:
[0140] f z (t) = f d1 (t)+f d2 (t),
[0141] Among them, f d1 (t) and f d2 (t) are the forces acting on the left and right suspensions at time t respectively.
[0142] The compensation force of the left and right suspensions can be calculated as:
[0143]
[0144] The required suspension system exploration command is calculated by the sum of the suspension compensation force and the learning feedback control force. The control action of the suspension system is expressed as:
[0145] a s1 =πs1 +f d1 +N s1 ,a s2 =π s2 +f d2 +N s2
[0146] Among them, a s1 with a s2 are the control actions of the left and right suspensions respectively; π s1 With π s2 are the exploration strategies of the left and right suspensions respectively; N s1 With N s2 are the random exploration noises of the left and right suspensions, respectively.
[0147] The control objectives of the suspension system mainly consider the ride comfort, roll stability, and suspension maneuverability during the autonomous driving path tracking process, while reducing the system control energy. Since the dimensions of the body acceleration, suspension dynamic travel, body roll acceleration, and suspension force of the suspension system are quite different, the various indicators in the reward process are normalized according to the dynamic response process of the passive suspension system to obtain the target reward function R n for:
[0148]
[0149] Among them, k a , k1, k2, k f1 , k f2 They are the penalty coefficients of vehicle speed acceleration, roll acceleration, left and right suspension travel, and suspension force; z 1nor 、z 2nor 、F 1nor 、F 2nor They are the normalized terms of vehicle speed acceleration, roll acceleration, left and right suspension travel and suspension force.
[0150] The constraints on the suspension travel are extended to the left and right suspensions, and both must be less than the maximum allowable limit travel to reduce the probability of hitting the suspension limit block, thereby ensuring good vehicle handling stability. That is, the constraints on the suspension travel need to meet the following requirements:
[0151]
[0152] Among them, R s1 With R s2 are the constraints of the left and right suspension travel respectively; k d1 With k d2 are the penalty coefficients of the left and right suspension travel constraints respectively; (z s1 -z w1) max ,(z s2 -z w2 ) max They are the maximum limit values of the left and right suspension dynamic travel respectively.
[0153] During the learning and exploration process, the suspension force does not exceed the maximum allowable value of the actuator. The control constraints of the left and right suspensions are set as:
[0154] |F s1 |≤F s1max ,|F s2 ≤F s2max
[0155] Among them, F s1max With F s2max They are the maximum limit values of the left and right suspension control quantities respectively.
[0156] Finally, the reward function R of the vehicle suspension system is obtained s for:
[0157] R s =R n +R s1 +R s2
[0158] 3) Use deep neural networks to build intelligent agents for the steering system and suspension system, and realize the end-to-end learning process from the motion state of the steering system and suspension system to the exploration strategy.
[0159] Specifically, the intelligent agent is the learner and decision maker in reinforcement learning. Reinforcement learning can use various neural network methods to build intelligent agents with different structures. The present invention uses a deep deterministic policy gradient algorithm to build an intelligent agent for the steering system and suspension system. The intelligent agent includes a behavior network and an evaluation network. The behavior network structure of the steering system is as follows: Figure 4a As shown in Figure 1, it consists of a state input layer, three hidden layers and a control output layer. The state of the state input layer is 6-dimensional, including lateral velocity, yaw rate, lateral deviation, heading deviation, road curvature and the control action at the last moment. Each hidden layer consists of 100 neurons, and the control output layer is the exploration strategy of the steering system. The network structure of the evaluation network of the steering system is shown in Figure 1. Figure 4b As shown in the figure, it mainly includes a state input layer, a control input layer, three hidden layers and an output layer. Among them, the state input layer obtains the 6-dimensional state of the intelligent agent from the behavior network of the steering system, the control input layer is a 1-dimensional control action, each hidden layer consists of 100 neurons, and the output layer is a 1-dimensional action value function that evaluates the path tracking control behavior of the steering system. The three hidden layers are connected in series in sequence, the state input layer is connected to the first hidden layer, and the control input layer skips the first hidden layer and is connected to the second hidden layer for transmission.
[0160] The behavior network structure of the suspension system is as follows Figure 5a As shown in the figure, it mainly includes a state input layer, three hidden layers and a control output layer. Among them, the state input layer is 6-dimensional, including vehicle body displacement, vehicle roll angle, left and right suspension dynamic travel and left and right suspension control actions at the previous moment. Each hidden layer is set with 100 neurons, and the control output layer is the exploration strategy of the left and right suspensions. The control output layer scales the suspension force within the range of [-1,1] through the activation function. The evaluation network of the suspension system includes a state input layer, a control input layer, three hidden layers and an output layer. The structure is as follows: Figure 5b As shown in the figure. The state input layer of the evaluation network of the suspension system is consistent with the behavior network of the suspension system, and the control input layer is the 2D left and right suspension control action. The three hidden layers are connected in series, the control input layer is directly connected to the second hidden layer, and the state input layer is connected to the first hidden layer. Each hidden layer consists of 100 neurons, and the output layer is the action value function that evaluates the behavior of the suspension system.
[0161] The behavior network parameters are updated based on the policy gradient algorithm, which can select appropriate actions in the continuous action space. Its input is the current state s i and the state s at the next moment i+1 , the output is the current strategy π(s i |θ π ) and the next moment’s strategy π′(s i+1 |θ π′ ), the direction of the behavior network iteration is to maximize the expected value of the reward J(π) within a certain period, and the strategy gradient calculation process of the behavior network is as follows:
[0162]
[0163] Among them, N is the total number of data samples; s is the state; a is the action; θ Q represents the parameters of the online evaluation network; θ π Parameters representing the online behavior network; is the policy gradient of the behavior network; a Q() is the gradient of the action value function; is the gradient of the policy function.
[0164] The evaluation network is updated using the Q learning method, and its input is the current state s i 、Strategy π(s i |θ π ) and the state s at the next moment i+1 , strategy π′(s i+1 |θ π′ ), the output is the Q value of the evaluation network, and the update process of the evaluation network is as follows:
[0165]
[0166] Where L() is the loss function of the evaluation network; Q′() is the action value function of the target network; Q() is the action value function of the online network; r i is the reward value at the current moment; γ is the discount factor; θ π′ is the parameter of the target behavior network; θ Q′ Evaluate the parameters of the network for the target.
[0167] The parameters of the target behavior and target evaluation networks are updated uniformly through the soft update process. The target network parameters are updated slowly with a small update rate, so that an easy-to-converge and stable online network optimization gradient can be obtained during the training process. The update method is as follows:
[0168] θ Q′ =τθ Q +(1-τ)θ Q′
[0169] θ π′ =τθ π +(1-τ)θ π′
[0170] Among them, τ is the parameter update rate.
[0171] When the intelligent agent is being trained, if the vehicle exceeds a certain error range during the path tracking process, continuing to run will add too much invalid data. Therefore, when the vehicle tracking error exceeds 1m during the learning process, the intelligent agent training process will be terminated to start a new training cycle, effectively improving the convergence speed of the training learning process.
[0172] 4) A reinforcement learning environment is constructed by combining the dynamic model and the reward function module, and the coordinated control of the steering system and the suspension system is achieved through the reinforcement learning method of the intelligent agent of the steering system and the suspension system.
[0173] Specifically: See Figure 6Reinforcement learning learns control strategies in the process of interaction with unknown dynamic environments. It adopts the method of obtaining samples while learning and iterates repeatedly until the model converges, and judges the control performance of the steering system and suspension system. The vehicle steering system and suspension system dynamics model, tracking error model, road model and steering and suspension system reward modules together constitute the reinforcement learning environment. In the path tracking control process of the autonomous driving vehicle, the lateral and vertical motion processes of the vehicle in the path tracking process are controlled by outputting the control actions of the steering and suspension systems of the controlled vehicle. The two agents respectively observe the motion states of the steering system and suspension system of the reinforcement learning environment, including the lateral velocity, yaw angular velocity, lateral deviation, heading deviation, road curvature of the steering system, and the body displacement, vehicle roll angle, and suspension dynamic travel of the suspension system. According to the strategy in the agent, the control actions of the steering system are executed separately. p Control action of the suspension system s1 with a s2 After making a control decision, the motion state of the autonomous vehicle changes, and the control state enters the next control state. At the same time, the reward function module of the steering system is calculated by the lateral motion state s p (v y ,r,e y ,e ψ ,κ) calculates the reward value obtained after executing the steering control action strategy. The reward function module of the suspension system is based on the vertical motion state The reward value obtained after executing the suspension control action strategy is calculated, and the agent further updates the applied strategy based on the feedback reward value. This cycle continues until the agent obtains enough cumulative rewards to achieve coordinated control of the steering system and suspension system.
[0174] like Figure 7 , Figure 8 As shown, by comparing with the model predictive control algorithm and the single-agent learning algorithm, the method of the present invention effectively tracks the reference path and obtains the minimum tracking error, and can obtain better steering path tracking performance. At the same time, the method of the present invention effectively reduces the roll acceleration of the autonomous driving vehicle, can obtain a smaller roll stability area, and the improved suspension comfort performance is significantly better than the other two control strategies, while the improvement effect of the model predictive control algorithm only exceeds the worst single-agent control. These results show that the method of the present invention can effectively improve the ride comfort and stability during the tracking steering process while ensuring accurate and efficient tracking performance, and complete the coordinated control of the steering system and the suspension system.
[0175] The above is only a specific implementation of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantial changes to the present invention using this concept shall be deemed as an infringement of the protection scope of the present invention.
Claims
1. A method for coordinated control of steering and suspension of an autonomous driving vehicle based on reinforcement learning, characterized in that: The steps include: 1) Establish the dynamic model of the vehicle steering system and suspension system, and obtain the lateral and vertical vehicle motion state by controlling the steering angle and suspension force; 2) Construct the control actions and reward functions of the vehicle steering system and suspension system to form a reinforcement learning exploration strategy for the coordinated control of vehicle steering and suspension; 3) Use deep neural networks to build intelligent agents for steering and suspension systems; 4) A reinforcement learning environment is constructed by combining the dynamic model and the reward function module, and the coordinated control of the steering system and the suspension system is achieved through the reinforcement learning method of the intelligent agent of the steering system and the suspension system.
2. The method for coordinated control of steering and suspension of an autonomous driving vehicle based on reinforcement learning according to claim 1, characterized in that: In the step 1), the steering system dynamics model is constructed by combining the roll and yaw motion states in the two-degree-of-freedom steering system: Among them, m and m s are the total vehicle mass and the sprung mass of the vehicle respectively; h is the distance from the center of mass of the vehicle to the roll center; is the road inclination; g is the acceleration due to gravity; l f With l r are the distances from the center of mass to the front and rear axles, respectively; I z is the yaw moment of inertia; r is the yaw angular velocity; is the yaw angular acceleration; v x With v y are the longitudinal and lateral velocities, respectively; is the lateral acceleration; is the roll acceleration; F yf With F yr are the lateral forces of the front and rear axle tires respectively; κ is the road curvature; e ψ is the heading deviation; and are the rates of change of lateral deviation and heading deviation respectively; The forces of the front and rear suspension systems are added together and concentrated on an equivalent roll axis, and the dynamic model of the suspension system is constructed as follows: Among them, z s is the displacement of the vehicle body; is the vehicle body acceleration; a y is the lateral acceleration; I x is the roll moment of inertia; is the vehicle roll angle; z s1 With z s2 are the displacement of the left and right sides of the vehicle body respectively; z w1 With z w2 are the left and right wheel displacements respectively; and are the left and right wheel accelerations respectively; k s1 With k s2 are the left and right suspension stiffness coefficients respectively; m w1 With m w2 are the unsprung masses on the left and right sides respectively; z r1 With z r2 are the road excitation displacements on the left and right sides respectively; k t1 With k t2 are the stiffness coefficients of the left and right tires respectively; F s1 With F s2 The left and right suspensions are powered respectively; T w is the wheel track.
3. The method for coordinated control of steering and suspension of an autonomous driving vehicle based on reinforcement learning according to claim 1, characterized in that: In the step 2), feedforward steering is applied in the steering control exploration process of reinforcement learning, and the relationship between the feedforward steering angle of the vehicle and the side slip angle of the front and rear tires is obtained through the front and rear wheel side slip motion relationship: d f ′=Lκ-α f ′+a r ′ Among them, δ f ′ is the feedforward steering angle; L is the distance from the front axle to the rear axle of the vehicle; α f ′ and α r ′ are the feedforward tire slip angles of the front and rear wheels respectively, and the feedforward tire force is expressed as: Among them, F yf ′ and F yr ′ are the feedforward tire forces of the front and rear wheels respectively; The required steering exploration command is calculated by the sum of feedforward and learning feedback control, and the control action of the steering system is expressed as: a p =π p +d f ′+N p Among them, a p is the control action of the steering system; π p is the exploration strategy for reinforcement learning; N p Random exploration noise for reinforcement learning; Objective function R m Set to: R m =-(k e e y 2 +k ψ e ψ 2 +k δ d f 2 +k Δδ Dd f 2 ) Among them, Δδ f is the control increment of the steering angle; k e , k ψ , k δ , k Δδ They are the penalty coefficients for lateral deviation, heading deviation, steering angle control amount and control increment respectively; The constraints of the steering angle control amount and control increment are set as: Among them, R p1 With R p2 are the constraints for controlling the steering angle and its increment respectively; k δ′ With k Δδ′ are the penalty coefficients for controlling the steering angle and its incremental constraint respectively; δ fmax With Δδ fmax They are the maximum limit values for controlling the steering angle and its increment respectively; At the same time, the envelope performance of the vehicle's motion state is considered in the active constraint function, and a penalty is set for actions that exceed the vehicle's yaw and lateral motion constraints. The constraints are: Among them, R p3 With R p4 are the constraints of vehicle yaw and lateral motion respectively; k β With k r are the penalty coefficients for the sideslip angle and yaw rate constraints respectively; α s is the saturated tire slip angle; b1 and b2 are the constraint boundaries of the yaw rate; Combining the above objective function and constraint function, we get the reward function R of the vehicle steering system: p for: R p =R m +R p1 +R p2 +R p3 +R p4 4. The method for coordinated control of steering and suspension of an autonomous driving vehicle based on reinforcement learning according to claim 1, characterized in that: In step 2), a fast-response PID control strategy is used to compensate for the uncertainty in the suspension exploration process, and two PID controllers are used to control the vertical and roll motions of the suspension system respectively. The corresponding PID control design is: Among them, f z (t) are the control forces of vertical and rolling motion at time t respectively; e z (t) are the error inputs of vertical and rolling motion at time t respectively; K zp , K zi With K zd is the PID control coefficient of vertical motion; and is the PID control coefficient of the roll motion; The vertical and roll motion control forces of the suspension system are distributed to the controllable actuators of the left and right suspensions. The decoupled distribution control force relationship is expressed as: Among them, f d1 (t) and f d2 (t) are the power of the left and right suspensions at time t respectively; The compensation force of the left and right suspensions is calculated as: The required suspension system exploration command is calculated by the sum of the suspension compensation force and the learning feedback control force. The control action of the suspension system is expressed as: a s1 =π s1 +f d1 +N s1 ,a s2 =π s2 +f d2 +N s2 Among them, a s1 with a s2 are the control actions of the left and right suspensions respectively; π s1 With π s2 are the exploration strategies of the left and right suspensions respectively; N s1 With N s2 are the random exploration noises of the left and right suspensions, respectively; According to the dynamic response process of the passive suspension system, the various indicators in the reward process are normalized to obtain the target reward function R n for: Among them, k a , k1, k2, k f1 , k f2 They are the penalty coefficients of vehicle speed acceleration, roll acceleration, left and right suspension travel, and suspension force; z 1nor 、z 2nor 、F 1nor 、F 2nor They are the normalized terms of vehicle speed acceleration, roll acceleration, left and right suspension travel, and suspension force; The constraint on the suspension travel is extended to the left and right suspensions, and both must be less than the maximum allowable limit travel. That is, the constraint on the suspension travel must meet the following requirements: Among them, R s1 With R s2 are the constraints of the left and right suspension travel respectively; k d1 With k d2 are the penalty coefficients of the left and right suspension travel constraints respectively; (z s1 -z w1 ) max ,(z s2 -z w2 ) max They are the maximum limit values of the left and right suspension travel respectively; The control constraints of the left and right suspensions are set as: |F s1 |≤F s1max ,|F s2 |≤F s2max Among them, F s1max With F s2max They are the maximum limit values of the left and right suspension control quantities respectively; Finally, the reward function R of the vehicle suspension system is obtained s for: R s =R n +R s1 +R s2 5. The method for coordinated control of steering and suspension of an autonomous driving vehicle based on reinforcement learning according to claim 1, characterized in that: In the step 3), a deep deterministic policy gradient algorithm is used to construct an intelligent agent for the steering system and the suspension system, and the intelligent agent includes a behavior network and an evaluation network; the behavior network of the steering system consists of a state input layer, three hidden layers and a control output layer, wherein the control output layer is the exploration strategy of the steering system; the evaluation network of the steering system includes a state input layer, a control input layer, three hidden layers and an output layer; the three hidden layers are connected in series in sequence, and the control input layer is connected to the second hidden layer for transmission. The behavioral network of the suspension system includes a state input layer, three hidden layers and a control output layer. The control output layer is the exploration strategy of the left and right suspensions. The control output layer scales the suspension force within the range of [-1,1] through an activation function. The evaluation network of the suspension system includes a state input layer, a control input layer, three hidden layers and an output layer. The three hidden layers are connected in series, and the control input layer is connected to the second hidden layer. During the learning process, it is set that when the vehicle tracking error exceeds 1m, the agent training process is terminated.
6. The method for coordinated steering and suspension control of an autonomous driving vehicle based on reinforcement learning according to claim 5, characterized in that: The state input layer of the steering system's behavioral network is 6-dimensional, including lateral velocity, yaw angular velocity, lateral deviation, heading deviation, road curvature and the control action at the previous moment. Each hidden layer of the steering system's behavioral network consists of 100 neurons.
7. The method for coordinated control of steering and suspension of an autonomous driving vehicle based on reinforcement learning according to claim 6, characterized in that: The state input layer of the evaluation network of the steering system obtains the 6-dimensional state of the intelligent agent from the behavior network of the steering system. The control input layer of the evaluation network of the steering system is a 1-dimensional control action. Each hidden layer of the evaluation network of the steering system is composed of 100 neurons. The output layer of the evaluation network of the steering system is a 1-dimensional action value function that evaluates the path tracking control behavior of the steering system.
8. The method for coordinated steering and suspension control of an autonomous driving vehicle based on reinforcement learning according to claim 5, characterized in that: The state input layer of the suspension system's behavior network is 6-dimensional, including vehicle body displacement, vehicle roll angle, left and right suspension dynamic travel, and left and right suspension control actions at the previous moment. Each hidden layer of the suspension system's behavior network is set with 100 neurons. The state input layer of the suspension system's evaluation network is consistent with the suspension system's behavior network, and the control input layer is 2-dimensional left and right suspension control actions. Each hidden layer of the evaluation network of the suspension system is composed of 100 neurons, and the output layer of the evaluation network of the suspension system is an action value function for evaluating the behavior of the suspension system.
9. The method for coordinated control of steering and suspension of an autonomous driving vehicle based on reinforcement learning according to claim 1, characterized in that: In the step 4), a reinforcement learning environment is formed by the vehicle steering system and suspension system dynamics model, tracking error model, road model and steering and suspension system reward modules; In the path tracking control process of the autonomous driving vehicle, the lateral and vertical motion processes of the vehicle in the path tracking process are controlled by outputting the control actions of the steering and suspension systems of the controlled vehicle; The agents of the steering system and suspension system observe the motion states of the steering system and suspension system in the reinforcement learning environment respectively; according to the strategies in the agents, the control actions of the steering system and the suspension system are respectively executed. After making control decisions, the motion state of the autonomous driving vehicle changes, and the control state enters the next control state; at the same time, the reward function module of the steering system p (v y ,r,e y ,e ψ ,κ) calculates the reward value obtained after executing the steering control action strategy. The reward function module of the suspension system is based on the vertical motion state The reward value obtained after executing the suspension control action strategy is calculated, and the agent further updates the applied strategy based on the feedback reward value. The above steps are repeated continuously until the agent obtains enough cumulative rewards to achieve coordinated control of the steering system and suspension system.
10. The method for coordinated control of steering and suspension of an autonomous driving vehicle based on reinforcement learning according to claim 9, characterized in that: The motion states of the steering system and suspension system include the lateral velocity, yaw angular velocity, lateral deviation, heading deviation, road curvature of the steering system, and the body displacement, vehicle roll angle, and suspension travel of the suspension system.
Citation Information
Patent Citations
Man-vehicle cooperative steering control method based on reinforcement learning corner weight distribution
CN115062539A
Vehicle active front wheel steering and active suspension system coordination control method
CN116767180A
Improved PID (Proportion Integration Differentiation) transverse tracking control method for self-driving automobile based on reinforcement learning
CN118306410A
Vehicle path tracking coordination control method and system based on danger level judgment and reinforcement learning
CN118457640A
Cited By
Suspension system control method for adjusting reward in combination with RND, medium and electronic equipment
CN120156237A
Suspension control method, storage medium and electronic equipment
CN120156240A
Suspension Control Method, Storage Medium and Electronic Device
CN120156240B
Control method considering safety of intelligent suspension, medium and equipment
CN120269980A
Intelligent suspension control method considering suspension system time lag, medium and equipment
CN120269982A