A method for cooperative control of steering and suspension of an autonomous vehicle based on reinforcement learning

By employing a reinforcement learning-based steering and suspension coordinated control method, combined with dynamic models and deep neural networks, the challenges of lateral and longitudinal comfort and stability in autonomous vehicles during path tracking are addressed, achieving coordinated control of vehicle safety, comfort, and stability.

CN119975527BActive Publication Date: 2025-10-21EAST CHINA JIAOTONG UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411971777.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-10-21
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

Existing autonomous vehicles struggle to simultaneously ensure both lateral and longitudinal comfort and stability during path tracking, especially in complex environments where a single path tracking task is insufficient to achieve vehicle safety, comfort, and intelligent control.

Method used

A reinforcement learning-based steering and suspension coordinated control method is adopted. By constructing dynamic models of the steering and suspension systems and combining deep neural networks and PID control strategies, end-to-end coordinated control is achieved. The intelligent agents of the steering and suspension systems are used to coordinate the lateral and vertical motion states.

Benefits of technology

It improves the path tracking accuracy and ride comfort of autonomous vehicles in complex environments, reduces learning costs and control complexity, and achieves coordinated control of vehicle safety, comfort and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119975527B_ABST
    Figure CN119975527B_ABST
Patent Text Reader

Abstract

The application relates to a kind of automatic driving vehicle steering and suspension collaborative control method based on reinforcement learning, specifically comprising the following steps: 1, the dynamic model of vehicle steering system and suspension system is established, the lateral and vertical vehicle motion state is obtained by controlling steering angle and suspension force;2, the control action and reward function of vehicle steering system and suspension system are constructed, and the reinforcement learning exploration strategy of vehicle steering and suspension collaborative control is formed;3, the agent of steering system and suspension system is constructed;4, the reinforcement learning environment is constructed, and the collaborative control of steering system and suspension system is realized by the reinforcement learning method of multiple agents.The application comprehensively considers different plane environments and different learning tasks in vertical and horizontal directions, reduces the dimension of data input, and considers the optimal motion control strategy in different working conditions, improves the lateral roll stability and vertical smooth comfort by collaborative steering system and suspension system control under the premise of guaranteeing accurate path tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous vehicle motion control, and in particular to a method for coordinated steering and suspension control of autonomous vehicles based on reinforcement learning. Background Art

[0002] Autonomous driving technology for smart cars is a key approach to alleviating traffic congestion, reducing traffic accidents, and improving road and vehicle utilization. High-level autonomous driving requires a high degree of autonomous coordination between steering and other subsystems. Designing a path-following solution that enhances comfort and handling stability is a prerequisite for highly intelligent autonomous vehicles. Current lateral control methods primarily focus on improving the stability of autonomous vehicles, while research on ride comfort remains limited. As autonomous vehicles continue to develop towards higher intelligence, consideration of comfort and handling stability will further enhance their safety and user experience.

[0003] In path-following control for autonomous vehicles, underactuated processes only track the vehicle's path by controlling the steering angle, making it difficult to simultaneously track a reference path and maintain comfort and stability during lateral and longitudinal maneuvers. To achieve safe, comfortable, and intelligent autonomous driving control in complex driving environments, most research has employed methods such as yaw torque and active steering control to enhance the vehicle's lateral and longitudinal comfort and stability within a single path-following task. However, there is a certain constraint between tracking capability and comfort and stability in planar models. Summary of the Invention

[0004] The main purpose of the present invention is to overcome the above-mentioned defects of vehicle multi-system motion control and propose a reinforcement learning-based coordinated control method for steering and suspension of autonomous driving vehicles.

[0005] The present invention adopts the following technical solutions:

[0006] A method for coordinated steering and suspension control of an autonomous vehicle based on reinforcement learning includes the following steps:

[0007] 1) Establish dynamic models of the vehicle's steering and suspension systems, and obtain the lateral and vertical vehicle motion states by controlling the steering angle and suspension forces;

[0008] 2) Constructing control actions and reward functions for the vehicle steering and suspension systems, and forming a reinforcement learning exploration strategy for coordinated control of vehicle steering and suspension;

[0009] 3) Using deep neural networks to build intelligent agents for the steering and suspension systems, achieving an end-to-end learning process from steering and suspension system motion states to exploration strategies;

[0010] 4) A reinforcement learning environment is constructed by combining the dynamic model and the reward function module, and the coordinated control of the steering system and the suspension system is achieved through the reinforcement learning method of the intelligent agents of the steering system and the suspension system.

[0011] In step 1), the steering system dynamics model is constructed by combining the roll and yaw motion states in the two-degree-of-freedom steering system model:

[0012]

[0013] Among them, m and m s are the total vehicle mass and the sprung mass of the vehicle, respectively; h is the distance from the center of mass of the vehicle to the roll center; is the road inclination; g is the acceleration due to gravity; l f With l r are the distances from the center of mass to the front and rear axles respectively; I z is the yaw moment of inertia; r is the yaw angular velocity; is the yaw angular acceleration; v x With v y are the longitudinal and lateral velocities, respectively; is the lateral acceleration; is the roll acceleration; F yf With F yr are the lateral forces of the front and rear axle tires respectively; κ is the road curvature; e ψ is the heading deviation; and are the rates of change of lateral deviation and heading deviation, respectively.

[0014] The magic formula model is used to describe the relationship between tire lateral force and sideslip angle. The tire model is expressed as:

[0015] F y =Dsin[Carctan(Bx-E(Bx-arctan(Bx)))]+S v

[0016] x=α+S h

[0017] D=A1F z 2 +A2F z

[0018] B=A3sin(A4arctan(A5F z )) / CD

[0019] E=A6F z 2 +A7F z +A8

[0020] Among them, F y is the tire lateral force; B is the curve stiffness factor; C is the curve shape factor; D is the curve peak factor; E is the curve curvature factor; α is the tire side slip angle; S h With S v are the horizontal and vertical offsets of the curve respectively; F z is the vertical load of the tire; A1~A8 are the identification parameters of the tire model.

[0021] Furthermore, the front and rear tire slip angles can be expressed as:

[0022]

[0023] Among them, α f With α r are the front and rear tire slip angles respectively; δ f is the front wheel steering angle.

[0024] At this point, a vehicle steering system dynamics model considering suspension roll has been established, which can obtain the lateral vehicle motion state by controlling the steering angle.

[0025] The forces of the front and rear suspension systems are added together and concentrated on an equivalent roll axis. The dynamic model of the suspension system is constructed as follows:

[0026]

[0027] Among them, z s is the vehicle body displacement; is the vehicle body acceleration; a y is the lateral acceleration; I x is the roll moment of inertia; is the vehicle roll angle; z s1 With z s2 are the displacements of the left and right sides of the vehicle body respectively; z w1 With z w2 are the left and right wheel displacements respectively; and are the left and right wheel accelerations respectively; k s1 With k s2 are the left and right suspension stiffness coefficients respectively; m w1 With m w2 are the unsprung masses on the left and right sides respectively; z r1 With z r2 are the road excitation displacements on the left and right sides respectively; k t1 With k t2 are the stiffness coefficients of the left and right tires respectively; F s1 With F s2 Powers the left and right suspensions respectively; T w is the wheel track.

[0028] The road excitation displacement generated by the filtered white noise method is:

[0029]

[0030] Among them, z r (t) is the road surface excitation displacement at time t; is the rate of change of the road surface excitation displacement at time t; η and σ are constants related to the road surface grade; G q (n0) is the road surface power spectrum density at the reference spatial frequency n0, that is, the road surface roughness coefficient, which can be obtained through the road surface roughness classification standard; w(t) is a unit white noise with a mean of 0 and a variance of 1.

[0031] At this point, a vehicle suspension system dynamics model associated with the steering system dynamics model has been established, which can obtain the vertical vehicle motion state by controlling the suspension force.

[0032] In step 2), the present invention combines a priori feedforward steering to improve the efficiency of reinforcement learning exploration, and applies feedforward steering during the reinforcement learning steering control exploration process. Based on the front and rear wheel side slip motion relationship, the relationship between the vehicle's feedforward steering angle and the front and rear tire side slip angle can be described as:

[0033] δ f ′=Lκ-α f ′+α r '

[0034] Among them, δ f ′ is the feedforward steering angle; L is the distance from the front axle to the rear axle of the vehicle; α f ′ and α r ′ are the feedforward tire slip angles of the front and rear wheels, respectively, which can be obtained by the feedforward tire force, which is expressed as:

[0035]

[0036] Among them, F yf ′ and F yr ′ are the feedforward tire forces of the front and rear wheels respectively.

[0037] The steering exploration command required for reinforcement learning is calculated by the sum of feedforward and learning feedback control. The control action of the steering system is expressed as:

[0038] a p =π p +δ f ′+N p

[0039] Among them, a p is the control action of the steering system; π pis the exploration strategy for reinforcement learning; N p Random exploration noise for reinforcement learning.

[0040] In order to further improve the path tracking stability and accuracy in the reinforcement learning process, the vehicle's envelope constraint performance is considered in the reward function. m Set as:

[0041] R m =-(k e e y 2 +k ψ e ψ 2 +k δ δ f 2 +k Δδ Δδ f 2 )

[0042] Among them, Δδ f is the control increment of the steering angle; k e 、k ψ 、k δ 、k Δδ are the penalty coefficients for lateral deviation, heading deviation, steering angle control amount and control increment respectively.

[0043] The constraints of the steering angle control amount and control increment are set as:

[0044]

[0045] Among them, R p1 With R p2 are the constraints for controlling the steering angle and its increment respectively; k δ′ With k Δδ′ are the penalty coefficients for controlling the steering angle and its incremental constraint respectively; δ fmax and Δδ fmax They are the maximum limit values ​​for controlling the steering angle and its increment respectively.

[0046] Penalties are set for actions that exceed the vehicle's yaw and lateral motion constraints. The constraints are:

[0047]

[0048] Among them, R p3 With R p4 are the constraints for vehicle yaw and lateral motion respectively; k β With k r are the penalty coefficients for sideslip angle and yaw rate constraints respectively; α s is the saturated tire slip angle; b1 and b2 are the constraint boundaries of the yaw rate.

[0049] Combining the above objective function and constraint function, the reward function R of the vehicle steering system is obtained p for:

[0050] R p =R m +R p1 +R p2 +R p3 +R p4

[0051] To better extend the reinforcement learning method to complex suspension system environments, a fast-response PID control strategy is adopted to compensate for the uncertainty in the suspension exploration process. Two PID controllers are used to control the vertical and roll motions of the suspension system respectively. The control inputs are the vehicle acceleration and roll acceleration error terms, which are targeted at ride comfort and roll stability, respectively. The corresponding PID control design is:

[0052]

[0053] Among them, f z (t) with are the control forces for vertical and rolling motion at time t; e z (t) with are the error inputs of vertical and roll motion at time t respectively; K zp , K zi With K zd is the PID control coefficient of vertical motion; and is the PID control coefficient of the roll motion.

[0054] The control force relationship of the decoupled distribution can be expressed as:

[0055] f z (t) = f d1 (t)+f d2 (t),

[0056] Among them, f d1 (t) and f d2 (t) are the power of the left and right suspensions at time t respectively.

[0057] The compensation force of the left and right suspension can be calculated as:

[0058]

[0059] The required suspension system exploration command is calculated by the sum of the suspension compensation force and the learning feedback control force. The control action of the suspension system is expressed as:

[0060] a s1 =π s1 +f d1 +N s1 ,a s2 =π s2 +f d2 +N s2

[0061] Among them, a s1 with a s2 are the control actions of the left and right suspension respectively; π s1 and π s2 are the exploration strategies of the left and right suspensions respectively; N s1 With N s2 are the random exploration noises of the left and right suspensions, respectively.

[0062] Since the dimensions of the body acceleration, suspension travel, body roll acceleration and suspension force of the suspension system are quite different, the various indicators in the reward process are normalized according to the dynamic response process of the passive suspension system to obtain the target reward function R n for:

[0063]

[0064] Among them, k a 、 k1, k2, k f1 、k f2 They are the penalty coefficients for vehicle speed acceleration, roll acceleration, left and right suspension travel, and suspension force; z 1nor 、z 2nor 、F 1nor 、F 2nor They are the normalized terms of vehicle speed acceleration, roll acceleration, left and right suspension dynamic travel and suspension force respectively.

[0065] The constraints on suspension travel are extended to the left and right suspensions. The constraints on suspension travel must meet the following requirements:

[0066]

[0067] Among them, R s1 With R s2 are the constraints of the left and right suspension travel respectively; k d1 With k d2 are the penalty coefficients for the left and right suspension travel constraints respectively; (z s1 -z w1 ) max 、(z s2 -z w2 ) maxThey are the maximum limit values ​​of the left and right suspension dynamic travel respectively.

[0068] The control constraints of the left and right suspensions are set as:

[0069] |F s1 |≤F s1max ,|F s2 ≤F s2max

[0070] Among them, F s1max With F s2max They are the maximum limit values ​​of the left and right suspension control amounts respectively.

[0071] Finally, the reward function R of the vehicle suspension system is obtained s for:

[0072] R s =R n +R s1 +R s2

[0073] In step 3), the present invention uses a deep deterministic policy gradient algorithm to construct an intelligent agent for the steering system and suspension system, which includes a behavioral network and an evaluation network. The behavioral network of the steering system consists of a state input layer, three hidden layers and a control output layer. Among them, the state of the state input layer is 6-dimensional, including lateral velocity, yaw angular velocity, lateral deviation, heading deviation, road curvature and control action at the previous moment, each hidden layer is composed of 100 neurons, and the control output layer is the exploration strategy of the steering system. The evaluation network of the steering system mainly includes a state input layer, a control input layer, three hidden layers and an output layer. Among them, the state input layer obtains the 6-dimensional state of the intelligent agent from the behavioral network of the steering system, the control input layer is a 1-dimensional control action, each hidden layer is composed of 100 neurons, and the output layer is a 1-dimensional action value function for evaluating the path tracking control behavior of the steering system. The three hidden layers are connected in series in sequence, the state input layer is connected to the first hidden layer, and the control input layer skips the first hidden layer and is connected and transmitted with the second hidden layer.

[0074] The suspension system's behavioral network primarily consists of a state input layer, three hidden layers, and a control output layer. The state input layer is 6-dimensional and contains vehicle body displacement, roll angle, left and right suspension travel, and the previous left and right suspension control actions. Each hidden layer has 100 neurons. The control output layer represents the exploration strategy for the left and right suspensions. The control output layer uses an activation function to scale the suspension forces within the range [-1, 1]. The suspension system's evaluation network consists of a state input layer, a control input layer, three hidden layers, and an output layer. The state input layer of the suspension system's evaluation network is identical to the suspension system's behavioral network, while the control input layer represents the 2D left and right suspension control actions. The three hidden layers are connected in series, with the control input layer directly connected to the second hidden layer and the state input layer connected to the first hidden layer. Each hidden layer consists of 100 neurons, and the output layer is the action-value function that evaluates the suspension system's behavior.

[0075] When training an agent, if the vehicle's path tracking error exceeds a certain range, continued operation will result in excessive invalid data. Therefore, when the vehicle tracking error exceeds 1m during the learning process, the agent training process will be terminated and a new training cycle will be started, effectively improving the convergence speed of the training process.

[0076] In step 4), the vehicle steering system and suspension system dynamics model, tracking error model, road model and steering and suspension system reward modules together constitute the reinforcement learning environment. During the path tracking control process of the autonomous vehicle, the lateral and vertical motion processes of the vehicle during the path tracking process are controlled by outputting the control actions of the steering and suspension systems of the controlled vehicle. The two intelligent agents respectively observe the motion states of the steering system and suspension system of the reinforcement learning environment, including the lateral velocity, yaw angular velocity, lateral deviation, heading deviation, and road curvature of the steering system, and the body displacement, vehicle roll angle, and suspension dynamic travel of the suspension system. The control actions of the steering system are respectively executed according to the strategy in the intelligent agent. p Control action of the suspension system s1 with a s2 After making a control decision, the motion state of the autonomous vehicle changes, and the control state enters the next control state. At the same time, the reward function module of the steering system is transformed into the lateral motion state s p (v y ,r,e y ,e ψ ,κ) calculates the reward value obtained after executing the steering control action strategy. The reward function module of the suspension system is based on the vertical motion state The agent calculates the reward value obtained after executing the suspension control action strategy. The agent further updates the strategy based on the feedback reward value. This cycle continues until the agent accumulates enough rewards to achieve coordinated control of the steering and suspension systems.

[0077] From the above description of the present invention, it can be seen that compared with the prior art, the present invention has the following beneficial effects:

[0078] (1) By constructing lateral and vertical dynamic models of the steering and suspension systems, a multi-agent reinforcement learning method involving a steering agent and a suspension agent was designed to reduce the data input dimension. Lateral and vertical motion constraints were also derived and applied, effectively achieving efficient, safe, and comfortable multi-objective coordinated control of the two systems.

[0079] (2) To improve the exploration efficiency of the reinforcement learning method, the method of the present invention adds feedforward steering information to the steering system, reducing the workload of reinforcement learning feedback control and reducing the ineffective steering search process in the reinforcement learning tracking process. The PID compensation control process is added to the reinforcement learning of complex suspension systems, making it easier to search for the strategic actions of the coupled system, ensuring the efficiency of strategic learning and reducing the learning cost.

[0080] (3) The method of the present invention comprehensively considers different vertical and lateral planar environments and different learning tasks, and takes into account the optimal motion control strategy for different working conditions. Under the premise of ensuring accurate path tracking, the lateral roll stability and vertical smoothness and comfort are improved through the coordinated control of the steering system and the suspension system. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Figure 1 It is a flow chart of the method of the present invention;

[0082] Figure 2 is a schematic diagram of a dynamic model of a steering system of the present invention;

[0083] Figure 3 is a schematic diagram of the suspension system dynamics model of the present invention;

[0084] Figure 4a is a behavioral network diagram of the steering system of the present invention;

[0085] Figure 4b is a schematic diagram of an evaluation network of the steering system of the present invention;

[0086] Figure 5a is a behavioral network diagram of the suspension system of the present invention;

[0087] Figure 5b is a schematic diagram of an evaluation network of the suspension system of the present invention;

[0088] Figure 6is a schematic diagram of a reinforcement learning method for coordinated control of steering and suspension systems of the present invention;

[0089] Figure 7 2. It is a comparative schematic diagram of the path tracking of the present invention;

[0090] Figure 8 2 is a comparative diagram of the suspension control of the present invention.

[0091] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. DETAILED DESCRIPTION

[0092] The present invention is further described below through specific embodiments.

[0093] See also Figure 1 , the method of the present invention comprises the following steps:

[0094] 1) Establish a dynamic model of the vehicle steering system and suspension system, and obtain the lateral and vertical vehicle motion states by controlling the steering angle and suspension force.

[0095] Specifically, in order to consider the roll motion characteristics of the suspension in the steering system dynamics model, the vehicle's roll degree of freedom is increased. Further combined with the yaw motion state in the two-degree-of-freedom steering system model, such as Figure 2 As shown in Figure 2, the steering system dynamics model is constructed under the small angle assumption as follows:

[0096]

[0097] Among them, m and m s are the total vehicle mass and the sprung mass of the vehicle, respectively; h is the distance from the center of mass of the vehicle to the roll center; is the road inclination; g is the acceleration due to gravity; l f With l r are the distances from the center of mass to the front and rear axles respectively; I z is the yaw moment of inertia; r is the yaw angular velocity; is the yaw angular acceleration; v x With v y are the longitudinal and lateral velocities, respectively; is the lateral acceleration; is the roll acceleration; F yf With F yr are the lateral forces of the front and rear axle tires respectively; κ is the road curvature; e ψ is the heading deviation; and are the rates of change of lateral deviation and heading deviation, respectively.

[0098] The complex nonlinear mechanical behavior in the vehicle steering system dynamics model mainly comes from the interaction between the tire and the ground. This paper uses the magic formula model to describe the relationship between the tire lateral force and the sideslip angle. The tire model is expressed as:

[0099] F y =Dsin[Carctan(Bx-E(Bx-arctan(Bx)))]+S v

[0100] x=α+S h

[0101] D=A1F z 2 +A2F z

[0102] B=A3sin(A4arctan(A5F z )) / CD

[0103] E=A6F z 2 +A7F z +A8

[0104] Among them, F y is the tire lateral force; B is the curve stiffness factor; C is the curve shape factor; D is the curve peak factor; E is the curve curvature factor; α is the tire side slip angle; S h With S v are the horizontal and vertical offsets of the curve respectively; F z is the vertical load of the tire; A1~A8 are the identification parameters of the tire model.

[0105] Furthermore, the front and rear tire slip angles can be expressed as:

[0106]

[0107] Among them, α f With α r are the front and rear tire slip angles respectively; δ f is the front wheel steering angle.

[0108] At this point, a vehicle steering system dynamics model considering suspension roll has been established, which can obtain the lateral vehicle motion state by controlling the steering angle.

[0109] In order to establish a suspension system model associated with the roll characteristics in the steering system dynamics model, the forces of the front and rear suspension systems are added together and concentrated on an equivalent roll axis, such as Figure 3 As shown in Figure 2, the constructed suspension system dynamic model is:

[0110]

[0111] Among them, z s is the vehicle body displacement; is the vehicle body acceleration; a y is the lateral acceleration; I x is the roll moment of inertia; is the vehicle roll angle; z s1 With z s2 are the displacements of the left and right sides of the vehicle body respectively; z w1 With z w2 are the left and right wheel displacements respectively; and are the left and right wheel accelerations respectively; k s1 With k s2 are the left and right suspension stiffness coefficients respectively; m w1 With m w2 are the unsprung masses on the left and right sides respectively; z r1 With z r2 are the road excitation displacements on the left and right sides respectively; k t1 With k t2 are the stiffness coefficients of the left and right tires respectively; F s1 With F s2 Powers the left and right suspensions respectively; T w is the wheel track.

[0112] Road roughness is the main external interference factor affecting the dynamic characteristics of the suspension. Flat asphalt roads and washboard roads all exhibit significant discrete random road roughness. The road excitation displacement generated by the filtered white noise method is:

[0113]

[0114] Among them, z r (t) is the road surface excitation displacement at time t; is the rate of change of the road surface excitation displacement at time t; η and σ are constants related to the road surface grade; G q (n0) is the road surface power spectrum density at the reference spatial frequency n0, that is, the road surface roughness coefficient, which can be obtained through the road surface roughness classification standard; w(t) is a unit white noise with a mean of 0 and a variance of 1.

[0115] At this point, a vehicle suspension system dynamics model associated with the steering system dynamics model has been established, which can obtain the vertical vehicle motion state by controlling the suspension force.

[0116] 2) Construct the control actions and reward functions of the vehicle steering system and suspension system to form a reinforcement learning exploration strategy for the coordinated control of vehicle steering and suspension.

[0117] Specifically, during the autonomous driving path tracking process of the steering system, disordered learning exploration in a high-dimensional continuous state space is prone to falling into local optimization. This invention combines a priori feedforward steering to improve the efficiency of reinforcement learning exploration. By applying feedforward steering during the reinforcement learning steering control exploration process, it reduces the workload of reinforcement learning feedback control and reduces the number of ineffective steering searches. Based on the relationship between the front and rear wheel side slip motions, the relationship between the vehicle's feedforward steering angle and the front and rear tire side slip angles can be described as:

[0118] δ f ′=Lκ-α f ′+α r '

[0119] Among them, δ f ′ is the feedforward steering angle; L is the distance from the front axle to the rear axle of the vehicle; α f ′ and α r ′ are the feedforward tire slip angles of the front and rear wheels, respectively, which can be obtained by the feedforward tire force, which is expressed as:

[0120]

[0121] Among them, F yf ′ and F yr ′ are the feedforward tire forces of the front and rear wheels respectively.

[0122] Feedforward steering predicts steering commands in advance based on road curvature and vehicle speed information, while reinforcement learning feedback control adjusts steering input based on tracking error. The required steering exploration command is calculated by the sum of feedforward and learning feedback control. The control action of the steering system is expressed as:

[0123] a p =π p +δ f ′+N p

[0124] Among them, a p is the control action of the steering system; π p is the exploration strategy for reinforcement learning; N p Random exploration noise for reinforcement learning.

[0125] In order to further improve the path tracking stability and accuracy in the reinforcement learning process, the vehicle's envelope constraint performance is considered in the reward function, and the reward function of the vehicle steering system path tracking is constructed with reference to the model predictive control objective. In the process of designing the target value function, it is hoped that the reinforcement learning agent can minimize the tracking error and heading error as much as possible, and at the same time complete the steering system path tracking behavior more smoothly. Therefore, the objective function R m Set as:

[0126] R m=-(k e e y 2 +k ψ e ψ 2 +k δ δ f 2 +k Δδ Δδ f 2 )

[0127] Among them, Δδ f is the control increment of the steering angle; k e 、k ψ 、k δ 、k Δδ are the penalty coefficients for lateral deviation, heading deviation, steering angle control amount and control increment respectively.

[0128] The control state constraints are imposed in the reward function to actively constrain the steering angle and its increment within a reasonable range. The constraints on the steering angle control amount and the control increment are set as:

[0129]

[0130] Among them, R p1 With R p2 are the constraints for controlling the steering angle and its increment respectively; k δ′ With k Δδ′ are the penalty coefficients for controlling the steering angle and its incremental constraint respectively; δ fmax and Δδ fmax They are the maximum limit values ​​for controlling the steering angle and its increment respectively.

[0131] At the same time, the envelope performance of the vehicle's motion state is further considered in the active constraint function, and penalties are set for actions that exceed the vehicle's yaw and lateral motion constraints. The constraints are:

[0132]

[0133] Among them, R p3 With R p4 are the constraints for vehicle yaw and lateral motion respectively; k β With k r are the penalty coefficients for sideslip angle and yaw rate constraints respectively; α s is the saturated tire slip angle; b1 and b2 are the constraint boundaries of the yaw rate.

[0134] Combining the above objective function and constraint function, the reward function R of the vehicle steering system is obtained p for:

[0135] Rp =R m +R p1 +R p2 +R p3 +R p4

[0136] The larger range of motion exploration in the vehicle suspension system increases the difficulty of exploring strategies using reinforcement learning. To better extend reinforcement learning methods to complex suspension system environments, a fast-response PID control strategy is employed to compensate for the uncertainty in the suspension exploration process. Two PID controllers are used to control the vertical and roll motions of the suspension system, respectively. The control inputs are the vehicle acceleration and roll acceleration error terms, targeting ride comfort and roll stability, respectively. The corresponding PID control design is:

[0137]

[0138] Among them, f z (t) with are the control forces for vertical and rolling motion at time t; e z (t) with are the error inputs of vertical and roll motion at time t respectively; K zp , K zi With K zd is the PID control coefficient of vertical motion; and is the PID control coefficient of the roll motion.

[0139] In order to distribute the vertical and roll motion control forces of the suspension system to the controllable actuators of the left and right suspensions, the decoupled distribution control force relationship can be expressed as:

[0140] f z (t) = f d1 (t)+f d2 (t),

[0141] Among them, f d1 (t) and f d2 (t) are the power of the left and right suspensions at time t respectively.

[0142] The compensation force of the left and right suspension can be calculated as:

[0143]

[0144] The required suspension system exploration command is calculated by the sum of the suspension compensation force and the learning feedback control force. The control action of the suspension system is expressed as:

[0145] a s1 =πs1 +f d1 +N s1 ,a s2 =π s2 +f d2 +N s2

[0146] Among them, a s1 with a s2 are the control actions of the left and right suspension respectively; π s1 and π s2 are the exploration strategies of the left and right suspensions respectively; N s1 With N s2 are the random exploration noises of the left and right suspensions, respectively.

[0147] The control objectives of the suspension system mainly consider the ride comfort, roll stability, and suspension maneuverability during the autonomous driving path tracking process, while reducing the system control energy. Due to the large dimensional difference between the body acceleration, suspension dynamic travel, body roll acceleration, and suspension force of the suspension system, the various indicators in the reward process are normalized according to the dynamic response process of the passive suspension system to obtain the target reward function R n for:

[0148]

[0149] Among them, k a 、 k1, k2, k f1 、k f2 They are the penalty coefficients for vehicle speed acceleration, roll acceleration, left and right suspension travel, and suspension force; z 1nor 、z 2nor 、F 1nor 、F 2nor They are the normalized terms of vehicle speed acceleration, roll acceleration, left and right suspension dynamic travel and suspension force respectively.

[0150] The constraints on suspension travel are extended to both left and right suspensions. Both must be less than the maximum allowable travel limit to reduce the probability of hitting the suspension limit blocks, thereby ensuring good vehicle handling stability. In other words, the constraints on suspension travel must meet the following requirements:

[0151]

[0152] Among them, R s1 With R s2 are the constraints of the left and right suspension travel respectively; k d1 With k d2 are the penalty coefficients for the left and right suspension travel constraints respectively; (z s1 -z w1) max 、(z s2 -z w2 ) max They are the maximum limit values ​​of the left and right suspension dynamic travel respectively.

[0153] During the learning and exploration process, the suspension force does not exceed the maximum allowable value of the actuator. The control quantity constraints of the left and right suspensions are set as:

[0154] |F s1 |≤F s1max ,|F s2 ≤F s2max

[0155] Among them, F s1max With F s2max They are the maximum limit values ​​of the left and right suspension control amounts respectively.

[0156] Finally, the reward function R of the vehicle suspension system is obtained s for:

[0157] R s =R n +R s1 +R s2

[0158] 3) Use deep neural networks to build intelligent agents for the steering and suspension systems, and realize an end-to-end learning process from the motion state of the steering and suspension systems to the exploration strategy.

[0159] Specifically, the intelligent agent is the learner and decision maker in reinforcement learning. Reinforcement learning can use various neural network methods to build intelligent agents with different structures. The present invention uses a deep deterministic policy gradient algorithm to build an intelligent agent for the steering system and suspension system. The intelligent agent includes a behavior network and an evaluation network. The behavior network structure of the steering system is as follows: Figure 4a As shown in Figure 1, it consists of a state input layer, three hidden layers, and a control output layer. The state of the state input layer is 6-dimensional, including lateral velocity, yaw rate, lateral deviation, heading deviation, road curvature, and the control action at the previous moment. Each hidden layer consists of 100 neurons, and the control output layer is the exploration strategy of the steering system. The network structure of the evaluation network of the steering system is shown in Figure 1. Figure 4b As shown in the figure, the system primarily consists of a state input layer, a control input layer, three hidden layers, and an output layer. The state input layer obtains the agent's 6-dimensional state from the steering system's behavioral network. The control input layer provides 1-dimensional control actions. Each hidden layer consists of 100 neurons. The output layer is a 1-dimensional action-value function that evaluates the steering system's path-tracking control behavior. The three hidden layers are connected in series, with the state input layer connected to the first hidden layer and the control input layer skipping the first hidden layer and connecting to the second hidden layer.

[0160] The behavioral network structure of the suspension system is as follows Figure 5a As shown in the figure, it mainly consists of a state input layer, three hidden layers and a control output layer. Among them, the state input layer is 6-dimensional, including vehicle body displacement, vehicle roll angle, left and right suspension dynamic travel and left and right suspension control action at the previous moment. Each hidden layer is set with 100 neurons. The control output layer is the exploration strategy of the left and right suspensions. The control output layer scales the suspension force within the range of [-1,1] through the activation function. The evaluation network of the suspension system consists of a state input layer, a control input layer, three hidden layers and an output layer. The structure is as follows Figure 5b As shown in the figure, the state input layer of the suspension system's evaluation network is consistent with the suspension system's behavioral network. The control input layer represents the two-dimensional left and right suspension control actions. The three hidden layers are connected in series, with the control input layer directly connected to the second hidden layer and the state input layer connected to the first hidden layer. Each hidden layer consists of 100 neurons, and the output layer is the action-value function that evaluates the suspension system's behavior.

[0161] The behavior network parameters are updated based on the policy gradient algorithm, which can select appropriate actions in the continuous action space. Its input is the current state s i and the state s at the next moment i+1 , the output is the current strategy π(s i |θ π ) and the next moment’s strategy π′(s i+1 |θ π′ ), the direction of the behavior network iteration is to maximize the expected value of the reward J(π) within a certain period. The policy gradient calculation process of the behavior network is as follows:

[0162]

[0163] Where N is the total number of data samples; s is the state; a is the action; θ Q represents the parameters of the online evaluation network; θ π Parameters representing the online behavioral network; is the policy gradient of the behavior network; a Q() is the gradient of the action-value function; is the gradient of the policy function.

[0164] The update of the evaluation network adopts the Q learning method, and its input is the current state s i 、Strategyπ(s i |θ π ) and the state s at the next moment i+1 , strategy π′(s i+1 |θ π′ ), the output is the Q value of the evaluation network, and the update process of the evaluation network is as follows:

[0165]

[0166] Where L() is the loss function of the evaluation network; Q′() is the action value function of the target network; Q() is the action value function of the online network; r i is the reward value at the current moment; γ is the discount factor; θ π′ is the parameter of the target behavior network; θ Q′ Evaluate the parameters of the network for the target.

[0167] The parameters of the target behavior and target evaluation networks are updated uniformly through a soft update process. The target network parameters are updated slowly at a small update rate, so that an easy-to-converge and stable online network optimization gradient can be obtained during the training process. The update method is as follows:

[0168] θ Q′ =τθ Q +(1-τ)θ Q′

[0169] θ π′ =τθ π +(1-τ)θ π′

[0170] Where τ is the parameter update rate.

[0171] When training an agent, if the vehicle's path tracking error exceeds a certain range, continued operation will result in excessive invalid data. Therefore, when the vehicle tracking error exceeds 1m during the learning process, the agent training process will be terminated and a new training cycle will be started, effectively improving the convergence speed of the training process.

[0172] 4) A reinforcement learning environment is constructed by combining the dynamic model and the reward function module, and the coordinated control of the steering system and the suspension system is achieved through the reinforcement learning method of the intelligent agents of the steering system and the suspension system.

[0173] Specifically: See Figure 6Reinforcement learning learns control strategies in the process of interaction with unknown dynamic environments. It adopts the method of learning while obtaining samples to iterate and repeat until the model converges, and judges the control performance of the steering system and suspension system. The vehicle steering system and suspension system dynamics model, tracking error model, road model and steering and suspension system reward modules together constitute the reinforcement learning environment. In the path tracking control process of the autonomous driving vehicle, the lateral and vertical motion processes of the vehicle in the path tracking process are controlled respectively by outputting the control actions of the steering and suspension systems of the controlled vehicle. The two intelligent agents respectively observe the motion states of the steering system and suspension system of the reinforcement learning environment, including the lateral velocity, yaw angular velocity, lateral deviation, heading deviation, road curvature of the steering system, and the body displacement of the suspension system, vehicle roll angle, and suspension dynamic travel. The control actions of the steering system are executed according to the strategy in the intelligent agent. p Control action of the suspension system s1 with a s2 After making a control decision, the motion state of the autonomous vehicle changes, and the control state enters the next control state. At the same time, the reward function module of the steering system is transformed into the lateral motion state s p (v y ,r,e y ,e ψ ,κ) calculates the reward value obtained after executing the steering control action strategy. The reward function module of the suspension system is based on the vertical motion state The agent calculates the reward value obtained after executing the suspension control action strategy. The agent further updates the strategy based on the feedback reward value. This cycle continues until the agent accumulates enough rewards to achieve coordinated control of the steering and suspension systems.

[0174] like Figure 7 、 Figure 8 As shown, by comparing with the model predictive control algorithm and the single-agent learning algorithm, the method of the present invention effectively tracks the reference path and obtains the minimum tracking error, which can achieve better steering path tracking performance. At the same time, the method of the present invention effectively reduces the roll acceleration of the autonomous driving vehicle, can obtain a smaller roll stability area, and the improved suspension comfort performance is significantly better than the other two control strategies, while the improvement effect of the model predictive control algorithm only exceeds the worst single-agent control. These results show that the method of the present invention can effectively improve the ride comfort and stability during the tracking steering process while ensuring accurate and efficient tracking performance, and complete the coordinated control of the steering system and the suspension system.

[0175] The above is only a specific implementation of the present invention, but the design concept of the present invention is not limited to this. Any non-substantial changes to the present invention using this concept shall be deemed as an infringement of the protection scope of the present invention.

Claims

1. A method for coordinated steering and suspension control of an autonomous vehicle based on reinforcement learning, characterized in that: The steps include: 1) Establish dynamic models of the vehicle's steering and suspension systems, and obtain the lateral and vertical vehicle motion states by controlling the steering angle and suspension forces; 2) Constructing control actions and reward functions for the vehicle's steering and suspension systems, and forming a reinforcement learning exploration strategy for coordinated control of vehicle steering and suspension. 3) Using deep neural networks to build intelligent agents for steering and suspension systems; 4) Combining the dynamics model with the reward function module to build a reinforcement learning environment, and achieving coordinated control of the steering and suspension systems through reinforcement learning methods of the steering and suspension system agents; In step 1), the steering system dynamics model is constructed by combining the roll and yaw motion states in the two-degree-of-freedom steering system: ; ; ; ; Among them, m and m s are the total vehicle mass and the sprung mass of the vehicle, respectively; h is the distance from the center of mass of the vehicle to the roll center; φ r is the road inclination; g is the acceleration due to gravity; l f With l r are the distances from the center of mass to the front and rear axles respectively; I z is the yaw moment of inertia; r is the yaw angular velocity; is the yaw angular acceleration; v x With v y are the longitudinal and lateral velocities, respectively; is the lateral acceleration; is the roll acceleration; F yf With F yr are the lateral forces on the front and rear axle tires respectively; k is the road curvature; e ψ is the heading deviation; and are the rates of change of lateral deviation and heading deviation respectively; The forces acting on the front and rear suspension systems are added together and concentrated on an equivalent roll axis, and the dynamic model of the suspension system is constructed as follows: ; ; ; ; ; Among them, z s is the vehicle body displacement; is the vehicle body acceleration; a y is the lateral acceleration; I x is the roll moment of inertia; φ is the vehicle roll angle; z s1 With z s2 are the displacements of the left and right sides of the vehicle body respectively; z w1 With z w2 are the left and right wheel displacements respectively; and are the left and right wheel accelerations respectively; k s1 With k s2 are the left and right suspension stiffness coefficients respectively; m w1 With m w2 are the unsprung masses on the left and right sides respectively; z r1 With z r2 are the road excitation displacements on the left and right sides respectively; k t1 With k t2 are the stiffness coefficients of the left and right tires respectively; F s1 With F s2 Powers the left and right suspensions respectively; T w is the wheel track; In step 2), feedforward steering is applied during the steering control exploration process of reinforcement learning. The relationship between the feedforward steering angle of the vehicle and the front and rear tire slip angles is obtained through the front and rear wheel slip motion relationship: ; in, is the feedforward steering angle; L is the distance from the front axle to the rear axle of the vehicle; and are the feedforward tire slip angles of the front and rear wheels respectively, and the feedforward tire force is expressed as: ; in, and are the feedforward tire forces for the front and rear wheels respectively; The required steering exploration command is calculated by the sum of feedforward and learning feedback control, and the control action of the steering system is expressed as: ; Among them, a p is the control action of the steering system; π p is the exploration strategy for reinforcement learning; N p Random exploration noise for reinforcement learning; Objective function R m Set as: ; in, is the control increment of the steering angle; k e 、k ψ 、k δ 、k ∆δ are the penalty coefficients for lateral deviation, heading deviation, steering angle control amount and control increment respectively; The constraints of the steering angle control amount and control increment are set as: ; Among them, R p1 With R p2 are the constraints for controlling the steering angle and its increment respectively; k δ′ and are the penalty coefficients for controlling the steering angle and its incremental constraint respectively; δ fmax and They are the maximum limit values ​​for controlling the steering angle and its increment respectively; At the same time, the envelope performance of the vehicle's motion state is considered in the active constraint function, and penalties are set for actions that exceed the vehicle's yaw and lateral motion constraints. The constraints are: ; Among them, R p3 With R p4 are the constraints for vehicle yaw and lateral motion respectively; k β With k r are the penalty coefficients for sideslip angle and yaw rate constraints respectively; α s is the saturated tire slip angle; b1 and b2 are the constraint boundaries of the yaw rate; Combining the above objective function and constraint function, the reward function R of the vehicle steering system is obtained p for: 。 2. The method for coordinated steering and suspension control of an autonomous driving vehicle based on reinforcement learning according to claim 1, characterized in that: In step 2), a fast-response PID control strategy is used to compensate for the uncertainty in the suspension exploration process. Two PID controllers are used to control the vertical and roll motions of the suspension system respectively. The corresponding PID control design is: ; ; Among them, f z (t) and f φ (t) are the control forces for vertical and rolling motion at time t; e z (t) and e φ (t) are the error inputs of vertical and roll motion at time t, respectively; K zp , K zi With K zd is the PID control coefficient of vertical motion; K φp , K φi With K φd is the PID control coefficient of the roll motion; The vertical and roll motion control forces of the suspension system are distributed to the controllable actuators of the left and right suspensions. The decoupled distribution control force relationship is expressed as: ; Among them, f d1 (t) and f d2 (t) are the power of the left and right suspensions at time t respectively; The compensation force of the left and right suspensions is calculated as: ; The required suspension system exploration command is calculated by the sum of the suspension compensation force and the learning feedback control force. The control action of the suspension system is expressed as: ; Among them, a s1 with a s2 are the control actions of the left and right suspension respectively; π s1 and π s2 are the exploration strategies of the left and right suspensions respectively; N s1 With N s2 are the random exploration noises of the left and right suspensions, respectively; According to the dynamic response process of the passive suspension system, the various indicators in the reward process are normalized to obtain the target reward function R n for: ; Among them, k a 、k φ , k1, k2, k f1 、k f2 They are the penalty coefficients for vehicle speed acceleration, roll acceleration, left and right suspension travel, and suspension force; are the normalized terms of vehicle speed acceleration, roll acceleration, left and right suspension travel, and suspension force; The constraints on suspension travel are extended to the left and right suspensions, and both must be less than the maximum allowable limit travel. That is, the constraints on suspension travel must meet the following requirements: ; Among them, R s1 With R s2 are the constraints of the left and right suspension travel respectively; k d1 With k d2 are the penalty coefficients for the left and right suspension travel constraints respectively; They are the maximum limit values ​​of the left and right suspension travel respectively; The control constraints of the left and right suspensions are set as: ; Among them, F s1max With F s2max are the maximum limit values ​​of the left and right suspension control quantities respectively; Finally, the reward function R of the vehicle suspension system is obtained s for: 。 3. The method for coordinated steering and suspension control of an autonomous driving vehicle based on reinforcement learning according to claim 1, characterized in that: In step 3), a deep deterministic policy gradient algorithm is used to construct an intelligent agent for the steering system and suspension system. The intelligent agent includes a behavior network and an evaluation network. The behavior network of the steering system consists of a state input layer, three hidden layers, and a control output layer, wherein the control output layer is the exploration strategy of the steering system. The evaluation network of the steering system includes a state input layer, a control input layer, three hidden layers, and an output layer. The three hidden layers are connected in series, and the control input layer is connected to the second hidden layer for transmission. The suspension system's behavioral network consists of a state input layer, three hidden layers, and a control output layer. The control output layer defines the exploration strategy for the left and right suspensions and uses an activation function to scale the suspension forces within the range [-1, 1]. The suspension system's evaluation network consists of a state input layer, a control input layer, three hidden layers, and an output layer. The three hidden layers are connected in series, with the control input layer connected to the second hidden layer. During the learning process, it is set that when the vehicle tracking error exceeds 1m, the agent training process is terminated.

4. The method for coordinated steering and suspension control of an autonomous driving vehicle based on reinforcement learning according to claim 3, characterized in that: The state input layer of the steering system's behavioral network is 6-dimensional, including lateral velocity, yaw angular velocity, lateral deviation, heading deviation, road curvature and the control action at the previous moment. Each hidden layer of the steering system's behavioral network consists of 100 neurons.

5. The method for coordinated steering and suspension control of an autonomous driving vehicle based on reinforcement learning according to claim 4, characterized in that: The state input layer of the steering system's evaluation network obtains the intelligent agent's 6-dimensional state from the steering system's behavior network. The control input layer of the steering system's evaluation network is a 1-dimensional control action. Each hidden layer of the steering system's evaluation network consists of 100 neurons. The output layer of the steering system's evaluation network is a 1-dimensional action-value function that evaluates the steering system's path tracking control behavior.

6. The method for coordinated steering and suspension control of an autonomous driving vehicle based on reinforcement learning according to claim 3, characterized in that: The state input layer of the suspension system's behavioral network is 6-dimensional, including vehicle body displacement, vehicle roll angle, left and right suspension travel, and the left and right suspension control actions at the previous moment. Each hidden layer of the suspension system's behavioral network has 100 neurons. The state input layer of the suspension system's evaluation network is consistent with the suspension system's behavioral network, and the control input layer is 2-dimensional left and right suspension control actions. Each hidden layer of the suspension system evaluation network consists of 100 neurons, and the output layer of the suspension system evaluation network is the action value function for evaluating the suspension system behavior.

7. The method for coordinated steering and suspension control of an autonomous driving vehicle based on reinforcement learning according to claim 1, characterized in that: In step 4), a reinforcement learning environment is formed by the vehicle steering and suspension system dynamics models, the tracking error model, the road model, and the steering and suspension system reward modules. During the path tracking control process of the autonomous vehicle, the lateral and vertical motion processes of the vehicle during the path tracking process are respectively controlled by outputting control actions of the steering and suspension systems of the controlled vehicle. The intelligent agents of the steering system and suspension system respectively observe the motion states of the steering system and suspension system in the reinforcement learning environment; according to the strategies in the intelligent agents, the control actions of the steering system and the suspension system are respectively executed. After making control decisions, the motion state of the autonomous driving vehicle changes, and the motion state enters the next motion state; at the same time, the reward function module of the steering system is used to calculate the lateral motion state s p (v y , r, e y , e ψ , κ) calculates the reward value obtained after executing the steering control action strategy. The reward function module of the suspension system is calculated by the vertical motion state s s (z s φ z s1 -z w1 z s2 -z w2 ) Calculate the reward value obtained after executing the suspension control action strategy, and the agent further updates the applied strategy based on the feedback reward value. Repeat the above steps until the agent obtains enough cumulative rewards to achieve coordinated control of the steering system and suspension system.

8. The method for coordinated steering and suspension control of an autonomous driving vehicle based on reinforcement learning according to claim 7, characterized in that: The motion states of the steering system and suspension system include the lateral velocity, yaw rate, lateral deviation, heading deviation, and road curvature of the steering system, and the body displacement, vehicle roll angle, and suspension travel of the suspension system.

Citation Information

Patent Citations

  • Man-vehicle cooperative steering control method based on reinforcement learning corner weight distribution

    CN115062539A

  • Vehicle active front wheel steering and active suspension system coordination control method

    CN116767180A