A coordinated control method for vehicle stability based on reinforcement learning
By using a reinforcement learning-based approach, combined with a linear two-degree-of-freedom model and sliding mode control, the active rear-wheel steering and yaw moment are coordinated, solving the vehicle handling stability problem under extreme conditions and improving vehicle stability and safety.
Patent Information
- Application Number
- CN202411644012.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-11-18
AI Technical Summary
Under extreme conditions, when tire force approaches saturation, existing technologies struggle to effectively combine active rear-wheel steering with yaw moment control, thus failing to improve vehicle handling stability and safety.
A reinforcement learning-based approach is adopted to obtain the desired yaw rate and center of gravity sideslip angle through a linear two-degree-of-freedom model. The additional yaw torque and active rear wheel steering angle are obtained by combining sliding mode control. The agent is trained through reinforcement learning to achieve optimal cooperative control.
It achieves precise tracking of the ideal yaw rate, improves vehicle stability control margin, enhances tire utilization, and ensures vehicle stability and safety.
Smart Images

Figure CN119408527B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of electric vehicle control technology, specifically relating to a vehicle stability coordination control method based on reinforcement learning. Background Technology
[0002] The development of new energy vehicles is accelerating, boasting unique advantages such as zero pollution, zero emissions, simple structure, and ease of control. Distributed drive directly mounts the motor inside or beside each wheel, driving the wheels directly with independent torque control for each wheel. This fully leverages the advantages of fast motor response and precise torque adjustment, providing greater freedom for vehicle dynamics control.
[0003] Distributed drive electric vehicles can control the output torque of four drive motors to meet the longitudinal movement of the vehicle, and can also achieve functions such as additional lateral yaw moment and active rear-wheel steering by changing the torque of some drive motors. At low speeds, active rear-wheel steering can provide a steering angle opposite to the front wheel steering angle, reducing the turning radius and increasing the vehicle's agility. At high speeds, active rear-wheel steering can provide a steering angle in the same direction as the front wheel steering angle to prevent skidding. However, under extreme conditions, tire force tends to saturate, at which point the tires cannot provide sufficient lateral force to maintain vehicle stability, and the active rear-wheel steering system cannot be used to improve vehicle handling stability. In this case, a yaw moment can be generated by changing the longitudinal force of the tires to control the vehicle and maintain stability. However, as a complex nonlinear system, how to combine active rear-wheel steering with yaw moment control to improve the vehicle's maneuverability and stability during steering, thereby ensuring vehicle safety, is a problem that urgently needs to be solved. Summary of the Invention
[0004] This invention addresses the shortcomings of existing technologies by providing a reinforcement learning-based vehicle stability coordination control method that can coordinate the control of active rear wheel steering and yaw moment to ensure vehicle stability and safety.
[0005] This invention provides the following technical solution:
[0006] A reinforcement learning-based coordinated control method for vehicle stability is provided, comprising:
[0007] Based on the vehicle's structural parameters, speed, and front wheel steering angle, a linear two-degree-of-freedom model is used to obtain the desired yaw rate ω. d And the expected centroid side slip angle β d ;
[0008] Based on the desired yaw rate ω d Actual yaw rate, desired centroid sideslip angle β dThe additional yaw moment ΔM and the active rear wheel steering angle δ are obtained by sliding mode control of the actual center of gravity sideslip angle. r ;
[0009] Based on the established clustering diagram of vehicle state parameters under the specified working conditions and the current vehicle state parameters, the current vehicle state type is obtained; the state type is one of stable state, critically stable state, and unstable state.
[0010] Construct a reinforcement learning agent and reward function, and continuously train the agent until it outputs an additional yaw moment ΔM and an active rear wheel steering angle δ. r Each has its own optimal cooperative control coefficient to achieve coordinated control of the electric vehicle; the reward function includes steady-state reward, yaw rate error and centroid sideslip angle error reward, critical steady-state reward and anti-shake reward.
[0011] Optionally, the vehicle's structural parameters, speed, and front wheel steering angle are obtained using a linear two-degree-of-freedom model to determine the desired yaw rate ω. d And the expected centroid side slip angle β d middle:
[0012] The linear two-degree-of-freedom model is in the form of:
[0013]
[0014] Where k1 is the stiffness of the front axle, k2 is the stiffness of the rear axle, β is the sideslip angle of the vehicle's center of gravity, and v x Let ω be the longitudinal vehicle speed, a be the distance from the vehicle's center of gravity to the front axle, b be the distance from the vehicle's center of gravity to the rear axle, and ω be the longitudinal speed. exp The desired yaw rate is the output of the linear two-degree-of-freedom model, where δ is the standard front wheel steering angle, m is the vehicle mass, and a is the yaw rate. y For the lateral acceleration of the vehicle, I z Let be the vehicle's moment of inertia. ω exp The derivative;
[0015] The desired yaw rate ω output by the linear two-degree-of-freedom model exp for:
[0016]
[0017] Where K is the vehicle stability factor;
[0018] The desired yaw rate ω d And the expected centroid side slip angle β d for:
[0019]
[0020] Where μ is the adhesion coefficient, g is the gravitational acceleration, and β d Let sgn(δ) be the expected centroid sideslip angle, and let sgn(δ) be the judgment function indicating the sign of the front wheel steering angle. When the front wheel steering angle is positive, sgn(δ) is 1, and when the front wheel steering angle is negative, sgn(δ) is -1.
[0021] Optionally, the standard front wheel steering angle δ is the average wheel steering angle of the two front wheels of the vehicle; the wheel steering angle of the two front wheels and the longitudinal vehicle speed v x The vehicle's moment of inertia I is obtained through sensors. z Based on the vehicle's desired torque T d Obtain; the desired torque T of the vehicle d for:
[0022]
[0023] Where θ is the accelerator pedal opening, and T imax This represents the peak torque of the i-th hub motor.
[0024] Optionally, the step of determining the desired yaw rate ω d Actual yaw rate, desired centroid sideslip angle β d The additional yaw moment ΔM and the active rear wheel steering angle δ are obtained by sliding mode control of the actual center of gravity sideslip angle. r , specifically including:
[0025] Based on the desired yaw rate ω d Based on the actual yaw rate, design the sliding surface of the yaw torque controller to obtain the additional yaw torque ΔM;
[0026]
[0027] Based on the expected centroid sideslip angle β d Based on the actual sideslip angle of the center of gravity, the sliding surface of the active rear wheel controller is designed to obtain the active rear wheel steering angle δ. r ;
[0028]
[0029] Where Iz is the vehicle's moment of inertia, and c ω Here, represents the relative weighting coefficient between error and the rate of change of error; a is the distance from the vehicle's center of gravity to the front axle; b is the distance from the vehicle's center of gravity to the rear axle; and C... f For the stiffness of the vehicle's front wheels, C r The stiffness of the vehicle's rear wheels, v x For longitudinal vehicle speed, and e ω and e βThe derivative, e ω For the deviation of the control variable of yaw rate, e β The deviation is the control variable for the centroid sideslip angle, K0 is the rate of approach constant, and s ω This is the sliding mode shear function.
[0030] Optionally, based on the established vehicle state parameter clustering diagram under the set working condition and the current vehicle state parameters, the current vehicle state type is obtained. In the vehicle state parameter clustering diagram under the set working condition, there are a total of 5 cluster centers. Among them, the middle cluster is the stable state, the clusters located on the left and right edges of the middle cluster are the unstable states, and the clusters located between the stable state and the unstable state are the critical stable state.
[0031] The vehicle state parameters include: lateral velocity, lateral acceleration, yaw rate, center of gravity sideslip angle, and tire sideslip angle.
[0032] Optionally, the process involves constructing a reinforcement learning agent and a reward function, and continuously training the agent until it outputs an additional yaw moment ΔM and an active rear wheel steering angle δ. r Their respective optimal collaborative control coefficients include:
[0033] Define the state s of the agent. t for:
[0034] s t ={ω, Δω, β, Δβ, a y}
[0035] Where Δω is the difference between the yaw rate and the actual yaw rate, Δβ is the difference between the desired sideslip angle and the actual sideslip angle, ω is the actual yaw rate, β is the actual sideslip angle, and a y It is lateral acceleration;
[0036] Define action a t The synergistic relationship between the rear wheel steering angle and the additional yaw moment is expressed as follows:
[0037] a t = {κ, 1-κ}, κ∈[0, 1]
[0038] Where κ is the active rear wheel steering angle δ r The coordination control coefficient, 1-κ is the coordination coefficient of the additional yaw moment ΔM;
[0039] Define reward function r t for:
[0040] r t =a1r s +a2r e +a3r cs +a4rc
[0041] Where, r s For steady-state rewards, r e As a reward for yaw rate error and center of mass sideslip angle error, r cs For the critical steady-state reward, r c For anti-shake rewards, a1, a2, a3, and a4 are respectively r s 、r e 、r cs and r c The reward coefficient.
[0042] Optionally, the steady-state reward r s This includes rewards for the vehicle's current state and rewards for predicted future states, as shown in the following formula:
[0043] r s =r s1 +r s2 +r s3 +r s4
[0044]
[0045] Where, r s1 For the reward of the car's current state, r s2 、r s3 and r s4 t1, t2, t3, and t4 represent the predicted car state rewards for the first, second, and third time points, respectively; Dis1 is the Euclidean distance from the vehicle state parameters to the cluster centers; t1, t2, t3, and t4 are respectively the r... s1 、r s2 、r s3 and r s4 The adjustment coefficient; q is the cluster category to which the current vehicle belongs; q=1 indicates that the current vehicle is in a stable state, q=2 or q=4 indicates that the current vehicle is in a critically stable state, and q=3 and q=5 indicate that the current vehicle is in an unstable state.
[0046] r s2 、r s3 and r s4 The corresponding clustering categories are obtained based on the predicted vehicle state parameters and the clustering diagram of vehicle state parameters under the set operating conditions.
[0047] The predicted vehicle state parameters are: based on the current vehicle state parameters and the additional yaw moment ΔM and active rear wheel steering angle δ of the control input. r Based on the linear two-degree-of-freedom model, predict r s2 、r s3 and r s4The vehicle state parameters at the corresponding time.
[0048] Optionally, the yaw rate error and the center of mass sideslip angle error are rewarded r. e for:
[0049] r e =-a ω Δω 2 -a β Δβ 2 +W
[0050] Where W is a constant, a ω and a β These are the adjustment coefficients for Δω and Δβ, respectively;
[0051] The critical stable state reward r cs for:
[0052]
[0053] Where D1 is the dot product of the vector formed by the vehicle's current state parameters and the parameters of the cluster center to which the vehicle belongs, and the vector formed by the second cluster center to the first cluster center; D1 is the length of the vector formed by the second cluster center to the first cluster center; D2 is the dot product of the vector formed by the vehicle's current state parameters and the parameters of the vehicle's cluster center, and the vector formed by the fourth cluster center to the first cluster center. Let be the length of the vector formed by the fourth cluster center and the first cluster center; n is the adjustment coefficient of the dot product, and k p The adjustment coefficient for the critical steady-state reward;
[0054] The anti-shake reward r c for:
[0055] r c =-Δδ r 2 -ΔM z 2
[0056] Where, Δδ r ΔM is the difference between the current active rear wheel steering angle and the previous angle; z The difference between the yaw moment at the current moment and the moment at the previous moment is applied.
[0057] Compared with the prior art, the beneficial effects of the present invention are:
[0058] This invention applies reinforcement learning algorithms to vehicle stability coordination control, enabling precise tracking of the ideal yaw rate. It fully leverages the advantages of both active rear-wheel steering and direct yaw moment control methods, improving vehicle stability control margin, significantly increasing tire utilization, and ensuring vehicle stability, comfort, and appropriate control. Attached Figure Description
[0059] Figure 1 This is a control principle diagram of the vehicle stability coordination control method based on reinforcement learning of the present invention;
[0060] Figure 2 This is a K-means clustering diagram provided by the present invention;
[0061] Figure 3 This is a diagram showing the reinforcement learning training results of the present invention. Detailed Implementation
[0062] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.
[0063] It should be noted that the term "comprising" and any variations thereof in the specification and claims of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or device.
[0064] like Figure 1 As shown, a vehicle stability coordination control method based on reinforcement learning is provided, including the following steps:
[0065] S1: Based on the vehicle's structural parameters, speed, and front wheel steering angle, use a linear two-degree-of-freedom model to obtain the desired yaw rate ω. d And the expected centroid side slip angle β d .
[0066] The structural parameters of a vehicle include the stiffness of the front axle, the stiffness of the rear axle, the distance from the vehicle's center of gravity to the front axle, the distance from the vehicle's center of gravity to the rear axle, and the vehicle's mass.
[0067] Specifically, the desired torque of the vehicle is obtained based on the accelerator pedal opening, and the desired yaw rate and desired center of gravity sideslip angle are calculated using a linear two-degree-of-freedom model based on the actual vehicle speed and front wheel steering angle; the accelerator pedal opening is obtained by a sensor.
[0068] The vehicle's desired torque T d for:
[0069]
[0070] Where θ is the accelerator pedal opening, and T imax This represents the peak torque of the i-th hub motor.
[0071] The linear two-degree-of-freedom model is in the form of:
[0072]
[0073] In the formula, k1 is the stiffness of the front axle, k2 is the stiffness of the rear axle, β is the sideslip angle of the vehicle's center of gravity, and v x Let ω be the longitudinal vehicle speed, a be the distance from the vehicle's center of gravity to the front axle, n be the distance from the vehicle's center of gravity to the rear axle, and ω be the longitudinal speed. exp The desired yaw rate is the output of a linear two-degree-of-freedom model, where δ is the standard front wheel steering angle, m is the vehicle mass, and v is the yaw rate. y For lateral vehicle speed, a y Let I be the lateral acceleration of the vehicle, and Iz be the moment of inertia of the vehicle. ω exp The derivative of .
[0074] The standard front wheel steering angle δ is the average wheel steering angle of the two front wheels of the vehicle; the wheel steering angle δ of the two front wheels of the vehicle f and longitudinal vehicle speed v x Obtained by sensors; Vehicle rotational inertia I z Based on the vehicle's desired torque T d Obtain.
[0075] The desired yaw rate ω output by the linear two-degree-of-freedom model exp for:
[0076]
[0077] Where K is the vehicle stability factor.
[0078] Taking into account the influence of the adhesion coefficient μ, the lateral acceleration of the vehicle during driving... Should meet When the vehicle is in a steady state, the sideslip angle β is small, and the lateral acceleration can be expressed as:
[0079]
[0080] Considering the actual capabilities of electric wheel drive vehicles during operation, a 5% stability margin is applied; and since the desired sideslip angle of an electric wheel drive vehicle is 0, the final desired yaw rate and desired sideslip angle can be obtained:
[0081]
[0082] Where μ is the adhesion coefficient, g is the gravitational acceleration, and β d Let sgn(δ) be the expected centroid sideslip angle, and let sgn(δ) be the judgment function indicating the sign of the front wheel steering angle. When the front wheel steering angle is positive, sgn(δ) is 1, and when the front wheel steering angle is negative, sgn(δ) is -1.
[0083] S2: Based on the desired yaw rate ω d Actual yaw rate, desired centroid sideslip angle β d The additional yaw moment ΔM and the active rear wheel steering angle δ are obtained by sliding mode control of the actual center of gravity sideslip angle. r .
[0084] Specifically, S2.1: Using the desired yaw rate and the actual yaw rate of the vehicle, design the sliding surface of the yaw moment controller, and then solve for the additional yaw moment ΔM that ensures stability.
[0085] S2.2: Based on the desired centroid sideslip angle β d Based on the actual sideslip angle of the center of gravity, the sliding surface of the active rear wheel controller is designed to obtain the active rear wheel steering angle δ. r .
[0086] Step S2.1 specifically includes the following steps:
[0087] The equation for the vehicle yaw force with ΔM added is:
[0088]
[0089] Among them, C f For the stiffness of the vehicle's front wheels, C r This refers to the stiffness of the vehicle's rear wheels.
[0090] The control amount of the additional yaw moment is determined based on the deviation between the center of gravity sideslip angle and the yaw rate. The control variable deviation of the sliding mode controller is:
[0091] e ω =(ω-ω d ), e β =(ω-ω β );
[0092] In the formula, e is the control variable deviation, ω is the actual yaw rate, and β is the actual centroid sideslip angle;
[0093] The sliding mode switching function is defined as follows:
[0094]
[0095] In the formula, c ω c is the relative weighting coefficient between error and the rate of change of error.ω >0.
[0096]
[0097] In the formula, s ω For sliding mode shear function, The sliding mode index convergence rate;
[0098] The sliding mode index convergence rate satisfies:
[0099]
[0100] In the formula, K0 is the rate of approach constant, which indicates the rate at which the state point of the system approaches the sliding surface; to eliminate chattering, the saturation function sat(S) is chosen to replace the sign function, and both the boundary layer thickness and the rate of approach exponent are positive numbers;
[0101] The saturation function satisfies:
[0102]
[0103] In the formula, k e Let k be the boundary layer thickness. e =0.5.
[0104] Then there is
[0105]
[0106] Calculate the expected additional yaw moment using the sliding mode index convergence formula:
[0107]
[0108] Among them, I z Let c be the vehicle's moment of inertia. ω Here, represents the relative weighting coefficient between error and the rate of change of error; a is the distance from the vehicle's center of gravity to the front axle; b is the distance from the vehicle's center of gravity to the rear axle; and C... f For the stiffness of the vehicle's front wheels, C r The stiffness of the vehicle's rear wheels, v x For longitudinal vehicle speed, and e ω and e β The derivative, e ω For the deviation of the control variable of yaw rate, e β The deviation is the control variable for the centroid sideslip angle, K0 is the rate of approach constant, and s ω This is the sliding mode shear function.
[0109] When calculating the active rear wheel steering angle δ rAt this time, the first thing to do is to determine the added rear wheel steering angle δ. r The equation for the lateral sway force of a vehicle is:
[0110]
[0111] In the formula, δ f This refers to the steering angle of the front wheels.
[0112] Calculating the active rear wheel steering angle δ r The same sliding mode method is used to calculate the additional yaw moment ΔM; therefore, the steering angle δ of the active rear wheel is... r for:
[0113]
[0114] S3: Based on the established clustering diagram of vehicle state parameters under the set working conditions and the current vehicle state parameters, obtain the current vehicle state type; the state type is one of stable state, critically stable state and unstable state.
[0115] The cluster diagram of vehicle state parameters under the given operating conditions has a total of 5 cluster centers. The middle cluster is the stable state, the clusters on the left and right edges of the middle cluster are the unstable states, and the clusters between the stable and unstable states are the critical stable states.
[0116] Vehicle status parameters include: lateral velocity, lateral acceleration, yaw rate, center of gravity sideslip angle, and tire sideslip angle.
[0117] Specifically, a clustering diagram of vehicle state parameters under a given operating condition was established. Taking into account both vehicle input parameters and various state parameters characterizing the vehicle's lateral stability, a specific operating condition was selected in CarSim, and simulations were performed at speeds between 10-100 km / h. The obtained state parameter data was then clustered offline into five classes: Class 1 represents stable states, Classes 2 and 4 represent critically stable states, and Classes 3 and 5 represent unstable states. Five cluster centers were obtained from this clustering.
[0118] As an option, such as Figure 2 As shown, the clustering diagram obtained when the simulation test condition is the dual line-changing (DLC) condition is as follows: the green and orange areas represent the third and fifth clusters, respectively, which are unstable states; the yellow and purple areas represent the second and fourth clusters, respectively, which are critically stable states; and the blue area represents the first cluster, which is a stable state. When performing offline clustering, it is essential to ensure that the condition settings are the same as those used in step S4.
[0119] S4: Construct a reinforcement learning agent and reward function, and continuously train the agent until it outputs the additional yaw moment ΔM and the active rear wheel steering angle δ. rEach has its own optimal cooperative control coefficient to achieve coordinated control of the electric vehicle; the reward function includes steady-state reward, yaw rate error and centroid sideslip angle error reward, critical steady-state reward and anti-shake reward.
[0120] The reinforcement learning algorithm in this application can be described as including vehicle states s t Cooperative control action a t and instant rewards t The Markov decision process. Specifically, the agent, based on the current vehicle state s t Take an action a t To coordinate the rear wheel steering angle and add yaw moment, update the vehicle to a new state. t+1 According to the vehicle status s at the next moment. t+1 Generates an instant reward r t (s t a t Through iterative training, the agent can obtain the optimal policy network. To obtain the maximum cumulative return. In summary, the Markov decision process can be represented as follows: MDP = {S, A, P} s Let S be the state space, A be the action space, and P be the action space. s λ is the state transition function; r is the reward function; λ is the discount factor to ensure the convergence of the reward function.
[0121] Define the state s of the agent t for:
[0122] s t ={ω, Δω, β, Δβ, a y}
[0123] Where, Δω=ω-ω d The difference between the yaw rate and the actual yaw rate is Δβ = β - β. d ω is the difference between the desired and actual sideslip angles, β is the actual yaw rate, and a is the actual sideslip angle. y It is lateral acceleration;
[0124] Define action a t The synergistic relationship between the rear wheel steering angle and the additional yaw moment is expressed as follows:
[0125] a t = {κ, 1-κ}, κ∈[0, 1]
[0126] Where κ is the active rear wheel steering angle δ r The cooperative control coefficient is 1-κ, which is the cooperative coefficient of the additional yaw moment ΔM.
[0127] The corrected rear wheel steering angle and additional yaw moment are shown below:
[0128]
[0129] The reward is a numerical value returned to the agent by the environment after the agent performs an action. The reward function serves as a pointer for control optimization in deep reinforcement learning, and the definition of the reward directly affects the effect of reinforcement learning.
[0130] Define reward function r t for:
[0131] r t =a1r s +a2r e +a3r cs +a4r c
[0132] Where, r s For steady-state rewards, r e As a reward for yaw rate error and center of mass sideslip angle error, r cs For the critical steady-state reward, r c For anti-shake rewards, a1, a2, a3, and a4 are respectively r s 、r e 、r cs and r c The reward coefficient.
[0133] To ensure that the cars remain in a stable state, i.e., the first cluster, the stable state reward r s for:
[0134] r s =r s1 +r s2 +r s3 +r s4
[0135]
[0136]
[0137] Where, r s1 For the reward of the car's current state, r s2 、r s3 and r s4 t1, t2, t3, and t4 represent the predicted car state rewards for the first, second, and third time points, respectively; Dis1 is the Euclidean distance from the vehicle state parameters to the cluster centers; t1, t2, t3, and t4 are respectively the r... s1、 r sz 、r s3 and r s4The adjustment coefficient is q; q is the cluster category to which the current vehicle belongs; q=1 indicates that the current vehicle is in a stable state, q=2 or q=4 indicates that the current vehicle is in a critically stable state, and q=3 and q=5 indicate that the current vehicle is in an unstable state.
[0138] steady-state reward r s This includes rewards for the vehicle's current state and rewards for predicted future states. Incorporating the predicted future state reward into the reward function can prevent sudden changes in the coordination coefficient caused by abrupt shifts in the vehicle's state parameters at the next time step.
[0139] r s2 r s3 and r s4 The corresponding clustering categories are obtained based on the predicted vehicle state parameters and the clustering diagram of vehicle state parameters under the set operating conditions;
[0140] The predicted vehicle state parameters are: based on the current vehicle state parameters and the additional yaw moment ΔM and active rear wheel steering angle δ of the control input. r Based on the linear two-degree-of-freedom model, predict r s2 r s3 and r s4 The vehicle state parameters at the corresponding time.
[0141] Specifically, the linear two-degree-of-freedom model is transformed into the following form:
[0142]
[0143] Discretize the above form:
[0144]
[0145] Where β′ and ω′ are the sideslip angle and yaw rate of the center of mass after differentiation, respectively.
[0146] The discrete form of the linear two-degree-of-freedom model, the current vehicle state parameters, and the additional yaw moment ΔM and active rear wheel steering angle δ of the control input are used. r To determine the vehicle state parameters for the predicted future few moments.
[0147] Yaw rate error and center of mass sideslip angle error bonus r e for:
[0148] r e =-a ω Δω 2 -a β Δβ 2 +W
[0149] Among them, aω and a β Δω and Δβ are the adjustment coefficients, respectively; W is a constant, and when Δω 2 When W = 0.5, it is less than or equal to 0.01; otherwise, it is less than or equal to 0.
[0150] When the car is in a critically stable state, i.e., in the second or fourth cluster, we hope to improve the car's state from the critically stable state to a stable state. Therefore, a critically stable state reward r is introduced. cs for:
[0151]
[0152] Where D1 is the dot product of the vector formed by the vehicle's current state parameters and the parameters of the cluster center to which the vehicle belongs, and the vector formed by the second cluster center to the first cluster center; D1 is the length of the vector formed by the second cluster center to the first cluster center; D2 is the dot product of the vector formed by the vehicle's current state parameters and the parameters of the vehicle's cluster center, and the vector formed by the fourth cluster center to the first cluster center. Let be the length of the vector formed by the fourth cluster center and the first cluster center; n is the adjustment coefficient of the dot product, and k p The adjustment coefficient for the critical steady-state reward;
[0153] During control, to ensure ride comfort, the changes in rear wheel steering angle and additional yaw moment between the previous and current moments should not be too large; the smaller the values, the less the car vibrates. (Anti-vibration bonus r) c for:
[0154] r c =-Δδ r 2 -ΔM z 2
[0155] Where, Δδ r ΔM is the difference between the current active rear wheel steering angle and the previous angle; z The difference between the yaw moment at the current moment and the moment at the previous moment is applied.
[0156] As an option, such as Figure 3 As shown, the graph illustrates the training curves after 100 iterations under DLC conditions with a vehicle speed of 90 km / h and a coefficient of friction of 0.5. The light blue line represents the reward value obtained at each learning step, and the dark blue line represents the average reward. Figure 3 It can be seen that the average reward of this application is stable, and the agent remains stable during the training process.
[0157] The active rear wheel steering controller outputs the active rear wheel steering angle δ in step S2. rAnd the active rear wheel steering angle δ output in step S4 r The optimal cooperative control coefficient is used to perform active rear-wheel steering control on the vehicle; the direct yaw moment controller performs torque distribution based on the additional yaw moment ΔM output in step S2 and the optimal cooperative control coefficient of the additional yaw moment ΔM output in step S4. The torque distribution method refers to the existing technology, thereby obtaining the driving or braking torque of the four wheels so as to control the four wheels separately.
[0158] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present invention.
[0159] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A vehicle stability coordination control method based on reinforcement learning, characterized in that, include: Based on the vehicle's structural parameters, speed, and front wheel steering angle, a linear two-degree-of-freedom model is used to obtain the desired yaw rate ω. d And the expected centroid side slip angle β d ; Based on the desired yaw rate ω d Actual yaw rate, desired centroid sideslip angle β d The additional yaw moment ΔM and the active rear wheel steering angle δ are obtained by sliding mode control of the actual center of gravity sideslip angle. r ; Based on the established clustering diagram of vehicle state parameters under the specified working conditions and the current vehicle state parameters, the current vehicle state type is obtained; the state type is one of stable state, critically stable state, and unstable state. Construct a reinforcement learning agent and reward function, and continuously train the agent until it outputs an additional yaw moment ΔM and an active rear wheel steering angle δ. r Each has its own optimal cooperative control coefficient to achieve coordinated control of the electric vehicle; the reward function includes steady-state reward, yaw rate error and centroid sideslip angle error reward, critical steady-state reward and anti-shake reward.
2. The vehicle stability coordination control method based on reinforcement learning according to claim 1, characterized in that, The vehicle's structural parameters, speed, and front wheel steering angle are obtained using a linear two-degree-of-freedom model to determine the desired yaw rate ω. d And the expected centroid side slip angle β d middle: The linear two-degree-of-freedom model is in the following form: Where k1 is the stiffness of the front axle, l2 is the stiffness of the rear axle, β is the sideslip angle of the vehicle's center of gravity, and v x Let ω be the longitudinal vehicle speed, a be the distance from the vehicle's center of gravity to the front axle, b be the distance from the vehicle's center of gravity to the rear axle, and ω be the longitudinal speed. exp The desired yaw rate is the output of the linear two-degree-of-freedom model, where δ is the standard front wheel steering angle, m is the vehicle mass, and a is the yaw rate. y For the lateral acceleration of the vehicle, I z Let be the vehicle's moment of inertia. For ω exp The derivative; The desired yaw rate ω output by the linear two-degree-of-freedom model exp for: Where K is the vehicle stability factor; The desired yaw rate ω d And the expected centroid side slip angle β d for: Where μ is the adhesion coefficient, g is the gravitational acceleration, and β d Let sgn(δ) be the expected centroid sideslip angle, and let sgn(δ) be the judgment function indicating the sign of the front wheel steering angle. When the front wheel steering angle is positive, sgn(δ) is 1, and when the front wheel steering angle is negative, sgn(δ) is -1.
3. The vehicle stability coordination control method based on reinforcement learning according to claim 2, characterized in that, The standard front wheel steering angle δ is the average wheel steering angle of the two front wheels of the vehicle; the wheel steering angle of the two front wheels and the longitudinal vehicle speed v x The vehicle's moment of inertia I is obtained through sensors. z Based on the vehicle's desired torque T d Obtain; the desired torque T of the vehicle d for: Where θ is the accelerator pedal opening, and T imax This represents the peak torque of the i-th hub motor.
4. The vehicle stability coordination control method based on reinforcement learning according to claim 1, characterized in that, The desired yaw rate ω d Actual yaw rate, desired centroid sideslip angle β d The additional yaw moment ΔM and the active rear wheel steering angle δ are obtained by sliding mode control of the actual center of gravity sideslip angle. r Specifically, it includes: Based on the desired yaw rate ω d Based on the actual yaw rate, design the sliding surface of the yaw torque controller to obtain the additional yaw torque ΔM; Based on the expected centroid sideslip angle β d Based on the actual sideslip angle of the center of gravity, the sliding surface of the active rear wheel controller is designed to obtain the active rear wheel steering angle δ. r ; Among them, I z Let c be the vehicle's moment of inertia. ω Here, represents the relative weighting coefficient between error and the rate of change of error; a is the distance from the vehicle's center of gravity to the front axle; b is the distance from the vehicle's center of gravity to the rear axle; and C... f For the stiffness of the vehicle's front wheels, C r The stiffness of the vehicle's rear wheels, v x For longitudinal vehicle speed, and e ω and e β The derivative, e ω For the deviation of the control variable of yaw rate, e β The deviation is the control variable for the centroid sideslip angle, K0 is the rate of approach constant, and s ω is the sliding mode shear function; sat(·) is the saturation function.
5. The vehicle stability coordination control method based on reinforcement learning according to claim 1, characterized in that, Based on the established clustering diagram of vehicle state parameters under set operating conditions and the current vehicle state parameters, the current vehicle state type is obtained. There are a total of 5 cluster centers in the cluster diagram of vehicle state parameters under the given working condition. The cluster in the middle is the stable state, the clusters on the left and right edges of the middle cluster are the unstable states, and the clusters between the stable and unstable states are the critical stable states. The vehicle state parameters include: lateral velocity, lateral acceleration, yaw rate, center of gravity sideslip angle, and tire sideslip angle.
6. The vehicle stability coordination control method based on reinforcement learning according to claim 1, characterized in that, The process involves constructing a reinforcement learning agent and a reward function, and continuously training the agent until it outputs an additional yaw moment ΔM and an active rear wheel steering angle δ. r Their respective optimal collaborative control coefficients include: Define the state s of the agent. t for: s t ={ω,Δω,β,Δβ,a y } Where Δω is the difference between the yaw rate and the actual yaw rate, Δβ is the difference between the desired sideslip angle and the actual sideslip angle, ω is the actual yaw rate, β is the actual sideslip angle, and a y It is lateral acceleration; Define action a t The synergistic relationship between the rear wheel steering angle and the additional yaw moment is expressed as follows: a t ={k,1-k},k∈[0,1] Where k is the active rear wheel steering angle δ r The coordination control coefficient, where 1-k is the coordination coefficient of the additional yaw moment ΔM; Define reward function r t for: r t =a1r s +a2r e +a3r cs +a4r c Where, r s For steady-state rewards, r e As a reward for yaw rate error and center of mass sideslip angle error, r cs For the critical steady-state reward, r c For anti-shake rewards, a1, a2, a3, and a4 are respectively r s r e r cs and r c The reward coefficient.
7. The vehicle stability coordination control method based on reinforcement learning according to claim 6, characterized in that, The steady-state reward r s This includes rewards for the vehicle's current state and rewards for predicted future states, as shown in the following formula: r s =r s1 +r s2 +r s3 +r s4 Where, r s1 For the reward of the car's current state, r s2 r s3 and r s4 t1, t2, t3, and t4 represent the predicted car state rewards for the first, second, and third time points, respectively; Dis1 is the Euclidean distance from the vehicle state parameters to the cluster centers; t1, t2, t3, and t4 are respectively the r... s1 r s2 r s3 and r s4 The adjustment coefficient; q is the cluster category to which the current vehicle belongs; q=1 indicates that the current vehicle is in a stable state, q=2 or q=4 indicates that the current vehicle is in a critically stable state, and q=3 and q=5 indicate that the current vehicle is in an unstable state. r s2 r s3 and r s4 The corresponding clustering categories are obtained based on the predicted vehicle state parameters and the clustering diagram of vehicle state parameters under the set operating conditions; The predicted vehicle state parameters are: based on the current vehicle state parameters and the additional yaw moment ΔM and active rear wheel steering angle δ of the control input. r Based on the linear two-degree-of-freedom model, predict r s2 r s3 and r s4 The vehicle state parameters at the corresponding time.
8. The vehicle stability coordination control method based on reinforcement learning according to claim 6, characterized in that, The yaw rate error and the center of mass sideslip angle error are rewarded r e for: r e =-a ω Give 2 -a β Db 2 +W Where W is a constant, a ω and a β These are the adjustment coefficients for Δω and Δβ, respectively; The critical stable state reward r cs for: Where D1 is the dot product of the vector formed by the vehicle's current state parameters and the parameters of the cluster center to which the vehicle belongs, and the vector formed by the second cluster center to the first cluster center; D1 is the length of the vector formed by the second cluster center to the first cluster center; D2 is the dot product of the vector formed by the vehicle's current state parameters and the parameters of the vehicle's cluster center, and the vector formed by the fourth cluster center to the first cluster center. Let be the length of the vector formed by the fourth cluster center and the first cluster center; n is the adjustment coefficient of the dot product, and k p The adjustment coefficient for the critical steady-state reward; The anti-shake reward r c for: r c =-Dδ r 2 -DM z 2 Where, Δδ r ΔM is the difference between the current active rear wheel steering angle and the previous angle; z The difference between the yaw moment at the current moment and the moment at the previous moment is applied.
Citation Information
Patent Citations
Rear wheel active steering control method and device
CN115092118A
AFS and DYC cooperative control method for eight-wheel distributed electric drive vehicle
CN115431790A