Four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning

Through the four-wheel steering vehicle anti-roll control method with deep reinforcement learning, the problems of strong model dependence, multi-objective conflict and low training efficiency in traditional methods are solved, and robust anti-roll control and real-time requirements under complex operating conditions are achieved, taking into account the driver's intentions.

CN120348281APending Publication Date: 2025-07-22NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510402498.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The traditional vehicle anti-roll control method relies on the vehicle dynamics model, and there are problems such as model mismatch, multi-objective conflict and limited control margins. In addition, reinforcement learning is low in training efficiency and poor real-time performance in vehicle control, making it difficult to integrate driver input signals.

Method used

The four-wheel steering vehicle anti-roll control method based on deep reinforcement learning is adopted. By constructing a reinforcement learning framework driven by hybrid model and data, the layered network structure is separated from stability decision-making and execution control, and combined with virtual and real transfer learning, the coordinated optimization of the four-wheel steering angle and braking force is achieved.

Benefits of technology

It improves the robustness of anti-roll control under complex operating conditions, reduces model dependence, improves training efficiency and generalization capabilities, meets real-time requirements, takes into account the driver's intentions, and reduces human-machine conflicts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120348281A_ABST
    Figure CN120348281A_ABST
Patent Text Reader

Abstract

The invention discloses an anti-rollover control method for a four-wheel steering vehicle based on deep reinforcement learning, and the method achieves the collaborative optimization of a four-wheel steering angle and a braking force through constructing a reinforcement learning frame of hybrid model driving and data driving. A hierarchical network structure is adopted to separate stability decision and execution control, and the generalization ability of the algorithm is improved in combination with virtual-real transfer learning. The anti-rollover control robustness under the complex working condition is remarkably improved, and the method can be applied to an active safety system of an intelligent driving vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle active safety control, and particularly relates to a four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning. Background Art

[0002] With the rapid development of intelligent driving technology, vehicle rollover prevention control has become one of the core challenges in enhancing active safety. Traditional methods mainly rely on rule-based controllers (such as PID) or model predictive control (MPC), which calculate stability control quantities through predefined vehicle dynamics models. However, traditional methods have the following problems: First, traditional methods are highly dependent on vehicle dynamics models, and in practical applications, factors such as load changes and tire pressure fluctuations can lead to model mismatches. Especially in the non-linear region of tires (such as when approaching the adhesion limit), the control error increases significantly, even leading to misjudgments. For example, in the scenario of emergency obstacle avoidance on a low-adhesion road surface, traditional MPC may cause improper braking force distribution due to the failure to correct the tire force model in real time, increasing the risk of rollover. Second, most traditional methods only focus on a single index (such as roll angle or load transfer ratio), and it is difficult to balance stability and maneuverability. For example, the electronic stability program (ESP) can suppress roll through unilateral braking intervention, but it will cause vehicle speed loss and steering response lag, conflicting with the driver's intention. Third, traditional methods usually only control front-wheel steering or rear-wheel follow-up steering, and do not fully utilize the potential of four-wheel independent steering. In high-speed lane changes or complex working conditions (such as braking in a curve), the lack of dynamic coordinated distribution of four-wheel lateral forces results in limited control margin. These problems are difficult to be effectively solved in traditional technical solutions, and reinforcement learning has been tried to be applied to vehicle control due to its adaptive ability, but there are still certain bottlenecks: First, model-free reinforcement learning requires a large amount of interaction data for training, and the high-dimensional state space of vehicle dynamics (such as yaw rate, tire slip rate, etc.) leads to slow policy convergence and low training efficiency; Second, the joint control of four-wheel steering angles and braking forces needs to generate high-dimensional actions in a continuous space, and existing network architectures are prone to falling into local optima and difficult to meet the real-time requirement (usually a decision delay of <50 ms); Finally, existing reinforcement learning schemes do not effectively integrate driver input signals (such as steering wheel angle and torque), which may cause interference between machine decisions and human operations. For example, when the driver makes an emergency steering, if the RL controller does not adjust the strategy in time, it may cause overcorrection. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning to address the deficiencies in the background art. By constructing a hybrid model-driven and data-driven reinforcement learning framework, the collaborative optimization of the four-wheel steering angle and braking force is completed, and a hierarchical network structure is adopted to separate stability decision-making and execution control. Combining virtual-real transfer learning to improve the generalization ability of the algorithm, the robust dynamic rollover prevention control of four-wheel steering vehicles under complex working conditions is realized.

[0004] The present invention adopts the following technical solutions to solve the above technical problems:

[0005] A four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning, comprising the following steps:

[0006] Step 1), obtain vehicle state data and environmental information in real time and construct a state vector; the state vector at time t is where φ is the roll angle of the vehicle, is the roll angular velocity of the vehicle, a y is the lateral acceleration of the vehicle, δ dirver is the steering wheel angle input by the driver, μ road is the road surface adhesion coefficient; LTR is the lateral load transfer ratio of the vehicle, where F z,r , F z,l are the vertical loads of the right and left wheels; v x is the longitudinal vehicle speed;

[0007] Step 2), extract temporal features according to the state vectors at the current moment and the previous n moments; obtain the state vectors s t , s t-1 , …, s t-n at time t, t - 1, …, t - n respectively, where n is the preset sliding window length, and input them into the TCN network to extract the temporal feature f t = TCN(s t , s t-1 ,..., s t-n );

[0008] Step 3), establish and train a hierarchical reinforcement learning decision network, and use the trained hierarchical reinforcement learning decision network to calculate the target steering angles and target braking torques of the four wheels according to the temporal feature f t ;

[0009] The hierarchical reinforcement learning decision network includes an upper-layer decision network and a lower-layer execution network;

[0010] The upper-layer decision network adopts a network structure based on the PPO algorithm and is used to calculate the stability target, and its input is the temporal feature ft The output is the desired roll angle φ for the stability target des and the desired yaw rate

[0011] The lower - layer execution network adopts a network structure based on the SAC algorithm to calculate the target steering angles and target braking torques of the four wheels, and its input is f t 、φ des 、 The output is the target steering angles and target braking torques of the four wheels;

[0012] Step 4), compare the LTR with the preset maximum load transfer rate threshold;

[0013] Step 4.1), if the LTR is greater than the preset maximum load transfer rate threshold, trigger the MPC fallback controller;

[0014] Step 4.2), if the LTR is less than or equal to the preset maximum load transfer rate threshold;

[0015] Step 4.2.1), compare the target steering angles of the four wheels with the preset minimum steering angle threshold θ1 respectively. If there is a wheel with a target steering angle less than θ1, modify the target steering angle of this wheel to θ1;

[0016] Step 4.2.2), compare the target steering angles of the four wheels with the preset maximum steering angle threshold θ2 respectively. If there is a wheel with a target steering angle greater than θ2, modify the target steering angle of this wheel to θ2;

[0017] Step 4.2.3), control the four wheels of the vehicle through the SBW system so that the actual steering angles of the four wheels are equal to the target steering angles of the four wheels respectively;

[0018] Step 4.2.4), compare the target braking torques of the four wheels with the preset minimum torque threshold N1 respectively. If there is a wheel with a target braking torque less than N1, modify the target braking torque of this wheel to L1;

[0019] Step 4.2.5), compare the target braking torques of the four wheels with the preset maximum torque threshold N2 respectively. If there is a wheel with a target braking torque greater than N2, modify the target braking torque of this wheel to N2;

[0020] Step 4.2.6), control the four wheels of the vehicle through the ESP system so that the actual braking torques of the four wheels are equal to the target braking torques of the four wheels respectively.

[0021] As a further optimization solution of the four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning of the present invention, the training steps of the hierarchical reinforcement learning decision network are as follows:

[0022] Step 3.1), construct a simulation environment based on a high-fidelity vehicle dynamics model;

[0023] Step 3.2), train the upper-layer decision network;

[0024] Step 3.2.1), pre-train the upper-layer decision network. The pre-training is dominated by the model drive. Based on the vehicle dynamics model, the vehicle rollover dynamics model, and the tire magic formula, a virtual environment is constructed to guide the upper-layer decision network to learn the roll stability control strategy according to the vehicle dynamics principle, and the pre-trained upper-layer decision network is obtained;

[0025] Step 3.2.2), transfer the pre-trained upper-layer decision network to the simulation environment for training. The training data includes 70% simulation data and 30% historical real-vehicle data. Fine-tune the upper-layer decision network to improve its robustness to data beyond the model range;

[0026] Step 3.3), place the lower-layer execution network in the simulation environment for pre-training, and guide the lower-layer execution network to learn how to convert the stability target of the upper-layer decision network into the target steering angles and target braking torques of the four wheels. The training data at this stage includes 50% simulation data and 50% historical real-vehicle data. Simulate the real road conditions through noise injection and sensor delay, and inject extreme working conditions generated by real road data for training to optimize the control robustness in complex scenarios;

[0027] Step 3.4), jointly fine-tune the upper-layer decision network and the lower-layer execution network;

[0028] Step 3.4.1), fix the parameters of the upper-layer decision network and optimize the parameters of the lower-layer execution network;

[0029] Step 3.4.2), unfreeze the parameters of the upper-layer decision network, and jointly optimize the parameters of the upper-layer decision network and the lower-layer execution network through end-to-end gradient backpropagation;

[0030] Step 3.5), the lower-layer execution network performs online learning and updating;

[0031] After the real vehicle is deployed, collect real vehicle data through the edge computing device, and dynamically update the parameters of the lower-layer execution network according to the collected real vehicle data. At this time, the parameters of the upper-layer decision network are frozen.

[0032] As a further optimization solution of the four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning of the present invention, the TCN network in step 2) includes an input layer, a first residual block, a second residual block, a third residual block, and an output layer connected in sequence. The input layer is used to receive a time series state vector s of length n t 、s t-1 、…、s t-n , with a dimension of n×d s , d s = 8; the first residual block, the second residual block, and the third residual block have the same structure, each containing 2 layers, and each layer adopts a structure of causal dilated convolution, weight normalization, ReLU activation, and Dropout connected in sequence. Moreover, the dilation coefficients of the first residual block, the second residual block, and the third residual block are 1, 2, and 4 in sequence, and the kernel size is 3; the output layer adopts a structure of global average pooling and fully connected layer connected in sequence, and is used to output a feature vector

[0033] As a further optimization solution of the four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning of the present invention, the upper-level decision network includes a policy network and a value network;

[0034] The policy network includes an input layer, a first fully connected layer, a second fully connected layer, a third fully connected layer, and an output layer connected in sequence; the input layer is used to input f t ; the first fully connected layer, the second fully connected layer, and the third fully connected layer all adopt Swish activation, and the number of neurons is 256, 128, and 64 respectively; the output layer generates a policy distribution through Softmax;

[0035] The value network includes an input layer, a fourth fully connected layer, a fifth fully connected layer, a sixth fully connected layer, and an output layer connected in sequence. The input layer is used to input f t ; the fourth fully connected layer, the fifth fully connected layer, and the sixth fully connected layer all adopt Swish activation, and the number of neurons is 256, 128, and 64 respectively; the output layer uses linear activation to output a state value scalar.

[0036] As a further optimization solution of the four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning of the present invention, the lower-level execution network includes an Actor network and a Critic network;

[0037] The Actor network includes an input layer, a seventh fully connected layer, an eighth fully connected layer, a ninth fully connected layer, and an output layer connected in sequence. The input layer is used to input f t 、φ des 、 The total dimension is 34; the seventh, eighth, and ninth fully connected layers all use Tanh activation, and the number of neurons is 128, 64, and 32 respectively; the output layer is used to output the action mean and variance;

[0038] The Critic network includes an input layer, an LSTM layer, a tenth fully connected layer, and an output layer connected in sequence. The input layer is used to input f t , φ des , the target steering angles and target braking torques of the current four wheels, with a total dimension of 42; the LSTM layer is used to process temporal dependencies; the tenth fully connected layer all uses Tanh activation; the output layer is used to output the current state-action value score Q(s,a).

[0039] As a further optimization scheme of the four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning of the present invention, the method of migrating the pre-trained upper decision-making network to the simulation environment in step 3.2.2) is parameter freezing and feature space alignment.

[0040] As a further optimization scheme of the four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning of the present invention, the roll dynamics equation of the vehicle roll dynamics model in step 3.2.1) is as follows:

[0041]

[0042] Among them, I x is the moment of inertia of the sprung mass of the vehicle about the roll center, m s is the sprung mass of the vehicle, h is the height of the vehicle's center of mass, k φ is the equivalent roll stiffness of the vehicle, c φ is the equivalent roll damping of the vehicle, λ is the correction coefficient, and b is the offset coefficient.

[0043] As a further optimization scheme of the four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning of the present invention, when pre-training the upper decision-making network in step 3.2.1) and migrating the pre-trained upper decision-making network to the simulation environment for training in step 3.2.2), the reward function is set as follows:

[0044]

[0045] Among them, is the model-driven term, k1 and k2 are the weight coefficients of the roll angle and yaw rate respectively. This reward function focuses on suppressing the oscillations of the roll angle and yaw rate. is the data-driven term, k3 is the human-machine cooperation weight coefficient, and δ actual is the actual steering wheel angle of the vehicle.

[0046] As a further optimization solution for the four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning of the present invention, the dynamic weight α(t) in the reward function in step 3.2.1) is set to α(t) = 1.0, and the dynamic weight α(t) in the reward function in step 3.2.2) is set to α(t) = 0.5×μ road .

[0047] As a further optimization solution for the four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning of the present invention, the reward function during the pre-training of the lower-level execution network in step 3.3) is set as follows:

[0048]

[0049] Wherein, is the model-driven term, LTR safe = 0.6 is the critical safety load transfer coefficient, (Δδ i / Δt) is the instantaneous change rate of the four-wheel steering angle, and k4 and k5 are respectively the load transfer and sudden steering penalty weight coefficients, is the data-driven term, and k6 and k7 are respectively the energy consumption penalty coefficient and the upper-layer target tracking weight coefficient.

[0050] The present invention can achieve dynamic rollover prevention control of four-wheel steering vehicles under various complex working conditions. By deeply integrating technologies such as model-driven, data-driven, hierarchical control, and transfer learning, it solves the core pain points of strong model dependence, multi-objective conflicts, and poor generalization ability in traditional rollover prevention control, and has significant technological advancement and engineering implementation value, which is specifically reflected in the following aspects:

[0051] 1. The present invention adopts a hybrid model-driven and data-driven reinforcement learning framework, and realizes the complementarity of physical prior and data-driven through a hybrid reward function, action space constraint, and staged training. It not only avoids the "blind exploration" of pure data-driven, greatly improves the training efficiency, but also can compensate for the errors caused by model simplification, and greatly improves the control effect of the rollover prevention controller in the non-linear region (such as the tire force saturation region). Compared with the traditional MPC controller, it improves the control effect of vehicle rollover prevention under non-linear working conditions while reducing the model dependence.

[0052] 2. The present invention adopts a hierarchical reinforcement learning framework and uses a phased training and collaborative optimization strategy for training. While ensuring the stability of the upper-layer stability control strategy, it can quickly adapt to the dynamic changes of the real vehicle, and greatly improves the training efficiency. By designing a human-machine collaboration and safety guarantee control mechanism (weight fusion and guaranteed MPC controller) in the reward function and control process, the driving intention of the driver is taken into account while ensuring the safety of vehicle control. Compared with the traditional MPC controller, the degree of conflict between human and machine driving rights during real vehicle control is greatly reduced.

[0053] 3. The present invention applies transfer learning technology throughout the training process, reduces the differences between the model data and simulation data, and between the simulation data and real vehicle data during training, improves the generalization ability of the system for unseen data, and can adapt to the driving styles of different drivers through online learning. In addition, the parallel processing of TCN significantly reduces the decision-making delay, meeting the real-time requirements of real vehicle control. Compared with the traditional MPC controller, the patent solution has outstanding advantages in generalization and adaptability, meets the real-time requirements, and has significant technological advancement and engineering implementation value. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 is the flow chart of the control algorithm of the present invention;

[0055] Figure 2 is the framework diagram of the reinforcement learning training of the present invention;

[0056] Figure 3 is the schematic diagram of the four-wheel steering angle distribution strategy of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0057] The technical solution of the present invention will be further described in detail below with reference to the drawings:

[0058] The present invention can be implemented in many different forms and should not be considered limited to the embodiments described herein. On the contrary, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present invention to those skilled in the art. In the drawings, components are enlarged for clarity.

[0059] As Figure 1 shown, the present invention discloses a four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning, which is characterized by including the following steps:

[0060] Step 1), obtain vehicle state data and environmental information in real time, and construct a state vector; the state vector at time t is where φ is the roll angle of the vehicle, is the roll angle speed of the vehicle, a y is the lateral acceleration of the vehicle, δ dirverThe steering wheel angle input by the driver, μ road is the road surface adhesion coefficient; LTR is the vehicle lateral load transfer ratio, wherein, F z,r 、F z,l are the right wheel vertical load and the left wheel vertical load; v x is the longitudinal vehicle speed;

[0061] Step 2), extract the time series features according to the state vectors at the current moment and the previous n moments; obtain the state vectors s t 、s t-1 、…、s t-n at the moments t, t - 1, …, t - n respectively, n is the preset sliding window length, and input them into the TCN network to extract the time series feature f t =TCN(s t ,s t-1 ,...,s t-n );

[0062] The TCN network in the step 2) includes an input layer, a first residual block, a second residual block, a third residual block and an output layer connected in sequence. The input layer is used to receive the time series state vectors s t 、s t-1 、…、s t-n with the dimension of n×d s , d s =8; the first residual block, the second residual block and the third residual block have the same structure and both contain 2 layers. Each layer adopts a structure of causal dilated convolution, weight normalization, ReLU activation and Dropout connected in sequence, and the dilation coefficients of the first residual block, the second residual block and the third residual block are 1, 2, 4 in sequence, and the kernel size is 3; the output layer adopts a structure of global average pooling and fully connected layer connected in sequence to output the feature vector

[0063] Step 3), establish and train a hierarchical reinforcement learning decision network, and use the trained hierarchical reinforcement learning decision network to calculate the target steering angles and target braking torques of the four wheels according to the time series feature f t ;

[0064] The hierarchical reinforcement learning decision network includes an upper - layer decision network and a lower - layer execution network;

[0065] The upper - layer decision network adopts a network structure based on the PPO algorithm and is used to calculate the stability target. Its input is the time series feature f t , and the output is the expected roll angle φ des of the stability target and the expected yaw rate

[0066] The upper - layer decision - making network includes a policy network and a value network;

[0067] The policy network includes an input layer, a first fully - connected layer, a second fully - connected layer, a third fully - connected layer, and an output layer connected in sequence; the input layer is used to input f t ; The first fully - connected layer, the second fully - connected layer, and the third fully - connected layer all use the Swish activation, and the number of neurons is 256, 128, and 64 respectively; the output layer generates a policy distribution through Softmax;

[0068] The value network includes an input layer, a fourth fully - connected layer, a fifth fully - connected layer, a sixth fully - connected layer, and an output layer connected in sequence, and the input layer is used to input f t ; The fourth fully - connected layer, the fifth fully - connected layer, and the sixth fully - connected layer all use the Swish activation, and the number of neurons is 256, 128, and 64 respectively; the output layer uses linear activation to output the state - value scalar.

[0069] The lower - layer execution network adopts a network structure based on the SAC algorithm to calculate the target steering angles and target braking torques of the four wheels, and its input is f t , φ des , The output is the target steering angles and target braking torques of the four wheels;

[0070] The lower - layer execution network includes an Actor network and a Critic network;

[0071] The Actor network includes an input layer, a seventh fully - connected layer, an eighth fully - connected layer, a ninth fully - connected layer, and an output layer connected in sequence, and the input layer is used to input f t , φ des , with a total dimension of 34; the seventh fully - connected layer, the eighth fully - connected layer, and the ninth fully - connected layer all use Tanh activation, and the number of neurons is 128, 64, and 32 respectively; the output layer is used to output the action mean and variance;

[0072] The Critic network includes an input layer, an LSTM layer, a tenth fully - connected layer, and an output layer connected in sequence, and the input layer is used to input f t , φ des , the current target steering angles and target braking torques of the four wheels, with a total dimension of 42; the LSTM layer is used to process temporal dependencies; the tenth fully - connected layer all uses Tanh activation; the output layer is used to output the current state - action value score Q(s,a).

[0073] Step 4), compare the LTR with a preset maximum load transfer rate threshold;

[0074] Step 4.1), if the LTR is greater than the preset maximum load transfer rate threshold, trigger the MPC fallback controller;

[0075] Step 4.2), if the LTR is less than or equal to the preset maximum load transfer rate threshold;

[0076] Step 4.2.1), compare the target steering angles of the four wheels with the preset minimum steering angle threshold θ1 respectively. If there is a wheel with a target steering angle less than θ1, modify the target steering angle of this wheel to θ1;

[0077] Step 4.2.2), compare the target steering angles of the four wheels with the preset maximum steering angle threshold θ2 respectively. If there is a wheel with a target steering angle greater than θ2, modify the target steering angle of this wheel to θ2;

[0078] Step 4.2.3), control the four wheels of the vehicle through the SBW system so that the actual steering angles of the four wheels are equal to the target steering angles of the four wheels respectively;

[0079] Step 4.2.4), compare the target braking torques of the four wheels with the preset minimum torque threshold N1 respectively. If there is a wheel with a target braking torque less than N1, modify the target braking torque of this wheel to L1;

[0080] Step 4.2.5), compare the target braking torques of the four wheels with the preset maximum torque threshold N2 respectively. If there is a wheel with a target braking torque greater than N2, modify the target braking torque of this wheel to N2;

[0081] Step 4.2.6), control the four wheels of the vehicle through the ESP system so that the actual braking torques of the four wheels are equal to the target braking torques of the four wheels respectively.

[0082] As Figure 2 shown, the training steps of the hierarchical reinforcement learning decision network are as follows:

[0083] Step 3.1), construct a simulation environment based on a high-fidelity vehicle dynamics model (such as CarSim / Simulink);

[0084] Step 3.2), train the upper-layer decision network;

[0085] Step 3.2.1), pre-train the upper-layer decision network. The pre-training is dominated by model driving. Based on the vehicle dynamics model, vehicle roll dynamics model, and tire magic formula, construct a virtual environment to guide the upper-layer decision network to learn the roll stability control strategy according to the vehicle dynamics principle, and obtain the pre-trained upper-layer decision network;

[0086] Step 3.2.2), transfer the pre-trained upper-level decision-making network to the simulation environment for training. The training data includes 70% simulation data and 30% historical real vehicle data. Fine-tune the upper-level decision-making network to improve its robustness to data beyond the model range;

[0087] In step 3.2.2), the method of transferring the pre-trained upper-level decision-making network to the simulation environment is parameter freezing and feature space alignment.

[0088] In the said step 3.2.1), the roll dynamics equation of the vehicle rollover dynamics model is as follows:

[0089]

[0090] Among them, I x is the moment of inertia of the sprung mass of the vehicle about the roll center, m s is the sprung mass of the vehicle, h is the height of the vehicle's center of mass, k φ is the equivalent roll stiffness of the vehicle, c φ is the equivalent roll damping of the vehicle, λ is the correction coefficient, and b is the offset coefficient.

[0091] When pre-training the upper-level decision-making network in step 3.2.1) and transferring the pre-trained upper-level decision-making network to the simulation environment for training in step 3.2.2), the reward function is set as follows:

[0092]

[0093] Among them, is the model-driven term, k1 and k2 are the weight coefficients of the roll angle and yaw rate respectively. This reward function focuses on suppressing the oscillation of the roll angle and yaw rate. is the data-driven term, k3 is the human-machine collaboration weight coefficient, and δ actual is the actual steering wheel angle of the vehicle. This reward function focuses on rewarding the control strategy that matches the driver's steering intention.

[0094] In step 3.2.1), the dynamic weight α(t) in the reward function is set as α(t) = 1.0 (completely model-dependent), and in step 3.2.2), the dynamic weight α(t) in the reward function is set as α(t) = 0.5×μ road .

[0095] Furthermore, in the pre-training stage, initialize the policy network with the expert trajectory (such as the MPC optimal control sequence) generated by the vehicle dynamics model to accelerate convergence.

[0096] Step 3.3), place the lower - layer execution network in the simulation environment for pre - training, guiding the lower - layer execution network to learn how to convert the stability target of the upper - layer decision network into the target steering angles and target braking torques of the four wheels. The training data at this stage includes 50% simulation data and 50% historical real - vehicle data. Simulate the real road conditions through noise injection and sensor delay, and inject extreme working conditions (emergency lane change, unilateral low adhesion, strong cross - wind) generated from real road - collected data for training to optimize the control robustness in complex scenarios;

[0097] When pre - training the lower - layer execution network in Step 3.3), the reward function is set as:

[0098]

[0099] Among them, is the model - driven term, LTR safe = 0.6 is the critical safety load transfer coefficient, (Δδ i / Δt) is the instantaneous change rate of the four - wheel steering angles, k4 and k5 are the load transfer and sudden - steering penalty weight coefficients respectively. This term focuses on suppressing load transfer and sudden steering, is the data - driven term, k6 and k7 are the energy - consumption penalty coefficient and the upper - layer target - tracking weight coefficient respectively. This term focuses on rewarding actions that can reduce energy consumption and track the upper - layer target action strategy.

[0100] When pre - training the lower - layer execution network in Step 3.3), the dynamic weight β(t) in the reward function is set as β(t)=0.3 + 0.4×β s , β s is the sideslip angle of the vehicle's center of mass. The larger the sideslip angle, the higher the model weight.

[0101] Step 3.4), jointly fine - tune the upper - layer decision network and the lower - layer execution network;

[0102] Step 3.4.1), fix the parameters of the upper - layer decision network and optimize the parameters of the lower - layer execution network;

[0103] Step 3.4.2), unfreeze the parameters of the upper - layer decision network and jointly optimize the parameters of the upper - layer decision network and the lower - layer execution network through end - to - end gradient backpropagation;

[0104] Step 3.5), the lower - layer execution network performs online learning and updates;

[0105] After the real vehicle is deployed, real vehicle data is collected through edge computing devices, and the parameters of the lower-level execution network are dynamically updated according to the collected real vehicle data. At this time, the parameters of the upper-level decision-making network are frozen; because the upper-level network is responsible for generating the global stability target and needs to maintain policy stability, while the lower-level network needs to quickly adapt to the dynamic changes of the real vehicle (such as tire pressure fluctuations and driver habits).

[0106] The update formula for the online learning update of the lower-level execution network is as follows:

[0107]

[0108] where θ is the parameter of the lower-level execution network, η is the dynamic learning rate, which is adjusted according to the data confidence,[[]] is the reinforcement learning loss function, which includes the policy gradient and the value function error,[[]] is the state prediction loss based on the dynamic model, and λ is the model loss weight.[[]]

[0109] Furthermore, a human-machine collaboration enhancement term R human =-k8||T driver -T expected || 2 and an environment adaptation term R adapt =-k9||μ estimated -μ actual || 2 are added to the reward function in the online learning stage. Here, T driver and T expected are the actual driver's steering wheel torque and the desired steering wheel torque, and μ estimated and μ actual are the estimated road adhesion coefficient and the actual road adhesion coefficient. The purpose of adding the above two reward functions is to perceive the driver's steering torque intention and correct the road adhesion estimation error in real time;[[]]

[0110] As Figure 3 shown, when generating the four-wheel corner control command, the lower-level execution network follows the strategy of "same direction at low speed, opposite direction at high speed" to generate the equivalent front and rear wheel corner, and calculates the four-wheel corner adjustment amount based on the equivalent front and rear wheel corner; the relationship formula between the equivalent front and rear wheel corners is as follows:

[0111]

[0112] In the formula, δ r ′ is the equivalent rear wheel corner, δ′ f is the equivalent front wheel corner, ρ is the steering gain coefficient, λ1 is the corner distribution ratio coefficient, and λ2 is the yaw rate compensation coefficient, where ρ and λ2 are obtained by real vehicle calibration;[[]]

[0113] λ1 is calculated by the following formula:

[0114]

[0115] m is the vehicle mass, l f and l r are the distances from the vehicle's center of mass to the front and rear axles, γ1 and γ2 are the equivalent cornering stiffnesses of the vehicle, and L is the vehicle's axle length.

[0116] Those skilled in the art of this technology can understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with their meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as such here.

[0117] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning, characterized in that It includes the following steps: Step 1), obtain vehicle status data and environmental information in real time and construct a state vector; the state vector at time t is where φ is the roll angle of the vehicle, φ is the angular velocity of the vehicle's roll, a y is the lateral acceleration of the vehicle, δ dirver is the steering wheel angle input by the driver, μ road is the road surface adhesion coefficient; LTR is the vehicle's lateral load transfer ratio, where F z,r 、F z,l are the vertical loads on the right and left wheels; v x is the longitudinal vehicle speed; Step 2), extract temporal features according to the state vectors at the current moment and the previous n moments; respectively obtain the state vectors s at time t, t-1,..., t-n t , s t-1 , …, s t-n , where n is the preset sliding window length, and input them into the TCN network to extract the temporal feature f t = TCN(s t , s t-1 ,..., s t-n ); Step 3), establish and train a hierarchical reinforcement learning decision network, and use the trained hierarchical reinforcement learning decision network to calculate the target steering angles and target braking torques of the four wheels according to the temporal feature f t ​ The hierarchical reinforcement learning decision-making network includes an upper-layer decision-making network and a lower-layer execution network; The upper-layer decision-making network adopts a network structure based on the PPO algorithm and is used to calculate the stability target. Its input is the time-series feature f t , and the output is the expected roll angle φ des of the stability target and the expected yaw rate The lower - layer execution network adopts a network structure based on the SAC algorithm to calculate the target steering angles and target braking torques of the four wheels, and its input is f t , φ des , The output is the target steering angles and target braking torques of the four wheels; Step 4), compare the LTR with a preset maximum payload transfer rate threshold; Step 4.1), if the LTR is greater than the preset maximum payload transfer rate threshold, trigger the MPC fallback controller; Step 4.2), if the LTR is less than or equal to the preset maximum payload transfer rate threshold; Step 4.2.1), compare the target steering angles of the four wheels with a preset minimum steering angle threshold θ1 respectively. If there is a wheel with a target steering angle less than θ1, modify the target steering angle of this wheel to θ1; Step 4.2.2), compare the target steering angles of the four wheels with a preset maximum steering angle threshold θ2 respectively. If there is a wheel with a target steering angle greater than θ2, modify the target steering angle of this wheel to θ2; Step 4.2.3), control the four wheels of the vehicle through the SBW system so that the actual steering angles of the four wheels are equal to the target steering angles of the four wheels respectively; Step 4.2.4), compare the target braking torques of the four wheels with a preset minimum torque threshold N1 respectively. If there is a wheel with a target braking torque less than N1, modify the target braking torque of this wheel to L1; Step 4.2.5), compare the target braking torques of the four wheels with a preset maximum torque threshold N2 respectively. If there is a wheel with a target braking torque greater than N2, modify the target braking torque of this wheel to N2; Step 4.2.6), control the four wheels of the vehicle through the ESP system so that the actual braking torques of the four wheels are equal to the target braking torques of the four wheels respectively.

2. The four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning according to claim 1, characterized in that The training steps of the hierarchical reinforcement learning decision-making network are as follows: Step 3.1), construct a simulation environment based on a high-fidelity vehicle dynamics model; Step 3.2), train the upper-layer decision-making network; Step 3.2.1), pre-train the upper-layer decision-making network. The pre-training is dominated by model driving. Based on the vehicle dynamics model, the vehicle roll dynamics model, and the tire magic formula, a virtual environment is constructed to guide the upper-layer decision-making network to learn the roll stability control strategy according to the vehicle dynamics principle, and obtain the pre-trained upper-layer decision-making network; Step 3.2.2), transfer the pre-trained upper-layer decision-making network to the simulation environment for training. The training data includes 70% simulation data and 30% historical real-vehicle data. Fine-tune the upper-layer decision-making network to improve its robustness to data beyond the model range; Step 3.3), place the lower-layer execution network in the simulation environment for pre-training, guiding the lower-layer execution network to learn how to convert the stability target of the upper-layer decision-making network into the target steering angles and target braking torques of the four wheels. The training data at this stage includes 50% simulation data and 50% historical real-vehicle data. Simulate the real road conditions through noise injection and sensor delay, and inject extreme working conditions generated by real road data for training to optimize the control robustness in complex scenarios; Step 3.4), jointly fine-tune the upper-layer decision-making network and the lower-layer execution network; Step 3.4.1), fix the parameters of the upper-level decision-making network and optimize the parameters of the lower-level execution network; Step 3.4.2), unfreeze the parameters of the upper-level decision-making network, and jointly optimize the parameters of the upper-level decision-making network and the lower-level execution network through end-to-end gradient backpropagation; Step 3.5), the lower-level execution network performs online learning and update; After in-vehicle deployment, collect in-vehicle data through edge computing devices, and dynamically update the parameters of the lower-level execution network according to the collected in-vehicle data. At this time, the parameters of the upper-level decision-making network are frozen.

3. The four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning according to claim 1, wherein The TCN network in step 2) includes an input layer, a first residual block, a second residual block, a third residual block, and an output layer connected in sequence. The input layer is used to receive the time series state vector s of length n t , s t-1 , …, s t-n , with a dimension of n×d s , d s = 8; the first residual block, the second residual block, and the third residual block have the same structure, each containing 2 layers, and each layer adopts a structure of causal dilated convolution, weight normalization, ReLU activation, and Dropout connected in sequence. The dilation coefficients of the first residual block, the second residual block, and the third residual block are 1, 2, and 4 in sequence, and the kernel size is 3; the output layer adopts a structure of global average pooling and fully connected layer connected in sequence to output the feature vector 4. The four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning according to claim 1, wherein, The upper-level decision-making network includes a policy network and a value network; The policy network includes an input layer, a first fully connected layer, a second fully connected layer, a third fully connected layer, and an output layer connected in sequence; the input layer is used to input f t ; the first fully connected layer, the second fully connected layer, and the third fully connected layer all use the Swish activation, and the number of neurons is 256, 128, and 64 respectively; the output layer generates a policy distribution through Softmax; The value network includes an input layer, a fourth fully-connected layer, a fifth fully-connected layer, a sixth fully-connected layer, and an output layer that are connected in sequence. The input layer is used to input f t ; The fourth fully-connected layer, the fifth fully-connected layer, and the sixth fully-connected layer all use Swish activation, and the number of neurons is 256, 128, and 64 respectively; The output layer uses linear activation to output the state value scalar.

5. The four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning according to claim 1, characterized in that The lower-level execution network includes an Actor network and a Critic network; The Actor network includes an input layer, a seventh fully connected layer, an eighth fully connected layer, a ninth fully connected layer, and an output layer connected in sequence. The input layer is used to input f t , φ des , with a total dimension of 34. The seventh fully connected layer, the eighth fully connected layer, and the ninth fully connected layer all use Tanh activation, and the number of neurons is 128, 64, and 32 respectively. The output layer is used to output the action mean and variance; The Critic network includes an input layer, an LSTM layer, a tenth fully connected layer, and an output layer connected in sequence. The input layer is used to input f t , φ des , the target steering angles and target braking torques of the current four wheels, with a total dimension of 42; the LSTM layer is used to process temporal dependencies; The tenth fully connected layer all uses Tanh activation; the output layer is used to output the current state-action value score Q(s,a).

6. The four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning according to claim 2, characterized in that, The method of migrating the pre-trained upper-level decision-making network to the simulation environment in Step 3.2.2) is parameter freezing and feature space alignment.

7. The four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning according to claim 2, wherein The roll dynamics equation of the vehicle roll dynamics model in Step 3.2.1) is as follows: Among them, I x is the moment of inertia of the sprung mass of the vehicle about the roll center, m s is the sprung mass of the vehicle, h is the height of the vehicle's center of mass, k φ is the equivalent roll stiffness of the vehicle, c φ is the equivalent roll damping of the vehicle, λ is the correction coefficient, and b is the offset coefficient.

8. The four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning according to claim 7, characterized in that, When pre-training the upper-level decision-making network in Step 3.2.1) and training the pre-trained upper-level decision-making network migrated to the simulation environment in Step 3.2.2), the reward function is set as follows: Among them, is the model-driven term, where k1 and k2 are the roll angle and yaw rate weight coefficients respectively. This reward function focuses on suppressing the oscillations of the roll angle and yaw rate. is the data-driven term, k3 is the human-machine collaboration weight coefficient, and δ actual is the actual steering wheel angle of the vehicle.

9. The four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning according to claim 8, wherein In step 3.2.1), the dynamic weight α(t) in the reward function is set to α(t) = 1.0, and in step 3.2.2), the dynamic weight α(t) in the reward function is set to α(t) = 0.5×μ road .

10. The four-wheel steering vehicle rollover prevention control method based on deep reinforcement learning according to claim 8, wherein, The reward function when pre-training the lower-level execution network in Step 3.3) is set to: Among them, is the model-driven term, LTR safe = 0.6 is the critical safety load transfer coefficient, (Δδ i / Δt) is the instantaneous change rate of the four-wheel steering angles, k4 and k5 are the load transfer and sudden steering penalty weight coefficients respectively, is the data-driven term, k6 and k7 are the energy consumption penalty coefficient and the upper-layer target tracking weight coefficient respectively.

Citation Information

Cited By

  • Vehicle rollover prevention control method and device based on angle module, equipment and medium

    CN120886817A

  • Wheeled robot control method, device and equipment and storage medium

    CN121535741A