Limit collision avoidance control method and system for four-wheel steering vehicle
Through the four-wheel steering vehicle extreme collision avoidance control method, combined with reinforcement learning and model prediction control, the steering ratio is dynamically adjusted, and the front and rear wheel steering angles are optimized, which solves the problem of insufficient dynamics of four-wheel steering in extreme collision avoidance scenarios, and improves the vehicle's collision avoidance ability and stability.
Patent Information
- Application Number
- CN202510590788.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-22
AI Technical Summary
The existing four-wheel steering control strategy has insufficient dynamic utilization ability in extreme collision avoidance scenarios, especially when combined with intelligent control methods such as reinforcement learning, it has failed to fully improve collision avoidance performance.
A four-wheel steering vehicle extreme collision avoidance control method is constructed, through real-time acquisition of vehicle status and environmental information, combined with reinforcement learning and model prediction control, dynamically adjust the front and rear wheel steering ratio, establish a two-degree of freedom dynamic model, optimize the front and rear wheel steering angles, and realize real-time tracking of the extreme collision avoidance trajectory.
It significantly improves the maximum passing speed of smart vehicles in extreme collision avoidance scenarios, improves emergency collision avoidance capabilities, and improves the maneuverability and stability of the vehicle.
Smart Images

Figure CN120348290A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of active vehicle safety, and particularly to a method and system for extreme collision avoidance control of a four-wheel steering vehicle. Background Art
[0002] With the rapid development of autonomous driving technology, the active safety collision avoidance ability of vehicles under extreme working conditions has received increasing attention. Among them, four-wheel steering technology has gradually become one of the research hotspots due to its outstanding advantages in improving vehicle maneuverability and stability. Through a systematic analysis of existing academic papers, patent documents and related resources, the background art of four-wheel steering in the field of vehicle collision avoidance is mainly reflected in the following aspects:
[0003] First, the technical advantages of four-wheel steering systems and their applications in actual working conditions have received extensive attention. Compared with traditional two-wheel steering systems, four-wheel steering systems can significantly reduce the turning radius of vehicles and improve the stability of vehicles when driving at high speeds, making them more suitable for emergency collision avoidance scenarios. For example, existing research has shown that when a vehicle equipped with a four-wheel steering system performs an emergency obstacle avoidance, compared with a vehicle with only front-wheel steering, a later collision avoidance reaction time can be achieved, and the safety can be significantly improved. Therefore, four-wheel steering technology has gradually been applied to scenarios such as emergency collision avoidance of autonomous vehicles and avoidance of dynamic obstacles in urban traffic.
[0004] Second, the current research on four-wheel steering collision avoidance control strategies presents diverse characteristics. In path planning and trajectory tracking control, model predictive control (MPC) has been widely applied to four-wheel steering systems due to its high efficiency in constraint handling and real-time optimization capabilities. For example, existing research has proposed an emergency collision avoidance strategy for four-wheel steering vehicles based on MPC, which further ensures the motion stability of the vehicle through an adaptive control method. In addition, the linear quadratic regulator (LQR) method has also been used to calculate the ideal steering angle to achieve precise lateral control of the vehicle. The effectiveness of these methods is usually verified through MATLAB / Simulink and CarSim simulation platforms.
[0005] In addition, reinforcement learning, as an intelligent control method that has emerged in recent years, has gradually been applied to the field of vehicle collision avoidance control. At present, the research on reinforcement learning in vehicle control mostly focuses on traditional two-wheel steering systems. For example, existing literature has used deep reinforcement learning techniques to imitate human driving behavior and achieve the task of collision avoidance with static obstacles. However, the application of reinforcement learning technology in four-wheel steering systems is still in its infancy, and there is relatively little related research. Nevertheless, a small amount of existing literature has shown that reinforcement learning has great potential in the path tracking control of four-wheel independent steering vehicles. For example, existing research has introduced the group intelligence experience replay (GER) and two-stream information bottleneck (TIB) methods to optimize the deep reinforcement learning framework, significantly reducing the path tracking error of four-wheel independent steering vehicles and improving the convergence and generalization ability of the algorithm.
[0006] In summary, four-wheel steering technology has significant handling and stability advantages, which can improve the maneuverability with respect to obstacles and the stability of the vehicle itself, and has great application prospects in the field of vehicle collision avoidance. However, the existing research on four-wheel steering control strategies has not yet involved extreme collision avoidance, and there is still much room for improvement in the dynamic utilization ability under extreme working conditions, especially in further improving the collision avoidance performance by combining intelligent control methods (such as reinforcement learning), which urgently requires in-depth research and exploration. Summary of the Invention
[0007] The present invention aims to solve at least one of the technical problems in the related art to some extent.
[0008] The present invention proposes a method for extreme collision avoidance control of a four-wheel steering vehicle to fully exploit the potential of the four-wheel steering system, maximize the maximum passing speed for tracking the extreme collision avoidance trajectory, and significantly improve the emergency collision avoidance ability of intelligent vehicles.
[0009] Another object of the present invention is to propose a system for extreme collision avoidance control of a four-wheel steering vehicle.
[0010] To achieve the above object, on the one hand, the present invention proposes a method for extreme collision avoidance control of a four-wheel steering vehicle, including:
[0011] Obtaining in real time perception information including the state information of the intelligent vehicle itself and the external environment information;
[0012] Based on the perception information and in combination with a pre-stored typical collision avoidance working condition library or an online real-time trajectory planning algorithm, generating an extreme collision avoidance trajectory suitable for extreme collision avoidance scenarios to output the extreme collision avoidance trajectory and vehicle state information;
[0013] Constructing the state space and action space of the reinforcement learning algorithm according to the extreme collision avoidance trajectory and vehicle state information, and designing a reward function to train the network parameters in real time through the reinforcement learning algorithm and dynamically adjust the steering ratio parameters;
[0014] Based on the extreme collision avoidance trajectory, vehicle state information, and the dynamically adjusted steering ratio parameter, a two-degree-of-freedom dynamic model of the vehicle is constructed. The optimized front wheel steering angle and the calculated rear wheel steering angle are output through the model predictive control algorithm;
[0015] The optimized front wheel steering angle, the calculated rear wheel steering angle, vehicle state information, and the trained network parameters are input into the in-vehicle computer, and an action instruction is generated and sent to the front and rear wheel steering systems of the vehicle after the extreme collision avoidance control is triggered, so as to realize the extreme collision avoidance control of the intelligent vehicle.
[0016] The extreme collision avoidance control method for a four-wheel steering vehicle in an embodiment of the present invention may further have the following additional technical features:
[0017] In an embodiment of the present invention, perception information including the state information of the intelligent vehicle itself and the external environment information is obtained in real time, including:
[0018] The perception information of the vehicle is measured in real time through a millimeter-wave radar, a camera, an inertial measurement unit, a vehicle speed sensor, a steering wheel angle sensor, and front and rear wheel angle sensors installed on the vehicle; wherein, the perception information includes the longitudinal speed u, the lateral speed v, the yaw angular velocity r, and the lateral acceleration a y , the actual steering angle δ of the front wheel f , the actual steering angle δ of the rear wheel rt , and the lateral tracking error e between the vehicle and the planned target collision avoidance trajectory.
[0019] In an embodiment of the present invention, the constructed reinforcement learning algorithm includes a state space, an action space, a reward function, a steering ratio stability constraint, and a training method, and the process is as follows:
[0020] The state space adopted is:
[0021] S = [e, u, r, a y , I is
[0022] e is the tracking error between the current vehicle and the target collision avoidance trajectory; u is the longitudinal vehicle speed; r is the yaw angular velocity of the vehicle; a y is the lateral acceleration of the vehicle; I is is a parameter for judging whether the vehicle exceeds the dynamic stability boundary;
[0023] The action space adopted is the real-time adjustment amount Δβ of the front and rear wheel steering ratio parameter β.
[0024] In an embodiment of the present invention, the reward function considers two action modes: one is to give trajectory tracking compensation when the tire enters the non-linear region during the front wheel angle tracking of the target path, corresponding to the reward function R1:
[0025] R1 = λ1·e 2
[0026] Second, when the lateral acceleration is too large, it rotates in the same direction for instability compensation, corresponding to the reward function R2:
[0027] R2 = λ2·I is ·a y
[0028] It is also considered to minimize the intervention of rear-wheel steering in normal operating conditions, corresponding to the reward function R3:
[0029] R2 = λ3Δβ 2
[0030] In the formula, λ1, λ2, and λ3 are negative parameters, which are used to weigh the trajectory accuracy, steering smoothness, and vehicle stability respectively; the total reward function is:
[0031] R = R1 + R2 + R3
[0032] The reinforcement learning training process uses the non-conservative dynamics safety boundary as the prior constraint in the training process.
[0033] In an embodiment of the present invention, a two-degree-of-freedom dynamic model of the vehicle is constructed based on the limit collision avoidance trajectory, vehicle state information, and dynamically adjusted steering ratio parameters. The optimized front-wheel steering angle and the calculated rear-wheel steering angle output by the model predictive control algorithm include:
[0034] Construct a two-degree-of-freedom dynamic model of the vehicle to describe the dynamic responses of the vehicle's lateral motion and yaw motion respectively; the vehicle lateral dynamics equation is expressed as follows:
[0035]
[0036] The vehicle yaw dynamics equation is expressed as follows:
[0037]
[0038] Among them, m is the vehicle mass; I z is the moment of inertia of the vehicle about the vertical axis; u is the vehicle's longitudinal speed; v is the vehicle's lateral speed; a and b respectively represent the distances from the vehicle's center of mass to the front and rear axles; C f and C r are the lateral tire stiffnesses of the front and rear wheels; δ f and δ r are the steering angles of the front and rear wheels:
[0039] δ r,k = β·δ f,k
[0040] The specific real-time update formula for the steering ratio is as follows:
[0041] β t = β t-1 + Δβ
[0042] In a control step, taking the vehicle state measured at the current moment as the starting point, a prediction range of N steps in length is set. Then, the vehicle's dynamic model is simplified into a linear discrete form to obtain a discrete model of the state space:
[0043] X k+1 = AX k + BU k
[0044] where X k is the vehicle state vector at the k-th step within the prediction time domain, including lateral deviation, yaw angle deviation, lateral velocity, and yaw rate; U k is the control input at the k-th step within the prediction time domain, that is, the front wheel steering angle δ f,k ; A and B are matrices solved according to the dynamic formula and the discrete sampling time;
[0045] Design an optimization problem to be achieved by minimizing the cost function:
[0046]
[0047] is the lateral tracking error at the k-th step within the prediction time domain, is the change amount of the front wheel steering angle, and w1 and w2 are preset weight coefficients; at the same time, constraints are set on the magnitude and change rate of the front wheel steering angle:
[0048]
[0049] where δ f,max is the maximum allowable steering angle of the front wheel, is the recommended maximum change rate of the steering angle of the front wheel steering actuator;
[0050] After defining the discrete model and the constraint conditions, the problem of solving the front wheel steering angle is transformed into a standard quadratic programming problem; among the optimized control input sequences obtained by solving, the first one is selected as the final input.
[0051] In an embodiment of the present invention, after the reinforcement learning network parameters converge and stabilize in the simulation environment training, they are imported into the vehicle-mounted computer and used for real-time online inference of the actual vehicle; when the triggering condition of the extreme collision avoidance control is satisfied, according to the real-time state, the optimized steering ratio adjustment amount Δβ is output by the trained policy network, and the final steering ratio decision is executed in combination with the amplitude and change rate constraints.
[0052] To achieve the above object, on the other hand, the present invention provides a four-wheel steering vehicle extreme collision avoidance control system, including:
[0053] A vehicle sensing system that supports the extreme collision avoidance function, which is used to obtain perception information including the state information of the intelligent vehicle itself and the external environment information in real time;
[0054] An upper-layer collision avoidance trajectory planning module, which is used to generate an extreme collision avoidance trajectory applicable to the extreme collision avoidance scenario based on the perception information and in combination with a pre-stored typical collision avoidance working condition library or an online real-time trajectory planning algorithm, so as to output the extreme collision avoidance trajectory and vehicle state information;
[0055] A front and rear wheel steering ratio decision module based on reinforcement learning, which is used to construct the state space and action space of the reinforcement learning algorithm according to the extreme collision avoidance trajectory and vehicle state information, and design a reward function to train the network parameters in real time through the reinforcement learning algorithm, and dynamically adjust the steering ratio parameters;
[0056] A steering angle calculation module based on a model predictive controller, which is used to construct a two-degree-of-freedom dynamic model of the vehicle based on the extreme collision avoidance trajectory, vehicle state information, and dynamically adjusted steering ratio parameters, and output the optimized front wheel steering angle and the calculated rear wheel steering angle through the model predictive control algorithm;
[0057] After triggering the control, it sends instructions to the actuator module, which is used to input the optimized front wheel steering angle, the calculated rear wheel steering angle, vehicle state information, and the trained network parameters into the in-vehicle computer, and generate action instructions and send them to the front and rear wheel steering systems of the vehicle after the extreme collision avoidance control is triggered, so as to realize the extreme collision avoidance control of the intelligent vehicle.
[0058] The four-wheel steering vehicle extreme collision avoidance control method and system of the embodiments of the present invention construct a four-wheel steering extreme collision avoidance control architecture, establish an MPC trajectory tracking control strategy based on a two-degree-of-freedom vehicle dynamic model, and dynamically adjust the front and rear wheel steering ratios based on reinforcement learning to realize the extreme collision avoidance control of the vehicle, that is, to maximize the maximum passing vehicle speed for tracking the extreme collision avoidance trajectory.
[0059] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0060] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0061] Figure 1 is a flowchart of the four-wheel steering vehicle extreme collision avoidance control method according to the embodiments of the present invention;
[0062] Figure 2 is the architecture diagram of the four-wheel steering vehicle dynamics model adopted according to the embodiment of the present invention;
[0063] Figure 3 is the structure diagram of the four-wheel steering vehicle extreme collision avoidance control system according to the embodiment of the present invention. Detailed implementation manners
[0064] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.
[0065] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0066] The following describes a four-wheel steering vehicle extreme collision avoidance control method and system according to an embodiment of the present invention with reference to the drawings.
[0067] Figure 1 is the flowchart of the four-wheel steering vehicle extreme collision avoidance control method according to the embodiment of the present invention. As Figure 1 shown, the method includes:
[0068] S1, obtaining perception information including the state information of the intelligent vehicle itself and the external environment information in real time;
[0069] S2, generating an extreme collision avoidance trajectory applicable to the extreme collision avoidance scenario based on the perception information and in combination with a pre-stored typical collision avoidance working condition library or an online real-time trajectory planning algorithm, so as to output the extreme collision avoidance trajectory and vehicle state information;
[0070] S3, constructing a state space and an action space of a reinforcement learning algorithm according to the extreme collision avoidance trajectory and vehicle state information, and designing a reward function to train network parameters in real time through the reinforcement learning algorithm, and dynamically adjusting the steering ratio parameter;
[0071] S4, constructing a two-degree-of-freedom dynamics model of the vehicle based on the extreme collision avoidance trajectory, vehicle state information, and dynamically adjusted steering ratio parameter, and outputting an optimized front wheel steering angle and a calculated rear wheel steering angle through a model predictive control algorithm;
[0072] S5. Input the optimized front-wheel steering angle, the calculated rear-wheel steering angle, vehicle state information, and the trained network parameters into the on-vehicle computer, and generate an action command after the extreme collision avoidance control is triggered and send it to the front and rear wheel steering systems of the vehicle to achieve the extreme collision avoidance control of the intelligent vehicle.
[0073] In an embodiment of the present invention, a four-wheel steering extreme collision avoidance control architecture is constructed, an MPC trajectory tracking control strategy based on a two-degree-of-freedom vehicle dynamics model is established, and the front and rear wheel steering ratios are dynamically adjusted based on reinforcement learning to achieve vehicle extreme collision avoidance control, that is, to maximize the maximum passing speed of tracking the extreme collision avoidance trajectory.
[0074] Further, the four-wheel steering extreme collision avoidance control architecture includes: calculating the front-wheel steering angle δ of the vehicle in real time through a model predictive controller (MPC). f Dynamically determining the steering ratio parameter β through a reinforcement learning algorithm, and calculating the rear-wheel steering angle δ in real time. r According to the key formula:
[0075] δ r = β·δ f
[0076] Realize the coordinated steering control of the front and rear wheels of the vehicle.
[0077] Further, the MPC trajectory tracking control strategy based on the two-degree-of-freedom vehicle dynamics model includes: establishing a vehicle lateral and yaw dynamics model; using this dynamics model to construct a cost function including the trajectory tracking error and the change amount of the front-wheel steering angle, and obtaining the front-wheel steering angle control amount by solving the minimization problem of this cost function; at the same time, imposing the tire force limit and the front and rear wheel steering relationship constraints to ensure the trajectory tracking stability.
[0078] Further, the steering ratio decision method based on reinforcement learning includes: constructing the state space and action space of the reinforcement learning algorithm; establishing a vehicle real-time dynamics feedback model as the reinforcement learning algorithm environment; designing a reward function including the trajectory tracking error, vehicle stability, and steering action smoothness; training the network parameters in real time through the reinforcement learning algorithm to dynamically determine the steering ratio β, so as to realize the intelligent strategy of reverse steering of the front and rear wheels when the vehicle tracking error is large and the same-direction steering when the vehicle is critically unstable.
[0079] The state space includes but is not limited to the following variables: trajectory tracking error, vehicle lateral offset, yaw angular velocity, lateral acceleration.
[0080] The action space includes the steering ratio adjustment amount Δβ, and its specific expression is:
[0081] β new = β old +Δβ
[0082] The design of the reward function includes the immediate reward R i and the terminal reward R t . The design principle of the immediate reward R i is as follows: comprehensively consider the vehicle trajectory tracking accuracy, vehicle motion stability, and smoothness of the steering action. Among them, the vehicle tracking error and the adjustment range of the steering ratio need to be penalized to improve the trajectory tracking accuracy and smooth control effect, and the vehicle stability index is used to reflect the degree of vehicle instability and corresponding penalties are given to ensure vehicle safety. The terminal reward R t is a one-time reward given according to the terminal state at the end of each training episode, including: the non-collision reward item R T1 , the penalty item R T2 for instability (the sideslip angle of the vehicle's center of mass is greater than the preset threshold), and the penalty item R T3 for collision or exceeding the lateral road boundary.
[0083] The training method will combine the prior knowledge of vehicle dynamics to improve the exploration efficiency of the reinforcement learning algorithm. The network parameters of the reinforcement learning algorithm are trained in a simulation environment, and the trained network parameters are imported into an in-vehicle computer in a real environment; after the extreme collision avoidance control is triggered, the in-vehicle computer sends action instructions to the front and rear wheel steering systems of the vehicle to achieve the extreme collision avoidance control of the intelligent vehicle.
[0084] In summary, the present invention constructs a four-wheel steering extreme collision avoidance control architecture, designs a precise trajectory tracking control strategy and a dynamic reinforcement learning steering ratio decision method, realizes the ability of an intelligent vehicle to safely and efficiently track the extreme collision avoidance trajectory under emergency conditions, and fills the technical gap in the field of four-wheel steering extreme collision avoidance control.
[0085] In an embodiment of the present invention, the core idea is to incorporate the front and rear wheel steering coordination of the vehicle into the same closed-loop control system, realize the real-time tracking of the emergency collision avoidance trajectory through model predictive control (MPC), and at the same time use the reinforcement learning (RL) algorithm to dynamically adjust the front and rear wheel steering ratio parameters, so as to balance flexibility and stability in high-risk scenarios and maximize the collision avoidance success rate of the vehicle. The overall system framework of this embodiment can be summarized into the following key steps:
[0086] The upper-layer perception-decision in the present invention is used to obtain the state information of the intelligent vehicle itself and the external environment information in real time, and plan the collision avoidance trajectory of the vehicle under extreme conditions accordingly. Specifically, this step includes:
[0087] First, through the millimeter-wave radar, camera, inertial measurement unit (IMU), vehicle speed sensor, steering wheel angle sensor, and front and rear wheel angle sensors installed on the vehicle, the longitudinal speed u, lateral speed v, yaw rate r, and lateral acceleration a of the vehicle are measured in real timey , the actual steering angle δ of the front wheels f , the actual steering angle δ of the rear wheels rt , and the lateral tracking error e between the vehicle and the planned target collision avoidance trajectory. Among them, the millimeter-wave radar and the camera are mainly responsible for detecting information such as the position, size, and relative speed of road obstacles ahead, and fusing the obstacle information with the vehicle state information to form a real-time understanding of the front environment.
[0088] Secondly, based on the above real-time perception information, combined with a pre-stored typical collision avoidance working condition library or an online real-time trajectory planning algorithm, an ideal collision avoidance path suitable for extreme collision avoidance scenarios is quickly generated. In the specific process, considering factors such as the current speed and dynamic characteristics of the vehicle, the position, size, and motion state of the obstacle, a collision avoidance trajectory coordinated in the transverse and longitudinal directions is calculated through an optimization method.
[0089] During the actual operation process, the vehicle state information and the planned extreme collision avoidance trajectory are transmitted in real time to the subsequent trajectory tracking control module and related modules for steering ratio decision-making through Ethernet and CAN FD buses after data fusion, preprocessing, and synchronization, as the key input data for the subsequent modules to perform real-time trajectory tracking control and reinforcement learning dynamic decision-making.
[0090] Furthermore, the present invention designs a steering ratio decision-making step based on reinforcement learning. Specifically, the reinforcement learning algorithm constructed in this embodiment includes a state space, an action space, a reward function, a steering ratio stability constraint, and an efficient training method. The specific implementation process is as follows:
[0091] The state space adopted in this embodiment is:[[]]
[0092] S = [e, u, r, a y , I is
[0093] e is the tracking error between the current vehicle and the target collision avoidance trajectory; u is the longitudinal vehicle speed; r is the yaw angular velocity of the vehicle; a y is the lateral acceleration of the vehicle; I is is a parameter for judging whether the vehicle exceeds the dynamic stability boundary, and is assigned 1 when exceeded.
[0094] The action space adopted in this embodiment is the real-time adjustment amount Δβ of the front and rear wheel steering ratio parameters β.
[0095] In this embodiment, the reward function considers two action modes. One is to give trajectory tracking compensation when the tire enters the non-linear region during the front wheel angle tracking of the target path, so as to improve the trajectory tracking ability. The corresponding reward function R1 is:[[]]
[0096] R1 = λ1 · e 2
[0097] Second, when the lateral acceleration is too large, it rotates in the same direction to compensate for instability, corresponding to the reward function R2:
[0098] R2 = λ2· is ·a y
[0099] In addition, it should also be considered to reduce the intervention of rear-wheel steering as much as possible in normal working conditions, corresponding to the reward function R3:
[0100] R2 = λ3Δβ 2
[0101] In the formula, λ1, λ2 and λ3 are negative parameters, which are used to balance the trajectory accuracy, steering smoothness and vehicle stability respectively. The total reward function is:
[0102] R = R1 + R2 + R3
[0103] First, the Soft-Actor-Critic algorithm is applied in the offline simulation environment to enable the reinforcement learning agent to perform policy evaluation and policy update. In this embodiment, the non-conservative dynamic safety boundary (such as the side slip angle and yaw rate diagram) is used as the prior constraint in the training process during the reinforcement learning training process.
[0104] After the reinforcement learning network parameters converge and stabilize in the simulation environment training, they are imported into the vehicle-mounted computer and used for real-time online inference of the actual vehicle. When the limit collision avoidance control trigger condition is satisfied, in this step, the optimized steering ratio adjustment amount Δβ is output by the trained policy network according to the real-time state, and the final steering ratio decision is executed in combination with the above amplitude and change rate constraints.
[0105] Furthermore, the MPC trajectory tracking control based on the two-degree-of-freedom vehicle dynamics model proposed by the present invention is used to achieve precise trajectory tracking control of the intelligent vehicle in the limit collision avoidance scenario. In the specific implementation process, first, a two-degree-of-freedom dynamics model of the vehicle is constructed to respectively describe the dynamic responses of the lateral motion and yaw motion of the vehicle. As Figure 2 shown, the vehicle lateral dynamics equation is expressed as follows:
[0106]
[0107] The vehicle yaw dynamics equation is expressed as follows:
[0108]
[0109] Among them, m is the vehicle mass; I z is the moment of inertia of the vehicle about the vertical axis; u is the longitudinal speed of the vehicle; v is the lateral speed of the vehicle; a and b respectively represent the distances from the vehicle center of mass to the front and rear axles; C f and Cr are the cornering stiffnesses of the front and rear tires; δ f and δ r are the steering angles of the front and rear wheels:
[0110] δ r,k = β·δ f,k
[0111] The specific real-time update formula for the steering ratio is as follows:
[0112] β t = β t-1 + Δβ
[0113] Based on the above dynamic model, in this step, the model predictive control (MPC) method is adopted to achieve real-time and efficient tracking control of the planned collision avoidance trajectory. In a control step, the present invention takes the vehicle state measured at the current moment as the starting point and sets a prediction horizon with a length of N steps. Then, the dynamic model of the vehicle is simplified into a linear discrete form to obtain a discrete model of the state space:
[0114] X k+1 = AX k + BU k
[0115] where X k is the vehicle state vector at the k-th step within the prediction horizon, including lateral deviation, yaw angle deviation, lateral velocity, yaw rate, etc.; U k is the control input at the k-th step within the prediction horizon, that is, the front wheel steering angle δ f,k ; A and B are matrices solved according to the aforementioned dynamic formula and discrete sampling time.
[0116] The objective of the present invention is to find a set of optimal control commands to enable the vehicle to closely follow the preset trajectory while ensuring driving smoothness and comfort. To this end, the present invention designs an optimization problem and realizes it by minimizing the following cost function:
[0117]
[0118] is the lateral tracking error at the k-th step within the prediction horizon, is the change amount of the front wheel steering angle, and w1 and w2 are preset weight coefficients. At the same time, considering the hardware limitations of the actual vehicle, the present invention sets constraint conditions on the magnitude and change rate of the front wheel steering angle:
[0119]
[0120] where δ f,max is the maximum allowable steering angle of the front wheel, Recommended maximum rate of change of steering angle for front wheel steering actuator.
[0121] After the above discrete model and constraint conditions are defined, this step transforms the problem of solving the front wheel steering angle into a standard quadratic programming problem. The first of the optimized control input sequences is selected as the final input.
[0122] Finally, based on the steering angle control of the front and rear wheels, and the real-time status information of the vehicle. At the same time, it also receives the trained reinforcement learning network parameters, which have been pre-imported into the on-board computer. After the extreme collision avoidance control is triggered, action instructions are generated according to these parameters and sent to the front and rear wheel steering systems of the vehicle, thereby realizing the extreme collision avoidance control of the intelligent vehicle. In addition, it is also responsible for monitoring the actual steering execution of the vehicle. If deviations or abnormalities are found, adjustment signals will be generated and fed back to the upstream module to optimize the control strategy. Finally, the overall operation status report of the system is output to ensure the safety and efficiency of the entire system.
[0123] The extreme collision avoidance control method for a four-wheel steering vehicle according to an embodiment of the present invention utilizes four-wheel steering technology, and through a control strategy that integrates model predictive control (MPC) and reinforcement learning (RL), dynamically adjusts the steering ratio of the front and rear wheels in real time to achieve coordinated control of the front and rear wheels. Under emergency collision avoidance conditions, the trajectory tracking error can be significantly reduced, the dynamic stability of the vehicle can be maximized, and the maximum passing speed of the collision avoidance trajectory can be improved. The reinforcement learning algorithm is used to optimize the steering ratio parameters in real time, and the steering relationship between the front and rear wheels is flexibly adjusted according to the real-time state of the vehicle and the trajectory tracking error. When the trajectory error is large, the front and rear wheels are reversed to improve the maneuverability of the vehicle; when the vehicle is close to instability, it automatically turns to the same direction of steering to ensure the stability and safety of the vehicle, reflecting high intelligence and adaptability.
[0124] In order to implement the above embodiment, Figure 3 As shown, this embodiment also provides a four-wheel steering vehicle extreme collision avoidance control system 10, including:
[0125] The vehicle sensing system 10 supporting the extreme collision avoidance function is used to obtain perception information including the intelligent vehicle's own state information and external environment information in real time;
[0126] The upper collision avoidance trajectory planning module 21 is used to generate an extreme collision avoidance trajectory suitable for extreme collision avoidance scenarios based on the perception information and in combination with a pre-stored typical collision avoidance condition library or an online real-time trajectory planning algorithm, so as to output the extreme collision avoidance trajectory and vehicle status information;
[0127] The front and rear wheel steering ratio decision-making module 22 based on reinforcement learning is used to construct the state space and action space of the reinforcement learning algorithm according to the extreme collision avoidance trajectory and vehicle state information, and design a reward function to train the network parameters in real time through the reinforcement learning algorithm, and dynamically adjust the steering ratio parameters;
[0128] The steering angle calculation module 23 based on the model predictive controller is used to construct a two-degree-of-freedom dynamic model of the vehicle based on the extreme collision avoidance trajectory, vehicle state information, and dynamically adjusted steering ratio parameters, and output the optimized front wheel steering angle and the calculated rear wheel steering angle through the model predictive control algorithm;
[0129] The trigger control module 24 that sends instructions to the actuator after triggering is used to input the optimized front wheel steering angle, the calculated rear wheel steering angle, vehicle state information, and trained network parameters into the in-vehicle computer, and generate action instructions and send them to the front and rear wheel steering systems of the vehicle after the extreme collision avoidance control is triggered, so as to realize the extreme collision avoidance control of the intelligent vehicle.
[0130] Furthermore, the vehicle sensing system 10 that supports the extreme collision avoidance function is also used for:
[0131] Real-time measurement of vehicle perception information through millimeter wave radar, camera, inertial measurement unit, vehicle speed sensor, steering wheel angle sensor, and front and rear wheel angle sensors installed on the vehicle; among them, the perception information includes longitudinal speed u, lateral speed v, yaw rate r, and lateral acceleration a y The actual steering angle δ of the front wheel f The actual steering angle δ of the rear wheel rt , and the lateral tracking error e between the vehicle and the planned target collision avoidance trajectory.
[0132] Furthermore, the constructed reinforcement learning algorithm includes a state space, an action space, a reward function, a steering ratio stability constraint, and a training method, and the process is as follows:
[0133] The state space adopted is:
[0134] S = [e, u, r, a y , I is
[0135] e is the tracking error between the current vehicle and the target collision avoidance trajectory; u is the longitudinal vehicle speed; r is the yaw rate of the vehicle; a y is the lateral acceleration of the vehicle; I is is a parameter for judging whether the vehicle exceeds the dynamic stability boundary;
[0136] The action space adopted is the real-time adjustment amount Δβ of the front and rear wheel steering ratio parameter β.
[0137] Furthermore, the reward function considers two modes of action: one is to compensate for trajectory tracking when the tire enters the non-linear region during the previous wheel angle tracking of the target path, corresponding to the reward function R1:
[0138] R1 = λ1·e 2
[0139] The other is to perform instability compensation by turning in the same direction when the lateral acceleration is too large, corresponding to the reward function R2:
[0140] R2 = λ2·I is ·a y
[0141] It also considers minimizing the intervention of rear-wheel steering in normal working conditions, corresponding to the reward function R3:
[0142] R2 = λ3Δβ 2
[0143] In the formula, λ1, λ2 and λ3 are negative parameters, which are used to balance trajectory accuracy, steering smoothness and vehicle stability respectively; the total reward function is:
[0144] R = R1 + R2 + R3
[0145] The reinforcement learning training process uses a non-conservative dynamic safety boundary as a prior constraint during the training process.
[0146] It can be understood that during the actual operation process, after the vehicle state information and the planned extreme collision avoidance trajectory are fused, preprocessed and synchronized, they are transmitted in real time to the subsequent trajectory tracking control module and steering ratio decision module through Ethernet and CAN FD bus, as the key input data for the subsequent modules to perform real-time trajectory tracking control and reinforcement learning dynamic decision-making.
[0147] According to the extreme collision avoidance control system for four-wheel steering vehicles of the embodiment of the present invention, through a control strategy that combines model predictive control (MPC) and reinforcement learning (RL), the front and rear wheel steering ratios are dynamically adjusted in real time to achieve coordinated control of the front and rear wheels. In the emergency collision avoidance working condition, the trajectory tracking error can be significantly reduced, the dynamic stability of the vehicle is maximized, and the maximum passing speed of the collision avoidance trajectory is increased. The reinforcement learning algorithm is used to optimize the steering ratio parameters in real time, and the front and rear wheel steering relationships are flexibly adjusted according to the real-time state of the vehicle and the trajectory tracking error. When the trajectory error is large, reverse steering of the front and rear wheels is realized to improve the maneuverability of the vehicle; when the vehicle is close to instability, it automatically turns to the same direction to ensure the stable and safe operation of the vehicle, reflecting high intelligence and self-adaptability.
[0148] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0149] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
Claims
1. A method for controlling extreme collision avoidance of a four-wheel steering vehicle, characterized in that, Including: Obtaining in real time perception information including the state information of the intelligent vehicle itself and the external environment information; Based on the perception information and in combination with a pre-stored typical collision avoidance working condition library or an online real-time trajectory planning algorithm, generating a limit collision avoidance trajectory applicable to an extreme collision avoidance scenario to output the limit collision avoidance trajectory and vehicle state information; Constructing the state space and action space of the reinforcement learning algorithm according to the limit collision avoidance trajectory and vehicle state information, and designing a reward function to train the network parameters in real time through the reinforcement learning algorithm and dynamically adjust the steering ratio parameter; Constructing a two-degree-of-freedom dynamic model of the vehicle based on the limit collision avoidance trajectory, vehicle state information, and the dynamically adjusted steering ratio parameter, and outputting the optimized front wheel steering angle and the calculated rear wheel steering angle through the model predictive control algorithm; Inputting the optimized front wheel steering angle, the calculated rear wheel steering angle, vehicle state information, and the trained network parameters into an in-vehicle computer, and generating an action instruction and sending it to the front and rear wheel steering systems of the vehicle after the limit collision avoidance control is triggered to achieve the limit collision avoidance control of the intelligent vehicle.
2. The method according to claim 1, characterized in that, Obtaining in real time perception information including the state information of the intelligent vehicle itself and the external environment information, including: The perception information of the vehicle is measured in real time by millimeter-wave radar, camera, inertial measurement unit, vehicle speed sensor, steering wheel angle sensor, and front and rear wheel angle sensors installed on the vehicle; wherein, the perception information includes longitudinal speed u, lateral speed v, yaw rate r, and lateral acceleration a y , the actual steering angle δ of the front wheel f , the actual steering angle δ of the rear wheel rt , and the lateral tracking error e between the vehicle and the planned target collision avoidance trajectory.
3. The method according to claim 1, characterized in that The constructed reinforcement learning algorithm includes a state space, an action space, a reward function, a steering ratio stability constraint, and a training method, and the process is as follows: The state space adopted is: S = [e, u, r, a y , I is e is the tracking error between the current vehicle and the target collision avoidance trajectory; u is the longitudinal vehicle speed; r is the yaw angular velocity of the vehicle; a y is the lateral acceleration of the vehicle; I is is a parameter for judging whether the vehicle exceeds the dynamic stability boundary; The action space adopted is the real-time adjustment amount Δβ of the front and rear wheel steering ratio parameter β.
4. The method according to claim 3, characterized in that, The reward function considers two action modes: one is to give trajectory tracking compensation when the tire enters the non-linear region during the front wheel corner tracking the target path, corresponding to the reward function R1: R1 = λ1·e 2 The other is to give a co-rotating instability compensation when the lateral acceleration is too large, corresponding to the reward function R2: R2 = λ2·I is ·a y It is also considered to minimize the intervention of the rear wheel steering in normal working conditions, corresponding to the reward function R3: R2 = λ3Δβ 2 In the formula, λ1, λ2, and λ3 are negative parameters, which are respectively used to balance the trajectory accuracy, steering smoothness, and vehicle stability; the total reward function is: R = R1 + R2 + R3 The reinforcement learning training process uses a non-conservative dynamic safety boundary as a prior constraint during the training process.
5. The method according to claim 4, characterized in that Constructing a two-degree-of-freedom dynamic model of the vehicle based on the limit collision avoidance trajectory, vehicle state information, and the dynamically adjusted steering ratio parameter, and outputting the optimized front wheel steering angle and the calculated rear wheel steering angle through the model predictive control algorithm includes: Constructing a two-degree-of-freedom dynamic model of the vehicle to respectively describe the dynamic responses of the lateral motion and yaw motion of the vehicle; the vehicle lateral dynamic equation is expressed as follows: The vehicle yaw dynamic equation is expressed as follows: where m is the vehicle mass; I z is the moment of inertia of the vehicle about the vertical axis; u is the longitudinal speed of the vehicle; v is the lateral speed of the vehicle; a and b respectively represent the distances from the vehicle's center of mass to the front and rear axles; C f and C r are the cornering stiffnesses of the front and rear tires; δ f and δ r are the steering angles of the front and rear wheels: δ r,k = β · δ f,k The specific real-time update formula of the steering ratio is as follows: β t = β t-1 + Δβ In a control step, taking the vehicle state measured at the current moment as the starting point, setting a prediction range of length N steps, and then simplifying the dynamic model of the vehicle into a linear discrete form to obtain a discrete model of the state space: X k+1 = AX k + BU k Among them, X k is the vehicle state vector at the k-th step within the prediction horizon, including lateral deviation, yaw angle deviation, lateral velocity, and yaw rate; U k is the control input at the k-th step within the prediction horizon, that is, the front wheel steering angle δ f,k ; A and B are matrices obtained according to the dynamic formula and discrete sampling time; Designing an optimization problem to be achieved by minimizing the cost function: is the lateral tracking error at the k-th step within the prediction horizon, is the change in the front wheel steering angle, and w1 and w2 are preset weight coefficients; meanwhile, constraint conditions are set for the magnitude and change rate of the front wheel steering angle: where δ f,max is the maximum allowable steering angle of the front wheels, and is the recommended maximum angular change rate of the front-wheel steering actuator; After the discrete model and constraint conditions are defined, the problem of solving the front wheel corner is transformed into a standard quadratic programming problem; in the optimized control input sequence obtained by solving, the first one is selected as the final input.
6. The method according to claim 5, characterized in that After the reinforcement learning network parameters converge and stabilize during simulation environment training, they are imported into the vehicle-mounted computer and used for real-time online inference of the actual vehicle. When the triggering condition for extreme collision avoidance control is met, the optimized steering ratio adjustment Δβ is output by the trained policy network according to the real-time state, and the final steering ratio decision is executed in combination with amplitude and rate-of-change constraints.
7. A limit collision avoidance control system for a four-wheel steering vehicle, characterized in that, It includes: A vehicle sensing system that supports the extreme collision avoidance function, which is used to obtain perception information including the state information of the intelligent vehicle itself and the external environment information in real time; An upper-layer collision avoidance trajectory planning module, which is used to generate an extreme collision avoidance trajectory applicable to the extreme collision avoidance scenario based on the perception information and in combination with a pre-stored typical collision avoidance working condition library or an online real-time trajectory planning algorithm, so as to output the extreme collision avoidance trajectory and vehicle state information; A front and rear wheel steering ratio decision module based on reinforcement learning, which is used to construct the state space and action space of the reinforcement learning algorithm according to the extreme collision avoidance trajectory and vehicle state information, and design a reward function to train the network parameters in real time through the reinforcement learning algorithm and dynamically adjust the steering ratio parameters; A steering angle calculation module based on a model predictive controller, which is used to construct a two-degree-of-freedom dynamic model of the vehicle based on the extreme collision avoidance trajectory, vehicle state information, and dynamically adjusted steering ratio parameters, and output the optimized front wheel steering angle and the calculated rear wheel steering angle through the model predictive control algorithm; After triggering the control, send instructions to the actuator module, which is used to input the optimized front wheel steering angle, the calculated rear wheel steering angle, vehicle state information, and the trained network parameters into the vehicle-mounted computer, and generate action instructions and send them to the front and rear wheel steering systems of the vehicle after the extreme collision avoidance control is triggered, so as to realize the extreme collision avoidance control of the intelligent vehicle.
8. The system according to claim 7, wherein The vehicle sensing system that supports the extreme collision avoidance function is also used for: The perception information of the vehicle is measured in real time by millimeter-wave radar, camera, inertial measurement unit, vehicle speed sensor, steering wheel angle sensor, and front and rear wheel angle sensors installed on the vehicle; wherein, the perception information includes longitudinal speed u, lateral speed v, yaw angular velocity r, and lateral acceleration a y , the actual steering angle δ of the front wheel f , the actual steering angle δ of the rear wheel rt , and the lateral tracking error e between the vehicle and the planned target collision avoidance trajectory.
9. The system according to claim 8, wherein The constructed reinforcement learning algorithm includes a state space, an action space, a reward function, a steering ratio stability constraint, and a training method. The process is as follows: The state space adopted is: S = [e, u, r, a y , I is e is the tracking error between the current vehicle and the target collision avoidance trajectory; u is the longitudinal vehicle speed; r is the yaw angular velocity of the vehicle; a y is the lateral acceleration of the vehicle; I is is a parameter for judging whether the vehicle exceeds the dynamic stability boundary; The action space adopted is the real-time adjustment amount Δβ of the front and rear wheel steering ratio parameter β.
10. The system according to claim 9, wherein The reward function considers two action modes: one is to give trajectory tracking compensation when the tire enters the non-linear region during front wheel angle tracking of the target path, corresponding to the reward function R1: R1 = λ1·e 2 The other is to give co-rotating instability compensation when the lateral acceleration is too large, corresponding to the reward function R2: R2 = λ2·I is ·a y It also considers minimizing the intervention of rear wheel steering as much as possible in normal working conditions, corresponding to the reward function R3: R2 = λ3Δβ 2 In the formula, λ1, λ2, and λ3 are negative parameters, which are used to weigh trajectory accuracy, steering smoothness, and vehicle stability respectively; the total reward function is: R = R1 + R2 + R3 The reinforcement learning training process uses a non-conservative dynamic safety boundary as a prior constraint during the training process.