An emergency steering control strategy network model, training method, modeling method and simulation method for autonomous commercial vehicles
Through a multi-task division method based on deep reinforcement learning and a variable Gaussian safety field strategy, the problem of commercial vehicles being easily instable during emergency obstacle avoidance is solved, the control stability and obstacle avoidance safety are improved, and the reliability of simulation experiments is improved.
Patent Information
- Application Number
- CN202210748793.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-06-29
AI Technical Summary
Commercial vehicles are prone to instability and overturning problems during emergency obstacle avoidance, resulting in serious traffic accidents. The existing model-based control algorithm is computationally large and the simulation effect is not convincing.
The multi-task division method based on deep reinforcement learning is adopted to divide and train the decision-making and control tasks of commercial vehicles, combine the variable Gaussian safety field to improve the reward function, and dynamic modeling is performed in Matlab, and joint simulation is performed with Carla.
It improves the control stability and safety of commercial vehicles in emergency obstacle avoidance situations, ensures the reliability of simulation experiments, and greatly improves training efficiency.
Smart Images

Figure CN114925461B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of autonomous driving in artificial intelligence, and relates to an autonomous emergency steering control (AES) policy network model, a training method, a modeling method, and a simulation method for commercial vehicles based on deep reinforcement learning. Background Art
[0002] Automobiles have become an indispensable means of transportation in today's world. However, while vehicles bring mobility, they also bring risks. With the rapid development of artificial intelligence and automotive technology, people expect autonomous vehicles to take on more of the driver's burden and stress, thereby improving safety. Currently, the autonomous driving of commercial vehicles with fixed routes has the feasibility of being deployed. However, due to the characteristics of large size, heavy weight, and large blind spots in the driver's field of vision of commercial vehicles, problems such as instability and rollover are likely to occur when emergency avoidance of obstacles ahead, leading to serious traffic accidents. Therefore, how to solve the problem of easy rollover of commercial vehicles during emergency obstacle avoidance has become an important topic.
[0003] Currently, the computational complexity of model-based control algorithms (such as MPC) is very high. When MPC attempts to optimize the cost function of the control behavior for each control cycle, the huge computational complexity will lead to a long decision-making time, which is unsafe. At the same time, when conducting simulations in the Carla simulator, commercial vehicles do not have special dynamic characteristics, making the simulation results not very persuasive. Based on the original Carla simulation environment, how to construct a model with the dynamic characteristics of large size and high center of gravity of commercial vehicles is an urgent problem to be solved. Summary of the Invention
[0004] In view of the above problems, the present invention will propose an autonomous emergency steering control scheme for commercial vehicles based on deep reinforcement learning. Although the training computational complexity of deep reinforcement learning (DRL) is also relatively high, the weights in the reasoning process are relatively light. Compared with other methods, complex actions can be planned in a short time. The present invention adopts a multi-task partitioned reinforcement learning method to divide and train the decision-making and control tasks. At the same time, a variable Gaussian safety field is added as the basis for decision-making, and the reward function is improved according to the variable Gaussian safety field. Finally, the dynamic modeling of commercial vehicles is carried out in Matlab, and combined with Carla for joint simulation.
[0005] The purpose of the present invention is to provide a policy network model, a training method, a modeling method, and a simulation method for autonomous emergency steering control (AES) of commercial vehicles based on deep reinforcement learning, using a multi-task partitioned training method, and combining a variable Gaussian safety field model to improve the safety of decision-making. So as to complete autonomous emergency steering when there are obstacles ahead and the own vehicle cannot achieve the braking target, avoiding rear-end or collision accidents.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] Part 1: Modeling of Autonomous Commercial Vehicles
[0008] Step 1: Model the commercial vehicle in Matlab. The type of commercial vehicle designed in the present invention is a three-axle commercial vehicle.
[0009] The lateral motion equation of the vehicle is:
[0010]
[0011] Yaw motion equation:
[0012]
[0013] Where, m is the vehicle mass, u is the forward speed of the vehicle's center of mass, β is the sideslip angle of the vehicle's center of mass, ω r is the yaw angular velocity, B is the distance between the intersections of the center lines of the two kingpins with the ground, I z is the yaw inertia moment, δ i is the tire steering angle, i = 1, 2, 3 respectively represent the front, middle, and rear axles of the three-axle commercial vehicle, F yi is the lateral force of the ground on the tire, i = 1, 2, 3, 4, 5, 6 respectively represent the 6 tires of the three-axle commercial vehicle. a, b, c are the distances from the front, middle, and rear axles to the center of mass, F yij represents the lateral moment of the j-th axle on the i-th axle of the tire.
[0014] Calculate the vertical load of each wheel of the three-axle commercial vehicle, which is divided into two parts: static load and dynamic load:
[0015] (1) Static load
[0016] The mathematical equation of the vehicle dynamics model is:
[0017]
[0018] Where, z is the dynamic displacement of the vehicle vibration structure, M, C, K, f are the mass matrix, damping matrix, stiffness matrix, and external force column vector.
[0019] Then the mathematical equation of the three-axle commercial vehicle at rest under its own gravity is:
[0020] Kz j = G
[0021] Where, z j is the static displacement of the vehicle vibration structure, G is the gravity in the direction of each static displacement. Because the structural stiffness of the vehicle remains unchanged whether it vibrates or not within the linear range, the expression of K is as follows:
[0022]
[0023] where k 1 , k 2 , k 3 are the vertical stiffnesses of each suspension respectively.
[0024] The static displacement z j can be solved from the above mathematical equation, and then the ground normal reaction forces on the front axle, middle axle and rear axle are respectively:
[0025]
[0026]
[0027]
[0028] k t1 , k t2 , k t3 are the vertical stiffnesses of each tire respectively.
[0029] When stationary, the ground normal reaction forces caused by the vertical loads on each wheel are respectively denoted as F zi0 (i = 1 to 6), then:
[0030]
[0031] where F zf is the ground normal reaction force on the front wheel, F zm is the ground normal reaction force on the middle wheel, and F zr is the ground normal reaction force on the rear wheel.
[0032] (2) Dynamic load
[0033] Denote the distribution ratio coefficients of the vehicle gravity on each axle as t 1 , t 2 , t 3 , then t 1 = G f / G, t 2 = G m / G, t 3 = G r / G. The distribution ratio of the lateral inertia force on each axle is the same as that of the vehicle gravity on each axle. Therefore, (ma y ) f = ma y t 1 , (ma y ) m = ma y t 2 , (may ) r = ma y t 3 . Among them, G f , G m , G r respectively represent the separation of the vehicle gravity on the front axle, middle axle and rear axle, a y is the lateral acceleration, and ma y is the lateral inertial force.
[0034] Taking moments for different tires respectively, the vertical load of each tire can be obtained:
[0035]
[0036] Among them, h g is the height from the center of mass to the ground, and F zi1 (i = 1 - 6) is the vertical load on each wheel during turning.
[0037] Next, calculate the sideslip angle of each tire of the three-axle commercial vehicle. Using the base point method for finding the velocity of each point in a planar figure, with the vehicle's center of mass as the base point, the components of the velocity of each wheel center in the longitudinal and lateral directions can be obtained.
[0038] The sideslip angles of the tires to be obtained are respectively:
[0039]
[0040] Among them, ε i (i = 1 - 6) is the angle between u i (i = 1 - 6) and the ground (i.e., the x-axis), and u i (i = 1 - 6) is the center velocity of each tire of the vehicle. v is the lateral velocity of the vehicle's center of mass.
[0041] Part Two: Training and Simulation of the Policy Network Model
[0042] Step 2: Import the three-axle commercial vehicle model built in Matlab into the source file of the Carla simulation simulator to obtain a model with the dynamic characteristics of a commercial vehicle in Carla.
[0043] Step 3: Conduct training on the decision control model of an autonomous driving commercial vehicle based on multi-task reinforcement learning. The decomposed sub-tasks include lateral and longitudinal control tasks and decision-making tasks. Obtain the navigation points of the commercial vehicle through the temporal bird's-eye view and conduct training.
[0044] Lateral control training. Using the coordinates (x i , y i ) of the obtained navigation points, the heading deviation and the vehicle speed v and acceleration for controlling the vehicle as state variables:
[0045]
[0046] s lane_keep is the state variable obtained during the lane-keeping training of the intelligent agent.
[0047] The action is only the steering wheel angle a steer ∈[-1, 1]. Regarding the design of the reward function r for the lane-keeping task part, the present invention uses the lateral error x of the current coordinates of the vehicle lane_keep and the heading angle deviation 0 as evaluation indicators:
[0048]
[0049] λ 1 、λ 2 are the weights of the two parts of the reward function.
[0050] If the lateral deviation of the current position of the autonomous vehicle during training is greater than the set maximum lateral deviation threshold x 0max then end the iterative training of the current round and start the training of the next round.
[0051] Longitudinal control training. The longitudinal trajectory tracking control task uses the vehicle speed v of the current vehicle, the acceleration the vehicle speed v of the vehicle in front l 、acceleration the distance d from the vehicle in front and the desired vehicle speed v of the current vehicle des as state variables:
[0052]
[0053] s acc is the state variable obtained during the longitudinal following control training of the intelligent agent.
[0054] The output action a of the intelligent agent acc ∈[-1, 1], including the throttle action a throttle and the brake action a brake :
[0055]
[0056] For the longitudinal control task, the reward function is designed as:
[0057]
[0058] Among them, d is the real-time distance from the vehicle in front, d des is the desired distance from the vehicle in front, d safe is the safe distance from the vehicle ahead. When the distance between the intelligent vehicle and the vehicle ahead is less than the safe distance, the reward is -100 and the current interaction is stopped to start the next round of interaction. When conducting longitudinal training, the vehicle speed v of the vehicle ahead is randomly given in each round l and the desired vehicle speed v of the current vehicle des , so that the trained model can be generalized to more complex situations.
[0059] Decision task training. The decision task in a complex environment is trained based on the premise that both the lateral and longitudinal trajectory tracking control tasks can be well completed. The decision behaviors defined in the present invention are emergency braking and emergency steering. Compared with the lateral and longitudinal control tasks, the action space of the decision task is discrete.
[0060]
[0061] When a decision is 0, the decision module selects emergency braking; when a decision is 1, the decision module selects to make an emergency turn to the right.
[0062] Step 4: To achieve the stability of vehicle control and the safety of obstacle avoidance during emergency obstacle avoidance, the present invention designs the reward function of the decision using a variable Gaussian safety field. When the obstacle is outside the expansion domain, the vehicle can take braking measures; when the obstacle is in the expansion domain, the vehicle takes steering and lane-changing measures; when the obstacle is in the core domain and the restricted domain, a collision is likely to occur. The reward function r decision is designed as follows:
[0063]
[0064] Among them, d lon,min , d lon,mid , d lon,max are respectively the longitudinal safety distances of the core domain, restricted domain, and expansion domain of the recognizable Gaussian safety field. l v is the length of the vehicle model, w v is the width of the vehicle model, l′ v is the length of the vehicle model when the vehicle is moving, w v ′ is the width of the vehicle model when the vehicle is moving.
[0065] Among them:
[0066]
[0067] Among them, is the velocity vector of the vehicle movement, k v is the adjustment factor, and 0 < k v < 1 or -1 < k v< 0, whose sign corresponds to the front - rear direction of the movement, and ξ is the yaw angle of the vehicle.
[0068] In this paper, through real - vehicle measurement, the k v = 0.3.
[0069] Step 5: Conduct co - simulation in Carla.
[0070] Advantages of the present invention:
[0071] (1) Aiming at the emergency braking and steering problem of commercial vehicles, the present invention uses Matlab to build a model and conducts co - simulation with Carla, solving the problems in model - free reinforcement learning that cannot reflect the high center of gravity, easy rollover, and difficult braking due to large mass of commercial vehicles, and ensuring the reliability of the simulation experiment.
[0072] (2) The present invention uses a multi - task - divided reinforcement learning method, greatly improving the training efficiency. At the same time, a variable Gaussian safety field strategy is introduced to ensure that the vehicle control has high stability and obstacle - avoidance safety during decision - making and control. Description of the drawings
[0073] Figure 1 Flowchart of the present invention;
[0074] Figure 2 Matlab modeling diagram used in the present invention;
[0075] Figure 3 Variable Gaussian safety field model Detailed implementation manners
[0076] The present invention will be further described below with reference to the drawings.
[0077] The purpose of the present invention is to provide a commercial vehicle automatic emergency steering method based on deep reinforcement learning, using a multi - task - divided training method and combining a variable Gaussian safety field. It enables automatic emergency steering when there are obstacles in front and the vehicle itself cannot achieve the braking target, avoiding rear - end or collision accidents, as Figure 1 shown, and specifically includes the following steps:
[0078] As Figure 1As shown, commercial vehicle modeling is carried out in Matlab. Generally, based on the linear two-degree-of-freedom three-axle commercial vehicle model, it is only to improve the handling stability of the vehicle under general operating conditions, so the nonlinearity of the tires is not considered, and the control method adopted is the linear control method based on the linear vehicle state equation. When the vehicle is in extreme operating conditions such as high-speed emergency steering, the side slip characteristics of the tires will show obvious nonlinearity. The tire nonlinearity has an important impact on the steering characteristics and driving stability of the vehicle. Therefore, under such conditions, the nonlinear side slip characteristics of the tires cannot be ignored. So it is necessary to establish a nonlinear vehicle model of a three-axle commercial vehicle.
[0079] As Figure 2 shown:
[0080] The lateral motion equation of the vehicle is:
[0081]
[0082] Yaw motion equation:
[0083]
[0084] Among them, m is the vehicle mass, u is the forward speed of the vehicle center of mass, β is the side slip angle of the vehicle center of mass, ω r is the yaw angular velocity, B is the distance between the intersections of the center lines of the two kingpins with the ground, I z is the yaw inertia moment, δ i is the tire steering angle, i = 1, 2, 3 respectively represent the front, middle and rear axles of the three-axle commercial vehicle, F yi is the side slip force of the ground on the tire, i = 1, 2, 3, 4, 5, 6 respectively represent the 6 tires of the three-axle commercial vehicle. a, b, c are the distances from the front, middle and rear axles to the center of mass, F yij represents the side slip moment of the j-th axle on the i-th axle of the tire.
[0085] Calculate the vertical loads of each wheel of the three-axle commercial vehicle, which are divided into two parts: static load and dynamic load:
[0086] (1) Static load
[0087] The mathematical equation of the vehicle dynamics model is:
[0088] Among them, z is the dynamic displacement of the vehicle vibration structure, M, C, K, f are the mass matrix, damping matrix, stiffness matrix and external force column vector.
[0089] Then the mathematical equation of the three-axle commercial vehicle at rest under its own gravity is: Kz j = G.
[0090] Among them, z jis the static displacement of the vehicle vibration structure, and G is the gravity in the directions of each static displacement. Since the vehicle's structural stiffness remains unchanged whether it vibrates or not within the linear range, the expression of K is as follows:
[0091]
[0092] where k 1 , k 2 , k 3 are the vertical stiffnesses of each suspension respectively.
[0093] The static displacement z j can be solved from the above mathematical equation, and then the normal reaction forces of the ground on the front axle, middle axle and rear axle are respectively:
[0094]
[0095]
[0096]
[0097] where k t1 , k t2 , k t3 are the vertical stiffnesses of each tire respectively.
[0098] When stationary, the normal reaction forces of the ground caused by the vertical loads on each wheel are respectively denoted as F zi0 (i = 1 to 6), then:
[0099]
[0100] where F zf is the normal reaction force of the ground on the front wheel, F zm is the normal reaction force of the ground on the middle wheel, and F zr is the normal reaction force of the ground on the rear wheel.
[0101] (2) Dynamic load
[0102] The distribution ratio coefficients of the vehicle gravity on each axle are respectively denoted as t 1 , t 2 , t 3 , then t 1 = G f / G, t 2 = G m / G, t 3 = G r / G. The distribution ratio of the lateral inertial force on each axle is the same as the distribution ratio of the vehicle gravity on each axle. Therefore, (ma y ) f = ma y t1 , (ma y ) m = ma y t 2 , (ma y ) r = ma y t 3 . Among them, G f , G m , G r respectively represent the separation of the vehicle gravity on the front axle, middle axle and rear axle, a y is the lateral acceleration, and ma y is the lateral inertia force.
[0103] Taking moments for different tires respectively, the vertical loads of each tire can be obtained:
[0104]
[0105] Among them, h g is the height from the center of mass to the ground, and F zi1 (i = 1 to 6) is the vertical load on each wheel during turning.
[0106] Next, calculate the sideslip angles of each tire of the three-axle commercial vehicle. Using the base point method for finding the velocities of each point in a planar figure, with the vehicle center of mass as the base point, the components of the velocities of each wheel center in the longitudinal and lateral directions can be obtained.
[0107] The obtained sideslip angles of each tire are respectively:
[0108]
[0109] Among them, ε i (i = 1 to 6) is the angle between u i (i = 1 to 6) and the ground (i.e., the x-axis), and u i (i = 1 to 6) is the velocity of the center of each tire of the vehicle. v is the lateral velocity of the vehicle center of mass.
[0110] The designed vehicle parameters are as follows:
[0111]
[0112] Among them, k 1 , k 2 , k 3 are the vertical stiffnesses of each suspension, and k t1 , k t2 , k ti are the vertical dampings of each suspension.
[0113] Perform decision-making and lateral and longitudinal control training. The present invention uses a multi-task reinforcement learning method to train a vehicle. The decomposed sub-tasks include lateral and longitudinal control tasks and decision-making tasks. Among them, the actions of the lateral and longitudinal control tasks are the steering wheel angle, braking, and throttle, all of which belong to the continuous action space. The Actor-Critic algorithm is selected for training. The characteristic of the decision-making task is that each decision-making behavior can be regarded as an action in the discrete space. The purpose of the decision-making network is to select the most valuable decision-making behavior according to the current environmental information. Here, it is more appropriate to select the method based on the value-based DQN series.
[0114] Obtain state variables. To obtain the navigation point coordinates for lateral and longitudinal control training, the present invention uses an effective and concise temporal bird's-eye view as the state variable of the policy network, which greatly improves the learning of the policy network and ensures the safety of the output trajectory.
[0115] The generation of the temporal bird's-eye view includes the following two steps: (1) According to the perception module of the autonomous vehicle, obtain the surrounding environmental information, including dynamic and static obstacles and lane lines. Use the prediction module (both lstm and GCN networks are available) to obtain the position information of dynamic obstacles within the time range of 0 to t in the future; (2) Generate feature bird's-eye views in three dimensions: lateral, longitudinal, and time, from the information obtained by the perception module and prediction; end The time range of the three-dimensional temporal bird's-eye view matrix is (40, 400, 80). Among them, the first dimension 40 represents the lateral range of 10m on both sides of the reference line, and the lateral displacement interval is 0.5m; the second dimension 400 represents the range of 200m longitudinally forward with the vehicle itself as the origin, and the longitudinal displacement interval is 0.5m. The third dimension 80 represents the time range within the next 8s, and the time interval is 1s. Specifically, when the point [α, β, γ] in the temporal bird's-eye view matrix is -1, it means that there is an obstacle or an unnavigable area at this point in time and space; when the point [α, β, γ] in the temporal bird's-eye view matrix is 0, it means that this point is a navigable area in time and space; when the point [α, β, γ] in the temporal bird's-eye view matrix is 1, it means that this point is a point on the reference line.
[0116] To ensure the rationality of the coordinates of the navigation point, the current vehicle position is input into the policy network as a dynamic feature. The policy network π
[0117] (z, p) specifically includes two parts: a convolutional feature extraction network and a fully connected network. Among them, z is the input state variable of the policy network, including the temporal bird's-eye view matrix and the current position of the vehicle; p is the output of the policy network, that is, the navigation point p = (x θ , y i , i); θ are the weight and bias parameters of the network. The convolutional module of the policy network includes one convolutional layer and three fully connected layers. The convolutional layer Conv1 consists of convolutional kernels of size 2*2, the number of convolutional kernels is 9*32, the stride is stride = 1, and the activation function is ReLU; the first fully connected layer is the fully connected layer FC1 and the fully connected layer FC1-σ. The fully connected layer FC1 processes the output result of the flattened convolutional layer Conv1, with a size of 2*2*9*32, and the activation function is ReLU; the output of the fully connected layer FC1-σ is the historical trajectory information of the ego vehicle in the past few moments, with a size of 1024*1, and the activation function is ReLU; the second fully connected layer is the fully connected layer FC2, which processes the concatenated state quantities of the fully connected layer FC1 and the fully connected layer FC1-σ, with a size of 4096*1, and the activation function is ReLU; the third fully connected layer is the fully connected layer FC3, which processes the state quantity output by the fully connected layer FC2, with a size of 1024*1, and the activation function is Tanh. Finally, the fully connected layer FC3 outputs the state feature z.
[0118] Lateral control training. Using the coordinates (x i , y i ) of the obtained navigation points, the heading deviation and the vehicle speed v and acceleration of the controlled vehicle as state quantities:
[0119]
[0120] s lane_keep is the state quantity obtained during the lane keeping training of the agent.
[0121] The action is only the steering wheel angle a steer ∈[-1, 1]. Regarding the design of the reward function r lane_keep for the lane keeping task part, the present invention uses the lateral error x 0 of the current coordinates of the vehicle and the heading angle deviation as evaluation indicators:
[0122]
[0123] λ 1 and λ 2 are the weights of the two parts of the reward function.
[0124] If the lateral deviation of the current position of the autonomous vehicle during training is greater than the set maximum lateral deviation threshold x 0max then end the iterative training of the current round and start the training of the next round.
[0125] Longitudinal control training. The longitudinal trajectory tracking control task uses the vehicle speed v and acceleration The speed v of the vehicle ahead l , acceleration The distance d from the vehicle ahead and the desired speed v of the current vehicle des are state variables:
[0126]
[0127] s acc is the state variable obtained during the longitudinal following control training of the agent.
[0128] The output action a of the agent acc ∈[-1,1], including the throttle action a throttle and the brake action a brake :
[0129]
[0130] For the longitudinal control task, the reward function is designed as:
[0131]
[0132] where d is the real-time distance from the vehicle ahead, d des is the desired distance from the vehicle ahead, d safe is the safety distance from the vehicle ahead. When the distance between the intelligent vehicle and the vehicle ahead is less than the safety distance, the reward is -100 and the current interaction stops and the next round of interaction starts. During the longitudinal training, the speed v of the vehicle ahead l and the desired speed v of the current vehicle des are randomly given in each round so that the trained model can be generalized to more complex situations.
[0133] Decision task training. The decision task in a complex environment is trained based on the premise that both the lateral and longitudinal trajectory tracking control tasks can be well completed. The decision behaviors defined in the present invention are emergency braking and emergency steering. Compared with the lateral and longitudinal control tasks, the action space of the decision task is discrete.
[0134]
[0135] When a decision is 0, the decision module selects emergency braking; when a decision is 1, the decision module selects to make an emergency turn to the right.
[0136] To achieve the stability of vehicle control and the safety of obstacle avoidance during emergency obstacle avoidance, the present invention designs the reward function of the decision-making using a variable Gaussian safety field. When the obstacle is outside the expansion domain, the vehicle can brake by taking braking measures. When the obstacle is in the expansion domain, the vehicle takes steering and lane-changing measures. When the obstacle is in the core domain and the restricted domain, a high-probability collision will occur. The reward function for the decision-making part is as follows:
[0137]
[0138] Among them, d lon,min 、d lon,mid 、d lon,max are the longitudinal safety distances of the core domain, restricted domain, and expansion domain of the recognizable Gaussian safety field respectively. l v is the length of the vehicle model, w v is the width of the vehicle model, l′ v is the length of the vehicle model during vehicle movement, w v ′ is the width of the vehicle model during vehicle movement. Among them:
[0139]
[0140] Among them, is the velocity vector of vehicle movement, k v is the adjustment factor, and there is 0 < k v < 1 or -1 < k v < 0, and its sign corresponds to the front and back directions of movement. ξ is the yaw angle of the vehicle.
[0141] In this paper, through actual vehicle measurement, k v = 0.3.
[0142] As Figure 3 shown, the variable Gaussian safety field abstracts the static vehicle as a rectangle, with its length being l v , width being w v , and the risk center O(x 0 , y 0 ) being its geometric center. The static safety field is described using a two-dimensional Gaussian function:
[0143]
[0144] In the formula, C a is the field strength coefficient, a x and b y are functions related to the vehicle shape. The main control parameter of the shape of the static safety field is anisotropy:
[0145]
[0146] The parameter ε can be equivalently represented by the aspect ratio φ = a x / b y = l v / w v .
[0147] The direction of the safety field is a vector emitted from the risk center. The top-down projection of its equipotential lines is a series of ellipses. Each region represents a different risk state. Their sizes are related to the vehicle shape and motion state and can be determined based on the σ parameter of the Gaussian function. When the vehicle is in motion, the risk center will shift along with the vector and the new risk center becomes O'(x' 0 , y 0 ') and there is:
[0148]
[0149] where is the velocity vector of the vehicle's motion, k v is the adjustment factor and 0 < k v < 1 or -1 < k v < 0, and its sign corresponds to the forward and backward directions of the motion. β is the angle between the transfer vector and the x-axis.
[0150] Under the action of the risk center transfer, a virtual vehicle is formed, with a length of l' v , a width of w' v , and its geometric center is (x' 0 , y' 0 ), thus establishing its dynamic safety field:
[0151]
[0152] where a' x and b' y are parameters related to the vehicle shape and motion state. The new aspect ratio is expressed as φ' = a' x / b' y = l' v / w' v .
[0153] Obviously, the Gaussian safety field is variable. The aspect ratio of the virtual vehicle will change with the change of the vehicle's motion state, and thus significantly change the core domain, restricted domain, and extended domain of the Gaussian safety field. When necessary, the safety field can be moderately adjusted in terms of expansion and contraction by adding human driving characteristics such as the driving style factor. Figure 3 In
[0154] In summary, for the emergency braking and steering problem of commercial vehicles, the present invention uses Matlab to build a model and conducts joint simulation with Carla, solving the problems in model-free reinforcement learning that the high center of gravity, easy rollover, and large mass and difficult braking of commercial vehicles cannot be reflected, ensuring the reliability of the simulation experiment. The present invention uses a reinforcement learning method with multi-task partitioning, greatly improving the training efficiency. At the same time, a variable Gaussian safety field strategy is introduced to ensure that the vehicle has high safety during decision-making and control.
[0155] The series of detailed descriptions listed above are only specific descriptions of the feasible implementation manners of the present invention, and they are not intended to limit the protection scope of the present invention. Any equivalent manners or changes that do not depart from the technology created by the present invention should be included in the protection scope of the present invention.
Claims
1. A training method for an emergency steering control strategy network model of an autonomous commercial vehicle, including a network structure and related training methods, characterized in that, the network structure is as follows: Policy network p θ (z, p) consists of two parts: a convolutional feature extraction network and a fully connected network. Here, z is the input state quantity of the policy network, including the temporal bird's-eye view matrix and the current position of the ego vehicle; p is the output of the policy network, that is, the navigation point of the commercial vehicle p = (x i , y i ); θ is the weight and bias parameter of the network; the policy network specifically includes one convolutional layer and three fully connected layers. The convolutional layer Conv1 is composed of convolutional kernels of size 2*2, the number of convolutional kernels is 9*32, the stride stride = 1, and the activation function is ReLU; the first fully connected layer is the fully connected layer FC1 and the fully connected layer FC1-σ. The fully connected layer FC1 processes the output result of the flattened convolutional layer Conv1, with a size of 2*2*9*32, and the activation function is ReLU; the output of the fully connected layer FC1-σ is the historical trajectory information of the ego vehicle in the past few moments, with a size of 1024*1, and the activation function is ReLU; the second fully connected layer is the fully connected layer FC2, which processes the concatenated state quantity of the fully connected layer FC1 and the fully connected layer FC1-σ, with a size of 4096*1, and the activation function is ReLU; the third fully connected layer is the fully connected layer FC3, which processes the state quantity output by the fully connected layer FC2, with a size of 1024*1, and the activation function is Tanh. Finally, the fully connected layer FC3 outputs the state feature z; the training method is as follows: including lateral control training; specifically as follows: With the coordinates (x i , y i ) of the navigation point, the course deviation and the vehicle speed v and acceleration as state variables: s lane_keep The state variables obtained during the lane-keeping training of the agent; The action is the steering wheel angle a steer ∈[-1,1]. The design of the reward function for this part is based on the lateral error x of the vehicle's current coordinates 0 and the heading angle deviation as evaluation indicators: λ 1 and λ 2 are the weights of the two parts of the reward function; If the lateral deviation of the current position of the autonomous vehicle during training is greater than the set maximum lateral deviation threshold x 0max then end the iterative training of the current round and proceed to the next round of training; also including longitudinal control training, specifically as follows: The longitudinal trajectory tracking control task uses the vehicle speed v, acceleration of the current vehicle vehicle speed v of the vehicle ahead l , acceleration distance d to the vehicle ahead, and the desired vehicle speed v of the current vehicle des as state variables: s acc The state variables obtained during the longitudinal car-following control training of the agent; The output action a of the agent acc ∈[-1, 1], including the throttle action a throttle and the brake action a brake : For the longitudinal control task, the reward function is designed as: Among them, d is the real-time distance from the vehicle ahead, d des is the desired distance from the vehicle ahead, d safe is the safety distance from the vehicle ahead. When the distance between the intelligent vehicle and the vehicle ahead is less than the safety distance, the reward is -100, and at the same time, the current interaction is stopped and the next round of interaction is started. When conducting longitudinal training, the vehicle speed v l of the vehicle ahead and the desired vehicle speed v des of the current vehicle are randomly given in each round, so that the trained model can be generalized to more complex situations; also including decision-making behavior training; the decision-making behaviors include emergency braking and emergency steering, When a decision is 0, the decision-making module selects emergency braking; when a decision is 1, the decision-making module selects an emergency turn to the right.
2. The training method for an emergency steering control strategy network model of an autonomous commercial vehicle according to claim 1, characterized in that, the state quantity of the policy network is a three-dimensional time-series bird's-eye view; the size of the three-dimensional time-series bird's-eye view matrix is (40, 400, 80), where the first dimension 40 represents the lateral range of 10m on each side of the reference line, and the lateral displacement interval is 0.5m; the second dimension 400 represents the range of 200m longitudinally forward with the vehicle itself as the origin, and the longitudinal displacement interval is 0.5m, and the third dimension 80 represents the time range within the next 8s, and the time interval is 1s. When the point [α, β, γ] in the time-series bird's-eye view matrix is -1, it means that there is an obstacle or an inoperable area at this point in space-time. When the point [α, β, γ] in the time-series bird's-eye view matrix is 0, it means that this point is an operable area in space-time. When the point [α, β, γ] in the time-series bird's-eye view matrix is 1, it means that this point is a point on the reference line.
3. The training method for an emergency steering control strategy network model of an autonomous commercial vehicle according to claim 1, characterized in that, when the decision-making behavior is emergency braking to avoid obstacles, a variable Gaussian safety field is used to design the reward function of the decision. When the obstacle is outside the expansion domain, the vehicle can brake by taking braking measures. When the obstacle is in the expansion domain, the vehicle takes steering and lane-changing measures. When the obstacle is in the core domain and the restricted domain, a high-probability collision will occur. The reward function is designed as follows: Among them, d lon,min 、d lon,mid 、d lon,max are the longitudinal safety distances of the core region, restricted region, and extended region of the variable Gaussian safety field respectively, l v is the length of the vehicle model, w v is the width of the vehicle model, l′ v is the length of the vehicle model during vehicle movement, w′ v is the width of the vehicle model during vehicle movement; where: wherein, is the velocity vector of the vehicle movement, k v is the adjustment factor, and 0 < k v < 1 or -1 < k v < 0, whose sign corresponds to the front-back direction of the movement, and ξ is the yaw angle of the vehicle.
4. The training method for an emergency steering control strategy network model of an autonomous commercial vehicle according to claim 3, characterized in that, The variable Gaussian safety field abstracts a static vehicle as a rectangle with a length of l v , a width of w v , and the risk center O(x 0 , y 0 ) is its geometric center. The two-dimensional Gaussian function is used to describe its static safety field: where C a is the field strength coefficient, a x and b y are functions related to the vehicle shape, and the control parameter of the shape of the static safety field is anisotropy: The parameter ε is equivalently expressed by the aspect ratio φ = a x / b y = l v / w v .
5. The training method for an emergency steering control strategy network model of an autonomous commercial vehicle according to claim 3 or 4, characterized in that, The direction of the variable Gaussian safety field is a vector emitted from the risk center. The top-down projection of its equipotential lines is a series of ellipses. Each region represents a different risk state, and their sizes are related to the vehicle shape and motion state, and are determined based on the σ parameter of the Gaussian function. When the vehicle is in a moving state, the risk center will move along with the vector and transfer. The new risk center becomes O′(x′ 0 , y′ 0 ), and there is: Among them, is the velocity vector of vehicle movement, k v is the adjustment factor, and 0 < k v < 1 or -1 < k v < 0, and its sign corresponds to the front-back direction of the movement. β is the angle between the transfer vector and the x coordinate axis; Under the action of the risk center transfer, a virtual vehicle is formed, with a length of l′ v , a width of w′ v , and its geometric center is (x′ 0 , y′ 0 ), thus establishing its dynamic safety field: where a′ x and b′ y are parameters related to the vehicle's shape and motion state, and the new aspect ratio is expressed as φ′ = a′ x / b′ y = l′ v / w′ v .
6. An emergency steering control simulation method for an autonomous commercial vehicle, characterized in that, use Matlab to establish an autonomous commercial vehicle model, and then import the autonomous commercial vehicle model into the Carla simulator, and use the policy network model trained by the training method described in claim 1 in the Carla simulator to simulate the autonomous commercial vehicle model.
Citation Information
Patent Citations
Intelligent vehicle neural network dynamics model, reinforcement learning network model and automatic driving training method thereof
CN113420368A
Automatic driving track planning system and method based on space-time aerial view and strategy gradient algorithm
CN114407925A