Autonomous Driving Vehicle Control Method Based on Depth Prediction Network and Deep Reinforcement Learning
By adopting deep prediction network and deep reinforcement learning methods in autonomous vehicles, the problems of reward value setting and delay impact are solved, and the safe and efficient autonomous driving control of the vehicle on the highway is achieved.
Patent Information
- Application Number
- CN202211316067.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-10-26
AI Technical Summary
In autonomous vehicles, deep reinforcement learning algorithms are difficult to effectively set reward values, resulting in the vehicle becoming radical or conservative, affecting the traffic efficiency; at the same time, the complexity of data transmission delay and vehicle trajectory prediction also limits the application of the algorithm.
Using an autonomous vehicle control method based on deep prediction network and deep reinforcement learning, a reward function and loss function are designed until the preset reward value or training round number is reached.
It realizes safe and efficient autonomous driving control of vehicles on highways, reduces the impact of time delay on the algorithm, and improves the accuracy and traffic efficiency of vehicle trajectory prediction.
Smart Images

Figure CN115629608B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent traffic control, and mainly relates to a control method for an autonomous vehicle based on a deep prediction network and deep reinforcement learning. Background Art
[0002] Autonomous driving is divided into five levels according to the degree of automation. In recent years, certain progress has been made in the field of autonomous driving, but there is still a certain gap from fully autonomous driving at level 5. Currently, the research on autonomous driving still focuses on different scenarios, and different algorithm designs will be adopted for different scenarios. Among them, following, lane-changing and overtaking on highways are important research scenarios.
[0003] Vehicles driving on highways need to balance safety and efficiency. Currently, some research has applied deep reinforcement learning technology to the control of autonomous vehicles. However, as a core issue of deep reinforcement learning, how to set the reward value has become one of the main problems in this research. Excessive encouragement for acceleration will train aggressive vehicles that will ignore the risk of collision in order to obtain higher speed rewards, while excessive punishment for collisions will make the vehicles conservative and reluctant to increase speed, resulting in low traffic efficiency of the vehicles. At the same time, considering the particularity of the autonomous driving problem, establishing the physical control from the upper-layer decision actions to the lower layer is also a major factor restricting the application of deep reinforcement learning in autonomous driving.
[0004] In the scenario of autonomous driving, the perception of the vehicle's surrounding environment comes from its own sensors and the interaction with surrounding vehicles. Therefore, the time delay of data transmission is an objective and non-negligible problem. Since the speed of vehicles on highways is generally relatively fast, the impact of time delay is also relatively large. If the time delay is not compensated, the decisions made by the deep reinforcement learning algorithm based on the lagged state signals may also be lagged, which reduces the safety and reliability of the algorithm. Predicting the trajectory of the vehicle in front of the controlled vehicle to compensate for the time delay is an intuitive solution. The biggest problem with this solution is that the trajectory of the vehicle is not only related to its own historical trajectory but also related to the trajectories of surrounding vehicles. The same historical trajectory may result in different future trajectories due to different surrounding environments. How to characterize the impact of the interaction of the surrounding environment on vehicle trajectory prediction is one of the important research difficulties of this solution. Summary of the Invention
[0005] The present invention precisely aims at the problems existing in the prior art and provides an autonomous vehicle control method based on a depth prediction network and deep reinforcement learning. First, a discrete controller at the upper layer corresponding to the control signals at the vehicle's lower layer is defined; hyperparameters are set, a depth prediction network and a double deep Q network based on an encoder-decoder framework are built, and the weights of each network are initialized; then, deep reinforcement learning training is carried out on the controlled vehicle, a reward function is designed, and the weights of the network are iteratively updated until the reward value obtained by the controlled vehicle in a round of driving behavior reaches a preset level or the number of training rounds reaches a preset value; the historical data collected is preprocessed, the data and labels are determined according to the time delay situation, the feature data of the vehicle in front of the controlled vehicle is transformed into graph data to provide a training set and a validation set for the depth prediction network, and training is carried out on the training set until the loss function value on the validation set no longer decreases; finally, the trained depth prediction network and double deep Q network are deployed into the controlled vehicle, the position and speed of the vehicle in front of the controlled vehicle are predicted by the depth prediction network, the prediction information together with the speed and position information of the controlled vehicle is expanded into a column vector, which is used as the input of the double deep Q network to obtain the action with the highest action value at the current moment, and then this discrete action at the upper layer is mapped into the physical control signal at the lower layer through a reference value and a proportional controller to achieve the autonomous driving control of the vehicle. To achieve the above object, the technical solution adopted by the present invention is: an autonomous vehicle control method based on a depth prediction network and deep reinforcement learning, comprising the following steps:
[0006] S1: Define a discrete controller at the upper layer corresponding to the control signals at the vehicle's lower layer. The discrete controller at the upper layer includes at least five action instructions: left lane change, keep, right lane change, acceleration, and deceleration. Map the discrete action at the upper layer into the physical control signal at the lower layer, and the physical control signal at the lower layer is obtained by a proportional controller;
[0007] S2: Set hyperparameters, build a depth prediction network and a double deep Q network based on an encoder-decoder framework, and initialize the weights of each network;
[0008] S3: Carry out deep reinforcement learning training on the controlled vehicle, design a reward function, the controlled vehicle interacts with the environment to obtain a reward value, and iteratively update the weights of the network until the reward value obtained by the controlled vehicle in a round of driving behavior reaches a preset level or the number of training rounds reaches a preset value, and terminate the training;
[0009] S4: Preprocess the historical data collected, determine the data and labels according to the time delay situation, transform the feature data of the vehicle in front of the controlled vehicle into graph data to provide a training set and a validation set for the depth prediction network, and carry out training on the training set until the loss function value on the validation set no longer decreases, and terminate the training;
[0010] S5: Deploy the trained depth prediction network and the double deep Q network to the controlled vehicle. Use the depth prediction network to predict the position and speed of the vehicle in front of the controlled vehicle. Unfold the prediction information together with the speed and position information of the controlled vehicle into a column vector, which is used as the input of the double deep Q network to obtain the action with the highest action value at the current moment. Then map this upper-layer discrete action into a lower-layer physical control signal through a reference value and a proportional controller to achieve the autonomous driving control of the vehicle.
[0011] As an improvement of the present invention, in the step S1, the kinematic model of the vehicle is defined as:
[0012]
[0013]
[0014]
[0015]
[0016] β = tan -1 (1 / 2tanδ),
[0017] where (x, y) are the coordinates of the vehicle in the Frenet coordinate system; v is the forward speed of the vehicle; ψ is the heading angle; β is the slip angle at the center of gravity; a is the acceleration command; δ is the front-wheel steering command;
[0018] a and δ are obtained by the proportional controller. The specific expression of the acceleration control amount a is:
[0019] a = K p (v r - v),
[0020] where v r is the desired speed; v is the current speed;
[0021] When the instruction is to accelerate, the given speed reference value v r is greater than the current speed value, and K p is the controller gain;
[0022] The specific expression of the front-wheel steering command δ is:
[0023]
[0024]
[0025] ψ r = ψ L + Δψ r ,
[0026]
[0027] v lat,r = -K p,lat Δ lat ,
[0028] where ψ L is the lane orientation; v lat,r is the lateral velocity command; Δψ r is the required heading change corresponding to the controlled vehicle; K p,lat and K p,ψ are controller gains; obtained from the action mapping of the upper-layer reinforcement learning controller,
[0029] When the command is to change lanes, the lateral position Δ lat of the lane centerline will change according to the command requirements;
[0030] The desired speed v r and the lateral position Δ lat of the vehicle relative to the lane centerline are both obtained from the action mapping of the upper-layer reinforcement learning controller.
[0031] As an improvement of the present invention, in step S2, in the depth prediction network of the constructed encoder-decoder framework, the encoder is formed by stacking convolutional long short-term memory network (convlstm) modules, and the input of the encoder is the historical trajectories of all vehicles to be predicted in front of the controlled vehicle; the decoder consists of a channel attention mechanism and a fully connected layer; in the double deep Q network, both the main network and the target network are composed of two layers of fully connected layers, its input is the current state, which is composed of the positions and speeds of the controlled vehicle and the nearest vehicle in front of it, and the output is the upper-layer discretized action of the controlled vehicle.
[0032] As another improvement of the present invention, in step S2, assuming that the number of all vehicles to be predicted in front of the controlled vehicle is m, where m is a hyperparameter that needs to be adjusted in deep reinforcement learning, collect the position and speed information of the vehicles in front, that is, collect as the observation information of the i-th vehicle in front, where and represent the longitudinal coordinate and lateral coordinate of the i-th vehicle in front at time t, and represent the longitudinal speed and lateral speed of the i-th vehicle in front at time t, respectively.
[0033] As another improvement of the present invention, in step S3, a double deep Q network is used for the training of deep reinforcement learning, and the reward function is set as:
[0034]
[0035] where k is the speed coefficient; v max and v min are respectively the maximum and minimum values of the speed of the controlled vehicle; represents the normalized tangential speed, which is used as a reward for the controlled vehicle to maintain a high speed along the lane.
[0036] As another improvement of the present invention, in the step S4, the road information of the vehicle ahead is converted into graph data of C×H×W, where H is the height of the graph data, representing the number of lanes; W is the length of the graph data, which is obtained by discretizing the lane at intervals of l meters l is the average speed of other vehicles ahead, and L is the maximum possible distance between the vehicle ahead and the controlled vehicle during the observation sequence and the prediction process;
[0037] If the observation period is δ seconds, the length of the historical observation sequence for prediction is s, and the time delay is d seconds, then L = X+(δs + d)*v m , where X is the maximum distance of the field of view of the controlled vehicle, and v m is the maximum speed of other vehicles ahead; C is the characteristic length; if there is no vehicle in the discretized grid, the values are all 0 numerically, and if there is a vehicle, it is represented by the normalized vector, where L y represents the lane width, is the relative position of the i-th vehicle at time t in this grid, and are respectively the maximum values of the lateral speed and longitudinal speed of the vehicle, is the relative speed value of the i-th vehicle at time t.
[0038] As another improvement of the present invention, the loss function in the step S4 is a weighted mean square error function, and the specific expression is:
[0039]
[0040] where, is the predicted value of the trajectory of the vehicle ahead described by the graph data, that is, the output value of the depth prediction network with X i as the input, Q∈W×W is the weight matrix, and b is the size of the batch training.
[0041] As a further improvement of the present invention, in the depth prediction network of the step S5, among the graph data of the controlled vehicle, the feature quantity of the relative speed of the longitudinal movement whose predicted value denotes the predicted value of the relative position of the i-th vehicle at time t obtained from the depth prediction network, and the predicted value of the position of the i-th vehicle at time t is Use denotes the predicted value of the relative speed of the i-th vehicle at time t obtained from the depth prediction network, and the predicted value of the speed of the i-th vehicle at time t is
[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0043] (1) The present invention transforms the trajectories of historical vehicles into graph data, which can characterize the influence of interactions between vehicles on vehicle trajectory prediction. At the same time, the output of the prediction network is the trajectories of all vehicles to be predicted ahead, reducing the overall model parameters and saving training and computational costs. Compared with equipping each vehicle with a separate prediction network, the computational cost is reduced, and the prediction accuracy of vehicle lateral movement is improved.
[0044] (2) The present invention introduces a channel attention mechanism into the prediction network of the encoder-decoder framework, enhancing the model's attention to features with greater influence in trajectory prediction and improving the prediction accuracy.
[0045] (3) The present invention uses a deep reinforcement learning algorithm as the vehicle controller and obtains the predicted value of the current state using a depth prediction network, compensating for the time delay and reducing the adverse effects brought by the time delay. Brief Description of the Drawings
[0046] Figure 1 is a flowchart of the steps of the autonomous vehicle control method based on the depth prediction network and deep reinforcement learning of the present invention;
[0047] Figure 2 is a block diagram of the algorithm structure of the autonomous vehicle control method based on the depth prediction network and deep reinforcement learning of the present invention;
[0048] Figure 3 is a training curve graph of vehicle lateral movement when a long short-term memory network is equipped for each vehicle separately;
[0049] Figure 4 is a comparison graph of training curves with and without the channel attention mechanism of the depth prediction network proposed in Embodiment 2 of the present invention;
[0050] Figure 5 is a schematic diagram of the performance of the controlled vehicle after compensating for the time delay using the depth prediction network proposed by the present invention. Detailed Embodiments
[0051] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0052] Example 1
[0053] An autonomous vehicle control method based on a depth prediction network and deep reinforcement learning, as Figure 1 shown, includes the following steps:
[0054] Step S1: Define the control signals of the upper-layer discrete controller corresponding to the vehicle's lower layer;
[0055] Define the control signals of the upper-layer discrete controller corresponding to the vehicle's lower layer. The upper-layer deep reinforcement learning discrete controller defines five action commands: left lane change, keep, right lane change, acceleration, and deceleration. These five commands are respectively mapped to the lower-layer controller and reflected as reference values. First, consider the vehicle's kinematic model as:
[0056] · x = v cos(ψ + β), (1)
[0057]
[0058]
[0059]
[0060] β = tan -1 (1 / 2 tan δ), (5)
[0061] where (x, y) are the coordinates of the vehicle in the Frenet coordinate system, v is the vehicle's forward speed, ψ is the heading angle, β is the slip angle at the center of gravity, a is the acceleration command, δ is the front-wheel steering command, a and δ are obtained by a proportional controller, and the specific expression of the acceleration control quantity a is:
[0062] a = K p (v r - v), (6)
[0063] where v r is the desired speed, obtained by mapping the actions of the upper-layer reinforcement learning controller, v is the current speed. When the command is acceleration, the given speed reference value v r is greater than the current speed value, and K p is the controller gain;
[0064] The specific expression of the front-wheel steering command δ is:
[0065]
[0066]
[0067] ψ r = ψ L + Δψr , (9)
[0068]
[0069] v lat,r = -K p,lat Δ lat , (11)
[0070] where ψ is the heading of the current controlled vehicle, ψ L is the lane orientation, v lat,r is the lateral velocity command, Δψ r is the required heading change corresponding to the controlled vehicle, K p,lat and K p,ψ are controller gains, which need to be adjusted according to the actual situation, Δ lat is the lateral position of the vehicle relative to the center line of the lane, which is obtained by mapping the actions of the upper-level reinforcement learning controller. When the command is to change lanes, the lateral position Δ lat of the center line of the lane will change according to the command requirements.
[0071] S2: Set hyperparameters, build a deep prediction network and a double deep Q network based on the encoder-decoder framework, and initialize the weights of each network using the He initialization method;
[0072] As Figure 2 shown, build a deep prediction network with an encoder-decoder framework. The encoder part is formed by stacking convlstm modules of convolutional long short-term memory networks. The input of the encoder is the historical trajectories of all vehicles to be predicted in front of the controlled vehicle. The convlstm module can effectively extract the interactions between vehicles and participate in the prediction of the future trajectories of vehicles as the spatial information of the vehicles;
[0073] The decoder consists of a channel attention mechanism and a fully connected layer. The channel attention mechanism is introduced to assign attention weights to the features of different channels, enabling the decoder to focus on the more important channel features, and then aggregating the features through the fully connected layer to obtain the predicted values of the current positions of all vehicles in front of the controlled vehicle;
[0074] Build a double deep Q network. Both the main network and the target network are composed of two layers of fully connected layers. Its input is the current state, which consists of the positions and speeds of the controlled vehicle and the m nearest vehicles in front of it, that is, S t = (S e , S i ), represents the position and speed information of the controlled vehicle, where and represent the longitudinal coordinate and lateral coordinate of the controlled vehicle at time t, and respectively represent the longitudinal speed and lateral speed of the controlled vehicle at time t; represent the position and speed information of the m closest vehicles in front of the controlled vehicle, where and represent the longitudinal coordinate and lateral coordinate of the i-th vehicle in front at time t, and respectively represent the longitudinal speed and lateral speed of the i-th vehicle in front at time t. If there are less than m vehicles within the field of view, they are filled with zeros. The output is the upper-layer discretized action of the controlled vehicle.
[0075] S3: Conduct deep reinforcement learning training on the controlled vehicle, design a reward function, have the controlled vehicle interact with the environment to obtain a reward value, and iteratively update the weights of the network until the reward value obtained by the controlled vehicle in a round of driving behavior reaches a preset level or the number of training rounds reaches a preset value, then terminate the training;
[0076] Use a double deep Q-network for deep reinforcement learning training, and set the reward function as:
[0077]
[0078] where k is the speed coefficient. The larger k is, the more it encourages the controlled vehicle to maintain a high speed. v max and v min are the maximum and minimum values of the speed of the controlled vehicle respectively. Therefore represents the normalized tangential speed, which is used as the reward for the controlled vehicle to maintain a high speed along the lane.
[0079] S4: Preprocess the collected historical data, determine the data and labels according to the time delay situation, convert the feature data of the vehicles in front of the controlled vehicle into graph data to provide a training set and a validation set for the depth prediction network, and conduct training on the training set until the loss function value on the validation set no longer decreases, then terminate the training;
[0080] Use a depth prediction network based on graph data to compensate for the impact of time delay. The depth prediction network uses convlstm as the feature extractor. Therefore, the input data needs to be converted into graph data before training can be carried out. It specifically includes the following two steps:
[0081] Step 4.1: Collect the position and speed information of the vehicles in front as observation data, that is, collect as the observation information of the i-th vehicle in front; collect the observation data of the m closest vehicles in front of the controlled vehicle. m is a hyperparameter that needs to be adjusted in deep reinforcement learning;
[0082] First, convert the road information within a certain range in front of the controlled vehicle into graph data of C×H×W. Here, H is the height of the graph data, which represents the number of lanes; W is the length of the graph data, obtained by discretizing the lanes at intervals of l meters per grid. l is the average speed of other vehicles in front, and L is the maximum possible distance of the vehicle in front from the controlled vehicle during the observation sequence and prediction process. It is related to the delay size, observation period, and the length of the observation sequence. If the observation period is δ seconds, the length of the historical observation sequence for prediction is s, and the delay is d seconds, then L = X+(δs + d)*v m , where X is the maximum distance of the controlled vehicle's field of view, and v m is the maximum speed of other vehicles in front; C is the characteristic length. Considering the use of vehicle position and speed information, C = 4. If there is no vehicle in the discretized grid, the values are all 0 numerically. If there is a vehicle, it is represented by the normalized vector, where L y represents the lane width. Therefore, is the relative position of the i-th vehicle at time t in this grid, and are the maximum values of the vehicle's lateral speed and longitudinal speed respectively. Therefore, is the relative speed value of the i-th vehicle at time t;
[0083] Step 4.2: After converting all historical record data into graph data, prepare the input data X i and output label Y i of the network according to the historical observation sequence length s and delay time d, and split the training set and validation set according to 7:3. Select the weighted mean square error function as the loss function, and the specific expression is:
[0084]
[0085] where, is the predicted value of the trajectory of the vehicle in front described by graph data, that is, the output value of the depth prediction network with X i as the input. Q ∈ W×W is the weight matrix, and b is the size of the batch training volume
[0086] S5: Deploy the trained depth prediction network and double deep Q network into the controlled vehicle. Predict the position and speed of the vehicle in front of the controlled vehicle through the depth prediction network, expand the prediction information together with the speed and position information of the controlled vehicle into a column vector, use it as the input of the double deep Q network, obtain the action with the highest action value at the current moment, and then map this upper-layer discrete action into a lower-layer physical control signal through a reference value and a proportional controller to achieve the automatic driving control of the vehicle.
[0087] Deploy the trained depth prediction network and the double deep Q-network to the controlled vehicle. Due to latency, the controlled vehicle cannot obtain the position and speed information of other vehicles ahead at the current moment. Therefore, the depth prediction network is required to predict the position and speed information of the vehicles ahead at the current moment:
[0088] First, convert the latest s observations of the m vehicles closest to the controlled vehicle ahead into graph data and input it into the trained depth prediction network to obtain the predicted values of the current positions and speeds of the m vehicles ahead, which are also described by graph data. In the graph data, the characteristic quantity of the relative speed of longitudinal motion of the predicted value The grid with a predicted value greater than 0.4 is considered to have a vehicle, and according to its relative position and relative speed characteristics, it is converted into general position and speed information; use to represent the predicted value of the relative position of the i-th vehicle at time t obtained from the depth prediction network. Then, the predicted value of the position of the i-th vehicle at time t is Use to represent the predicted value of the relative speed of the i-th vehicle at time t obtained from the depth prediction network. Then, the predicted value of the speed of the i-th vehicle at time t is Expand the restored predicted information together with the speed and position information of the controlled vehicle into a column vector, which is used as the input of the deep Q-network to obtain the action with the highest action value at the current moment. Then, map this upper-layer discrete action into a lower-layer physical control signal through a reference value and a proportional controller. The controlled vehicle can complete normal lane-changing and overtaking actions in the presence of latency, thus completing the control of autonomous driving.
[0089] Embodiment 2
[0090] In this embodiment, the car-following, lane-changing, and overtaking scenarios on a highway in the highway_env environment are used for simulation verification. Assume that the observation and control frequencies of the vehicle are both 1 Hz. The vehicle observes its own state in real time, but the observation of the vehicles ahead depends on interaction and there is a latency of 1 second. The average speed of other vehicles ahead is 20 m / s, the maximum speed is 23 m / s, and the field of view of the controlled vehicle is 180 meters ahead
[0091] The objective of the present invention is to predict the positions and speeds of other vehicles ahead 1 second later based on the historical trajectories of the vehicles ahead, compensate for the influence brought by latency, and use the predicted values of the positions and speeds of other vehicles ahead together with the real-time position and speed information of the controlled vehicle as the input state of deep reinforcement learning. Calculate the optimal upper-layer discrete action in the current state using the deep Q-network, and map this action into a lower-layer physical control signal to achieve lane-changing and overtaking of the controlled vehicle.
[0092] Step S1: Define the control signals of the upper-layer discrete controller corresponding to the vehicle's lower layer
[0093] According to equations (1)-(11), map the upper-layer discrete actions into the lower-layer physical control signals. The upper-layer discrete control actions give the target values, while the lower-layer physical control signals are obtained by a proportional controller. The controller parameters are set as follows: K p = 1.87, K p,lat = 1.67 and K p,ψ = 5.
[0094] Considering a series of actions such as overtaking, the maximum speed of the controlled vehicle is set to 30 m / s, and the minimum speed is 20 m / s. The set values of the speed are divided into three gears within this range. When the discrete action gives an acceleration command, the set value of the speed increases by one gear until the maximum value. When a deceleration command is given, the speed is set to decrease by one gear until the minimum value.
[0095] Step S2: Hyperparameter setting and data initialization
[0096] Considering the observation frequency of 1 Hz and the time delay of 1 s, select the length of the predicted historical sequence as s = 5. Therefore, the width of the graph data to be constructed is Set the scenario to 4 lanes. Therefore, the height of the graph data is 4.
[0097] Set the hyperparameters of the depth prediction network. The input data format of the encoder is batch×5×4×4×16, which contains 2 layers of convlstm modules with the number of channels being 64 and 128 respectively. The decoder consists of a channel attention mechanism and a fully connected layer. The channel attention mechanism adopts the global average pooling method. All activation functions use the Relu function. The initial values of the model weights use the He initialization method. The optimization method of the model selects stochastic gradient descent with a learning rate of 0.001, and adaptively decays based on the change of the loss function on the validation set with a decay factor of 0.5. The batch training size is 3, and the number of training epochs is 500. The early stopping condition is that in 20 consecutive training epochs, the loss function value on the validation set no longer decreases.
[0098] Set the hyperparameters of the double deep Q network. The action value network adopts a two-layer fully connected structure with the number of neurons being 128 and 256 respectively. The activation function uses the Relu function. The discount factor γ = 0.8, the learning rate is 0.001, the update frequency of the target network is 50 steps, and the total learning time step is 10 6 , the batch training size is 32, and the buffer size is 15000.
[0099] Finally, use the highway scenario in the highway_env environment to simulate the vehicle control situation.
[0100] The initial weights of all networks follow a normal distribution with a mean of 0 and a standard deviation of 0.1, and the initial value of the bias is 0.01.
[0101] Step S3: Deep reinforcement learning training
[0102] The controlled vehicle is trained with deep reinforcement learning. The reward value is set according to Equation (12), the speed reward weight k = 0.4 is selected, and the hyperparameter m is adjusted according to the training results. Finally, the hyperparameter m = 9 is determined.
[0103] Step S4: Preprocess the collected historical data, determine the data and labels according to the time delay situation, convert the feature data of the vehicle in front of the controlled vehicle into graph data to provide a training set and a validation set for the deep prediction network, and perform training on the training set until the loss function value on the validation set no longer decreases, then terminate the training.
[0104] Collect historical data, including the position and speed information of the controlled vehicle and the m vehicles in front of it, and convert it into graph data of 4×4×16 as the training data of the deep prediction network. Set the loss function according to Equation (13), where Q is a 16×16 weight matrix and is selected as:
[0105]
[0106] Considering that the signs of the end of training include reaching the preset number of training rounds or triggering the early stopping condition, after the training ends, the deep prediction network model is maintained and fixed to compensate for the impact of time delay. Figure 3 It is the training curve of the lateral movement of the vehicle when a long short-term memory network is equipped for each vehicle separately. Figure 4 It is the comparison of the training curves of the deep prediction network with and without the channel attention mechanism proposed by the present invention. Figure 4 Compared with Figure 3 It can be seen from the comparison that if a long short-term memory network is equipped for each vehicle separately, the loss function value of the validation set in the training curve of the lateral movement will be much larger than that of the training set. This is because a single long short-term memory network cannot capture the interaction between vehicles and cannot represent the impact of vehicle interaction on the vehicle trajectory, and this impact is mainly reflected in the lateral movement (lane change). The prediction method based on graph data proposed by the present invention includes a convolutional layer, which can extract spatial information and depict the impact of the interaction between vehicles on the vehicle trajectory. Therefore, the loss function values of the training set and the validation set are similar. At the same time, Figure 4 It shows that after introducing the channel attention mechanism, the loss function value on the validation set is smaller and the prediction accuracy is improved.
[0107] Step S5: Deploy the trained depth prediction network and the double deep Q network into the controlled vehicle. Assume that the remaining vehicles all follow the intelligent driver model, and conduct simulations in the highway scenario of the highway_env environment. The controlled vehicle predicts the speed and position of the vehicle ahead at the current moment based on 5 historical observations, compensates for the impact caused by a 1-second time delay, aggregates the predicted values with its own real-time speed and position information as the state of the current controlled vehicle, obtains the upper-layer discrete action with the highest action value from the deep Q network, and then maps it into the underlying physical control signal. Figure 5 It shows that the controlled vehicle (black, marked in the figure) can judge the real-time positions of other vehicles (white) in the presence of a 1-second time delay and complete a series of actions such as lane changing and overtaking.
[0108] In summary, for the control method of an autonomous driving vehicle based on a depth prediction network and deep reinforcement learning of the present invention, considering that in practical applications, the observation of the position and speed of other vehicles ahead by the controlled vehicle may have a time delay, a depth prediction network is trained for the controlled vehicle, and the position and speed of the vehicle ahead at the current moment are estimated using the trajectories of the vehicles ahead at historical moments, so as to compensate for the impact caused by the time delay. The present invention converts the historical trajectories of the vehicles ahead of the controlled vehicle into graph data, extracts the interaction information between vehicles from the convlstm module, depicts the impact of the interaction on vehicle prediction, introduces a channel attention mechanism to improve the prediction accuracy of vehicle trajectories, and uses the output of the prediction network as the input of the controller to reduce the adverse impact of the time delay. At the same time, the underlying vehicle control is mapped into the upper-layer discrete action, and the double deep Q network is used to control the controlled vehicle, enabling the vehicle to accurately complete actions such as lane changing and overtaking, and realizing the control of autonomous driving.
[0109] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. An autonomous vehicle control method based on a depth prediction network and deep reinforcement learning, characterized in that, it includes the following steps: S1: Define the control signals of the vehicle's lower layer corresponding to the upper-layer discrete controller. The upper-layer discrete controller includes at least five action instructions: left lane change, keep, right lane change, acceleration, and deceleration. Map the upper-layer discrete actions into the physical control signals of the lower layer, and the physical control signals of the lower layer are obtained by a proportional controller; Define the kinematic model of the vehicle as: β = tan -1 (1 / 2tanδ), where (x, y) are the coordinates of the vehicle in the Frenet coordinate system; v is the forward speed of the vehicle; ψ is the heading angle; β is the slip angle at the center of gravity; a is the acceleration command; δ is the front-wheel steering command; a and δ are obtained by the proportional controller, and the specific expression of the acceleration control amount a is: a = K p (v r - v), where v r is the desired speed; v is the current speed; When the instruction is to accelerate, the given speed reference value v r is greater than the current speed value, K p is the controller gain; the specific expression of the front-wheel steering command δ is: ψ r = ψ L + Δψ r , v lat,r = -K p,lat Δ lat , Among them, ψ L is the lane orientation; v lat,r is the lateral speed command; Δψ r is the required heading change corresponding to the controlled vehicle; K p,lat and K p,ψ are controller gains; obtained from the action mapping of the upper-layer reinforcement learning controller. When the command is to change lanes, the lateral position Δ lat of the lane centerline will change according to the command requirements; The desired speed v r and the lateral position Δ of the vehicle relative to the center line of the lane lat are both obtained by the action mapping of the upper-layer reinforcement learning controller; S2: Set hyperparameters, build a depth prediction network and a double deep Q network based on the encoder-decoder framework, and initialize the weights of each network; S3: Perform deep reinforcement learning training on the controlled vehicle, design a reward function, let the controlled vehicle interact with the environment to obtain a reward value, and iteratively update the weights of the network until the reward value obtained by the controlled vehicle in a round of driving behavior reaches a preset level or the number of training rounds reaches a preset value, and terminate the training; S4: Preprocess the collected historical data, determine the data and labels according to the time delay situation, convert the feature data of the vehicle in front of the controlled vehicle into graph data, provide a training set and a validation set for the depth prediction network, and train on the training set until the loss function value on the validation set no longer decreases, and terminate the training; S5: Deploy the trained depth prediction network and double deep Q network into the controlled vehicle, predict the position and speed of the vehicle in front of the controlled vehicle through the depth prediction network, expand the prediction information together with the speed and position information of the controlled vehicle into a column vector, and use it as the input of the double deep Q network to obtain the action with the highest action value at the current moment. Then map this upper-layer discrete action into the physical control signal of the lower layer through a reference value and a proportional controller to achieve autonomous driving control of the vehicle.
2. The autonomous vehicle control method based on a depth prediction network and deep reinforcement learning according to claim 1, characterized in that: In step S2, in the depth prediction network with the encoder-decoder framework built, the encoder is formed by stacking convolutional long short-term memory network (Convlstm) modules, and the input of the encoder is the historical trajectories of all vehicles to be predicted in front of the controlled vehicle; the decoder consists of a channel attention mechanism and a fully connected layer; In the double deep Q network, both the main network and the target network are composed of two layers of fully connected layers. Its input is the current state, which is composed of the position and speed of the controlled vehicle and the vehicle closest to it in front, and the output is the upper-layer discretized action of the controlled vehicle.
3. The autonomous vehicle control method based on a depth prediction network and deep reinforcement learning according to claim 2, characterized in that: In the step S2, assume that the number of all vehicles to be predicted in front of the controlled vehicle is m, where m is a hyperparameter that needs to be adjusted by deep reinforcement learning. Collect the position and speed information of the vehicles in front, that is, collect as the observation information of the i-th vehicle in front, where and represent the longitudinal coordinate and the lateral coordinate of the i-th vehicle in front at the moment t, and respectively represent the longitudinal speed and the lateral speed of the i-th vehicle in front at the moment t.
4. The autonomous vehicle control method based on a depth prediction network and deep reinforcement learning according to claim 1, characterized in that: In step S3, a double deep Q network is used for deep reinforcement learning training, and the reward function is set as: where k is the speed coefficient; v max and v min are the maximum and minimum values of the speed of the controlled vehicle, respectively; represents the normalized tangential speed, which is used as a reward for the controlled vehicle to maintain a high speed along the lane.
5. The method for controlling an autonomous vehicle based on a depth prediction network and deep reinforcement learning according to claim 4, characterized in that: In the step S4, the road information of the vehicle ahead is converted into a graph data of C×H×W, where H is the height of the graph data, representing the number of lanes; W is the length of the graph data, which is obtained by discretizing the lanes at one-meter intervals. l is the average speed of other vehicles ahead, and L is the maximum possible distance between the vehicle ahead and the controlled vehicle during the observation sequence and prediction process. If the observation period is δ seconds, the length of the historical observation sequence for prediction is s, and the time delay is d seconds, then L = X + (δs + d)*v m , where X is the maximum distance of the field of view of the controlled vehicle, and v m is the maximum speed of other vehicles ahead; C is the characteristic length; if there is no vehicle in this discretized grid, the values are all 0 numerically, and if there is a vehicle, it is represented by the normalized vector, where L y represents the lane width, is the relative position of the i-th vehicle at time t in this grid, and are the maximum values of the lateral speed and longitudinal speed of the vehicle respectively, is the relative speed value of the i-th vehicle at time t.
6. The method for controlling an autonomous vehicle based on a depth prediction network and deep reinforcement learning according to claim 5, characterized in that: the loss function in the step S4 is a weighted mean square error function, and the specific expression is: Among them, is the predicted value of the front vehicle trajectory described by graph data, that is, the output value of the depth prediction network with X i as the input, Q ∈ W × W is the weight matrix, and b is the size of the batch training volume.
7. The method for controlling an autonomous vehicle based on a depth prediction network and deep reinforcement learning according to claim 6, characterized in that: In the depth prediction network of step S5, among the graph data of the controlled vehicle, the feature quantity of the longitudinal motion relative speed predicted value A grid with a value greater than 0.4 is considered to have a vehicle, and according to its relative position and relative speed characteristics, it is converted into general position and speed information; use to represent the predicted value of the relative position of the i-th vehicle at time t obtained from the depth prediction network, then the predicted value of the position of the i-th vehicle at time t is Use to represent the predicted value of the relative speed of the i-th vehicle at time t obtained from the depth prediction network, then the predicted value of the speed of the i-th vehicle at time t is
Citation Information
Patent Citations
Automatic driving overtaking decision-making method based on reinforcement learning under opposite double lanes
CN110969848A