An autonomous driving controller and training method based on variational autoencoder and reinforcement learning
By combining the variational autoencoder and the reinforced learning network, the variational autoencoder extracts potential variable features and constructs additional rewards, the problem of low learning efficiency caused by excessive state space in autonomous driving is solved, the exploration rate and learning efficiency of the agent are improved, and more efficient autonomous driving control is achieved.
Patent Information
- Application Number
- CN202110124110.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-01-29
AI Technical Summary
When the existing technology directly applies reinforcement learning in autonomous driving, it faces the problem of huge state space and low exploration rate of the agent for unfamiliar state space, resulting in low learning efficiency.
Using a combination of variational autoencoder (VAE) and reinforcement learning network (RL-net), the latent variable autoencoder extracts potential variable features as the state quantity of reinforcement learning network through variational autoencoder, and uses its loss function to construct additional rewards to improve the exploration rate and learning efficiency of the agent in the strange state space.
It effectively solves the problem of low learning efficiency caused by excessive state space in reinforcement learning in autonomous driving, improves the ability and learning efficiency of the agent to explore unfamiliar state spaces, and enhances the control effect of autonomous driving.
Smart Images

Figure CN112801273B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving vehicles, and particularly relates to an autonomous driving controller and a training method based on a variational autoencoder and reinforcement learning. Background Art
[0002] As a major research content in the field of intelligent transportation control, intelligent vehicles integrate a variety of modern electronic information technologies. With the increasing demand for the intelligence and safety of modern vehicles in current society, intelligent driving has become a hot issue and technological frontier that various countries in the world are competing to research in the transportation field. As a rapidly developing machine learning algorithm, more and more experts and scholars apply reinforcement learning to the field of intelligent driving.
[0003] Reinforcement learning is a rapidly developing machine learning method that emphasizes selecting an action based on the current environmental state so that the action can achieve the maximum expected reward. It is a trial-and-error learning method. During the learning process, through the stimulation of rewards, it can gradually make actions that maximize the expected reward. Traditional intelligent driving control methods mainly include feedback control, feedforward-feedback control, fuzzy control, sliding mode control strategy, single-point preview strategy, model predictive control, optimal control, etc. However, the above control methods have many limitations. They may have good control effects under specific working conditions, but overall control effects are not good under mixed complex working conditions, and at the same time, they rely on the accuracy of perception-end information. Different from traditional control methods, reinforcement learning can be widely applied to various complex traffic scenarios through continuous interaction and learning with the environment, and can even surpass experienced drivers in the end.
[0004] However, due to the vastness of environmental information and the complexity of traffic scenarios during the autonomous driving process, directly applying the method of reinforcement learning to the field of autonomous driving has a series of problems such as a huge state quantity space and low learning efficiency caused by the low exploration rate of the intelligent agent in the unfamiliar state space. Summary of the Invention
[0005] The present invention proposes an autonomous driving training method based on a variational autoencoder and reinforcement learning, including designing an algorithm model controller part and a training process part. The algorithm model controller includes two parts: a variational autoencoder (VAE) and a reinforcement learning network (RL-net).
[0006] Further, the variational autoencoder (VAE) includes an encoder and a decoder. The input of the encoder is the environmental state quantity s with time series information t , and the output is the latent variable z t ; the input of the decoder is the latent variable feature z t, the output is the predicted feature at the next moment. The input of the reinforcement learning network (RL-net) is the latent variable feature z t and the real-time reward r t , and the output is the specific action a t .
[0007] Furthermore, the real-time reward r t includes the reward r' of the environmental real-time feedback t and the additional reward B(s t ) constructed by the loss function of the variational autoencoder. The specific expression is:
[0008] r t = r' t + B(s t )
[0009] where the expression of the additional reward B(s t ) is:
[0010] B(s t ) = -γlog(p(s t ))
[0011] γ is the scaling factor; p(s t ) is the probability of the state quantity s t ; -log(p(s t )) is the information quantity of the state quantity, which represents the density of the state quantity s t .
[0012] Reinforcement learning is a trial-and-error learning method. The agent needs to continuously interact with the environment and continuously explore the unfamiliar state space to continuously improve the agent. In the area of the state space where the agent explores less, the state quantity s t is relatively sparse, the information quantity of the state quantity -log(p(s t )) is larger, the value of the additional reward B(s t ) will also be larger, and the corresponding real-time reward r t will also increase, which will encourage the agent to explore the relatively unfamiliar state space. When the entire state space is explored, the information quantity -log(p(s t )) of each state quantity in the state space tends to be equal, and the incentive exploration effect brought by the additional reward B(s t ) gradually disappears. Adding the additional reward B(s t ) can improve the ability of the reinforcement learning agent to explore the state space and the learning efficiency.
[0013] In the later stage of training the variational autoencoder (VAE), the final loss function of the variational autoencoder (VAE) can be expressed as:
[0014] loss = -ELBO ≈ -log(p(s t ))
[0015] Therefore, the loss function of the pre-trained variational autoencoder (VAE) can be used to approximate the information content of the state quantity, and further construct the additional reward B(s t ), to improve the exploration rate and learning rate of the reinforcement learning agent.
[0016] The present invention also proposes a training method for autonomous driving using the above controller, and the training method includes the following steps:
[0017] 1) Collect environmental information during the driving process in the way of manual driving;
[0018] 2) Process the collected data set, and use the environmental state quantity at each moment to construct the feature map at the current moment. In particular, each piece of data in the data set includes the environmental state quantity s of continuous n moments t and the feature map of the next moment in the future;
[0019] 3) Use the processed data set to pre-train the variational autoencoder (VAE), so that for every input of the environmental state quantity s of the previous n moments t predict the feature map of the next moment;
[0020] 4) Use the variational autoencoder (VAE) to extract the latent variable feature z t , as the state quantity of the reinforcement learning network (RL-net) to train the reinforcement learning network;
[0021] 5) Continuously collect data during the training of the reinforcement learning network (RL-net) to expand the data set for training the variational autoencoder. At the same time, use the additional reward B(s t ) constructed by the loss function of the variational autoencoder (VAE) to accelerate the training of the reinforcement learning network (RL-net).
[0022] 6) While training the reinforcement learning network (RL-net), fill the experience pool in real time to support the later network training. The experience in the experience pool is <s t , a t , r t , s t+1 >.
[0023] Furthermore, the encoder of the variational autoencoder (VAE) includes a convolutional module and a recurrent neural network module. The convolutional module processes the environmental state at each moment and extracts the feature f m . The encoder combines the continuous n-moment features f1, f2... f nProcessed into a time series feature group and input into the recurrent neural network module. The recurrent neural network module finally further extracts the latent variables from the time series feature group with features at n moments.
[0024] The decoder of the variational autoencoder (VAE) is a transposed convolution module, which reconstructs the features of the next moment from the latent variables extracted by the recurrent neural network.
[0025] Furthermore, use the variational autoencoder (VAE) to extract the latent variable features z t When doing so, the data reading process of the variational autoencoder is as follows:
[0026] 1) Preferably, process k RGB images obtained by the sensor at the current moment into k matrices of (w*h*3), where w is the width of the matrix, h is the height of the matrix, and 3 represents 3 channels;
[0027] 2) Merge the k matrices of (w*h*3) in the third dimension into a matrix of (w*h*(3*k));
[0028] 3) Pass the matrix of (w*h*(3*k)) through the convolution module to extract the features f at the current moment.
[0029] Advantages of the present invention:
[0030] (1) By using the variational autoencoder (VAE) to extract the surrounding environmental state information, the present invention can simultaneously extract the features of multiple sensors, fully extract the information of the surrounding environment, and avoid the loss of surrounding environmental information when training the reinforcement learning network (RL-net); at the same time, use the latent variables extracted by the variational autoencoder (VAE) for dimensionality reduction as the state quantity of the training reinforcement learning network (RL-net), which solves the problem of the overly large state space in the reinforcement learning part and improves the convergence speed of the reinforcement learning network.
[0031] (2) The encoder part of the variational autoencoder (VAE) of the present invention adopts the method of a convolution module plus a recurrent neural network, fully considering the influence of historical information on the future, and avoiding the loss of state quantity information caused by the dynamic characteristics of the ego vehicle and surrounding obstacles being hidden in the historical information.
[0032] (3) By using the loss function of the variational autoencoder (VAE) to construct an additional reward term B(s t ), when training the reinforcement learning network (RL-net), it accelerates the exploration of the agent in the unfamiliar state space, accelerates the exploration of the agent in the unfamiliar state space, and improves the exploration rate and learning efficiency of the agent.
[0033] (4) The present invention trains a variational autoencoder (VAE) by means of pre-training plus re-training. The dataset for pre-training is the dataset collected by manual driving, which improves the network convergence speed. At the same time, during the training of the reinforcement learning network (RL-net), the dataset is augmented and collected, avoiding the overfitting problem caused by the small capacity of the dataset for the variational autoencoder (VAE). Description of the Drawings
[0034] Figure 1 is the framework diagram of the autonomous driving algorithm model based on the variational autoencoder and reinforcement learning;
[0035] Figure 2 is the training flow chart of the autonomous driving based on the variational autoencoder and reinforcement learning;
[0036] Figure 3 is the network structure diagram of the corresponding variational autoencoder;
[0037] Figure 4 is the data reading flow chart of the corresponding variational autoencoder; Detailed Embodiment
[0038] The present invention will be further described below in conjunction with the description of the drawings and the detailed embodiment, but the protection scope of the present invention is not limited thereto.
[0039] Figure 1 is the framework diagram of the autonomous driving algorithm model based on the variational autoencoder and reinforcement learning. The method of the present invention includes two parts: a variational autoencoder (VAE) and a reinforcement learning network (RL-net), which are specifically as follows:
[0040] 1) The variational autoencoder (VAE) includes an encoder and a decoder.
[0041] The input of the encoder is the environmental state quantity s with time series information t , and the output is the latent variable z t ; the input of the decoder is the latent variable feature z t , and the output is the predicted feature of the next moment.
[0042] 2) Preferably, the reinforcement learning network (RL-net) is an actor-critic algorithm.
[0043] The input of the reinforcement learning network (RL-net) is the latent variable feature z t and the real-time reward r t , and the output is the specific action a t , including the steering wheel angle and the brake and throttle opening.
[0044] The real-time reward r t includes the reward r' feedback by the environment in real timet and an additional reward B(s constructed from the loss function of the variational autoencoder t ).
[0045] Figure 2 The flowchart for training autonomous driving based on variational autoencoder and reinforcement learning is as follows:
[0046] 1) By means of manual driving, use on-vehicle cameras, lidar, and GPS navigators to collect environmental information during the driving process;
[0047] 2) Process the collected dataset, and use the on-vehicle camera images, radar point cloud maps, and maps collected by the GPS navigator at each moment to construct a semantic segmentation bird's-eye view of the current moment. In particular, each piece of data in the dataset includes camera images, radar point cloud maps, maps at four consecutive moments, and the semantic segmentation bird's-eye view of the next moment in the future;
[0048] 3) Use the processed dataset to pre-train the variational autoencoder (VAE) so that the camera images, radar point cloud maps, and maps of the previous four moments can be input to predict the semantic segmentation bird's-eye view of the next moment;
[0049] 4) Use the latent variable features z extracted by the variational autoencoder (VAE) t as the state quantity of the reinforcement learning network (RL-net) to train the reinforcement learning network;
[0050] 5) Continuously collect data during the training of the reinforcement learning network (RL-net) to expand the dataset for training the variational autoencoder. At the same time, use the additional reward B(s t ) constructed from the loss function of the variational autoencoder (VAE) to accelerate the training of the reinforcement learning network (RL-net).
[0051] 6) While training the reinforcement learning network (RL-net), fill the experience pool in real time to support the later network training. The experience in the experience pool is <s t ,a t ,r t ,s t+1 >.
[0052] Figure 3 is the network structure diagram of the variational autoencoder, including an encoder and a decoder. Among them, the encoder specifically includes: a convolutional module and a recurrent neural network module, and the decoder is a transposed convolutional module. The convolutional module processes the front view camera images, radar point cloud maps, and maps at the m-th moment, and extracts features f m. Specifically, the convolution module processes the front view camera images, radar point cloud maps, and maps of four consecutive time instants each time, and the extracted features are f1, f2, f3, and f4 respectively. The encoder processes the features of four consecutive time instants into a temporal feature group and inputs it into the recurrent neural network module. The recurrent neural network module finally further extracts the latent variables from the temporal feature group with the features of four time instants. The deconvolution module extracts the latent variables and reconstructs them into the features of the next time instant. The features of the next time instant are the semantic segmentation bird's-eye view.
[0053] Preferably, the convolution module includes three convolutional layers, two pooling layers, and three fully connected layers, and the specific structure is as follows:
[0054] Name Size Output Input layer (256*256*3)*3 256*256*9 Convolutional layer Conv1 (3*3*9)*32, stride = 2 128*128*32 Pooling layer Pool1 (2*2), stride = 2 64*64*32 Convolutional layer Conv2 (3*3*32)*64, stride = 2 32*32*128 Pooling layer Pool2 (2*2), stride = 2 16*16*128 Convolutional layer Conv3 (3*3*128)*128, stride = 2 8*8*128 Fully connected layer FC (8*8*128)*512 1*1*512
[0055] The input layer combines three matrices of 256*256*3 into a matrix of 256*256*9; the convolutional layer Conv1 consists of a convolutional kernel of (3*3*9)*32 with a stride of stride = 2, and its input is the output of the input layer, which is a matrix of 256*256*9, and its output is a feature of 128*128*32; the pooling layer Pool1 consists of a pooling kernel of (2*2) with a stride of stride = 2, and its input is the output of the convolutional layer Conv1, which is a feature of 128*128*32, and its output is a feature of 64*64*32; the convolutional layer Conv2 consists of a convolutional kernel of (3*3*32)*64 with a stride of stride = 2, and its input is the output of the pooling layer Pool1, which is a feature of 64*64*32, and its output is a feature of 32*32*128; the pooling layer Pool2 consists of a pooling kernel of (2*2) with a stride of stride = 2, and its input is the output of the convolutional layer Conv2, which is a feature of 32*32*128, and its output is a feature of 16*16*128; the convolutional layer Conv3 consists of a convolutional kernel of (3*3*128)*128 with a stride of stride = 2, and its input is the output of the pooling layer Pool2, which is a feature of 16*16*128, and its output is a feature of 8*8*128; the size of the fully connected layer FC is (8*8*128)*512, and its input is the output of the convolutional layer Conv3, which is a feature of 8*8*128, and its output is a feature of 1*1*512, f.
[0056] Preferably, the recurrent neural network module is an LSTM long short-term memory network, and the input is a temporal feature group with the features of four time instants, which are 4 consecutive time instants of 1*1*512 features f extracted by the convolution module, and the output is a feature of 1*1*512.
[0057] The latent variable z tIt is further obtained by the recurrent neural network module through the fully connected layer FC-μ and the fully connected layer FC-σ. The fully connected layer FC-μ and the fully connected layer FC-σ are in a parallel structure, and the input is the feature extracted by the recurrent neural network module, which is a feature of 1*1*512. The output of the fully connected layer FC-μ is a feature of 1*1*512, and the output of the fully connected layer FC-σ is a feature of 1*1*512. The features extracted by the fully connected layer FC-μ and the fully connected layer FC-σ together constitute the latent variable z t .
[0058] The deconvolution module described above is the reverse structure of the convolution module. Its input is the feature of 1*1*512 sampled from the latent variable z t , and the output is the feature map of the next moment of 256*256*3.
[0059] Figure 4 It is the data reading flow chart of the variational autoencoder, and the specific process is as follows:
[0060] 1) Preferably, the front view camera image, the radar point cloud map, and the RGB image of the map at the current moment are processed into 3 matrices of (256*256*3);
[0061] 2) The 3 matrices of (256*256*3) are merged in the third dimension to form a matrix of (256*256*9);
[0062] 3) The matrix of (256*256*9) passes through the convolution module to extract the feature f at the current moment.
[0063] The series of detailed descriptions listed above are only specific descriptions of the feasible implementation modes of the present invention, and they are not intended to limit the protection scope of the present invention. Any equivalent modes or changes not departing from the technology created by the present invention should be included in the protection scope of the present invention.
Claims
1. An autonomous driving controller based on variational autoencoder and reinforcement learning, characterized in that It includes two parts: a variational autoencoder and a reinforcement learning network; the variational autoencoder includes an encoder and a decoder; the input of the encoder is the environmental state quantity s with temporal information t , and the output is the latent variable feature z t ; the input of the decoder is the latent variable feature z t , and the output is the predicted feature at the next moment; The input of the reinforcement learning network is the latent variable feature z t and the real-time reward r t , and the output is the specific action a t ; The real-time reward r t includes the reward r' with real-time environmental feedback t and the additional reward B(s t ), and the specific expression is: r t = r' t + B(s t ) Among them, the expression of the additional reward B(s t ) is as follows: B(s t ) = -γ log(p(s t )) γ is a scaling factor, -log(p(s t )) is the information content of the state quantity, p(s t ) represents the density of the state quantity s t ; The encoder includes a convolutional module and a recurrent neural network module. The convolutional module processes the front view camera image, radar point cloud map, and map at the m-th moment, and extracts the feature f m , and the convolutional module processes the front view camera image, radar point cloud map, and map of four consecutive moments each time. The extracted features are f1, f2, f3, and f4 respectively. The features of four consecutive moments are processed into a time series feature group and input into the recurrent neural network module; the recurrent neural network module finally further extracts the latent variable from the time series feature group with the features of four moments; The described convolutional module includes three convolutional layers, two pooling layers, and three fully connected layers, and the specific structure is as follows: The input layer combines three matrices of 256*256*3 into a matrix of 256*256*9; the convolutional layer Conv1 consists of convolutional kernels of (3*3*9)*32 with a stride of stride = 2. Its input is the output of the input layer, which is a matrix of 256*256*9, and its output is a feature of 128*128*32; the pooling layer Pool1 consists of pooling kernels of (2*2) with a stride of stride = 2. Its input is the output of the convolutional layer Conv1, which is a feature of 128*128*32, and its output is a feature of 64*64*32; the convolutional layer Conv2 consists of convolutional kernels of (3*3*32)*64 with a stride of stride = 2. Its input is the output of the pooling layer Pool1, which is a feature of 64*64*32, and its output is a feature of 32*32*128; the pooling layer Pool2 consists of pooling kernels of (2*2) with a stride of stride = 2. Its input is the output of the convolutional layer Conv2, which is a feature of 32*32*128, and its output is a feature of 16*16*128; the convolutional layer Conv3 consists of convolutional kernels of (3*3*128)*128 with a stride of stride = 2. Its input is the output of the pooling layer Pool2, which is a feature of 16*16*128, and its output is a feature of 8*8*128; the size of the fully connected layer FC is (8*8*128)*512. Its input is the output of the convolutional layer Conv3, which is a feature of 8*8*128, and its output is a feature f of 1*1*512; The described recurrent neural network module is an LSTM long short-term memory network. The input is a time-series feature group with features at four moments, which are 4 consecutive moments of features f of 1*1*512 extracted by the convolutional module, and the output is a feature of 1*1*512; The potential variable z t is further extracted by the recurrent neural network module through the fully connected layer FC-μ and the fully connected layer FC-σ: The fully connected layer FC-μ and the fully connected layer FC-σ are in a parallel structure, and the input is the feature extracted by the recurrent neural network module, which is a 1*1*512 feature. The output of the fully connected layer FC-μ is a 1*1*512 feature, and the output of the fully connected layer FC-σ is a 1*1*512 feature. The features extracted by the fully connected layer FC-μ and the fully connected layer FC-σ together constitute the potential variable z t ; The decoder is a transposed convolution module, and its input is the feature of 1*1*512 sampled from the latent variable z t The output is the feature map of the next moment of 256*256*3; the feature map of the next moment is the bird's-eye view of semantic segmentation.
2. The autonomous driving controller based on variational autoencoder and reinforcement learning according to claim 1, wherein The described action a t includes the steering wheel angle and the brake and throttle opening degrees.
3. A method for training an autonomous driving of an autonomous driving controller based on a variational autoencoder and reinforcement learning as claimed in claim 1, characterized in that, It includes the following steps: S1 Use on-vehicle cameras, lidar, and GPS navigators to collect environmental information during driving; S2 Process the collected data set. Use the on-vehicle camera images, radar point cloud maps, and maps collected by the GPS navigator at each moment to construct a semantic segmentation bird's-eye view of the current moment. Each piece of data in the data set includes camera images, radar point cloud maps, maps, and semantic segmentation bird's-eye views of the next moment in the future at four consecutive moments; S3 uses the data set processed by S2 to pre-train a variational autoencoder (VAE) so that the environmental state quantity s of the first n moments is input each time t to predict the feature map of the next moment; the environmental state quantity s t includes camera pictures, radar point cloud maps, and maps, and the feature map of the next moment is a semantic segmentation bird's-eye view; S4 uses a variational autoencoder (VAE) to extract the latent variable feature z t , and uses it as the state quantity of the reinforcement learning network (RL-net) to train the reinforcement learning network; During the process of training the reinforcement learning network (RL-net), S5 continuously collects data to expand the dataset for training the variational autoencoder. Meanwhile, an additional reward B(s t ) constructed using the loss function of the variational autoencoder (VAE) is used to accelerate the training of the reinforcement learning network (RL-net); While training the reinforcement learning network (RL-net), S6 fills the experience pool in real time to support the subsequent network training. The experiences in the experience pool are <s t ,a t ,r t ,s t+1 >.
4. The autonomous driving training method of an autonomous driving controller based on a variational autoencoder and reinforcement learning according to claim 3, characterized in that, The implementation of S3 includes the following: S3.1 Process k RGB images obtained by sensors at the current moment into k matrices of (w*h*3); S3.2 Merge the k matrices of (w*h*3) in the third dimension into a matrix of (w*h*(3*k)); S3.3 Pass the matrix of (w*h*(3*k)) through the convolutional module to extract the feature f at the current moment.