A deep learning-based automatic driving field-of-view vehicle trajectory prediction method

CN115512323BActive Publication Date: 2026-09-11NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211219292.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-08
Publication Date
2026-09-11
Estimated Expiration
2042-10-08

Smart Images

  • Figure CN115512323B_ABST
    Figure CN115512323B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of image processing, and particularly relates to a kind of automatic driving field of vision outside vehicle trajectory prediction method based on deep learning.The present application fully exploits the characteristics of small feature targets in automatic driving trajectory prediction task, and proposes a kind of deep learning neural network training method.The method combines codec architecture, dilated convolution network and self-attention mechanism, and the deep learning neural network training process is completed through the three learnable neural networks of encoder, self-attention unit and decoder.Compared with the traditional method, the present application can more effectively capture the vehicles outside the field of vision that may suddenly rush into the key area, and help the unmanned vehicle to make corresponding decisions quickly.The experimental results show that, compared with the existing method, the deep learning neural network training method proposed by the present application can improve the recall rate of dangerous vehicles outside the field of vision under different thresholds, reduce the omission probability and false positive rate, and greatly reduce the required time.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] [Technical Field] This invention discloses a method for predicting vehicle trajectories outside the field of view in autonomous driving based on deep learning, belonging to the field of image processing technology.

[0002] [Background Technology] The Society of Automotive Engineers (SAE) defines five levels of autonomous driving. Lower levels offer basic driver assistance features, while higher levels are suitable for vehicles that do not require any human interaction. Currently, no practical products can fully realize Level 4 or higher autonomous driving. Machine learning and artificial intelligence are undoubtedly the core technologies for achieving autonomous driving. In real-world driving scenarios, vehicles not visible to the driver are very common due to obstacle obstruction and limited sensor range. "Out-of-view vehicles" refer to vehicles that have not yet entered the driver's field of vision, either currently or historically, but will enter the driver's field of vision in the near future and affect planning decisions. The lack of prediction for out-of-view vehicles threatens the safety of planning decisions and can lead to traffic accidents.

[0003] Convolutional Neural Networks (CNNs) have demonstrated powerful performance in image feature extraction and are widely used in tasks such as classification and segmentation. In recent years, generative models based on CNNs have provided new approaches for low-dimensional object detection and tracking. The U-Net architecture, based on CNNs and featuring a U-shaped connection, has been widely applied in biomedical image segmentation and enhancement tasks. The UNet architecture can achieve relatively accurate segmentation using only a small number of training images. Unlike CNNs that output image class labels, the UNet architecture outputs pixel labels. However, the number of patches is directly proportional to the number of max-pooling layers required, making localization accuracy and the acquisition of contextual information a trade-off. Max-pooling can lead to the loss of spatial relationships between target pixels and their neighbors, while a small number of patches reduces the visible range of local information.

[0004] In autonomous driving, the task of detecting vehicles outside the driver's field of vision aims to accurately predict vehicles that may suddenly appear outside the driver's field of vision within a critical area K in a short period of time. To simplify the problem, this invention defines the critical area K as a rectangle:

[0005] K={(x, y)|w1≤x≤w2, l1≤y≤l2, x, y∈Z}

[0006] Where w1, w2, l1, and l2 are vehicle parameters, which are set to -25, 25, -15, and 35 respectively in this invention. This invention uses a dynamic coordinate system, always with the center point of the vehicle body as the origin. This invention uses the pixel occupancy image G at time t. t This indicates the currently occupied pixel area:

[0007]

[0008] Where P t It is the pixel area occupied by the vehicle body, A t This indicates the area that is currently inaccessible for driving. The earliest occupancy map of the entire driving area is represented as follows:

[0009]

[0010] Where t+Δt is the time step between t and t+T, the goal of this invention is to find the earliest time step in which a pixel is occupied and derive a prediction P(x,y) that is earlier and as accurate as possible than the ground reality G(x,y).

[0011] Due to the complexity of real-time road conditions, traditional algorithms are not entirely effective at capturing them. In recent years, many methods and applications combining deep neural networks have been proposed. Among them, the most attention has been paid to applying generative models combined with random sampling to multimodal autonomous driving tasks, but these tasks rarely consider vehicles outside the field of view.

[0012] [Summary of the Invention] To address the problem of trajectory prediction for vehicles outside the field of view in autonomous driving vehicle-to-everything (V2X) environments, this invention proposes a deep learning neural network training algorithm that combines a Transformer encoder, a dilated convolutional network, and a self-attention unit.

[0013] Assumption It is an input top-view rasterized real-time traffic information radar image. The purpose of this invention is to utilize feature-coded maps. Generate a predicted trajectory image y. Feature pyramids have varying resolutions at different scales. Since the input images in this invention are captured at different heights, the required feature sizes also differ. Therefore, this invention introduces a pyramid structure.

[0014] Because the processed images output at different scales, this invention needs to extract features at different scales, including vehicle position, motion trajectory, and drivable area information. The encoder has four stages, similar to a CNN backbone network. This invention employs a stepped feature sampling method, which saves GPU memory used during training. Each slice is 2×2×3 pixels in size. This invention uses the output feature map of the previous stage as the input of the next stage, with strides of 2, 4, 8, and 16 pixels, thus forming the feature pyramid of this invention. The size of the feature map generated in each stage is:

[0015]

[0016] To address the problem of task-driven feature extraction, this invention proposes a method that combines the advantages of Transformer and dilated convolution. Figure 1The network model shown divides the solution to the out-of-view vehicle trajectory prediction problem into three parts: encoder, decoder, and self-attention unit.

[0017] In this invention, the encoder is divided into two parts: a Transformer with a pyramid and a dilated convolutional network, with the function of maximizing the extraction of contextual information. The Transformer module consists of four stages, each of which comprises two regularization layers, one multi-head attention layer, two fully connected layers, one activation function layer, and one Dropout layer. The input to the sampler is the raw radar image. The output is a feature map, the purpose of which is to fully extract background information within the driving environment.

[0018] Unlike traditional convolutional algorithms, this invention uses hybrid dilated convolution to further extract contextual information, shorten algorithm runtime, and conserve computational resources. The dilated convolution module consists of 6 convolutional layers, 6 regularization layers, and 1 Dropout layer. Each convolutional layer is a fully connected layer with 1024 units, using the ReLU activation function. This invention employs a sawtooth wave structure with a dilated convolution stride combination of (2, 4, 9, 2, 4, 9). Combinations like (2, 4, 8) in Dilated-UNet, which are not coprime, would lead to grid artifacts and negatively impact training performance, and are therefore undesirable. The final output is the feature map F. m The dimensions are (h, w, n), where n is the number of channels (usually 3), and h and w are the height and width of the image. Dilated convolution operator* d It can be represented as:

[0019]

[0020] After dilation, the coverage field length of the convolutional kernel is a = k + (k-1)(d-1), where k is the kernel length, which is 3 in this invention. For example, when the dilation rate is 4, the coverage field length is 3 + 2 × 4 = 11. This invention uses hybrid dilated convolution to solve the receptive field problem. Hybrid dilated convolution aims to ensure that the final receptive field completely covers a square region after a series of convolution operations. Each layer uses arbitrary dilation rates, and using a combination of dilation rates with a sawtooth wave shape allows for natural integration with the original network without adding other modules while maintaining the top layer's receptive field. During training, to penalize predictions P(x,y) that lag behind the ground reality G(x,y), the loss function for decision velocity is defined as:

[0021]

[0022] The Count function increments the count by 1 when the condition is true.

[0023] The model consists of an encoder, a decoder, and self-attention units. The encoder first projects and reshapes the input into a tensor of size [256, 4, 4], and then the decoder recovers it into an RGB image of size [128, 128, 3]. The encoder has 4 Transformer stages, and the dilated convolutional network has a kernel size of 3×3 and a stride of 1. The number of hidden features in the feature pyramid is 64, 32, 16, and 8, respectively. The decoder consists of 3 deconvolutional layers, 3 activation function layers, 3 regularization layers, 3 residual convolutional modules, and 1 output layer. The decoder's kernel size is also 3×3, with a stride of 1.

[0024] To avoid a solution where all vehicles are stationary—which meets the problem-solving conditions but completely fails to reflect the actual ground conditions—a challenging loss function is defined:

[0025] L2=∑ (x,y)∈K Count(P(x, y) = 0)

[0026] This invention uses mean squared error (MSE) to evaluate the model's error because the output occupancy map of this invention is image-level. This invention defines a loss function L for single scene reconstruction. MSE Calculate the L2 norm between the actual ground conditions and the predicted values.

[0027] Meanwhile, to make the model more focused on predicting vehicles outside the field of view, a pixel occupancy map O(x, y) of known vehicles outside the field of view is introduced to define the loss function for vehicles that mistakenly enter the driving area of ​​vehicles outside the field of view:

[0028]

[0029] The self-attention unit is designed to encode spatial information onto a feature map, improving the network's focus on vehicles outside its field of view. This module consists of two regularization layers, two Dropout layers, and one activation layer. In this invention, the ReLU function used for activation is replaced with a GELU function, given a feature map F. m The dimension is (h, w, n), where n is the number of channels (usually 3), and h and w are the height and width of the image. Input feature map F m The data is fed into two CNN branches to generate the value matrix K and the query Q, respectively. The present invention then focuses the attention on F. m The final output F′ is generated using a skip connection. m The output is:

[0030]

[0031] The loss function for the entire model is:

[0032] L = L MSE+δ1L1+L2+δ3L3

[0033] δ1 and δ3 are weight parameters, and in actual training, both parameters are set to 1500.

[0034] [Advantages and Positive Effects of the Invention] Compared with the prior art, the present invention has the following advantages and positive effects:

[0035] 1. This invention proposes a prediction algorithm for the trajectory of vehicles outside the field of view in an autonomous driving vehicle-to-everything (V2X) environment, based on Dilated-UNet. The encoder selection, dilation rate of the dilated convolution, and the pyramid gradient settings are all optimal for the specific task, significantly reducing training time while improving the ability to capture vehicles outside the field of view.

[0036] 2. This invention decomposes the training process into three parts: an encoder, a decoder, and a self-attention unit. Each part, except the encoder, consists of a fully connected network or a convolutional neural network, and each part can be learned independently. The purpose of choosing a Transformer with dilated convolutional networks as the encoder is to enhance the extraction of contextual information.

[0037] 3. This invention tests the proposed method on the large open-source autonomous driving dataset nuScenes, with the corresponding application scenario being road condition prediction in autonomous driving tasks. Experiments show that the proposed algorithm can efficiently capture vehicles outside the field of view that may cause traffic hazards at extremely low sampling rates. Compared with existing methods, it significantly improves the prediction accuracy of possible trajectories of vehicles outside the field of view, while greatly reducing the time required for network training.

[0038] [Attached Image Description] Figure 1 This is a diagram of the training model structure for a prediction algorithm for vehicle trajectories outside the field of view in an autonomous driving vehicle network environment, as proposed in this invention.

[0039] Figure 2 This is a partial schematic diagram illustrating the capture effect of the present invention, the Dilated-UNet model, and the Trajectory++ model on vehicles outside the field of view;

[0040] [Detailed Description of Embodiments] To make the implementation schemes and advantages of the present invention clearer, the present invention will be described in more detail below with reference to the accompanying drawings and examples.

[0041] (1) The Transformer encoder consists of four stages of Transformer modules with a pyramid mechanism. The input to the encoder is the original image. The output is the encoded feature map. This invention divides the image into 2×2×3 (RGB3 channels) blocks to make it suitable for Transformer processing, and uses a stepped feature pyramid for downsampling at each stage. Therefore, the feature map size obtained after each Transformer encoding layer is:

[0042]

[0043] (2) The dilated convolutional network has 6 layers, using a sawtooth wave structure with dilated convolution strides of (2, 4, 9, 2, 4, 9), and a dilated convolution operator * d It can be represented as:

[0044]

[0045] After dilation, the coverage field length of the convolutional kernel is a = k + (k-1)(d-1), where k is the kernel length, which is 3 in this invention. This invention uses hybrid dilated convolution to solve the receptive field problem. Hybrid dilated convolution aims to ensure that the final receptive field completely covers a square region after a series of convolution operations. The Transformer module, dilated convolutional network, and feature pyramid network are merged into an encoder. The purpose of using the pyramid mechanism in this invention is to improve training and judgment speed. Therefore, a penalty is applied to predictions P(x,y) that are later than the ground reality G(x,y). This loss function is defined as:

[0046]

[0047] (3) The model consists of an encoder and a decoder. The input to the model is an image of size 128. The model first encodes spatial information into the input image through the encoder and reshapes it into a tensor of size [256, 4, 4] as the encoded feature map. Then, the decoder restores it to an RGB image of size [128, 128, 3]. The kernel size of the decoder is still 3×3, and the stride is 1. At the same time, to avoid the situation where all vehicles are stationary, which conforms to the solution of the problem but does not conform to the actual road surface problem, a loss function is defined to ensure that the model has no stationary solution. The loss function is:

[0048] L2=∑ (x,y)∈K Count(P(x, y) = 0)

[0049] (4) This invention uses mean squared error (MSE) to evaluate the model's error because the output occupancy map of this invention is image-level. This invention defines a loss function for single scene reconstruction and calculates the L2 norm between the predicted value and the ground reality. Simultaneously, to make the model more focused on predicting vehicles outside the field of view, the earliest occupancy map O(x, y) of vehicles outside the field of view is introduced to define the loss function for vehicles mistakenly entering the driving area of ​​vehicles outside the field of view.

[0050]

[0051] (5) The purpose of the self-attention unit is to encode spatial information onto the feature map, thereby improving the network's focus on vehicles outside its field of view. This invention replaces the ReLU function used for activation with a GELU function, given the feature map F. m The dimension is (h, w, n), where n is the number of channels (usually 3), and h and w are the height and width of the image. Input feature map F m After being fed into two CNN branches to generate the value matrix K and query Q respectively, the generated attention is focused on F. m The final output F′ is generated using a skip connection. m The output is:

[0052]

[0053] The loss function for the entire model is:

[0054] L = L MSE +δ1L1+L2+δ3L3

[0055] δ1 and δ3 are weight parameters, and both are set to 1500 during actual training.

[0056] The hardware configuration for the simulation experiment of this invention is as follows: AMD EPYC 7551P, 63G memory, 8 cores; the graphics card used is NVIDIA Quadro RTX3090 GPU.

[0057] The simulation experiment software configuration of this invention is as follows: Linux operating system, Python simulation language, and PyTorch 1.11 software library.

[0058] In the simulation experiments, the nuScenes dataset was used. This dataset collected data from 1000 scenes in densely trafficked areas such as Boston and Singapore using vehicles equipped with one roof-mounted rotating radar, five long-range radar sensors, and six cameras. Each scene was annotated at a frequency of 2 Hz, with a length of 20 seconds, containing up to 23 semantic object classes and high-resolution maps with 11 annotation layers. This invention follows the official benchmark of the nuScenes prediction challenge to segment the dataset. The training set contains 32,186 prediction scenes, and the validation set contains 8,560 prediction scenes. Its annotation count is more than 7 times higher than that of the KITTI autonomous driving dataset.

[0059] The invention uses the ADAM optimizer to train the compressed sensing network with a learning rate of 0.0001. In contrast, several other models also use ADAM as the optimizer with the same learning rate.

[0060] To ensure judgment speed, this invention defines a delay ratio (DR) standard:

[0061]

[0062] Where s represents the current scene, S is the set of all scenes, the Count function increments the count by 1 when the condition is true, and K s It is a subset of K containing vehicles outside the field of view. Another standard is MSE, which is the mean squared error.

[0063] The model's ability to detect vehicles outside its field of view is evaluated by the recall rate, which is calculated as follows:

[0064]

[0065] in,

[0066]

[0067] It is a subset of S containing vehicles outside the field of view, and α is the selected threshold. It is the set of locations predicted by motion. It is the set of pixels occupied by vehicles outside the field of view, which is a subset of the set of out-of-view vehicle occupancy maps mentioned above. IoU is the intersection-union ratio of the predicted pixel occupancy set and the actual occupancy set.

[0068] Table 1 Missing rate and mean squared error for each model

[0069] Physical 6.53 26.70 PPP 6.74 13.20 Trajectory++ 19.87 15.96 Dilated-UNet 1.36 10.60 This invention 1.18 9.78

[0070] Table 2. Out-of-field vehicle recall rates for each model at different thresholds.

[0071]

[0072] Table 1 shows the missing rate and mean squared error of different models in capturing dangerous vehicles outside the field of view. The comparative models used in this invention are the Physical discriminative model built into the nuScenes dataset, the graph-structured recursive model Trajectory++ which predicts future trajectories using past agent trajectories as input, the Occupied Map Sequence (PPP) which fuses LiDAR and map feature predictions, and the Dilated-UNet model, which was the first to propose a solution to this problem. As can be seen from the data in Table 1, the present invention has the lowest latency and mean squared error, i.e., the lowest error rate.

[0073] Table 2 summarizes the recall rates of different models for vehicles outside the field of view. The intersection-over-union (IoU) thresholds were set to 0.3, 0.5, and 0.7, respectively. It can be seen that the present invention significantly outperforms Dilated-UNet and Trajectory++ in capturing vehicles outside the field of view. Specifically, when the threshold is 0.3, the recall rate of the present invention reaches 74.56%, which is very close to the actual road conditions, while the recall rate of Trajectory++ is only 11.55%, and that of Dilated-UNet is 60.31%. When the thresholds are 0.5 and 0.7, the recall rate of the present invention is more than 10% higher than that of Dilated-UNet and the traditional algorithm. As can be seen from the data in Table 2, the present invention significantly reduces training time while improving the capture of dangerous vehicles outside the field of view, nearly doubling the speed.

[0074] Figure 2 This demonstrates the processing results of the present invention, the Dilated-UNet model, and the Trajectory++ model on the input image. Darker stripes represent the possible trajectories of vehicles outside the field of view; the darker the color, the earlier the trajectories are likely to be occupied, indicating the more probable trajectory of vehicles outside the field of view. Figure 2 The results show that the present invention can achieve a very good recall effect, and the predicted trajectory is in good agreement with the actual ground conditions.

Claims

1. A trajectory prediction algorithm for vehicles outside the field of view in an autonomous vehicle-to-everything (V2X) environment, comprising three pre-trained neural networks: an encoder, a self-attention unit, and a decoder, characterized in that: (1) The encoder consists of a four-stage Transformer module and a dilated convolutional network. The Transformer modules are connected in a stacked manner. They contain fully connected layers with 1024 units. The fully connected layers use the GELU activation function. A multi-head attention mechanism is embedded inside the Transformer module. The multi-head attention mechanism has 8 heads, and each head has a dimension of 128. The output of the final layer of the encoder serves as the initial context of the decoder, providing global information about the source sequence. (2) A feature pyramid mechanism combining multiple resolutions is adopted. The feature pyramid in the encoder is as follows: the input image is divided into 2×2×3 blocks and processed by four Transformer stages. The feature map size is downsampled to 1 / 2, 1 / 4, 1 / 8 and 1 / 16 of the original size, forming a four-layer feature pyramid. The feature values ​​of each layer of the dilated convolutional network are set to 64, 32, 16 and 8, respectively. The feature map is gradually restored to the original size (128×128×3) through the deconvolution layer. Each layer fuses the encoder feature map of the corresponding scale. Finally, the output layer generates a predicted trajectory image with the same resolution as the input image. (3) The dilated convolutional network has 6 layers, with a kernel size of 3×3 and a stride of (2,4,9) for every three layers. The dilated convolutional network adopts a six-layer structure with a dilation rate combination of (2,4,9,2,4,9) to form a double-layer sawtooth wave shape. The dilation rates of the first three layers are 2, 4, and 9 respectively, and the last three layers repeat this combination to ensure that the receptive field completely covers the 128×128 region. (4) The self-attention unit consists of two CNN branches. The input encoded feature map is fed into the two branches and then activated by Softmax. Each CNN branch consists of two linear layers and two Dropout layers. The GELU function is used as the activation function in the middle and skip connections are adopted. (5) The decoder consists of 3 deconvolutional layers, 3 activation function layers, 3 regularization layers, 3 residual convolutional modules and 1 output layer. The kernel size of the decoder is 3×3 and the stride is 1. (6) The recall rate of vehicles outside the field of view is used as the evaluation criterion for the model’s ability to capture dangerous vehicles outside the field of view. The recall rate refers to the evaluation result when the intersection ratio of the output predicted trajectory and the ground reality is greater than 0.3, 0.5 and 0.7.

Citation Information

Patent Citations

  • Variable length coding table selection based on video block type for refinement coefficient coding

    CN101523919A

  • Method and system for solving video question-answering problem by utilizing specific target network based on graph

    CN111652357A