Urban road vehicle trajectory prediction method based on LSTM network and attention mechanism

By introducing LSTM network and attention mechanism in vehicle trajectory prediction, dynamically updating the positional relationship between vehicles and roads, the problem of inaccurate vehicle trajectory prediction in the prior art is solved, and higher accuracy and better generalization are achieved.

CN120032505APending Publication Date: 2025-05-23BEIJING INST OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411881173.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing vehicle trajectory prediction methods are difficult to update the prediction trajectory in a timely manner when facing complex urban road networks and highly interactive traffic states, resulting in inaccurate prediction results.

Method used

Using an LSTM network and attention mechanism method, by encoding vehicle information and road information, dynamically update the position relationship between the target vehicle and the road, use the attention mechanism to extract important information, and then predict the future motion trajectory of the vehicle.

Benefits of technology

It realizes more accurate vehicle trajectory prediction, can dynamically update the position relationship between vehicles and roads, and improves the prediction accuracy and scenario generalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032505A_ABST
    Figure CN120032505A_ABST
Patent Text Reader

Abstract

The invention discloses an urban road vehicle trajectory prediction method based on an LSTM network and an attention mechanism, and belongs to the field of vehicle trajectory prediction. The method comprises the following steps: predicting a track sequence of a vehicle and surrounding vehicles by using LSTM network coding, predicting a driving intention of a target vehicle by using a method based on an attention mechanism or a residual network, and embedding the intention into a coding vector of the target vehicle; using a multi-scale lane map to carry out convolutional coding on lane center and boundary information, and capturing environment information needing to be focused on through a dot product attention mechanism; and after the target vehicle state and the environment state of the fused intention are spliced, a predicted track is obtained through LSTM network decoding. According to the invention, the position relation between the target vehicle and the road can be dynamically updated, and more excellent trajectory prediction performance is shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a vehicle trajectory prediction method in the field of vehicle prediction, and in particular to an urban road vehicle trajectory prediction method based on an LSTM network and an attention mechanism. Background Art

[0002] The urban environment has a complex road network, diverse traffic participants, and rich traffic rules. These elements work together to affect traffic flow and vehicle trajectories. First, the complex road network and diverse road types in the urban environment directly affect the vehicle's driving speed, road selection, and driving behavior. For example, a vehicle can choose a lane with less traffic to speed up on a spacious multi-lane road, while it needs to follow the vehicle in front and drive slowly on a narrow single-lane road; secondly, high-density traffic conditions usually mean that vehicles and vehicles, and vehicles and pedestrians are close, which increases the possibility of mutual influence on each other's behavior. In this environment, emergency braking or lane change by a vehicle may force surrounding vehicles to respond quickly, thereby affecting the overall traffic flow and driving trajectory.

[0003] Current vehicle trajectory prediction methods usually use the time series information of the target vehicle, the interaction between the target vehicle and surrounding traffic participants, and the interaction between the target vehicle and the road as input to directly obtain the future trajectory of the target vehicle. However, when faced with urban roads with complex road networks and highly interactive traffic conditions, such methods cannot update the predicted trajectory in a timely manner, which has obvious limitations. Summary of the invention

[0004] The technical problem to be solved by the present invention is: to overcome the shortcomings of the prior art and provide a method for predicting vehicle trajectories on urban roads based on an LSTM network and an attention mechanism. The method inputs vehicle information and road information into a prediction network, and predicts the vehicle position at the next moment in sequence through the LSTM network and the attention mechanism. Compared with the existing trajectory prediction method, the method obtains a more accurate trajectory prediction result by dynamically updating the positional relationship between the target vehicle and the road.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] A method for predicting urban road vehicle trajectories based on an LSTM network and an attention mechanism, the method comprising the following steps:

[0007] Step S1, encoding the historical movement trajectory of the traffic participant to obtain a predicted vehicle coding vector and a surrounding vehicle coding vector, using the predicted vehicle coding vector to perform an attention mechanism extraction on the surrounding vehicle coding vector, to obtain the surrounding vehicle coding vector after attention extraction;

[0008] Step S2, encoding the road information to obtain a lane centerline encoding vector and a lane boundary line encoding vector, and using the predicted vehicle encoding vector obtained in step S1 to perform an attention mechanism extraction on the lane centerline encoding vector and the lane boundary line encoding vector, to obtain an attention-extracted lane centerline encoding vector and an attention-extracted lane boundary line encoding vector;

[0009] Step S3, obtaining predicted vehicle intention information, and fusing the predicted vehicle coding vector and the predicted vehicle intention information to obtain a predicted vehicle coding vector containing an intention label;

[0010] Step S4, concatenating the predicted vehicle coding vector containing the intention label, the surrounding vehicle coding vector after attention extraction, the lane centerline coding vector after attention extraction, and the lane boundary line coding vector after attention extraction to obtain a context vector output by the encoder;

[0011] Step S5, using a decoder based on the LSTM network to obtain the coordinates of the predicted vehicle at the next moment according to the context vector;

[0012] Step S6, add the predicted vehicle coordinates at the next moment obtained in step S5 to the prediction input, and move the observation window backward by one time step to obtain the historical motion trajectory of the new traffic participant, re-execute steps S1-S5 for the historical motion trajectory of the new traffic participant, and predict the coordinates of the predicted vehicle at the next moment, and repeat this cycle for t f times, and the length is t f The predicted future trajectory of the vehicle.

[0013] In step S1, the method for obtaining the predicted vehicle coding vector and the surrounding vehicle coding vector after attention extraction includes:

[0014] Step S11, assume that the traffic scene contains a predicted vehicle and n surrounding vehicles, with a total of n+1 vehicles, and the length of the historical motion trajectory of each vehicle is t h , contains the location information of the vehicle at each time point {x t ,y t}, integrate the trajectories of all traffic participants into a (n+1)×t h ×2 three-dimensional tensor;

[0015] Step S12: The three-dimensional tensor formed by integrating the trajectories of all traffic participants is input into the LSTM network. The LSTM network will encode each vehicle trajectory independently, and the historical motion trajectory of each vehicle will generate a set of latent vectors h t (t∈{1,2,...,t h}), the hidden vector of the last moment As the coding vector of the vehicle's historical trajectory, the predicted vehicle coding vector and the surrounding vehicle coding vectors are obtained;

[0016] Step S13, taking the predicted vehicle code vector as the query matrix Q, the surrounding vehicle code information as the key matrix K and the value matrix V, and using the scaled dot product attention to obtain the surrounding vehicle code vector after attention extraction, the formula is as follows:

[0017]

[0018] in, is the dimension of the key vector.

[0019] In step S2, the road information is divided into a lane centerline and an insurmountable lane boundary line. The lane centerline is a marking line located in the center of the lane. The points on the road centerline are arranged along the lane direction. A lane node is defined as a line segment formed by any two consecutive points on the road centerline. Its coordinates are the midpoint coordinates of the two endpoints. The direction vector is a vector pointing from the starting endpoint to the ending endpoint. Continuous lane sub-segments are represented in the form of {lane node coordinates, lane node direction}. According to the connection relationship between the lane centerlines, four adjacency matrices of the lane nodes can be derived, which respectively store the connection relationships of the front drive, rear drive, left neighbor, and right neighbor. The front-drive node and the rear-drive node are defined as the nearest forward or backward lane nodes on the same lane centerline, and the left-neighbor node and the right-neighbor node are defined as the nearest lane nodes in the left or right space. The insurmountable lane boundary line is a solid line, which clearly indicates that vehicles must not cross it to change lanes or overtake. Each subsegment on the boundary line is expressed in the form of {lane node coordinates, lane node direction}, and the forward adjacency matrix and the backward adjacency matrix are obtained in the same way as the lane centerline. Since the boundary lines on both sides of the lane are generally independent of each other, the adjacency matrix storing the left and right neighbor relationships is set to empty.

[0020] In step S2, the method of obtaining the lane center line coding vector after attention extraction and the lane boundary line coding vector after attention extraction includes:

[0021] Step S21, construct a lane map. According to the construction method of the lane centerline and the insurmountable lane boundary line, the map information is extracted into a directed graph containing multiple groups of vertices and edges, where the vertices are lane node features and the edges are the lines connecting the adjacent lane nodes in four directions. The corresponding adjacency matrix is ​​obtained and the lane node feature x is defined. i The formula is as follows:

[0022]

[0023] Among them, u i is the coordinate of the i-th lane node, and are the coordinates of the starting and ending endpoints corresponding to the i-th lane node, respectively, i is the i-th row vector of the node feature matrix X.

[0024] In step S22, two map networks (MapNet) are used to encode the lane centerline and the insurmountable lane boundary line respectively. MapNet contains four modules with the same structure. Each module contains a lane map convolution layer, a linear layer and a residual calculation. The lane map convolution layer represents a multi-scale lane map convolution LaneConv (1, 2, 4, 8, 16, 32) with an expansion size of C = 6. The multi-scale lane map convolution enables each lane node to aggregate the information of itself, the left neighbor, the right neighbor node, and the 1st, 2nd, ..., 32nd neighbor node along the front and rear directions of the lane. The formula is as follows:

[0025]

[0026] in, and A 前驱 and A 后驱 K c Power, k c =2 c-1 .

[0027] Step S23, taking the predicted vehicle coding vector as the query matrix Q, taking the lane centerline coding information and lane boundary line coding information as the key matrix K and value matrix V respectively, and using scaled dot product attention to obtain the lane centerline coding vector after attention extraction and the lane boundary line coding vector after attention extraction.

[0028] In step S3, the method for obtaining the predicted vehicle code vector containing the intention label includes:

[0029] Step S31, judging the road scene according to the vehicle location and map information, if the vehicle is at an intersection, using the vehicle behavior interaction network based on the attention mechanism to predict the vehicle intention; if the vehicle is on a straight road, using the intention network based on ResNet and CBAM to predict the vehicle intention; otherwise, the predicted vehicle intention is set to 0, and the predicted vehicle intention includes five categories: going straight, turning left, turning right, changing lanes to the left, and changing lanes to the right;

[0030] Step S32, pass the predicted vehicle coding vector into a fully connected layer, reduce the feature dimension of the predicted vehicle coding vector by one dimension, and then add the predicted vehicle intention to the last dimension of the predicted vehicle coding vector to obtain the predicted vehicle coding vector containing the intention label.

[0031] The structure of the vehicle behavior interaction network based on the attention mechanism in step S31 is:

[0032] The first layer: GRU encoder, which uses three layers of GRU modules to encode the historical movement trajectories of traffic participants. The calculation formula of the GRU encoder is as follows;

[0033] Z t =σ(W r ·[h t-1 ,x t ])

[0034] r t =σ(W z ·[h t-1 ,x t ])

[0035]

[0036] Among them, x t is the input vector, W r , W z , W is the weight matrix, and σ is the activation function.

[0037] The second layer: paired interaction units, which encode the predicted vehicle output by the GRU encoder h i 、Surrounding vehicle code h j and the connection feature c between the two vehicles i,j After concatenation, the interaction vector is obtained through the linear layer and activation layer, where the connection feature c i,j To predict the coordinate difference Δx, Δy and speed difference Δv between the vehicle and surrounding vehicles x ,Δv y ;

[0038] The third layer: multi-head attention layer, which uses a multi-head self-attention mechanism on the vectors output by paired interaction units to obtain the weighted interaction vector s i ;

[0039] The fourth layer: behavior decoding layer, which adjusts the weighted interaction vector s i and predicted vehicle code h i The concatenated vectors are then decoded and outputted through the linear layer and the softmax layer to obtain the predicted vehicle intention.

[0040] In step S31, the method for predicting vehicle intention based on the intention network of ResNet and CBAM includes:

[0041] Step S311, based on the scene information near the predicted vehicle, an RGB image is generated, i.e., an image for intention prediction, and the image information of the predicted vehicle, surrounding vehicles, and drivable area is stored in R, G, and B channels respectively. The specific contents are as follows:

[0042] The first step is to determine the image size based on scene information such as average vehicle speed and lane width:

[0043] Image length = road speed limit or average traffic speed × advance prediction time × 2 / length of each pixel;

[0044] Image width = width of three lanes centered on the current lane / width represented by each pixel.

[0045] The second step is to draw the predicted vehicle motion image in the R channel, use a rectangle to represent the predicted vehicle, store the current and historical positions of the expected vehicle, and set the upper left corner coordinates of the target vehicle at time t. and the lower right corner coordinates They are:

[0046]

[0047] Among them, x t ,y t is the vehicle position coordinate, w is the vehicle width, and l is the vehicle length.

[0048] At the same time, the color depth of the pixel is used to indicate the distance between the observation time and the current time. h In the observation sequence of , the vehicle filling formula at the tth observation time is:

[0049]

[0050] The third step is to generate the surrounding vehicle motion image in the G channel, which is consistent with the generation process of the B channel.

[0051] The fourth step is to draw the drivable area image in the B channel: the pixel value of the drivable area is the default value 0, and the value of the non-drivable area is set to 255. On this basis, the lane lines are drawn and the B channel value of each lane boundary is set to 255.

[0052] In the fifth step, the three-channel information is finally fused to obtain an image for intent prediction.

[0053] In step S312, ResNet-18 is used as the backbone network, and the residual module is used to realize the learning of deep features. On this basis, a lightweight attention module CBAM is introduced. CBAM is integrated before and after the residual block, that is, between the initial convolution layer and the first residual block and between the last residual block and the output processing layer. The output layer is a fully connected layer. After being processed by the softmax function, the predicted vehicle intention is obtained.

[0054] In step S6, during the process of iteratively predicting the future coordinates of the vehicle, the teacher forcing technique is used to speed up the network training. During the training process, the real predicted future trajectory of the vehicle is used as the input of the next time step instead of the predicted value of the prediction network. f The step serial training is converted to a batch size of t f Parallel training is performed to reduce the time complexity to O(1). In the middle and late stages of training, a hybrid teacher forcing technique is used to avoid excessive reliance of the prediction network on the true value by gradually reducing the proportion of the true value in the input data.

[0055] The neural network loss function is set as the weighted sum of the root mean square error and the end point displacement error, and the formula is as follows:

[0056] Loss = RMSE + 0.5 FDE

[0057]

[0058] Where N is the number of samples, Y (i) is the actual trajectory of the ith sample, is the predicted trajectory of the sample, is the end of the real trajectory, To predict the end of the trajectory.

[0059] Compared with the prior art, the urban road vehicle trajectory prediction method based on LSTM network and attention mechanism obtained by the present invention has the following advantages: by using an encoder that takes into account the predicted vehicle intention vector and dynamically updating the positional relationship between the predicted vehicle and the road, the trajectory prediction result can be made more accurate and have better scene generalization. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The accompanying drawings are used to provide a further understanding of the technical solution of the present application or the prior art, and constitute a part of the specification. Among them, the accompanying drawings expressing the embodiments of the present application are used together with the embodiments of the present application to explain the technical solution of the present application, but do not constitute a limitation on the technical solution of the present application.

[0061] Figure 1 A flow chart of a method for predicting vehicle trajectories on urban roads based on an LSTM network and an attention mechanism according to an embodiment of the present disclosure is shown;

[0062] Figure 2 A flow chart of a vehicle trajectory prediction algorithm according to an embodiment of the present disclosure is shown;

[0063] Figure 3 A schematic diagram of a map network structure according to an embodiment of the present disclosure is shown;

[0064] Figure 4A schematic diagram of an expanded lane graph convolution according to an embodiment of the present disclosure is shown;

[0065] Figure 5 A schematic diagram of a vehicle behavior interaction network based on an attention mechanism according to an embodiment of the present disclosure is shown;

[0066] Figure 6 A schematic diagram of a paired interaction unit structure according to an embodiment of the present disclosure is shown;

[0067] Figure 7 A schematic diagram of constructing an RGB image based on sequence information according to an embodiment of the present disclosure is shown;

[0068] Figure 8 A schematic diagram of an intent network based on ResNet and CBAM according to an embodiment of the present disclosure is shown;

[0069] Fig. 9 A schematic diagram showing the visualization of the results of a vehicle trajectory prediction algorithm according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0070] In order to describe the objectives, technical solutions and advantages of the present application more clearly and completely, the specific data conversion process will be further described in detail below in combination with the specific algorithm details.

[0071] The specific implementation of the present invention is described in detail below in conjunction with specific embodiments.

[0072] Example

[0073] like Figure 1-Figure 8 As shown, the urban road vehicle trajectory prediction method based on LSTM network and attention mechanism provided in this embodiment includes the following steps:

[0074] Step S1, encoding the historical movement trajectory of the traffic participant to obtain a predicted vehicle coding vector and a surrounding vehicle coding vector, using the predicted vehicle coding vector to perform an attention mechanism extraction on the surrounding vehicle coding vector, to obtain the surrounding vehicle coding vector after attention extraction;

[0075] Step S2, encoding the road information to obtain a lane centerline encoding vector and a lane boundary line encoding vector, and using the predicted vehicle encoding vector obtained in step S1 to perform an attention mechanism extraction on the lane centerline encoding vector and the lane boundary line encoding vector, to obtain an attention-extracted lane centerline encoding vector and an attention-extracted lane boundary line encoding vector;

[0076] Step S3, obtaining predicted vehicle intention information, and fusing the predicted vehicle coding vector and the predicted vehicle intention information to obtain a predicted vehicle coding vector containing an intention label;

[0077] Step S4, concatenating the predicted vehicle coding vector containing the intention label, the surrounding vehicle coding vector after attention extraction, the lane centerline coding vector after attention extraction, and the lane boundary line coding vector after attention extraction to obtain a context vector output by the encoder;

[0078] Step S5, using a decoder based on the LSTM network to obtain the coordinates of the predicted vehicle at the next moment according to the context vector;

[0079] Step S6, add the predicted vehicle coordinates at the next moment obtained in step S5 to the prediction input, and move the observation window backward by one time step to obtain the historical motion trajectory of the new traffic participant, re-execute steps S1-S5 for the historical motion trajectory of the new traffic participant, and predict the coordinates of the predicted vehicle at the next moment, and repeat this cycle for t f times, and the length is t f The predicted future trajectory of the vehicle.

[0080] like Figure 2 As shown, in step S1, the method for obtaining the predicted vehicle coding vector and the surrounding vehicle coding vector after attention extraction includes:

[0081] Step S11, assume that the traffic scene contains a predicted vehicle and n surrounding vehicles, with a total of n+1 vehicles, and the length of the historical motion trajectory of each vehicle is t h , contains the location information of the vehicle at each time point {x t ,y t}, integrate the trajectories of all traffic participants into a (n+1)×t h ×2 three-dimensional tensor;

[0082] Step S12: The three-dimensional tensor formed by integrating the trajectories of all traffic participants is input into the LSTM network. The LSTM network will encode each vehicle trajectory independently, and the historical motion trajectory of each vehicle will generate a set of latent vectors h t (t∈{1,2,...,t h}), the hidden vector of the last moment As the coding vector of the vehicle's historical trajectory, the predicted vehicle coding vector and the surrounding vehicle coding vectors are obtained;

[0083] Step S13, taking the predicted vehicle code vector as the query matrix Q, the surrounding vehicle code information as the key matrix K and the value matrix V, and using the scaled dot product attention to obtain the surrounding vehicle code vector after attention extraction, the formula is as follows:

[0084]

[0085] in, is the dimension of the key vector.

[0086] In step S2, the road information is divided into a lane centerline and an insurmountable lane boundary line. The lane centerline is a marking line located in the center of the lane. The points on the road centerline are arranged along the lane direction. A lane node is defined as a line segment formed by any two consecutive points on the road centerline. Its coordinates are the midpoint coordinates of the two endpoints. The direction vector is a vector pointing from the starting endpoint to the ending endpoint. Continuous lane sub-segments are represented in the form of {lane node coordinates, lane node direction}. According to the connection relationship between the lane centerlines, four adjacency matrices of the lane nodes can be derived, which respectively store the connection relationships of the front drive, rear drive, left neighbor, and right neighbor. The front-drive node and the rear-drive node are defined as the nearest forward or backward lane nodes on the same lane centerline, and the left-neighbor node and the right-neighbor node are defined as the nearest lane nodes in the left or right space. The insurmountable lane boundary line is a solid line, which clearly indicates that vehicles must not cross it to change lanes or overtake. Each subsegment on the boundary line is expressed in the form of {lane node coordinates, lane node direction}, and the forward adjacency matrix and the backward adjacency matrix are obtained in the same way as the lane centerline. Since the boundary lines on both sides of the lane are generally independent of each other, the adjacency matrix storing the left and right neighbor relationships is set to empty.

[0087] In step S2, the method of obtaining the lane center line coding vector after attention extraction and the lane boundary line coding vector after attention extraction includes:

[0088] Step S21, construct a lane map. According to the construction method of the lane centerline and the insurmountable lane boundary line, the map information is extracted into a directed graph containing multiple groups of vertices and edges, where the vertices are lane node features and the edges are the lines connecting the adjacent lane nodes in four directions. The corresponding adjacency matrix is ​​obtained and the lane node feature x is defined. i The formula is as follows:

[0089]

[0090] Among them, u i is the coordinate of the i-th lane node, and are the coordinates of the starting and ending endpoints corresponding to the i-th lane node, respectively, u is the i-th row vector of the node feature matrix X.

[0091] Step S22, two map networks (MapNet) are used to encode the lane centerline and the insurmountable lane boundary line, wherein MapNet contains four modules with the same structure, such as Figure 3As shown in , each module contains a lane map convolution layer, a linear layer and residual calculation. The lane map convolution layer represents a multi-scale lane map convolution LaneConv(1,2,4,8,16,32) with an expansion size of C=6, as shown in Figure 4 As shown in the figure, the multi-scale lane graph convolution enables each lane node to aggregate the information of itself, its left neighbor, its right neighbor, and the information of the 1st, 2nd, ..., 32nd neighbor nodes along the front and rear directions of the lane. The formula is as follows:

[0092]

[0093] in, and A 前驱 and A 后驱 K c Power, k c =2 c-1 .

[0094] Step S23, taking the predicted vehicle coding vector as the query matrix Q, taking the lane centerline coding information and lane boundary line coding information as the key matrix K and value matrix V respectively, and using scaled dot product attention to obtain the lane centerline coding vector after attention extraction and the lane boundary line coding vector after attention extraction.

[0095] In step S3, the method for obtaining the predicted vehicle code vector containing the intention label includes:

[0096] Step S31, judging the road scene according to the vehicle location and map information, if it is located at an intersection, using the vehicle behavior interaction network based on the attention mechanism to predict the vehicle intention; if it is located on a straight road, using the intention network based on ResNet and CBAM to predict the vehicle intention; otherwise, the predicted vehicle intention is set to 0. In order to reduce the impact of intention prediction errors on trajectory generation, the intention with a normalized confidence greater than 0.7 is regarded as a valid intention. The predicted vehicle intention includes five categories: going straight, turning left, turning right, changing lanes to the left, and changing lanes to the right;

[0097] Step S32, pass the predicted vehicle coding vector into a fully connected layer, reduce the feature dimension of the predicted vehicle coding vector by one dimension, and then add the predicted vehicle intention to the last dimension of the predicted vehicle coding vector to obtain the predicted vehicle coding vector containing the intention label.

[0098] like Figure 5 As shown, the structure of the vehicle behavior interaction network based on the attention mechanism in step S31 is:

[0099] The first layer: GRU encoder, which uses three layers of GRU modules to encode the historical movement trajectories of traffic participants. The calculation formula of the GRU encoder is as follows;

[0100] z t =σ(W r ·[h t-1 ,x t ])

[0101] r t =σ(W z ·[h t-1 ,x t ])

[0102]

[0103] Among them, x t is the input vector, W r , W z , W is the weight matrix, and σ is the activation function.

[0104] Second layer: pairwise interaction units, such as Figure 6 As shown, the predicted vehicle encoding h output by the GRU encoder i 、Surrounding vehicle code h j and the connection feature c between the two vehicles i,j After concatenation, the interaction vector is obtained through the linear layer and activation layer, where the connection feature c i,j To predict the coordinate difference Δx, Δy and speed difference Δv between the vehicle and surrounding vehicles x ,Δv y ;

[0105] The third layer: multi-head attention layer, which uses a multi-head self-attention mechanism on the vectors output by paired interaction units to obtain the weighted interaction vector s i ;

[0106] The fourth layer: behavior decoding layer, which adjusts the weighted interaction vector s i and predicted vehicle code h i The concatenated vectors are then decoded and outputted through the linear layer and the softmax layer to obtain the predicted vehicle intention.

[0107] In step S31, the method for predicting vehicle intention based on the intention network of ResNet and CBAM includes:

[0108] Step S311, as Figure 7 As shown, the image is generated based on the scene information near the predicted vehicle, that is, the image used for intention prediction. The R, G, and B channels are used to store the image information of the predicted vehicle, surrounding vehicles, and drivable area. The specific contents are as follows:

[0109] The first step is to determine the image size based on scene information such as average vehicle speed and lane width:

[0110] Image length = road speed limit or average traffic speed × advance prediction time × 2 / length of each pixel;

[0111] Image width = width of three lanes centered on the current lane / width represented by each pixel.

[0112] The second step is to draw the predicted vehicle motion image in the R channel, use a rectangle to represent the predicted vehicle, store the current and historical positions of the expected vehicle, and set the upper left corner coordinates of the target vehicle at time t. and the lower right corner coordinates They are:

[0113]

[0114] Among them, x t ,y t is the vehicle position coordinate, w is the vehicle width, and l is the vehicle length.

[0115] At the same time, the color depth of the pixel is used to indicate the distance between the observation time and the current time. h In the observation sequence of , the vehicle filling formula at the tth observation time is:

[0116]

[0117] The third step is to generate the surrounding vehicle motion image in the G channel, which is consistent with the generation process of the B channel.

[0118] The fourth step is to draw the drivable area image in the B channel: the pixel value of the drivable area is the default value 0, and the value of the non-drivable area is set to 255. On this basis, the lane lines are drawn and the B channel value of each lane boundary is set to 255.

[0119] In the fifth step, the three-channel information is finally fused to obtain an image for intent prediction.

[0120] Step S312, as Figure 8 As shown in the figure, ResNet-18 is used as the backbone network, and the residual module is used to realize the learning of deep features. On this basis, a lightweight attention module CBAM is introduced. CBAM is integrated before and after the residual block, that is, between the initial convolution layer and the first residual block and between the last residual block and the output processing layer. The output layer is a fully connected layer. After being processed by the softmax function, the predicted vehicle intention is obtained.

[0121] Among them, the CBAM module sequentially integrates channel attention and spatial attention and can be directly embedded in the convolutional neural network. The specific operations are:

[0122] First, the channel attention mechanism is given by the feature map Get the feature map with added channel attention, the formula is as follows:

[0123]

[0124] Among them, σ is the sigmoid activation function, W 0 and W 1 is the weight matrix of the fully connected layer, and are the average pooling features and maximum pooling features of the feature map on each channel, Represents element-wise multiplication.

[0125] Subsequently, the spatial attention mechanism processes the feature map with added channel attention to obtain the final output, as follows:

[0126]

[0127] Among them, σ is the sigmoid activation function, f 7×7 is a convolution kernel of size 7×7, and the square brackets represent vector concatenation. and They are the average pooling features and maximum pooling features of the feature map in space, Represents element-wise multiplication.

[0128] In step S6, during the process of iteratively predicting the future coordinates of the vehicle, the teacher forcing technique is used to speed up the network training. During the training process, the real predicted future trajectory of the vehicle is used as the input of the next time step instead of the predicted value of the prediction network. f The step serial training is converted to a batch size of t f Parallel training is performed to reduce the time complexity to O(1). In the middle and late stages of training, a hybrid teacher forcing technique is used to avoid excessive reliance of the prediction network on the true value by gradually reducing the proportion of the true value in the input data.

[0129] The neural network loss function is set as the weighted sum of the root mean square error and the end point displacement error, and the formula is as follows:

[0130] Loss = RMSE + 0.5 FDE

[0131]

[0132] Where N is the number of samples, Y (i) is the actual trajectory of the ith sample, is the predicted trajectory of the sample, is the end of the real trajectory, To predict the end of the trajectory.

[0133] This section describes the effect of the present invention in combination with the measured data experiment. In order to evaluate the performance of the proposed detection method, this example is experimented on the Argoverse2 dataset.

[0134] Dataset and parameter settings:

[0135] The Argoverse2 dataset includes perception datasets, motion prediction datasets, Lidar datasets, and map transformation datasets, which can be used to study tasks such as perception and prediction of autonomous vehicles. In terms of motion prediction, it provides more than 250,000 scenes, with an observation time of up to 11 seconds for each scene, of which the first 5 seconds are observation values, including the state of the target vehicle and surrounding traffic participants, and the last 6 seconds are prediction values, which only contain the motion state of the target vehicle. The training model parameters are as follows:

[0136] There are 120,203 sets of training data and 15,030 sets of test data;

[0137] History trajectory length t h =50, predicted trajectory length t f =60;

[0138] The number of LSTM network layers in the encoder part is 2, and the hidden layer dimension is 64;

[0139] The number of LSTM network layers in the decoder part is 2, and the hidden layer dimension is 256;

[0140] The batch size is 64;

[0141] The parameter of Adam optimizer is β 1 =0.9,β 2 =0.98,ε=10 -9 , use a cyclically changing learning rate during training:

[0142]

[0143] Among them, n step is the current training step number, waemup steps is the number of steps in the warm-up phase, waemup steps =4000.

[0144] Training settings: First, the teacher forcing technique is used for 5 rounds, and then the hybrid teacher forcing technique is used for 18 rounds. When the hybrid teacher forcing technique is used, the proportion of the true value in the input is 0.9, ..., 0.1, 0.09, 0.08, ..., 0.01, and finally the teacher forcing technique is not used and the training is carried out for 30 rounds.

[0145] Experimental results:

[0146] The prediction results of the proposed model for four prediction scenarios are shown in Fig. 9 As shown in the figure, the orange rectangle represents the target vehicle, the orange line is the historical trajectory of the target vehicle, the yellow line is the actual value of the target vehicle's future trajectory, and the green line is the predicted value of the future trajectory. In typical urban scenes such as intersections and straight roads, this method can obtain a relatively smooth predicted trajectory, and the predicted trajectory is close to the actual trajectory.

[0147] The prediction index results are shown in Table 1. It can be found that the prediction results of this method are significantly better than those of the CN and NN models, and the endpoint displacement error is smaller than that of the LSTM model.

[0148] Table 1 Comparison of results of vehicle trajectory prediction algorithm in the embodiment

[0149]

[0150] The above implementation examples are a preferred embodiment of the present invention, but the embodiments of the present invention are not limited to the above examples. Without violating the core spirit and principles of the present invention, any other changes, modifications, substitutions, combinations, and simplified operations should be regarded as equivalent alternatives and included in the protection scope of the present invention. This means that when implementing the present invention, corresponding adjustments can be made according to specific needs to better adapt to actual application scenarios.

[0151] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for predicting urban road vehicle trajectories based on LSTM network and attention mechanism, characterized in that The steps of the method include: Step S1, encoding the historical movement trajectory of the traffic participant to obtain a predicted vehicle coding vector and a surrounding vehicle coding vector, using the predicted vehicle coding vector to perform an attention mechanism extraction on the surrounding vehicle coding vector, to obtain the surrounding vehicle coding vector after attention extraction; Step S2, encoding the road information to obtain a lane centerline encoding vector and a lane boundary line encoding vector, and using the predicted vehicle encoding vector obtained in step S1 to perform an attention mechanism extraction on the lane centerline encoding vector and the lane boundary line encoding vector, to obtain an attention-extracted lane centerline encoding vector and an attention-extracted lane boundary line encoding vector; Step S3, obtaining predicted vehicle intention information, and fusing the predicted vehicle coding vector and the predicted vehicle intention information to obtain a predicted vehicle coding vector containing an intention label; Step S4, concatenating the predicted vehicle coding vector containing the intention label obtained in step S3, the surrounding vehicle coding vector after attention extraction obtained in step S1, the lane centerline coding vector after attention extraction obtained in step S2, and the lane boundary line coding vector after attention extraction, to obtain a context vector output by the encoder; Step S5, using a decoder based on the LSTM network, obtains the coordinates of the predicted vehicle at the next moment according to the context vector obtained in step S4; Step S6, add the predicted vehicle coordinates at the next moment obtained in step S5 to the prediction input, and move the observation window backward by one time step to obtain the historical motion trajectory of the new traffic participant, re-execute steps S1-S5 for the historical motion trajectory of the new traffic participant, and predict the coordinates of the predicted vehicle at the next moment, and repeat this cycle for t f times, and the length is t f The predicted future trajectory of the vehicle.

2. The method for predicting urban road vehicle trajectories based on LSTM network and attention mechanism according to claim 1, characterized in that: In step S1, the method for obtaining the predicted vehicle coding vector and the surrounding vehicle coding vector after attention extraction includes: Step S11, assume that the traffic scene contains a predicted vehicle and n surrounding vehicles, with a total of n+1 vehicles, and the length of the historical motion trajectory of each vehicle is t h , contains the location information of the vehicle at each time point {x t ,y t }, integrate the trajectories of all traffic participants into a (n+1)×t h ×2 three-dimensional tensor; Step S12: The three-dimensional tensor formed by integrating the trajectories of all traffic participants is input into the LSTM network. The LSTM network will encode each vehicle trajectory independently, and the historical motion trajectory of each vehicle will generate a set of latent vectors h t (t∈{1,2,...,t h }), the hidden vector of the last moment As the coding vector of the vehicle's historical trajectory, the predicted vehicle coding vector and the surrounding vehicle coding vectors are obtained; Step S13, taking the predicted vehicle code vector as the query matrix Q, the surrounding vehicle code information as the key matrix K and the value matrix V, and using the scaled dot product attention to obtain the surrounding vehicle code vector after attention extraction, the formula is as follows: Among them, d k is the second dimension of the key matrix, K T is the transpose of the key matrix K.

3. A method for predicting urban road vehicle trajectories based on LSTM network and attention mechanism according to claim 2, characterized in that: In step S2, the road information is divided into a lane centerline and an insurmountable lane boundary line. The lane centerline is a marking line located in the center of the lane. The points on the road centerline are arranged along the lane direction. A lane node is defined as a line segment formed by any two consecutive points on the road centerline. Its coordinates are the midpoint coordinates of the two endpoints. The direction vector is a vector pointing from the starting endpoint to the ending endpoint. The continuous lane sub-segments are represented in the form of {lane node coordinates, lane node direction}. According to the connection relationship between the lane centerlines, four adjacency matrices of the lane nodes are derived to store the predecessor, successor, left neighbor, and right neighbor connection relationships respectively, wherein the predecessor node and the successor node are defined as the nearest forward or backward lane nodes on the same lane centerline, and the left neighbor node and the right neighbor node are defined as the nearest lane nodes in the left or right space; The insurmountable lane boundary line is a solid line, which clearly indicates that vehicles are not allowed to cross it to change lanes or overtake. Each sub-segment on the boundary line is expressed in the form of {lane node coordinates, lane node direction}. According to the connection relationship between the lane boundary lines, the four adjacency matrices of the lane nodes are derived to store the predecessor, rear drive, left neighbor, and right neighbor connection relationships respectively. The predecessor node and the rear drive node are defined as the closest forward or backward lane nodes on the same lane boundary. Since the boundary lines on both sides of the lane are generally independent of each other, the adjacency matrix storing the left neighbor and right neighbor relationships is set to empty.

4. The method for predicting urban road vehicle trajectories based on LSTM network and attention mechanism according to claim 3, characterized in that: In step S2, the method of obtaining the lane center line coding vector after attention extraction and the lane boundary line coding vector after attention extraction includes: Step S21, construct a lane map, extract the map information into a directed graph containing multiple groups of vertices and edges, where the vertices are lane node features, and the edges are the lines connecting the adjacent lane nodes in four directions, and obtain the corresponding adjacency matrix, defining the lane node feature x i The formula is as follows: Among them, u i is the coordinate of the i-th lane node, and are the coordinates of the starting and ending endpoints corresponding to the i-th lane node, respectively, i is the i-th row vector of the node feature matrix X, MLP shape and MLP loc is a multilayer perceptron network; In step S22, two map networks are used to encode the lane centerline and the insurmountable lane boundary line respectively. MapNet contains four modules with the same structure. Each module contains a lane map convolution layer, a linear layer and a residual calculation. The lane map convolution layer represents a multi-scale lane map convolution LaneConv (1, 2, 4, 8, 16, 32) with an expansion size of C = 6. The multi-scale lane map convolution enables each lane node to aggregate the information of itself, the left neighbor, the right neighbor node and the 1st, 2nd, ..., 32nd neighbor node along the front and rear directions of the lane. The formula is as follows: Among them, A i and W i They are the adjacency matrix and weight matrix in a specific direction, X is the node feature matrix, and A 前驱 and A 后驱 K c Power, k c =2 c-1 ; Step S23, taking the predicted vehicle coding vector as the query matrix Q, taking the lane centerline coding information and lane boundary line coding information as the key matrix K and value matrix V respectively, and using scaled dot product attention to obtain the lane centerline coding vector after attention extraction and the lane boundary line coding vector after attention extraction.

5. The method for predicting urban road vehicle trajectories based on LSTM network and attention mechanism according to claim 4, characterized in that: In step S3, the method for obtaining the predicted vehicle code vector containing the intention label includes: Step S31, judging the road scene according to the vehicle location and map information, if the vehicle is at an intersection, using the vehicle behavior interaction network based on the attention mechanism to predict the vehicle intention; if the vehicle is on a straight road, using the intention network based on ResNet and CBAM to predict the vehicle intention; otherwise, the predicted vehicle intention is set to 0, and the predicted vehicle intention includes five categories: going straight, turning left, turning right, changing lanes to the left, and changing lanes to the right; Step S32, passing the predicted vehicle coding vector into a fully connected layer to reduce the feature dimension of the predicted vehicle coding vector by one dimension, and then adding the predicted vehicle intention to the last dimension of the predicted vehicle coding vector to obtain the predicted vehicle coding vector containing the intention label.

6. A method for predicting urban road vehicle trajectories based on LSTM network and attention mechanism according to claim 5, characterized in that: In step S31, the structure of the vehicle behavior interaction network based on the attention mechanism is: The first layer: GRU encoder, which uses three layers of GRU modules to encode the historical movement trajectories of traffic participants. The calculation formula of the GRU encoder is as follows; z t =σ(W r ·[h t-1 ,x t ]) r t =σ(W z ·[h t-1 ,x t ]) Among them, x t is the input vector, h t-1 is the hidden unit of the previous moment, h t is the hidden unit at the current moment, W r , W z , W is the weight matrix, σ is the activation function; The second layer: paired interaction units, which encode the predicted vehicle output by the GRU encoder h i 、Surrounding vehicle code h j and the connection feature c between the two vehicles i,j After concatenation, the interaction vector is obtained through the linear layer and activation layer, where the connection feature c i,j To predict the coordinate difference Δx, Δy and speed difference Δv between the vehicle and surrounding vehicles x ,Δv y ; The third layer: multi-head attention layer, which uses a multi-head self-attention mechanism on the vectors output by paired interaction units to obtain the weighted interaction vector s i ; The fourth layer: behavior decoding layer, which adjusts the weighted interaction vector s i and predicted vehicle code h i The concatenated vectors are then decoded and outputted through the linear layer and the softmax layer to obtain the predicted vehicle intention.

7. The method for predicting urban road vehicle trajectories based on LSTM network and attention mechanism according to claim 5, characterized in that: In step S31, the method for predicting vehicle intention based on the intention network of ResNet and CBAM includes: Step S311, based on the scene information near the predicted vehicle, an RGB image is generated, i.e., an image for intention prediction, and the image information of the predicted vehicle, surrounding vehicles, and drivable area is stored in R, G, and B channels respectively. The specific contents are as follows: The first step is to determine the image size based on scene information such as average vehicle speed and lane width: Image length = road speed limit or average traffic speed × advance prediction time × 2 / length of each pixel; Image width = width of three lanes centered on the current lane / width represented by each pixel; The second step is to draw the predicted vehicle motion image in the R channel, use a rectangle to represent the predicted vehicle, store the current and historical positions of the predicted vehicle, and set the upper left corner coordinates of the predicted vehicle at time t. and the lower right corner coordinates They are: Among them, x t ,y t is the vehicle position coordinate, w is the vehicle width, and l is the vehicle length; At the same time, the color depth of the pixel is used to indicate the distance between the observation time and the current time. h In the observation sequence of , the vehicle filling formula at the tth observation time is: The third step is to draw the motion image of the surrounding vehicles in the G channel, use multiple rectangles to represent the surrounding vehicles, store the current and historical positions of the surrounding vehicles, and set the upper left corner coordinates of the surrounding vehicle i at time t. and the lower right corner coordinates They are: in, is the position coordinate of the surrounding vehicle i, w i ′ is the width of the surrounding vehicle i, l i ′ is the length of the surrounding vehicle i; At the same time, the color depth of the pixel is used to indicate the distance between the observation time and the current time. h In the observation sequence of , the vehicle filling formula at the tth observation time is: The fourth step is to draw the drivable area image in the B channel: the pixel value of the drivable area is the default value 0, and the value of the non-drivable area is set to 255. On this basis, the lane lines are drawn and the B channel value of each lane boundary is set to 255; Step 5: Finally, fuse the three-channel information to obtain an image for intent prediction; In step S312, ResNet-18 is used as the backbone network, and the residual module is used to realize the learning of deep features. On this basis, a lightweight attention module CBAM is introduced. CBAM is integrated before and after the residual block, that is, between the initial convolution layer and the first residual block and between the last residual block and the output processing layer. The output layer is a fully connected layer. After being processed by the softmax function, the predicted vehicle intention is obtained.

8. The method for predicting urban road vehicle trajectories based on LSTM network and attention mechanism according to claim 1, characterized in that: In step S6, the teacher forcing technique is used to speed up the network training. During the training process, the real predicted future trajectory of the vehicle is used as the input of the next time step, and the t f The step serial training is converted to a batch size of t f Parallel training is performed to reduce the time complexity to O(1). In the middle and late stages of training, the hybrid teacher forcing technique is used to gradually reduce the proportion of true values ​​in the input data to avoid excessive reliance of the prediction network on the true values. f After predicting the future trajectory of the vehicle, it is compared with the actual trajectory of the vehicle, and then the loss function of the vehicle trajectory prediction network is calculated.

9. The method for predicting urban road vehicle trajectories based on LSTM network and attention mechanism according to claim 8, characterized in that: The loss function of the vehicle trajectory prediction network is set as the weighted sum of the root mean square error and the end point displacement error, and the formula is as follows: Loss = RMSE + 0.5 FDE Among them, N is the number of predicted vehicle samples, Y (i) is the actual trajectory of the i-th predicted vehicle sample, is the predicted trajectory of the predicted vehicle sample, is the end of the real trajectory, To predict the end of the trajectory.

Citation Information

Cited By

  • Vehicle multi-modal trajectory prediction method based on improved attention network

    CN120672802A

  • Vehicle lane changing track prediction method, device and system

    CN120828828A