Neural network vehicle trajectory prediction method based on self-attention mechanism

Through a neural network based on the self-attention mechanism, the problem that existing vehicle trajectory prediction methods are difficult to achieve long-term stable prediction is solved, and high accuracy and stability of vehicle trajectory prediction is achieved, and applicability is improved.

CN120039275APending Publication Date: 2025-05-27LIAONING POLYTECHNIC VOCATIONAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510183856.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing vehicle trajectory prediction methods are difficult to achieve long-term stable prediction, with large prediction errors and inability to synchronize with reality, and insufficient applicability.

Method used

Using a neural network based on the self-attention mechanism, by constructing a neural network model including input layer, encoder and decoder, the vehicle's historical trajectory information and relative distance and relative speed are input to the network to achieve accurate prediction of vehicle trajectory.

Benefits of technology

The accuracy of vehicle trajectory prediction is improved to reach more than 95%, and the prediction stability is maintained, with an error margin of less than 0.3. At the same time, combined with the vehicle's lane change behavior, the applicability is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120039275A_ABST
    Figure CN120039275A_ABST
Patent Text Reader

Abstract

The invention discloses a neural network vehicle trajectory prediction method based on a self-attention mechanism. The neural network vehicle trajectory prediction method comprises the steps of 1, collecting vehicle historical trajectory information, a relative distance between a vehicle and a target vehicle in a target lane and a relative speed between the vehicle and the target vehicle in the target lane; step 2, constructing a neural network based on a self-attention mechanism, and inputting the parameters obtained in the step 1 into the neural network to obtain a predicted trajectory of the vehicle; the encoder comprises a second input layer, a first embedded layer, a first multi-head probability sparse self-attention module, a first feed-forward layer, a first self-attention distillation module, a second multi-head probability sparse self-attention module, a second embedded layer and a first output layer, and the output of the first output layer is the lane changing state of the vehicle. The method has the characteristic of improving the prediction accuracy and applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent driving, and more specifically, the present invention relates to a neural network vehicle trajectory prediction method based on a self-attention mechanism. Background Art

[0002] In recent years, with the continuous growth of the economy, the living standards of the people have been greatly improved, and the ownership of automobiles has continued to climb. However, traditional fuel vehicles will cause a large amount of consumption of non-renewable resources such as petroleum fuel and other problems such as environmental pollution. The development of new energy vehicles powered by clean energy has gradually been taken seriously. New energy vehicles have the advantages of low noise and higher riding comfort.

[0003] In new energy vehicles, intelligent driving has become a research hotspot. However, due to the different driving styles of the driver group, the complex existing traffic environment, and various complex factors such as the interaction relationships between pedestrians and vehicles, in real driving, it is impossible to achieve fully autonomous driving of vehicles. In order to improve the degree of autonomous driving, accurately predicting the trajectory of vehicles is an inevitable problem. Existing vehicle trajectory prediction methods are often only short-term predictions, and the prediction time is generally less than 1 s, which cannot maintain long-term stable predictions, and the predicted trajectories often cannot be synchronized with reality, with large errors. On this basis, there are also ways to improve the prediction accuracy by combining historical states, interaction information, etc. using neural networks such as RNN or LSTM, but the effects are not ideal. Summary of the Invention

[0004] The purpose of the present invention is to design and develop a neural network vehicle trajectory prediction method based on a self-attention mechanism, which reduces the prediction error, maintains the prediction stability, and improves the applicability through the neural network of the self-attention mechanism.

[0005] The technical solution provided by the present invention is as follows:

[0006] A neural network vehicle trajectory prediction method based on a self-attention mechanism, comprising the following steps:

[0007] Step 1, collect the historical trajectory information of the vehicle, the relative distance between the vehicle and the target vehicle in the target lane, and the relative speed between the vehicle and the target vehicle in the target lane;

[0008] Step 2, construct a neural network based on a self-attention mechanism, and input the historical trajectory information of the vehicle, the relative distance between the vehicle and the target vehicle in the target lane, and the relative speed between the vehicle and the target vehicle in the target lane into the neural network to obtain the predicted trajectory of the vehicle;

[0009] Among them, the neural network based on the self-attention mechanism includes a first input layer, an encoder, and a decoder connected in sequence;

[0010] The first input layer satisfies:

[0011]

[0012] where is the output of the input layer at time t, α is the scalar projection coefficient, and l i is the vehicle trajectory information at time t after being dimensionally elevated to d model ; PE(·) is the position encoding, ψ is the number of features extracted, SE(·) is the timestamp, and L x is the length of the input to the encoder and decoder;

[0013] The encoder includes a second input layer, a first embedding layer, a first multi-head probabilistic sparse self-attention module, a first feed-forward layer, a first self-attention distillation module, a second multi-head probabilistic sparse self-attention module, a second embedding layer, and a first output layer. The output of the first output layer is the lane-changing state of the vehicle itself;

[0014] The decoder includes a third input layer, a third embedding layer, a third multi-head probabilistic sparse self-attention module, a full attention layer, a second feed-forward layer, a second self-attention distillation module, and a second output layer. The output of the second output layer is the predicted trajectory of the vehicle.

[0015] Preferably, the position encoding satisfies:

[0016]

[0017] where pos is the position in the sequence at time t, and its value range is [1, 2, …, d model , and d model is the total length of the time feature vector;

[0018] The timestamp satisfies:

[0019]

[0020] Preferably, the second input layer satisfies:

[0021] x χ = concat(P 0 (t), S M (t));

[0022] where x χ is the input of the encoder, P 0 (t) is the historical trajectory of the vehicle itself, and S M (t) is the interaction information of the vehicle itself;

[0023] The historical trajectory of the vehicle itself satisfies:

[0024]

[0025] where \(t\in(-mT,0)\);

[0026] The interaction information of the host vehicle satisfies:

[0027]

[0028] where \(S\) M (t) is the interaction information of the host vehicle at time \(t\), is the lateral relative distance between the host vehicle and vehicle No. M, is the longitudinal relative distance between the host vehicle and vehicle No. M, is the lateral relative speed between the host vehicle and vehicle No. M, is the longitudinal relative speed between the host vehicle and vehicle No. M, and M is the number of the target vehicle in the target lane.

[0029] Preferably, the embedding layer satisfies:

[0030]

[0031] where \(K\) δ is the output of the embedding layer, and \(w\) 1 is the weight of the first fully connected layer.

[0032] Preferably, the first multi-head probabilistic sparse self-attention module satisfies:

[0033] \(K\) ε = norm[prob(K δ , K δ , K δ ) + K δ ;

[0034] where prob is the probabilistic sparse self-attention layer and norm is the normalization layer;

[0035] The probabilistic sparse self-attention layer satisfies:

[0036]

[0037] where \(A\) is the probabilistic sparse self-attention, \(Q\) is the set of queries, \(K\) is the set of keys, \(V\) is the set of values, and \(d\) is the dimension of the query and the key.

[0038] Preferably, the first feed-forward layer satisfies:

[0039] \(K\) μ = norm[relu(K ε \(w\) 2 ) \(w\) 2 + K ε ;​​​​​​​​​​

[0040] Where relu is the activation function, and w 2 is the weight of the second fully connected layer.

[0041] Preferably, the first self-attention distillation module satisfies:

[0042] K η = MaxPool[ELU(Conv(K μ ))];

[0043] Where MaxPool is the max pooling layer, ELU is the exponential linear unit, and Conv is the one-dimensional convolutional layer.

[0044] Preferably, the first output layer satisfies:

[0045]

[0046] Where is the probability that the vehicle maintains the original lane, is the probability that the vehicle changes lanes to the left, is the probability that the vehicle changes lanes to the right.

[0047] Preferably, the third input layer satisfies:

[0048]

[0049] Where is the masked trajectory of the vehicle; Ltoken is the length of the labeled historical trajectory, and nT is the target trajectory segment to be predicted.

[0050] Preferably, the third embedding layer satisfies:

[0051]

[0052] Where w 5 is the weight of the fifth fully connected layer, and w 6 is the sixth fully connected layer;

[0053] The second output layer satisfies:

[0054] Θ = K τ w 7 ;

[0055] Where K τ is the output of the second self-attention distillation module, and w 7 is the output of the seventh fully connected layer.

[0056] The beneficial effects of the present invention:

[0057] A neural network vehicle trajectory prediction method based on self-attention mechanism designed and developed by the present invention accurately predicts the vehicle trajectory through a neural network based on self-attention mechanism, with an accuracy rate reaching over 95%, and can ensure the stability of trajectory prediction, with an error range less than 0.3 under long-term prediction. At the same time, it can also combine the lane-changing behavior of the vehicle to further improve the applicability. Description of the Drawings

[0058] Figure 1 It is a schematic diagram of the curve comparison of the average accuracy rates of the prediction method of the present invention with the HMM model and the LSTM model. Detailed Embodiment

[0059] The following further elaborates on the present invention in detail to enable those skilled in the art to implement it with reference to the text of the specification.

[0060] A neural network vehicle trajectory prediction method based on self-attention mechanism provided by the present invention includes:

[0061] Step 1: Collect the historical trajectory information of the vehicle, the relative distance between the vehicle itself and the target vehicle in the target lane, and the relative speed between the vehicle itself and the target vehicle in the target lane;

[0062] Among them, the historical trajectory information of the vehicle is collected through GPS, and the relative distance between the vehicle itself and the target vehicle in the target lane and the relative speed between the vehicle itself and the target vehicle in the target lane can both be obtained through sensors;

[0063] Step 2: Construct a neural network based on self-attention mechanism, and input the historical trajectory information of the vehicle, the relative distance between the vehicle itself and the target vehicle in the target lane, and the relative speed between the vehicle itself and the target vehicle in the target lane into the neural network based on self-attention mechanism to obtain the predicted trajectory of the vehicle;

[0064] Among them, the neural network based on self-attention mechanism includes a first input layer, an encoder, and a decoder;

[0065] The first input layer includes multiple convolutional layers, and the convolutional kernel sizes of the multiple convolutional layers are all 3, and the strides are all 1. Therefore, the first input layer satisfies:

[0066]

[0067] In the formula, is the output of the input layer at time t, α is a scalar projection coefficient, and its value is [0, 1], l i is the vehicle trajectory information at time t after being dimensionally elevated to d model ​PE(·) is the position encoding, ψ is the number of extracted features. Since the extracted features are weekly, monthly, and holiday features, thus, ψ = 1, 2, 3. SE(·) is the timestamp, and L x is the length of the input to the encoder and decoder. X = (x 1 , x 2 , …, x i , … x ψ ), where i = 1, 2, …, L x , …;

[0068] The position encoding is the sine-cosine position encoding, and specifically satisfies:

[0069]

[0070] In the formula, pos is the position in the sequence at time t, and its value range is [1, 2, …, d model , and d model is the total length of the time feature vector;

[0071] The timestamp satisfies:

[0072]

[0073] The encoder includes a second input layer, a first embedding layer, a first multi-head probabilistic sparse self-attention module, a first feed-forward layer, a first self-attention distillation module, a second multi-head probabilistic sparse self-attention module, a second embedding layer, and a first output layer, and the output is the lane-changing state of the vehicle itself;

[0074] The second input layer satisfies:

[0075] x χ = concat(P 0 (t), S M (t));

[0076] In the formula, x χ is the input to the encoder, Y 0 (t) is the historical trajectory of the vehicle itself, and S M (t) is the interaction information of the vehicle itself;

[0077] Among them, the trajectory coordinates of the vehicle itself at time t are Then, under the observation domain with a domain length of mt, the historical trajectory of the vehicle itself satisfies:

[0078]

[0079] In the formula, t ∈ (-mT, 0);

[0080] The interaction information of the vehicle itself satisfies:

[0081]

[0082] In the formula, S M (t) is the vehicle interaction information at time t, is the lateral relative distance between the vehicle and vehicle No. M, is the longitudinal relative distance between the vehicle and vehicle No. M, is the lateral relative speed between the vehicle and vehicle No. M, is the longitudinal relative speed between the vehicle and vehicle No. M, and M is the target vehicle number in the target lane;

[0083] The first embedding layer is the first fully connected layer. Therefore, the first embedding layer satisfies:

[0084]

[0085] In the formula, K δ is the output of the embedding layer, and w 1 is the weight of the first fully connected layer;

[0086] The first multi-head probability sparse self-attention module includes a residual connection and normalization. Therefore, the first multi-head probability sparse self-attention module satisfies:

[0087] K ε = norm[prob(K δ , K δ , K δ ) + K δ ;

[0088] In the formula, prob is the probability sparse self-attention layer, and norm is the normalization layer;

[0089] Among them, the probability sparse self-attention layer satisfies:

[0090]

[0091] In the formula, A is the probability sparse self-attention, Q is the set of queries, K is the set of keys, V is the set of values, and d is the dimension of the query and the key;

[0092] The first feed-forward layer includes an activation function layer, a residual connection, and a second fully connected layer. The feed-forward layer satisfies:

[0093] K μ = norm[relu(K ε w 2 )w 2 + K ε ;

[0094] In the formula, relu is the activation function, and w 2 is the weight of the second fully connected layer;

[0095] The first self-attention distillation module includes a one-dimensional convolutional layer, an activation function layer, and a pooling layer, and the first self-attention distillation module satisfies:

[0096] K η = MaxPool[ELU(Conv(K μ ))];

[0097] In the formula, MaxPool is the max pooling layer, ELU is the exponential linear unit, and Conv is the one-dimensional convolutional layer;

[0098] Among them, the stride of the max pooling layer is 2;

[0099] The second multi-head probabilistic sparse self-attention module includes a residual connection and normalization. Therefore, the second multi-head probabilistic sparse self-attention module satisfies:

[0100] K ε ' = norm[prob(K η , K η , K η ) + K η ;

[0101] The second embedding layer includes an activation function layer, a residual connection, and a third fully connected layer. Therefore, the second embedding layer satisfies:

[0102] K' μ = norm[relu(K ε 'w 3 )w 3 + K ε '];

[0103] In the formula, w 3 is the weight of the third fully connected layer;

[0104] The first output layer is the fourth fully connected layer. Therefore, the output layer satisfies:

[0105]

[0106] In the formula, is the probability that the vehicle maintains its original lane, is the probability that the vehicle makes a left lane change, is the probability that the vehicle makes a right lane change;

[0107] The loss function of the encoder is:

[0108]

[0109] In the formula, m is the number of samples, is the lane change type, is the independent encoding of the i-th sample label, is the probability that the i-th sample belongs to the th intention.

[0110] The decoder includes a third input layer, a third embedding layer, a third multi-head probabilistic sparse self-attention module, a full attention layer, a second feed-forward layer, a second self-attention distillation module, and a second output layer. The output of the encoder is input into the decoder to obtain the predicted trajectory of the vehicle;

[0111] The third input layer satisfies:

[0112]

[0113] In the formula, is the masked trajectory of the vehicle itself; Ltoken is the length of the label history trajectory, and nT is the target trajectory segment to be predicted;

[0114] In this embodiment, Ltoken is 3 seconds and nT is 5 seconds.

[0115] The third embedding layer includes a fifth fully connected layer and a sixth fully connected layer, that is, the third embedding layer satisfies:

[0116]

[0117] In the formula, w 5 is the weight of the fifth fully connected layer, and w 6 is the sixth fully connected layer;

[0118] The second multi-head probabilistic sparse self-attention module includes a residual connection and normalization. Therefore, the second multi-head probabilistic sparse self-attention module satisfies:

[0119] K λ = norm[prob(K γ , K γ , K γ ) + K γ ;

[0120] The full attention layer is a multi-head attention layer, that is, the full attention layer satisfies:

[0121] K θ = norm[mah(K λ , Π, Π) + K λ ;

[0122] In the formula, mah is the multi-head attention layer;

[0123] The second feed-forward layer is an activation function layer and a residual connection layer, that is, the second feed-forward layer satisfies:

[0124] K σ= norm[relu(K λ w 6 )w 6 +K θ ;

[0125] The second self-attention distillation module includes a one-dimensional convolutional layer, an activation function layer, and a pooling layer, and the second self-attention distillation module satisfies:

[0126] K τ = MaxPool[ELU(Conv(K σ ))];

[0127] Wherein, MaxPool is the maximum pooling layer, ELU is the exponential linear unit, and Conv is the one-dimensional convolutional layer;

[0128] The second output layer is the seventh fully connected layer, that is, the predicted trajectory of the vehicle is obtained. Therefore, the second output layer satisfies:

[0129] Θ = K τ w 7 .

[0130] The loss function of the decoder is:

[0131]

[0132] Wherein, L′ is the loss function of the decoder, Θ′ is the true vehicle trajectory,

[0133] In another embodiment, the model is trained through the public dataset highD, which includes information such as vehicle trajectory data, number of lanes, vehicle ID, vehicle speed, vehicle coordinates, and vehicle acceleration. After the model is trained, the sampling window is set to 5 seconds, and the input time series is sampled by sliding. The prediction method described in the present invention is compared with the HMM model and the LSTM model according to the average accuracy. The comparison results are as Figure 1 shown. It can be seen that when the prediction length is within 2s, the accuracy rates of the HMM model, the LSTM model, and the prediction method described in the present invention all show an upward trend, and the highest can reach more than 92%. However, in the trajectory prediction after 2s, the method described in the present invention can maintain the prediction accuracy rate above 95%, and reaches the highest point of the prediction accuracy rate when the prediction time is 5s. At this time, the prediction accuracy rates of the HMM model and the LSTM model are less than 90%. It can be seen that the prediction method described in the present invention can not only maintain long-term stability, but also maintain prediction accuracy.

[0134] A neural network vehicle trajectory prediction method based on self-attention mechanism designed and developed by the present invention accurately predicts the trajectory of a vehicle through a neural network based on self-attention mechanism, with an accuracy rate reaching over 95%, and can ensure the stability of trajectory prediction. At the same time, it can also combine the lane-changing behavior of the vehicle to further improve the applicability.

[0135] Although the embodiments of the present invention have been disclosed as above, it is not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the embodiments shown and described herein.

Claims

1. A neural network vehicle trajectory prediction method based on self-attention mechanism, characterized in that: The steps include: Step 1: Collect vehicle historical trajectory information, the relative distance between the vehicle and the target vehicle in the target lane, and the relative speed between the vehicle and the target vehicle in the target lane; Step 2: construct a neural network based on a self-attention mechanism, and input the vehicle historical trajectory information, the relative distance between the vehicle and the target vehicle in the target lane, and the relative speed between the vehicle and the target vehicle in the target lane into the neural network to obtain the predicted trajectory of the vehicle; Wherein, the neural network based on the self-attention mechanism includes a first input layer, an encoder and a decoder connected in sequence; The first input layer satisfies: In the formula, is the output of the input layer at time t, α is the scalar projection coefficient, l i To upgrade to d model Vehicle trajectory information at time t PE(·) is the position code, ψ is the number of extracted features, SE(·) is the timestamp, L x The length of the encoder and decoder inputs; The encoder includes a second input layer, a first embedding layer, a first multi-head probabilistic sparse self-attention module, a first feedforward layer, a first self-attention distillation module, a second multi-head probabilistic sparse self-attention module, a second embedding layer and a first output layer, wherein the output of the first output layer is the lane changing state of the vehicle; The decoder includes a third input layer, a third embedding layer, a third multi-head probabilistic sparse self-attention module, a full attention layer, a second feedforward layer, a second self-attention distillation module and a second output layer, and the output of the second output layer is the predicted trajectory of the vehicle.

2. The neural network vehicle trajectory prediction method based on self-attention mechanism as claimed in claim 1, characterized in that: The position encoding satisfies: Where pos is the position in the sequence at time t, and its value range is [1, 2, …, d model ],d model is the total length of the time feature vector; The timestamp satisfies:

3. The neural network vehicle trajectory prediction method based on self-attention mechanism as claimed in claim 2, characterized in that: The second input layer satisfies: x χ =concat(P 0 (t),S M (t)); In the formula, x χ is the input of the encoder, P 0 (t) is the historical trajectory of the vehicle, S M (t) is the interactive information of the vehicle; The historical trajectory of the vehicle satisfies: Where, t∈(-mT,0); The interactive information of the vehicle satisfies: In the formula, S M (t) is the vehicle interaction information at time t, is the lateral relative distance between the vehicle and vehicle M, is the longitudinal relative distance between the vehicle and vehicle M, is the lateral relative speed between the vehicle and vehicle M, is the longitudinal relative speed between the vehicle and vehicle M, and M is the target vehicle number in the target lane.

4. The neural network vehicle trajectory prediction method based on self-attention mechanism as claimed in claim 3, characterized in that: The embedding layer satisfies: In the formula, K δ is the output of the embedding layer, and w1 is the weight of the first fully connected layer.

5. The neural network vehicle trajectory prediction method based on self-attention mechanism as claimed in claim 4, characterized in that: The first multi-head probabilistic sparse self-attention module satisfies: K ε =norm[prob(K δ ,K δ ,K δ )+K δ ]; In the formula, prob is the probabilistic sparse self-attention layer, norm is the normalization layer; The probabilistic sparse self-attention layer satisfies: Where A is the probabilistic sparse self-attention, Q is the set of queries, K is the set of keys, V is the set of values, and d is the dimension of queries and keys.

6. The neural network vehicle trajectory prediction method based on self-attention mechanism as claimed in claim 5, characterized in that: The first feed-forward layer satisfies: K μ =norm[relu(K ε w2)w2+K ε ]; Where relu is the activation function and w2 is the weight of the second fully connected layer.

7. The neural network vehicle trajectory prediction method based on self-attention mechanism as claimed in claim 6, characterized in that: The first self-attention distillation module satisfies: K η =MaxPool[ELU(Conv(K μ ))]; In the formula, MaxPool is the maximum pooling layer, ELU is the exponential linear unit, and Conv is the one-dimensional convolution layer.

8. The neural network vehicle trajectory prediction method based on self-attention mechanism as claimed in claim 7, characterized in that: The first output layer satisfies: In the formula, is the probability that the vehicle maintains the original lane, is the probability of the vehicle changing lanes left, is the probability of the vehicle changing lanes right.

9. The neural network vehicle trajectory prediction method based on self-attention mechanism as claimed in claim 8, characterized in that: The third input layer satisfies: In the formula, is the masked trajectory of the vehicle; Ltoken is the length of the label history trajectory, and nT is the target trajectory segment with prediction.

10. The neural network vehicle trajectory prediction method based on self-attention mechanism as claimed in claim 9, characterized in that: The third embedding layer satisfies: Where w5 is the weight of the fifth fully connected layer, and w6 is the weight of the sixth fully connected layer; The second output layer satisfies: Θ=K τ w7; In the formula, K τ is the output of the second self-attention distillation module, and w7 is the output of the seventh fully connected layer.