A vehicle trajectory prediction method based on attention mechanism

Through the vehicle trajectory prediction method based on attention mechanism, the problem of difficulty in making global prediction of multiple agents in the prior art is solved, and accurate prediction of the future trajectory of multiple vehicles and the safety of autonomous driving is improved.

CN118953406BActive Publication Date: 2025-05-06JIANGSU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411107822.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-13
Publication Date
2025-05-06
Estimated Expiration
2044-08-13

AI Technical Summary

Technical Problem

The existing trajectory prediction technology is difficult to make global predictions of multiple agents and cannot effectively predict the interaction and coordination between multiple vehicles.

Method used

The vehicle trajectory prediction method based on attention mechanism is adopted, and data encoding and spatial division is performed by obtaining vehicle information and map information, and a feature extraction network and global feature interaction network are used, and multimodal trajectory prediction is performed in combination with lane screening results.

Benefits of technology

Accurate prediction of the future trajectory of all vehicles in the scene is achieved, can adapt to scenarios with different vehicle densities, reduce feature deviations due to scene transformation, and improve the safety of autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118953406B_ABST
    Figure CN118953406B_ABST
Patent Text Reader

Abstract

The present invention discloses a vehicle trajectory prediction method based on an attention mechanism, comprising: obtaining driving target vehicle and adjacent vehicle data and global scene data; dividing local scenes, encoding local scenes, and obtaining a set of vectorized entities in local scenes; extracting vehicle-lane interaction features for each entity in the scene; extracting vehicle-vehicle interaction features for each entity in the scene; extracting trajectory features for each entity in the scene; performing global feature interaction through a global feature interaction network to obtain globally fused features; obtaining candidate paths through a lane screening network; and performing multimodal trajectory prediction based on globally fused features and candidate paths. The present invention can adapt to various scenarios with different vehicle densities, and can reasonably allocate attention in both sparse lanes and dense lanes; reducing feature deviations caused by scene changes, helping to accurately predict trajectories and ensure the driving safety of intelligent vehicles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a vehicle trajectory prediction method based on an attention mechanism, and belongs to the technical field of intelligent driving. Background Art

[0002] With the development of the times, smart cars have become the forefront of transportation development, so the concept of autonomous driving has gradually entered people's field of vision. Just as human drivers need to continuously observe, analyze and predict the movement trajectories of surrounding vehicles and pedestrians, smart cars should also have the ability to perceive and predict the trajectories and intentions of surrounding vehicles, and on this basis plan a reasonable driving path to effectively avoid collisions with surrounding vehicles, pedestrians or obstacles, thereby ensuring driving safety. Therefore, trajectory prediction is a very important technology in autonomous driving.

[0003] Trajectory prediction is a technology that predicts the trajectory of a moving target in the next period of time based on the known trajectory and surrounding information of a moving target in a period of time. Existing research can be divided into two categories: traditional physics-based models and data-driven deep learning models. Traditional physics-based models usually use dynamics and kinematics as constraints, use methods such as Kalman filtering for prediction, and use probabilistic models to predict multiple possibilities in the future. However, this method can only predict trajectories in a short period of time and cannot meet the needs of actual driving.

[0004] At present, most cutting-edge trajectory prediction technologies are based on deep learning, which uses deep neural networks to learn large data sets and establish prediction models. Due to the subjectivity of human driving, the interaction between surrounding vehicles and the vehicle itself is a very important part in the neural network model. However, the current trajectory prediction technology can only perform trajectory prediction for a single agent, and cannot perform global prediction for multiple agents. Summary of the invention

[0005] Purpose of the invention: In view of the shortcomings of the prior art, the present invention provides a vehicle trajectory prediction method based on the attention mechanism. The present invention obtains vehicle information and map information as input, encodes the relevant information through data processing, passes through a spatial division network, and then passes through a feature extraction network, a global feature interaction network and trajectory prediction in turn, and finally obtains the future trajectories of all vehicles in the scene in combination with the lane screening results.

[0006] Technical solution: A vehicle trajectory prediction method based on attention mechanism, comprising the following steps:

[0007] Step 1: Obtain the driving target vehicle and adjacent vehicle data as well as the global scene data;

[0008] Step 2: The global scene data is spatially divided according to the spatial division network to obtain a local scene, the local scene is encoded, and the past trajectory of the vehicle and the lane line composed of lane sampling points are vectorized to obtain a set of vectorized entities in the local scene;

[0009] Step 3: Using the vehicle-lane interaction feature extraction network, extract the vehicle-lane interaction features for the entities in each scene according to the set of vectorized entities obtained in step 2;

[0010] Step 4: Extract vehicle-to-vehicle interaction features for each entity in the scene through the vehicle-to-vehicle interaction feature extraction network;

[0011] Step 5: Extract trajectory features of entities in each scene through the trajectory feature extraction network;

[0012] Step 6: Based on the local scene features extracted from steps 3 to 5, global feature interaction is performed through a global feature interaction network to obtain globally fused features;

[0013] Step 7: Process the lane information and obtain candidate paths through the lane screening network: The global scene data obtained in step 1 is classified through a multi-layer perceptron, and each lane is scored and sorted from high to low according to the score. The top M candidate lane lines and the hidden layer features h of each candidate lane line are selected. m ;

[0014] Step 8: Perform multimodal trajectory prediction based on the globally fused features obtained in step 6 and the candidate paths obtained in step 7.

[0015] Preferably, the step 1 is specifically:

[0016] The vehicle data includes the type, number, two-dimensional coordinates and heading angle of the vehicle;

[0017] The scene data includes the 2D coordinates of the lane sampling points and the lane numbers.

[0018] The vehicle data and its corresponding scene data are divided into training set, validation set and test set as a whole.

[0019] Preferably, the step 2 is specifically as follows:

[0020] The vehicle data in step 1 and its corresponding local scene data are encoded. The past trajectory of the vehicle and the lane line composed of lane sampling points are vectorized according to the data in step 1 to obtain a set of vectorized entities in the scene, and different vehicles are classified as semantic E a enter;

[0021] In order to ensure the translation invariance of the coordinates, the two-dimensional coordinates of the adjacent vehicle or lane sampling points are converted to the coordinates centered on the target vehicle. The specific process can be expressed as the following formula:

[0022]

[0023] in, It is based on the heading angle θ of the target vehicle’s final position i The resulting rotation matrix is is the two-dimensional coordinate of vehicle number i at time t, is the two-dimensional coordinate of vehicle number j at time t;

[0024] When performing space division, when any vehicle is selected as the target vehicle, the surrounding lanes are searched, and the forward direction of the target vehicle is taken as the positive direction. n vehicles are searched from the positive and negative directions on each lane respectively, which are used as the boundary of space division. If there is no vehicle in one direction of the current lane, the distance d is used as the boundary of space division.

[0025] Preferably, the step three is specifically as follows:

[0026] Extract vehicle-lane interaction features and extract the trajectory features h of the target vehicle through a multi-layer perceptron s As a query of the spatial attention mechanism; use a multi-layer perceptron to extract lane features h l ; The interaction between the lane centerline and the target vehicle is realized through the spatial attention mechanism, and finally the vehicle-lane interaction feature embedding h is obtained ls , this process is specifically expressed as:

[0027]

[0028] Among them, MLP is a standard multi-layer perceptron. is the trajectory sequence of the target vehicle; E a is the semantic attribute of the target vehicle; represents the lane features of the mth lane; is the position sequence of the mth lane; E l is the lane semantic attribute; are three learnable linear projection matrices; M is the number of lanes in the local area, softmax is the activation function; q s is the query vector obtained through the target vehicle trajectory features; k m is the key vector obtained by all lane features in the local scene; v m is a value vector obtained by passing through all lane features in the local scene.

[0029] Preferably, the step 4 is specifically as follows:

[0030] Extract target vehicles through multi-layer perceptron and adjacent vehicle characteristics The interaction between vehicles is realized through the spatial attention mechanism to obtain environmental features Then through the gating function The target vehicle features and environmental characteristics Fusion, finally get the vehicle-to-vehicle interaction feature embedding This process is specifically expressed as:

[0031]

[0032]

[0033] Among them, MLP is a standard multi-layer perceptron. is the trajectory vector of the target vehicle at time t, E i and E j are the semantic attributes of the target vehicle and the adjacent vehicles, represents the feature embedding of the target vehicle at time t, and also implies the interactive feature embedding h of the target vehicle and the lane ls , It contains the trajectory vector of the neighboring vehicle j and its position relative to the target vehicle, making the neighbor embedding spatially aware, j∈N j , N j is the number of neighboring vehicles in the local area, W gate , W self is a learnable linear projection matrix, d k is the dimension of k, representing the scaling factor, softmax and sigmoid are activation functions, ⊙ is the element-wise product, representing the multiplication of elements at the same position in the matrix; q i is the query vector obtained by the target vehicle trajectory features, and k ij and v ij are the key vector and value vector obtained through the features of adjacent vehicle trajectories.

[0034] Preferably, the step five is specifically as follows:

[0035] The result extracted by the multi-head attention mechanism in step 4 As the input value of the trajectory feature extraction network, it contains the features of adjacent vehicles and related lanes in the local scene, and is embedded as a feature after passing through the multi-layer perceptron. i ; Then, we learn through the spatial attention mechanism, deeply extract the features of the time series, and then obtain the final spatiotemporal embedding h through the multi-layer perceptroni , this process is specifically expressed as:

[0036]

[0037] Among them, MLP is a multi-layer perceptron. is a learnable linear projection matrix, and Mask is a mask. Its purpose is to focus only on the current moment and the previous trajectory in the attention mechanism, which can effectively avoid being affected by the true value in the feedforward process of training; q i , k i 、v i They are respectively the query vector, key vector and value vector obtained through the target vehicle trajectory features.

[0038] Preferably, the step six is ​​specifically as follows:

[0039] Since the extracted features in the local scene are all extracted from the coordinate vector centered on the target vehicle, the coordinate vector needs to be transformed when performing global feature interaction to parameterize the coordinate difference between vehicle i and vehicle j;

[0040] The global scene is regarded as a graph, and the vehicles are regarded as nodes. The global information interaction is realized through the graph attention mechanism, and the node features are updated layer by layer to obtain the globally fused features as output. This process is specifically expressed as follows:

[0041]

[0042] z l =f l W l

[0043]

[0044] Among them, MLP is a multi-layer perceptron. and are the coordinates of vehicle i and vehicle j at time t, θ ij is the heading angle difference between the two vehicles, c ij is the coordinate parameterized feature between the two vehicles, l is the number of layers of the graph attention mechanism, f is the input feature of the current node, is the node feature of the i-th vehicle at layer l, f i l+1 is the node feature of the i-th vehicle at the l+1 layer after updating, z is the mapped node feature, is the edge feature between node i and node j, is the calculated attention score, W l and W elis a learnable linear projection matrix, LeakyReLU, softmax and ReLU are activation functions, Represents the connection in dimension; k∈N i , N i is the set of neighboring vehicles of vehicle i.

[0045] Preferably, the step eight is specifically as follows:

[0046] The hidden layer feature h m Replicate K times and use the global features obtained in step 6 as the input of the decoder to obtain M×K target vehicle selected driving trajectories and their feasible probability estimates. Select the first K trajectories with the largest probability. When calculating the final loss function, only select the trajectory with the smallest difference from the true value for calculation to achieve multiple possibilities of the final predicted trajectory. For the loss function, including a regression task and a classification task, the specific calculation formula of the loss function is:

[0047] L=L reg +L cls

[0048]

[0049]

[0050] Among them, L reg represents the regression task loss, L cls represents the classification task loss, T represents the time step of the predicted future trajectory, and y t represents the true value of the trajectory, represents the predicted value of the kth trajectory at the future time t, represents the coordinates of candidate lane j, D lane represents the sum of the Euclidean distances between the true value and the nearest node of the candidate lane, p l Represents the scores of the M candidate lanes obtained in step 7, and softmax is the activation function.

[0051] Preferably, the layer normalization method is used in the vehicle-lane interaction feature extraction network, the vehicle-vehicle interaction feature extraction network, the trajectory feature extraction network, and the global feature interaction network to help the model converge, and the residual connection is used to avoid the gradient disappearance during the back propagation process. The formula is as follows:

[0052] h'=h out +h in

[0053]

[0054] Among them, h out is the output of this module, hin The input of this module, γ and δ are adjustable parameters learned autonomously during model training, μ is the mean of h', σ is the standard deviation of h', and ε is a constant used to stabilize the standard deviation;

[0055] Since real-world errors are difficult to avoid, the evaluation criteria use the minimum average displacement error minADE, the minimum final displacement error minFDE and the missing rate MR. The missing rate MR is defined as the probability that the minFDE of all predicted trajectories in the scene is greater than 2m. The formulas for minADE and minFDE are:

[0056]

[0057] Where T is the time step of the predicted future trajectory, y t represents the true value of the trajectory, Represents the predicted value of the k-th trajectory at the future time t.

[0058] Beneficial effects: The present invention flexibly divides the scene by using a space division network, avoiding the disadvantage that the model cannot be applied to different scenarios due to the bias of a single data set, and can adapt to various scenarios with different vehicle densities. It can reasonably allocate attention in both sparse lanes and dense lanes; through the rotation matrix and translation-invariant scene representation, the feature deviation caused by scene change is reduced, which can be widely used in the field of autonomous driving, helping to accurately predict the trajectory and ensure the driving safety of intelligent vehicles; through a three-layer feature extraction network, feature extraction is performed from three levels: lanes, adjacent vehicles, and the target vehicle's own trajectory. The global interaction features are obtained through the global feature interaction network, and information is gradually gathered at multiple scales for efficient modeling. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0060] Figure 1 is a flow chart of the method of the present invention;

[0061] Figure 2 It is a schematic diagram of the overall model structure of the present invention;

[0062] Figure 3 It is a schematic diagram of the space division of the present invention;

[0063] Figure 4 This is a network structure diagram for extracting vehicle-lane interaction features according to the present invention;

[0064] Figure 5 The vehicle-to-vehicle interaction feature extraction network of the present invention;

[0065] Figure 6 Schematic diagram of the attention mechanism of the present invention; Figure 6 (a) shows the calculation method of node weighted features; Figure 6 (b) Represents the updating method of node features. DETAILED DESCRIPTION

[0066] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0067] In the description of the present invention, it is necessary to understand that the terms "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.

[0068] In the present invention, unless otherwise clearly specified and limited, a first feature being "above" or "below" a second feature may include that the first and second features are in direct contact, or may include that the first and second features are not in direct contact but are in contact through another feature between them. Moreover, a first feature being "above", "above" and "above" a second feature includes that the first feature is directly above and obliquely above the second feature, or simply indicates that the first feature is higher in level than the second feature. A first feature being "below", "below" and "below" a second feature includes that the first feature is directly below and obliquely below the second feature, or simply indicates that the first feature is lower in level than the second feature.

[0069] like Figure 1 and Figure 2 As shown, a vehicle trajectory prediction method based on an attention mechanism includes the following steps:

[0070] Step 1: Obtain the driving target vehicle and adjacent vehicle data as well as the global scene data;

[0071] The vehicle data includes the type, number, two-dimensional coordinates and heading angle of the vehicle;

[0072] The scene data includes the two-dimensional coordinates of the lane sampling points and the lane numbers;

[0073] In this example, we choose the Argoverse dataset, which contains high-precision maps and 323,557 valuable real driving scenarios. Each scenario includes a 5-second 10Hz sequence, which is divided into a training set of 205,942 samples, a validation set of 39,472 samples, and a test set of 78,143 samples. The first 2 seconds of the test set samples are used as historical trajectories to predict the trajectory of the next 3 seconds. The specific input can be expressed as:

[0074] X={A,M}

[0075] A={S,E a}

[0076] S={S i},E a ={e i}

[0077]

[0078] M={L,E m}

[0079] L={(x k ,y k ),d k},E m ={e k}

[0080] Among them, A is vehicle information, M is map information, S is vehicle trajectory sequence, i is the vehicle serial number, is the two-dimensional position coordinate of vehicle i at time t, is the heading angle of vehicle i at time t, t∈(1,…,T), T is the final time step of the historical trajectory, e is the semantics of the point, L is the coordinate of the lane sampling point, and k is the sequence number of the lane sampling point.

[0081] Step 2: The global scene data is spatially divided according to the spatial division network to obtain a local scene, the local scene is encoded, and the past trajectory of the vehicle and the lane line composed of lane sampling points are vectorized to obtain a set of vectorized entities in the local scene;

[0082] The vehicle data in step 1 and its corresponding local scene data are encoded. The past trajectory of the vehicle and the lane line composed of lane sampling points are vectorized according to the data in step 1 to obtain a set of vectorized entities in the scene. Different vehicles are divided into four types: pedestrians, bicycles or motorcycles, small vehicles, and large vehicles as semantic E a enter;

[0083] In order to ensure the translation invariance of the coordinates, the two-dimensional coordinates of the adjacent vehicle or lane sampling points are converted to the coordinates centered on the target vehicle. The specific process can be expressed as the following formula:

[0084]

[0085] in, It is based on the heading angle θ of the target vehicle’s final position i The resulting rotation matrix is is the two-dimensional coordinate of vehicle number i at time t, p t j is the two-dimensional coordinate of vehicle number j at time t;

[0086] When performing space division, when any vehicle is selected as the target vehicle, the surrounding lanes are searched, and the forward direction of the target vehicle is taken as the positive direction. n vehicles are searched from the positive and negative directions on each lane respectively, which are used as the boundary of space division. If there is no vehicle in one direction of the current lane, the distance d is used as the boundary of space division.

[0087] like Figure 3 As shown, in this example, the target vehicle A1 is selected, and the six surrounding vehicles are selected as neighboring vehicles. When the target vehicle is A3, there are only three vehicles as neighboring vehicles. This processing method can flexibly segment the global scene, so that the target vehicle pays more attention to vehicles that have a greater impact on its own trajectory. It can adapt to both crowded lanes and sparse lanes, and has good generalization ability.

[0088] Step 3: Using the vehicle-lane interaction feature extraction network, extract the vehicle-lane interaction features for the entities in each scene according to the set of vectorized entities obtained in step 2;

[0089] like Figure 4 As shown in the figure, the vehicle-lane interaction features are extracted, and the trajectory features h of the target vehicle are extracted through a multi-layer perceptron. s As a query of the spatial attention mechanism; use a multi-layer perceptron to extract lane features h l ; The interaction between the lane centerline and the target vehicle is realized through the spatial attention mechanism, and finally the vehicle-lane interaction feature embedding h is obtained ls , this process is specifically expressed as:

[0090]

[0091] Among them, MLP is a standard multi-layer perceptron. is the trajectory sequence of the target vehicle; E a is the semantic attribute of the target vehicle; represents the lane features of the mth lane; is the position sequence of the mth lane; E l is the lane semantic attribute; are three learnable linear projection matrices; M is the number of lanes in the local area, softmax is the activation function; q s is the query vector obtained through the target vehicle trajectory features; k m is the key vector obtained by all lane features in the local scene; v m is a value vector obtained by passing through all lane features in the local scene.

[0092] Step 4: Extract vehicle-to-vehicle interaction features for each entity in the scene through the vehicle-to-vehicle interaction feature extraction network;

[0093] like Figure 5 As shown, the target vehicle is extracted through a multi-layer perceptron and adjacent vehicle characteristics The interaction between vehicles is realized through the spatial attention mechanism to obtain environmental features Then through the gating function The target vehicle features and environmental characteristics Fusion, finally get the vehicle-to-vehicle interaction feature embedding This process is specifically expressed as:

[0094]

[0095] Among them, MLP is a standard multi-layer perceptron. is the trajectory vector of the target vehicle at time t, E i and E j are the semantic attributes of the target vehicle and the adjacent vehicles, represents the feature embedding of the target vehicle at time t, and also implies the interactive feature embedding h of the target vehicle and the lane ls , It contains the trajectory vector of the neighboring vehicle j and its position relative to the target vehicle, making the neighbor embedding spatially aware, j∈N j , N j is the number of neighboring vehicles in the local area, W gate , W self is a learnable linear projection matrix, d k is the dimension of k, representing the scaling factor, softmax and sigmoid are activation functions, ⊙ is the element-wise product, representing the multiplication of elements at the same position in the matrix; q i is the query vector obtained by the target vehicle trajectory features, and kij and v ij are the key vector and value vector obtained through the features of adjacent vehicle trajectories.

[0096] Step 5: Extract trajectory features of entities in each scene through the trajectory feature extraction network;

[0097] The result extracted by the multi-head attention mechanism in step 4 As the input value of the trajectory feature extraction network, it contains the features of adjacent vehicles and related lanes in the local scene, and is embedded as a feature after passing through the multi-layer perceptron. i ; Then, we learn through the spatial attention mechanism, deeply extract the features of the time series, and then obtain the final spatiotemporal embedding h through the multi-layer perceptron i , this process is specifically expressed as:

[0098]

[0099] Among them, MLP is a multi-layer perceptron. is a learnable linear projection matrix, and Mask is a mask. Its purpose is to focus only on the current moment and the previous trajectory in the attention mechanism, which can effectively avoid being affected by the true value in the feedforward process of training; q i , k i 、v i They are respectively the query vector, key vector and value vector obtained through the target vehicle trajectory features.

[0100] The formula for the mask is:

[0101]

[0102] Where u represents the current time step, and v represents the time step involved in the calculation at each step.

[0103] Step 6: Based on the local scene features extracted from steps 3 to 5, global feature interaction is performed through a global feature interaction network to obtain globally fused features;

[0104] Since the extracted features in the local scene are all extracted from the coordinate vector centered on the target vehicle, the coordinate vector needs to be transformed when performing global feature interaction to parameterize the coordinate difference between vehicle i and vehicle j;

[0105] like Figure 6 As shown in (a, b), the weighted features of node i and node j and The attention scores of the two nodes are calculated through the activation function, and then the attention scores of all adjacent nodes of a node are added to update the node feature.

[0106] The global scene is regarded as a graph, and the vehicles are regarded as nodes. The global information interaction is realized through the graph attention mechanism, and the node features are updated layer by layer to obtain the globally fused features as output. This process is specifically expressed as follows:

[0107]

[0108] z l =f l W l

[0109]

[0110]

[0111] Among them, MLP is a multi-layer perceptron. and are the coordinates of vehicle i and vehicle j at time t, θ ij is the heading angle difference between the two vehicles, c ij is the coordinate parameterized feature between the two vehicles, l is the number of layers of the graph attention mechanism, f is the input feature of the current node, is the node feature of the i-th vehicle at layer l, f i l+1 is the node feature of the i-th vehicle at the l+1 layer after updating, z is the mapped node feature, is the edge feature between node i and node j, is the calculated attention score, W l and W el is a learnable linear projection matrix, LeakyReLU, softmax and ReLU are activation functions, Represents the connection in dimension; k∈N i , N i is the set of neighboring vehicles of vehicle i.

[0112] Step 7: Process the lane information and obtain candidate paths through the lane screening network: The global scene data obtained in step 1 is classified through a multi-layer perceptron, and each lane is scored and sorted from high to low according to the score. The top M candidate lane lines and the hidden layer features h of each candidate lane line are selected. m ;

[0113] Step 8: Perform multimodal trajectory prediction based on the globally fused features obtained in step 6 and the candidate paths obtained in step 7.

[0114] The hidden layer feature h m Replicate K times and use the global features obtained in step 6 as the decoder input to obtain M×K target vehicle selected driving trajectories and their feasible probability estimates. Select the first K trajectories with the largest probability. When calculating the final loss function, only select the trajectory with the smallest difference from the true value for calculation to achieve multiple possibilities of the final predicted trajectory. The final output result is represents the coordinates of the K trajectories of the i-th vehicle from time T to T+P. For the loss function, it includes a regression task and a classification task. The specific calculation formula of the loss function is:

[0115] L=L reg +L cls

[0116]

[0117] Among them, L reg represents the regression task loss, L cls represents the classification task loss, T represents the time step of the predicted future trajectory, and y t represents the true value of the trajectory, represents the predicted value of the kth trajectory at time t in the future, l d j represents the coordinates of candidate lane j, D lane represents the sum of the Euclidean distances between the true value and the nearest node of the candidate lane, p l Represents the scores of the M candidate lanes obtained in step 7, and softmax is the activation function.

[0118] The vehicle-lane interaction feature extraction network, vehicle-vehicle interaction feature extraction network, trajectory feature extraction network, and global feature interaction network all use layer normalization to help the model converge, and use residual connections to avoid gradient disappearance during back propagation. The formula is as follows:

[0119] h'=h out +h in

[0120]

[0121] Among them, h out is the output of this module, h in The input of this module, γ and δ are adjustable parameters learned autonomously during model training, μ is the mean of h', σ is the standard deviation of h', and ε is a constant used to stabilize the standard deviation;

[0122] Since real-world errors are difficult to avoid, the evaluation criteria use the minimum average displacement error minADE, the minimum final displacement error minFDE and the missing rate MR. The missing rate MR is defined as the probability that the minFDE of all predicted trajectories in the scene is greater than 2m. The formulas for minADE and minFDE are:

[0123]

[0124] Where T is the time step of the predicted future trajectory, y t represents the true value of the trajectory, Represents the predicted value of the k-th trajectory at the future time t.

[0125] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0126] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A vehicle trajectory prediction method based on attention mechanism, characterized in that: The following steps are involved: Step 1: Obtain the driving target vehicle and adjacent vehicle data as well as the global scene data; Step 2: The global scene data is spatially divided according to the spatial division network to obtain a local scene, the local scene is encoded, and the past trajectory of the vehicle and the lane line composed of lane sampling points are vectorized to obtain a set of vectorized entities in the local scene; Step 3: Using the vehicle-lane interaction feature extraction network, extract the vehicle-lane interaction features for the entities in each scene according to the set of vectorized entities obtained in step 2; Step 4: Extract vehicle-to-vehicle interaction features for each entity in the scene through the vehicle-to-vehicle interaction feature extraction network; Step 5: Extract trajectory features of entities in each scene through the trajectory feature extraction network; Step 6: Based on the local scene features extracted from steps 3 to 5, global feature interaction is performed through a global feature interaction network to obtain globally fused features; Step 7: Process the lane information and obtain candidate paths through the lane screening network: Perform a classification task on the global scene data obtained in step 1 through a multi-layer perceptron, score each lane and sort it from high to low according to the score, select the first M candidate lane lines, and the hidden layer features h of each candidate lane line m ; Step 8: Perform multimodal trajectory prediction based on the globally fused features obtained in step 6 and the candidate paths obtained in step 7.

2. The vehicle trajectory prediction method based on the attention mechanism according to claim 1 is characterized in that: The step 1 is specifically as follows: The vehicle data includes the type, number, two-dimensional coordinates and heading angle of the vehicle; The scene data includes the two-dimensional coordinates of the lane sampling points and the lane numbers; The vehicle data and its corresponding scene data are divided into training set, validation set and test set as a whole.

3. The vehicle trajectory prediction method based on the attention mechanism according to claim 1 is characterized in that: The step 2 is specifically as follows: The vehicle data in step 1 and its corresponding local scene data are encoded. The past trajectory of the vehicle and the lane line composed of lane sampling points are vectorized according to the data in step 1 to obtain a set of vectorized entities in the scene, and different vehicles are classified as semantic E a enter; In order to ensure the translation invariance of the coordinates, the two-dimensional coordinates of the adjacent vehicles are converted to coordinates centered on the target vehicle. The specific process can be expressed as the following formula: in, It is based on the heading angle θ of the target vehicle's final position i The resulting rotation matrix is is the two-dimensional coordinate of vehicle number i at time t, is the two-dimensional coordinate of vehicle number j at time t; When performing space division, when any vehicle is selected as the target vehicle, the surrounding lanes are searched, and the forward direction of the target vehicle is taken as the positive direction. n vehicles are searched from the positive and negative directions on each lane respectively, which are used as the boundary of space division. If there is no vehicle in one direction of the current lane, the distance d is used as the boundary of space division.

4. The vehicle trajectory prediction method based on the attention mechanism according to claim 1 is characterized in that: The step three is specifically as follows: Extract vehicle-lane interaction features and extract the trajectory features h of the target vehicle through a multi-layer perceptron s As a query of the spatial attention mechanism; use a multi-layer perceptron to extract lane features h l ; The interaction between the lane centerline and the target vehicle is realized through the spatial attention mechanism, and finally the vehicle-lane interaction feature embedding h is obtained ls , this process is specifically expressed as: Among them, MLP is a standard multi-layer perceptron. is the trajectory sequence of the target vehicle; E a is the semantic attribute of the target vehicle; represents the lane features of the mth lane; is the position sequence of the mth lane; E l is the lane semantic attribute; are three learnable linear projection matrices; M is the number of lanes in the local area, softmax is the activation function; q s is the query vector obtained through the target vehicle trajectory features; k m is the key vector obtained by all lane features in the local scene; v m is a value vector obtained by passing through all lane features in the local scene.

5. The vehicle trajectory prediction method based on the attention mechanism according to claim 1 is characterized in that: The step 4 is specifically as follows: Extract target vehicles through multi-layer perceptron and adjacent vehicle characteristics The interaction between vehicles is realized through the spatial attention mechanism to obtain environmental features Then through the gating function The target vehicle features and environmental characteristics Fusion, finally get the vehicle-to-vehicle interaction feature embedding This process is specifically expressed as: Among them, MLP is a standard multi-layer perceptron. is the trajectory vector of the target vehicle at time t, E i and E j are the semantic attributes of the target vehicle and the adjacent vehicles, represents the feature embedding of the target vehicle at time t, and also implies the interactive feature embedding h of the target vehicle and the lane ls , It contains the trajectory vector of the neighboring vehicle j and its position relative to the target vehicle, making the neighbor embedding spatially aware, j∈N j , N j is the number of neighboring vehicles in the local area, W gate , W self is a learnable linear projection matrix, d k is the dimension of k, representing the scaling factor, softmax and sigmoid are activation functions, ⊙ is the element-wise product, representing the multiplication of elements at the same position in the matrix; q i is the query vector obtained by the target vehicle trajectory features, and k ij and v ij are the key vector and value vector obtained through the features of adjacent vehicle trajectories.

6. The vehicle trajectory prediction method based on the attention mechanism according to claim 1 is characterized in that: The step five is specifically as follows: The result extracted by the multi-head attention mechanism in step 4 As the input value of the trajectory feature extraction network, it contains the features of adjacent vehicles and related lanes in the local scene, and is embedded as a feature after passing through the multi-layer perceptron. i ; Then, we learn through the spatial attention mechanism, deeply extract the features of the time series, and then obtain the final spatiotemporal embedding h through the multi-layer perceptron i , this process is specifically expressed as: Among them, MLP is a multi-layer perceptron. is a learnable linear projection matrix, and Mask is a mask. Its purpose is to focus only on the current moment and the previous trajectory in the attention mechanism, which can effectively avoid being affected by the true value in the feedforward process of training; q i , k i 、v i They are respectively the query vector, key vector and value vector obtained through the target vehicle trajectory features.

7. The vehicle trajectory prediction method based on the attention mechanism according to claim 1 is characterized in that: The step six is ​​specifically as follows: Since the extracted features in the local scene are all extracted from the coordinate vector centered on the target vehicle, the coordinate vector needs to be transformed when performing global feature interaction to parameterize the coordinate difference between vehicle i and vehicle j; The global scene is regarded as a graph, and the vehicles are regarded as nodes. The global information interaction is realized through the graph attention mechanism, and the node features are updated layer by layer to obtain the globally fused features as output. This process is specifically expressed as follows: Among them, MLP is a multi-layer perceptron. and are the coordinates of vehicle i and vehicle j at time t, θ ij is the heading angle difference between the two vehicles, c ij is the coordinate parameterized feature between the two vehicles, l is the number of layers of the graph attention mechanism, f is the input feature of the current node, is the node feature of the i-th vehicle at layer l, is the node feature of the i-th vehicle at the l+1 layer after updating, z is the mapped node feature, is the edge feature between node i and node j, is the calculated attention score, W l and W el is a learnable linear projection matrix, LeakyReLU, softmax and ReLU are activation functions, ⊕ represents the connection in dimension, k∈N i , N i is the set of neighboring vehicles of vehicle i.

8. The vehicle trajectory prediction method based on the attention mechanism according to claim 1 is characterized in that: The step eight is specifically as follows: The hidden layer feature h m Replicate K times and use the global features obtained in step 6 as the input of the decoder to obtain M×K target vehicle selected driving trajectories and their feasible probability estimates. Select the first K trajectories with the largest probability. When calculating the final loss function, only select the trajectory with the smallest difference from the true value for calculation to achieve multiple possibilities of the final predicted trajectory. For the loss function, including a regression task and a classification task, the specific calculation formula of the loss function is: L=L reg +L cls Among them, L reg represents the regression task loss, L cls represents the classification task loss, T represents the time step of the predicted future trajectory, and y t represents the true value of the trajectory, represents the predicted value of the kth trajectory at the future time t, represents the coordinates of candidate lane j, D lane represents the sum of the Euclidean distances between the true value and the nearest node of the candidate lane, p l Represents the scores of the M candidate lanes obtained in step 7, and softmax is the activation function.

9. The vehicle trajectory prediction method based on the attention mechanism according to claim 1, characterized in that: The vehicle-lane interaction feature extraction network, vehicle-vehicle interaction feature extraction network, trajectory feature extraction network, and global feature interaction network all use layer normalization to help the model converge, and use residual connections to avoid gradient disappearance during back propagation. The formula is as follows: h'=h out +h in Among them, h out is the output of the module, h in is the input of the module, γ and δ are adjustable parameters learned autonomously during model training, μ is the mean of h', σ is the standard deviation of h', and ε is a constant used to stabilize the standard deviation; Since real-world errors are difficult to avoid, the evaluation criteria use the minimum average displacement error minADE, the minimum final displacement error minFDE and the missing rate MR. The missing rate MR is defined as the probability that the minFDE of all predicted trajectories in the scene is greater than 2m. The formulas for minADE and minFDE are: Where T is the time step of the predicted future trajectory, y t represents the true value of the trajectory, Represents the predicted value of the k-th trajectory at the future time t.

Citation Information

Patent Citations

  • Urban scene-oriented vehicle trajectory prediction method and system, and storage medium

    CN115009275A

  • Multi-modal vehicle trajectory prediction method based on hierarchical order network

    CN116513240A