Multi-mode vehicle trajectory prediction method fusing vehicle motion and lane information

By adopting a CVAE-based model architecture in vehicle trajectory prediction, latent variable modeling of vehicle motion and lane information is solved, and the problem of insufficient feature extraction in the prior art is achieved, and higher prediction accuracy and diversity are achieved.

CN120032331APending Publication Date: 2025-05-23ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510001828.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When the prior art combines vehicle motion and lane information for multi-mode vehicle trajectory prediction, the potential correlation between data is not fully explored, especially when dealing with scene map changes caused by vehicle motion, feature extraction is insufficient, which limits the improvement of prediction accuracy.

Method used

Using a CVAE-based model architecture, the personalized characteristics and common characteristics of vehicle motion and lane information are modeled through multiple latent variables, and the interaction between surrounding lane changes and vehicles is generated to generate diversified trajectories.

Benefits of technology

The accuracy and diversity of vehicle trajectory prediction results are effectively improved. By potentially modeling the spatial-semantic relationship between the vehicle and the lane, the ELBO objective function structure of CVAE is improved, and the expression ability of latent variables is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032331A_ABST
    Figure CN120032331A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode vehicle trajectory prediction method fusing vehicle motion and lane information, and the method comprises the steps: representing the vehicle trajectory and map lane original data in a sequence form, converting the data into low-dimensional and easy-to-process features through employing an MLP network, and carrying out the feature extraction through employing an LSTM network. In addition, a GRU-based graph neural network and an attention mechanism are utilized to extract spatial relation features between vehicles and global features of scene lanes. In order to further mine potential relevance between vehicle tracks and scene map features, the method is based on a CVAE model architecture, and personalized features and public features of vehicle motion and lane information are modeled by using a plurality of latent variables. According to the method, the latent variable sample generated through CVAE modeling is combined with two tasks of vehicle trajectory multi-mode prediction and surrounding lane sequence prediction, so that the precision and diversity of vehicle trajectory prediction results are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning and autonomous driving, and specifically relates to a multi-mode vehicle trajectory prediction method that integrates vehicle motion and lane information. Background Art

[0002] Vehicle trajectory prediction is crucial in the safety planning of autonomous driving; the upstream tasks of the autonomous driving system provide vehicle observation trajectories, which contain vehicle motion information (such as position, speed, and acceleration) and scene map data, including the geometric structure and topological relationship of lanes. At the same time, vehicle motion also includes the mutual influence between vehicles based on relative position relationships, and these features cannot be directly reflected from the original data. How to accurately model the complex correlation between vehicle motion and lane information has become a key issue that needs to be solved in the model during data utilization. In addition, the uncertainty in the traffic environment increases the complexity of trajectory prediction. By predicting multiple possible driving paths of other traffic participants, the system can better improve the safety of autonomous driving. Therefore, designing a method that efficiently utilizes vehicle motion and lane information, collaboratively models the relationship between these two types of multi-source data, and realizes multi-modal vehicle trajectory prediction has become a key task at present.

[0003] At present, the research in this field mainly integrates vehicle motion and lane information through deep learning technology to improve the accuracy and robustness of vehicle trajectory prediction. The mainstream method usually extracts vehicle trajectory, traffic participant behavior and features in high-precision maps by constructing a neural network model based on time series data (such as recurrent neural network, graph neural network, etc.) or a framework based on attention mechanism. In addition, some studies have also proposed strategies for multi-source data fusion and uncertainty modeling to better adapt to complex dynamic environments. For example, the Chinese patent application with publication number CN114889638A proposes a trajectory prediction method and system for an autonomous driving system. The method combines the high-precision map of the current frame and the vehicle trajectory planning information of the previous frame to generate and correct the trajectory prediction result of the current frame to ensure that the generated predicted trajectory is within the drivable area. The US patent application with publication number US12001958B2 proposes a method for predicting by combining high-precision maps and vehicle trajectory information. The method uses a recurrent neural network (RNN) and a confidence map generation method to predict future trajectories, and generates a complete participant trajectory through vector field tracking technology, thereby assisting self-driving vehicles in environmental navigation. However, these methods usually fail to fully explore the potential correlations between data when fusing multi-source data. In particular, when dealing with scene map changes caused by vehicle motion, they may face the problem of insufficient feature extraction, which in turn limits the improvement of the final prediction accuracy.

[0004] In order to solve the above problems and better cope with the diversity challenges in trajectory prediction, some studies have introduced generative model methods, including Variational Autoencoder (VAE) and Conditional Variational Autoencoder (CVAE). In terms of vehicle trajectory prediction, the literature [N.Lee, W.Choi, P.Vernaza, CBCchoy, PHSTorr and M.Chandraker, "DESIRE: Distant Future Prediction in Dynamic Scenes with Interacting Agents" 2017IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp.2165~2174] generates diverse future trajectories by modeling multi-modal latent variable distributions, and the literature [H.Xing, J.Hu and Z.Zhang, "Multi-modal Vehicle Trajectory Prediction via Attention-based Conditional Variational Autoencoder" 2022 37th Youth Academic Annual Conference of Chinese Association of Automation (YAC), pp.1367~1373] proposes a conditional variational autoencoder method based on the attention mechanism. However, these methods still have limitations in modeling the distribution of latent variables, especially in modeling the complex interaction between vehicles and scenes and the high-order dependency characteristics of multi-modal latent variables, which limits the diversity and accuracy of trajectory prediction.

[0005] Therefore, a deep learning-based model framework is designed for the multi-source input data of vehicle trajectories and scene maps. By potentially modeling the spatial-semantic relationship between vehicles and lanes, combined with the changes in surrounding lanes and the interaction between vehicles, diversified trajectories are generated, which can effectively solve the problem of multi-modal vehicle trajectory prediction. Summary of the invention

[0006] In view of the above, the present invention provides a multimodal vehicle trajectory prediction method that integrates vehicle motion and lane information. The method is based on the CVAE model architecture and uses multiple latent variables to model the personalized characteristics and common characteristics of vehicle motion and lane information respectively, effectively improving the accuracy and diversity of vehicle trajectory prediction results.

[0007] A multi-mode vehicle trajectory prediction method integrating vehicle motion and lane information comprises the following steps:

[0008] (1) Obtain scene map data and vehicle historical observation data;

[0009] (2) Preprocessing the data obtained in step (1) to obtain a vehicle observation data time series and a lane space tensor sequence;

[0010] (3) Construct a multi-modal vehicle trajectory prediction model that integrates vehicle motion and lane information, including:

[0011] The feature extraction network uses MLP (Multilayer Perceptron) and LSTM (Long Short-Term Memory) to extract the temporal motion features of each vehicle and the geometric and semantic features of each lane at t = 0 and t = F from the vehicle observation data time series and lane space tensor sequence;

[0012] Multi-vehicle spatial relationship feature encoding network, for any vehicle i, first calculate the relative position of vehicle i and other vehicles based on the position information of vehicle i and other vehicles at time t=0 in the vehicle observation data time series, then use MLP to calculate the relationship features of vehicle i and other vehicles based on the node features of all vehicles and the motion features of vehicle i, and then use GRU (Gated Recurrent Unit) to update the node features of all vehicles based on the relationship features (the node features of the vehicle are initialized by the motion features); finally, determine the surrounding vehicles adjacent to vehicle i based on the relative position, and accumulate the node features of these surrounding vehicles to obtain the multi-vehicle relationship features of vehicle i;

[0013] The lane global feature encoding network calculates the geometric and semantic features of all lanes at time t=0 through the attention mechanism to obtain the scene lane global features;

[0014] Latent variable encoding network, using MLP to calculate the vehicle personalized feature latent variable z according to the motion characteristics of each vehicle and the multi-vehicle relationship characteristics VP The prior and posterior distribution parameters of the lane are calculated by using MLP according to the geometric and semantic features of each lane at t = 0 and t = F as well as the global features of the scene lane. CP The prior and posterior distribution parameters of each vehicle are calculated by MLP and PoE (Product of Experts) based on the motion characteristics and multi-vehicle relationship characteristics of each vehicle, the geometric and semantic characteristics of each lane at t = 0 and t = F, and the global characteristics of the scene lane.S The prior and posterior distribution parameters of ; define the latent variable z according to these distribution parameters VP 、z CP 、z S The corresponding prior and posterior distributions are obtained, and then the prior and posterior distributions of the three latent variables are sampled K times respectively to obtain the corresponding sample matrix, where K is a natural number not less than 1;

[0015] Vehicle trajectory multi-mode prediction network, for any vehicle i, according to z VP and z S The corresponding sample matrix uses MLP to restore the sample matrix of the latent variable of the vehicle i trajectory feature, and then the sample matrix is ​​combined with the motion features and multi-vehicle relationship features of vehicle i, the vehicle observation data at time t=0, and the global features of the scene lane as inputs. LSTM uses autoregression to predict the position of vehicle i from time t=1 to time t=F as the prediction result;

[0016] The vehicle future surrounding lane sequence prediction network, for any lane j around any vehicle i at time t = F, according to z CP and z S The corresponding sample matrix uses MLP to restore the sample matrix of lane j’s feature latent variables, and then the sample matrix and the geometric and semantic features of lane j at time t=0 and the global features of the scene lane are used as input to calculate the geometric and semantic attribute sequence of lane j as the prediction result through the multi-head attention mechanism;

[0017] (4) using the vehicle observation data time series and lane space tensor sequence in step (2) to train the multi-modal vehicle trajectory prediction model;

[0018] (5) Use the trained multi-modal vehicle trajectory prediction model to predict the trajectory of the target vehicle.

[0019] Furthermore, the scene map data includes a sequence of geometric (length and direction angle) and semantic attributes (whether there are stop lines, traffic lights) of all lanes in the scene, and the vehicle historical observation data includes motion information (position, speed, acceleration, heading angle) of all vehicles in the scene over a period of time in the past.

[0020] Furthermore, in the step (2), the historical vehicle observation data is converted into a time series to obtain a vehicle observation data time series with a length of H+F, and the vehicle observation data time series is divided into two parts with time t=0 as the timeline: a historical vehicle observation data series and a future vehicle observation data series, wherein H is the length of the historical vehicle observation data series and F is the length of the future vehicle observation data series; for the scene map data, it is converted into a tensor to obtain two sets of lane space tensor sequences corresponding to time t=0 and time t=F, whose dimensions include the number of lanes, the number of spatial sampling points of the lanes, and the amount of information (position, direction angle, semantic attributes) of the spatial sampling points.

[0021] Furthermore, the following loss function is used in step (4): Train a multi-modal vehicle trajectory prediction model;

[0022]

[0023] in: and are the reconstruction losses for the vehicle trajectory prediction task and lane sequence prediction task, TC and DW are the total correlation and component dimension in the KL (Kullback-Leibler) divergence regularization term, respectively. C , KLD , β 1 , β 2 are all weight coefficients.

[0024] Furthermore, the reconstruction loss The expression is as follows:

[0025]

[0026] in: and They represent the kth prediction result and the true value of the position of the nth vehicle at time t, respectively. N is the number of vehicles in the scene, || || 2 represents the L2 norm.

[0027] Furthermore, the reconstruction loss The expression is as follows:

[0028]

[0029] in: and They represent the kth prediction result and true value of the lth spatial sampling point of the mth lane around the nth vehicle at time t=F regarding the geometric and semantic attribute sequence, M is the number of lanes around each vehicle, L is the number of spatial sampling points of the lane, || ||2 represents the L2 norm.

[0030] Furthermore, the total correlation TC and the component dimension DW are expressed as follows:

[0031] TC=KL[q(z VP ,z S )||q(z VP )q(z S )]+KL[q(z CP ,z S )||q(z CP )q(z S )]

[0032] DW=KL[q(z VP )||p(z VP )]+KL[q(z CP )‖p(z CP )]+2KL[q(z S )‖p(z S )]

[0033] Where: KL[A||B] represents the KL divergence of A and B, A and B are the variables of KL divergence, q(z CP ,z S ) represents z CP With z S The joint posterior probability density function, q(z VP ,z S ) represents z VP With z S The joint posterior probability density function, q(z CP ) and p(z CP ) represent z CP The posterior probability density function and the prior probability density function, q(z S ) and p(z S ) represent z S The posterior probability density function and the prior probability density function, q(z VP ) and p(z VP ) represent z VP The posterior probability density function and the prior probability density function of .

[0034] Furthermore, the weight coefficient λ KLD The expression is as follows:

[0035]

[0036] Where: KLD (E) Weight coefficient λ in the Eth iteration of model training KLD The value of λ0 and λ 1 They are KLD The initial and final values ​​(given values) of , κ is the factor that controls the rate of change of the curve (given value), E c is the center point of the Sigmoid curve (given value), and E is a natural number.

[0037] Furthermore, in step (5), the historical observation data of the target vehicle and the geometric and semantic attribute sequences of the lanes around the target vehicle are preprocessed and input into the trained multi-modal vehicle trajectory prediction model, so as to predict and output K groups of motion trajectories of the target vehicle in the future.

[0038] Based on the above technical solution, the present invention has the following beneficial technical effects:

[0039] 1. The present invention takes into account the dynamic changes of the scene, further utilizes the vehicle motion and lane information, and proposes a vehicle's future surrounding lane sequence prediction task, which effectively improves the prediction performance through multi-task design.

[0040] 2. The present invention models the personalized features and common feature latent variables in vehicle motion and lane information respectively, solving the problem that the temporal correlation between vehicle motion trajectory and lane changes and the spatial-semantic relationship between vehicle position and lane geometry and semantic attributes in the two types of inputs are not fully utilized. At the same time, the structure of the ELBO (Evidence Lower Bound) objective function of CVAE is improved, the expressive ability of latent variables is enhanced, and thus the prediction accuracy of the model in vehicle trajectory prediction tasks is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 Schematic diagram of a multi-mode vehicle trajectory prediction model that integrates vehicle motion and lane information in the present invention.

[0042] Figure 2 Schematic diagram of the serialized expression of vehicle trajectory and map lane.

[0043] Figure 3 Schematic diagram of the feature extraction network structure based on MLP and LSTM.

[0044] Figure 4 Schematic diagram of the multi-vehicle spatial relationship feature encoding network structure based on GRU.

[0045] Figure 5 Schematic diagram of the lane global feature encoding network structure based on the attention mechanism.

[0046] Figure 6 Schematic diagram of the latent variable encoding network structure with personalized features and public features.

[0047] Figure 7 Schematic diagram of the vehicle trajectory multi-modal prediction network structure based on vehicle trajectory feature latent variables.

[0048] Figure 8 Schematic diagram of the surrounding lane sequence prediction network structure based on scene map feature latent variables.

[0049] Fig. 9 Schematic diagram of the training and testing process of the multi-mode vehicle trajectory prediction model of the present invention.

[0050] Fig.10 Schematic diagram of the error comparison results between the model of the present invention and the baseline algorithm on a public data set.

[0051] Fig.11 Schematic diagram of the module ablation experiment results of the model of the present invention on a public dataset. DETAILED DESCRIPTION

[0052] In order to describe the present invention more specifically, the technical solution of the present invention is described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0053] like Figure 1 As shown, the deep learning vehicle trajectory prediction model of the present invention that integrates vehicle motion and lane information consists of a serialized representation of vehicle trajectory and map lane, a feature extraction network based on MLP and LSTM, a multi-vehicle spatial relationship feature encoding network based on GRU, a lane-to-lane global feature encoding network based on an attention mechanism, a latent variable encoding network with personalized features and public features, a vehicle trajectory multi-mode prediction network, and a surrounding lane sequence prediction network. The solid connecting lines in the figure represent calculations that exist in both the training and prediction stages, and the dotted connecting lines represent calculations that only exist in the training process.

[0054] like Figure 2 As shown in the figure, the scene information consists of vehicle trajectory and map information, where the map information selects the lane line data in the high-precision map; the observation data of the entire scene covers the time period from t = -H + 1 to t = F, where t = 0 divides the observation data into historical sequence and future sequence. represents the coordinates, speed and heading angle of vehicle n at time t in the traffic scene, where t = -H+1, -H+2, ..., 0, 1, ..., F, H and F represent the length of the historical sequence and the future sequence respectively, and D V is the dimension size of vehicle features. Represents the historical state of vehicle n in a traffic scene, with the matrix represents the future state of the vehicle, where the superscript T represents transposition, and the historical states of all vehicles are expressed as follows:

[0055]

[0056] The future states of all vehicles are represented as follows:

[0057]

[0058] Where: N is the number of vehicles in the scene; the coordinates of vehicle n in the scene at time t = 0 are is the state vector The part related to the coordinate dimension Here Indicates taking The first and second dimensions in the , that is, the values ​​of the dimensions corresponding to the horizontal and vertical coordinates of the vehicle.

[0059] In the traffic scene, the lane information is composed of L points uniformly sampled on its center line. The scene information around vehicle n at time t = 0 is recorded as:

[0060]

[0061] in: represents the sampling point sequence of the mth surrounding lane of vehicle n at time t=0, represents the lth sampling point of the mth surrounding lane of vehicle n at time t = 0. The semantic attribute part of the sampling point defines two independent binary attributes: one uses 0 and 1 to indicate whether the lane has a stop line, and the other uses 0 and 1 to indicate whether the lane has a traffic light. C is the dimension of lane sampling points; the scene information of all vehicles at time t = 0 is recorded as:

[0062]

[0063] Mask Vector Indicates whether the mth lane feature of vehicle n at time t = 0 is filled with all zero vectors. If the mth lane feature is not filled, then otherwise The scene information of vehicle n at time t = F is recorded as:

[0064]

[0065] Where: The mth lane sequence of vehicle n at time t = F is recorded as The mask vector is The scene information of all vehicles at time t = F is recorded as:

[0066]

[0067] like Figure 3As shown in Figure 1, different feature extraction networks are used for the two types of input: vehicle motion and lane information. The network used for feature extraction of vehicle historical trajectory is MLP. VE With LSTM H , subscript VE and H They represent the vehicle trajectory state vector encoding and vehicle history sequence encoding respectively. Taking vehicle n as an example to illustrate the vehicle history motion feature extraction process, the MLP VE The state with t=-H+1,-H+2,…,0, a total of H time steps As input, it consists of a fully connected layer and an activation function. The activation function uses ReLU (Rectified Linear Unit) and the output is of dimension D 1 Vehicle feature code

[0068]

[0069] Where: W (1) For MLP VE The weight matrix of Represents the feature encoding of the vehicle at H time steps, MLP VE The output of the network at each time step t Feed into LSTM H The encoder encodes:

[0070]

[0071] Where: t=-H+1,-H+2,…,0,W (2) For LSTM H The weight matrix of LSTM H The input feature dimension and hidden state dimension are both D 1 The hidden state of the encoder at t = 0 is As the output of the historical trajectory extraction module, the historical trajectory of vehicle n is encoded The historical trajectory encoding of all vehicles is expressed as:

[0072]

[0073] The network used for feature extraction of the vehicle's future trajectory is MLP VE With LSTM F , MLP VE The same network used in historical trajectory feature extraction, LSTM F Subscript F It is used to encode the future history sequence of the vehicle. Taking vehicle n as an example, the process of extracting the historical motion features of the vehicle is explained. VEThe state of the network for t=1,2,…,F, a total of F time steps Encode to get Represents the vehicle feature encoding of F time steps. MLP VE The output of the network at each time step t Feed into LSTM F The encoder encodes:

[0074]

[0075] Where: t=1,2,…,F,W (3) For LSTM F The weight matrix of LSTM F The input feature dimension and hidden state dimension are both D 1 The hidden state of the encoder at t = F As the output of the historical trajectory extraction module, the future trajectory of vehicle n is encoded

[0076] In the feature extraction of lane space tensor sequence at time t=0 and t=F, the networks used are all MLP CE With LSTM C , subscript CE and C They represent the lane sampling point encoding and lane sequence encoding respectively, and the encoding process is introduced by taking the mth lane at time t=0 and t=F as an example.

[0077] MLP CE by and As input, it consists of a fully connected layer and an activation function. The activation function uses ReLU, and the dimension is D 2 Lane feature encoding and

[0078]

[0079] Where: W (4) For MLP CE The weight matrix of the vehicle n at t = 0 and t = F is expressed as and

[0080]

[0081] At t=0 and t=F, use LSTM C right and Encode the features of each sampling position l and After LSTM C Get the hidden layer state and

[0082]

[0083] Where: W (5) For LSTM C The weight matrix of LSTM C The input feature dimension and hidden state dimension are both D 1 l = the hidden layer state at L sampling positions and As the output of lane sequence feature extraction, the lane feature codes at time t = 0 and t = F are respectively denoted as and All lane feature codes of vehicle n are expressed as:

[0084]

[0085] The lane feature encoding of all vehicles is expressed as:

[0086]

[0087] like Figure 4 As shown in the figure, the multi-vehicle spatial relationship feature encoding network is divided into node feature calculation and GRU-based node feature update. The node feature calculation process is introduced by taking vehicle n as an example.

[0088] The coordinates of the two vehicles n and j in the scene at time t = 0 are and If two coordinates satisfy the relationship There is a mutual influence between vehicles n and j, where ρ V is the maximum vehicle correlation distance selected based on prior knowledge, using δ nj To express the above relationship:

[0089]

[0090] The relationship between vehicle n and other vehicles in the scene is expressed as

[0091] Use the vehicle historical time series characteristic encoding X = (x 1 ,x 2 ,…,x N ) T As the initial feature matrix in vehicle impact calculation in Define the relative position between vehicles n and j Then the relative position of vehicle n and other vehicles in the scene is expressed as:

[0092]

[0093] The initial feature vector of vehicle n Copy N times and relative position and the initial feature matrix G (0) Splicing into MLP nbr Input The subscript nbr It is used to encode the spatial relationship between multiple vehicles, and cat() is a vector concatenation operation. According to this definition, the feature transformation matrix The calculation of is as follows:

[0094]

[0095] in: represents the Kronecker product, Indicates that the concatenated vector is copied N times along a certain dimension, W (6) It is MLP nbr The weight matrix of MLP nbr It consists of a fully connected layer and an activation function, using ReLU as the activation function. To simplify the subsequent calculation instructions, Expanding along the first dimension gives the following representation:

[0096]

[0097] in: is the feature transformation vector of vehicles n and j, indicating The vector of the nth row.

[0098] Will Perform addition-based aggregation to obtain feature aggregation vector

[0099]

[0100] The above node feature calculation is applied to all vehicles in the scene to obtain the feature aggregation matrix A (0) :

[0101]

[0102] Initial feature matrix G (0) and the feature aggregation matrix A (0) As input, GRU is used nbr For G (0) Update, the updated feature matrix is ​​G (1) :

[0103] G (1) =GRU nbr (A (0) ,G (0) ; W (7) )

[0104] Where: G (1) represents the feature matrix of all vehicles after the first update, W (7) It is GRU nbr The feature matrix after W updates is denoted as G (W) , and the initial feature matrix G (0) Similarly, it expands as follows:

[0105]

[0106] in: represents the feature vector of vehicle j after W updates.

[0107] Using G (W) Calculate the feature vector of the influence of neighboring vehicles on vehicle n in the scene, that is, the feature encoding of the multi-vehicle spatial relationship This process is obtained by weighted summation:

[0108]

[0109] The multi-vehicle spatial relationship feature encoding of all vehicles is recorded as

[0110] like Figure 5 The figure shows the lane-to-lane global feature encoding network based on the attention mechanism. The lane feature encoding of vehicle n at time t=0 is obtained through the attention mechanism to obtain the lane-to-lane global feature encoding. First, the lane feature encoding M at time t=0 is used n,0 Feature encoding of the two lanes u and v in and Calculate the attention coefficient

[0111] Then, κ uv The normalized attention weight α is obtained by Softmax calculation uv :

[0112]

[0113] Finally, using the attention weight α uv The feature encodings of all lanes are weighted summed, and the mask information r is considered n,0 , and get the lane-to-lane global feature encoding

[0114]

[0115] The global feature encoding between lanes around all vehicles is

[0116] like Figure 6 The figure shows a latent variable encoding network with personalized features and common features, including a posterior network and a priori network. The posterior network is mainly used to approximate the conditional posterior distribution, including the vehicle personalized latent variable z VP The conditional posterior probability density function Map personalization latent variable z CP The conditional posterior probability density function and the common latent variable z S The conditional posterior probability density function These posterior networks parameterize the mean and variance of the conditional posterior distribution via a neural network and assume that the latent variable follows a normal distribution:

[0117]

[0118] Where: (μ VP ,σ VP )、(μ CP ,σ CP ) and (μ S ,σ S ) are z VP 、z CP and z S The mean and variance of the conditional posterior distribution.

[0119] The prior network is mainly used to approximate the conditional prior distribution, including the vehicle personalized latent variable z VP The conditional prior probability density function Map personalization latent variable z CP The conditional prior probability density function and the common latent variable z S The conditional prior probability density function These prior networks parameterize the mean and variance of the conditional prior distribution through a neural network and assume that the latent variable follows a normal distribution:

[0120]

[0121] Where: (μ′ VP ,σ′ VP ), (μ′ CP ,σ′ CP ) and (μ′ S ,σ S ′) are z VP 、z CP and z SThe mean and variance of the conditional prior distribution.

[0122] The parameters of the above normal distribution are calculated by different MLPs; in the implementation of the posterior network, the vehicle input personalized latent variable z VP The parameters of the conditional posterior distribution of are calculated as follows:

[0123] (μ VP ,σ VP )=MLP VP (cat(x n ,y n ,u n );W (8) )

[0124] Where: W (8) It is MLP VP The weight matrix of VP Represents the calculation of the posterior network parameters for vehicle personalized latent variables, MLP VP It consists of a fully connected layer and an activation function, and ReLU is used as the activation function.

[0125] Considering the lane mask, the lane feature pooling at t=0 and t=F is calculated as follows:

[0126]

[0127] in: and They are the lane pooling features around vehicle n at t=0 and t=F respectively.

[0128] In the implementation of the posterior network, the map input personalizes the latent variable z CP The parameters of the conditional posterior distribution of are calculated as follows:

[0129] (μ CP ,σ CP )=MLP CP (cat(v n,0 ,v n,F ,j n,0 );W (9) )

[0130] Where: W (9) It is MLP CP The weight matrix of CP Represents the calculation of the posterior network parameters for vehicle personalized latent variables, MLP CP It consists of a fully connected layer and an activation function, and ReLU is used as the activation function.

[0131] In the implementation of the posterior network, the two types of input common latent variables z SThe calculation of the parameters of the conditional posterior distribution of is divided into two steps:

[0132] 1. Calculate z using vehicle input and map input respectively S The posterior normal distribution parameters of are:

[0133] Assume z S In the conditions and The posterior distribution is a normal distribution, and its probability density function is Compute the parameters of the distribution from the vehicle input features:

[0134] (μ S1 ,σ S1 )=MLP S1 (cat(x n ,y n ,u n );W (10) )

[0135] Where: (μ S1 ,σ S1 ) is z S In the conditions and The mean and variance of the posterior distribution, W (10) It is MLP S1 The weight matrix of S1 Indicates for vehicle input and Calculation of the parameters of the lower posterior distribution, MLP S1 It consists of a fully connected layer and an activation function, and ReLU is used as the activation function.

[0136] Assume z S In the conditions and The posterior distribution is a normal distribution, and its probability density function is Compute the parameters of a distribution from map input features:

[0137] (μ S2 ,σ S2 )=MLP S2 (cat(v n,0 ,v n,F ,j n,0 );W (11) )

[0138] Where: (μ S2 ,σ S2 ) is z S exist and The mean and variance of the conditional posterior distribution, W (11) It is MLPS2 The weight matrix of S2 Indicates that it is used for map input and Calculation of the parameters of the lower posterior distribution, MLP S2 It consists of a fully connected layer and an activation function, and ReLU is used as the activation function.

[0139] 2. Update z using PoE-based method S The parameters of the posterior normal distribution are:

[0140] z S In the conditions and The posterior distribution is assumed to be normal, and its probability density function is Referring to the method in the paper [Y.Cao and JFDavid, "Generalized Product ofExperts for Automatic and Principled Fusion of Gaussian Process Predictions"ArXiv,vol.abs / 1410.7827,2014], the PoE-based calculation is used to obtain the estimation of the new posterior distribution parameters from the parameters of the above two types of conditional distributions:

[0141]

[0142] Where: I is the unit vector of the corresponding dimension.

[0143] In the implementation of the prior network, the vehicle input personalized latent variable z VP The conditional prior probability density function The parameters of are calculated as follows:

[0144] (μ′ VP ,σ′ VP )=MLP VPri (cat(x n ,u n );W (12) )

[0145] Where: W (12) It is MLP VPri The weight matrix of VPri Represents the calculation of the prior network parameters for vehicle personalized latent variables, MLP VPri It consists of a fully connected layer and an activation function, and ReLU is used as the activation function.

[0146] In the implementation of the prior network, the map input personalizes the latent variable z CP The conditional prior probability density function The parameters of are calculated as follows:

[0147] (μ′ CP ,σ′ CP )=MLP CPri (cat(v n,0 ,j n,0 );W (13) )

[0148] Where: W (13) It is MLP CPri The weight matrix of CPri Represents the calculation of prior network parameters for personalized latent variables used for map input, MLP CPri It consists of a fully connected layer and an activation function, and ReLU is used as the activation function.

[0149] In the implementation of the prior network, the latent variable z of the common features of the two types of inputs S The calculation of the parameters of the conditional prior distribution is divided into two steps:

[0150] 1. Calculate z using vehicle input and map input respectively S The prior normal distribution parameters of :

[0151] Assume z S In the conditions The prior distribution is normal distribution, and its probability density function is Compute the parameters of the distribution from the vehicle input features:

[0152] (μ′ S1 ,σ S ' 1 )=MLP S1Pri (cat(x n ,u n );W (14) )

[0153] Where: (μ′ S1 ,σ S ' 1 ) is z S In the conditions The mean and variance of the prior distribution, W (14) It is MLP S1Pri The weight matrix of S1Pri Indicates for vehicle input Calculation of prior distribution parameters, MLP S1Pri It consists of a fully connected layer and an activation function, and ReLU is used as the activation function.

[0154] Assume z S In the conditions The prior distribution is normal distribution, and its probability density function is Compute the parameters of a distribution from map input features:

[0155] (μ′ S2 ,σ S ' 2 )=MLP S2Pri (cat(v n,0 ,j n,0 );W (15) )

[0156] Where: (μ′ S2 ,σ S ' 2 ) is z S In the conditions The mean and variance of the prior distribution, W (15) It is MLP S2Pri The weight matrix of S2Pri Represents map input Calculation of prior distribution parameters, MLP S2Pri It consists of a fully connected layer and an activation function, and ReLU is used as the activation function.

[0157] 2. Update z using PoE-based method S The prior normal distribution parameters of :

[0158] z S In the conditions and The prior distribution is assumed to be normal distribution, and its probability density function is Use PoE-based calculations to get estimates of the prior distribution parameters:

[0159]

[0160] Where: I is the unit vector of the corresponding dimension.

[0161] The parameters obtained by the posterior network are reparameterized to obtain the corresponding distribution, from which K latent variable samples are generated. Taking vehicle n as an example, they are recorded as vehicle personalized posterior latent variable samples Sample of posterior latent variables for personalized surrounding maps and the common posterior latent variable sample Where D 3 Represents the dimension of the latent variable related to the vehicle's personalized features, D 4 Represents the dimension of the latent variable related to the personalized features of the scene map, D 5 The dimension of latent variables representing the common features between vehicles and maps. Sample of vehicle-specific posterior latent variables for all vehicles Sample of posterior latent variables for personalized surrounding maps and the common posterior latent variable The definition is as follows:

[0162]

[0163]

[0164] The parameters generated by the prior network are used to define the corresponding distribution, from which K latent variable samples are sampled. Taking vehicle n as an example, they are recorded as vehicle personalized prior latent variable samples Sample of prior latent variables for personalization of surrounding maps and a sample of common prior latent variables Sample of vehicle personalization prior latent variables for all vehicles Sample of prior latent variables for personalization of surrounding maps and common prior latent variables The definition is as follows:

[0165]

[0166] The training and prediction stages use the posterior latent variable samples and the prior latent variable samples as the output of the latent variable encoding network with personalized features and public features:

[0167]

[0168] in: is a sample of vehicle personalized latent variables, is a sample of map personalized latent variables, is a sample of common latent variables.

[0169] like Figure 7 As shown in Figure 2, the vehicle trajectory multi-modal prediction network based on vehicle trajectory feature latent variables includes two parts: latent variable recovery and trajectory prediction. The latent variable recovery network is implemented through MLP. V Implement latent variable samples, taking vehicle n as an example, and Fusion generates latent variable samples Used to more fully characterize vehicle trajectory characteristics:

[0170]

[0171] Where: W (16) It is MLP V The weight matrix of V Represents the latent variable recovery of vehicle trajectory features, MLP V It consists of a fully connected layer and an activation function, using ReLU as the activation function. To simplify the subsequent calculation instructions, Expanding along the first dimension gives the following representation:

[0172]

[0173] in: yes The k-th row vector of is used to obtain the k-th prediction result.

[0174] Other features used in prediction include: bicycle historical time series feature encoding x n , neighboring vehicles influence the characteristic vector u n and the global feature j of the scene map of vehicle n at time t = 0 n,0 In addition, from the vehicle history status Extract the state of vehicle n at time t = 0 Used to initialize the hidden state of LSTM.

[0175] Next, we take the prediction process of the kth trajectory of the nth vehicle as an example to describe the prediction process of the model in detail. The process includes the processing of input features, the state update of LSTM, and the final state prediction. The specific steps are as follows:

[0176] x n 、u n and j n,0 Splice to get the spliced ​​feature vector The vehicle state at t = 0 Through an MLP D Encode it to get the initial vehicle state representation vector Used to initialize LSTM D Input, MLP D and LSTM D Subscript D Represents the decoding for concatenating feature matrices. MLP D It consists of a fully connected layer and an activation function, using ReLU as the activation function. In each subsequent time step t=1,2,…,F, MLP is used D The prediction result of the kth trajectory at the previous moment Process and get the vehicle status code

[0177] The above process generates b n and And the last moment LSTM D The hidden state Input to the decoder module LSTM D , to update the hidden state at the next moment

[0178]

[0179] Where: W (17) It is LSTM D The weight matrix of .

[0180] Through MLP P The updated status is shown Convert to the predicted vehicle state at the next moment

[0181]

[0182] Where: W (18) It is MLP P The weight matrix of P Represents MLP for vehicle state prediction P It consists of a fully connected layer and an activation function, using ReLU as the activation function. Each predicted state of time step t=1,2,…,F is stacked along the time dimension to obtain the kth predicted trajectory of the nth vehicle

[0183]

[0184] The K prediction results of the nth vehicle are stacked according to the number of predicted trajectories to obtain the predicted trajectory of the nth vehicle:

[0185]

[0186] The prediction results for all vehicles constitute the final output:

[0187]

[0188] like Figure 8 As shown, the lane sequence prediction network includes the latent variable sample Z of lane characteristics C Restoration and lane sequence prediction. The lane feature latent variable restoration network consists of MLP C To achieve this, take the lanes around vehicle n as an example and convert the latent variable samples and Combine to generate latent variable samples

[0189]

[0190] Where: W (19) It is MLP C The weight matrix of C Represents the latent variable recovery for scene map features, MLP CIt consists of a fully connected layer and an activation function, and ReLU is used as the activation function.

[0191] Map feature j n,0 After being copied K times, the linear layer generates the query matrix for multi-head attention calculation. Encode the characteristics of the mth lane around vehicle n at time t=0 Copy K times and pass through two different linear layers to generate the key matrix for multi-head attention calculation Sum Matrix This process can be expressed as follows:

[0192]

[0193] Where: Linear Q 、Linear K and Linear V They represent the linear layers used to generate the query matrix, key matrix, and value matrix, respectively, and the weight matrices are W (20) , W (21) and W (22) ; These linear layers have the same structure, and the results generated by multiplying the input matrix with the learnable weight matrix are used in the subsequent multi-head attention mechanism.

[0194] Use the generated query vector, key vector, and value vector as the input for the Multihead Attention (MHA) calculation to get K prediction results for lane m at time t = F

[0195]

[0196] Where: W (23) is the weight matrix of MHA.

[0197] The prediction results of all M lanes constitute the lane sequence prediction results around vehicle n:

[0198]

[0199] The prediction results of surrounding lane sequences for all vehicles for:

[0200]

[0201] like Fig. 9 The figure shows the training and testing process of the network model. The training process obtains serialized representation data of vehicle trajectories and map lanes from the public dataset. and The encoding networks of personalized features and shared feature latent variables generate posterior probability density functions respectively and With the prior probability density function and And get the vehicle trajectory prediction result Lane sequence prediction results The objective function is calculated based on the results obtained from all training data during the training process, and one round of training is completed, and a total of E rounds of training are completed; the objective function of the multi-task is as follows:

[0202]

[0203] in: and is the reconstruction loss for the vehicle trajectory prediction task and lane sequence prediction task, TC and DW are the total correlation in the regularization term of KL divergence and the KL divergence of the component dimension, respectively; λ C yes The weight, λ KLD is the weight based on the KL divergence regularization term, β 1 and β 2 are the weights of TC and DW respectively.

[0204] In the objective function, The mean square error (MSE) is used and the specific calculation is as follows:

[0205]

[0206] In the formula: |||| 2 represents the L2 norm, and They represent the kth prediction result and true value of vehicle n at time t respectively.

[0207] In the objective function, The specific calculation is as follows:

[0208]

[0209] Where: and They represent the kth prediction result and the true value of the lth spatial sampling point in the mth lane around vehicle n at time t = F, and the smallest MSE is selected from the K prediction results as The final loss value.

[0210] In the objective function, TC is calculated as follows:

[0211]

[0212] In the objective function, DW is calculated as follows:

[0213]

[0214] During the training process, the λ in the objective function KLD Use the following Sigmoid-based function implementation:

[0215]

[0216] Where: 0 and λ 1 They are λ KLD The initial and final values ​​of the curve, κ is the variable that controls the rate of change of the curve, E c It is the center point of the Sigmoid curve.

[0217] The test process obtains serialized representation data of vehicle trajectories and map lanes from the public dataset and The vehicle trajectory prediction result is obtained based on the vehicle trajectory prediction model of multi-source information and dynamic interaction modeling.

[0218] The method of the present invention uses a variant of two commonly used performance indicators, average displacement error (Average Displacement Error, ADE) and final displacement error (FDE) ADE K and FDE K The experimental results were evaluated, including ADE K and FDE K The calculation of is as follows:

[0219]

[0220] Where: and They represent the kth prediction result and true value of vehicle n at time t, respectively. and They represent the kth prediction result and true value of vehicle n at time t=F respectively.

[0221] The data used in this invention comes from the nuScenes dataset provided by the literature [H.Caesar, V.Bankiti, A.H.Lang, S.Vora, V.E.Long, Q.Xu, A.Krishnan, Y.Pan, G.Baldan, O.Beijbom, "nuScenes: A Multimodal Dataset for Autonomous Driving" Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020]. The standard division of training, validation, and test data in the nuScenes benchmark is followed in the experiment. The experimental platform and environment are shown in Table 1. The various parameters in the experimental process refer to the literature [DP Kingma, J.Ba, "Adam: A Method for Stochastic Optimization" 3rd International Conference on Learning Representations (ICLR), May 7-9, 2015]. The detailed settings are shown in Table 2.

[0222] Table 1

[0223] Experimental environment Parameter information operating system Ubuntu 22.04.4LTS processor Intel(R)Xeon(R)Gold 5218CPU@2.30GHz Graphics GeForce RTX 3080Ti Memory 128G Deep Learning Frameworks Pytorch 1.13.0

[0224] Table 2

[0225]

[0226]

[0227] The method of the present invention is compared with four baseline algorithms: comparison scheme 1 is constant speed and heading angle; comparison scheme 2 is the Physics Oracle in the literature [T. Phan-Minh, E. C. Grigore, F. A. Boulton, O. Beijbom and E. M. Wolff, "CoverNet: Multimodal Behavior Prediction Using Trajectory Sets" 2020 IEEE / CV F Conference on Computer Vision and Pattern Recognition (CVPR), pp. 14062-14071]; comparison scheme 3 is the literature [H. Cui, V. Radosavljevic, F. -C. Chou, T. -H. Lin, T. Nguyen, T. -K. Huang, J. Schneider, N. Djuric, "Multimodal Trajectory Predictions for Autonomous Driving Using Deep Convolutional Networks" 2019 International Conference on Robotics and Automation (ICRA), pp.2090-2096, 2019]; Comparison scheme 4 is MultiPath in the literature [Y.Chai, B.Sapp, M.Bansal, D.Anguelov, "MultiPath: Multiple Probabilistic Anchor Trajectory Hypotheses for Behavior Prediction" Proceedings of the Conference on Robot Learning, vol.100, pp.86-99, 2020]. Constant speed and heading is a physics-based model that calculates future trajectories while maintaining constant speed and heading of the vehicle's current state. Based on the vehicle's current state (speed, acceleration, and heading), Physics Oracle calculates the minimum average point-by-point Euclidean distance of the predicted trajectories generated by the following four models: (i) constant speed and heading, (ii) constant speed and heading rate, (iii) constant acceleration and heading, and (iv) constant acceleration and heading rate. MTP processes a rasterized representation of the scene and the target vehicle state using a CNN to generate a fixed number of trajectories (patterns) and their associated probabilities.The model uses a weighted sum of regression loss and classification loss during training. The Multipath model uses fixed anchor points obtained from the training set to represent the pattern and outputs the residual relative to the anchor point in its regression branch.

[0228] Fig.10 The performance evaluation results are given when K is 1, 5 and 10. When K = 1, the present invention performs well in ADE 1 Indicators and FDE 1 The prediction accuracy of the indicator is better than all other schemes, indicating that the model can more accurately capture the key behavior patterns in the vehicle trajectory distribution, such as common driving paths and typical turning characteristics. 5 The prediction accuracy of the indicators is comparable to that of the best performing comparison scheme 4, and is better than the other three comparison schemes. This shows that when the model generates multiple predicted trajectories, it can more effectively capture possible trajectories close to the true trajectory, reflecting the ability to accurately model key behavior patterns. 5 The prediction accuracy of the proposed method is better than all other methods, which further shows that the model has a significant advantage in the overall accuracy of multimodal trajectory prediction and can more comprehensively fit the potential driving mode and trajectory distribution of the vehicle. When K = 10, the experimental results are similar to those when K = 5. The proposed method has a better prediction accuracy than all other methods in FDE. 10 The performance of the indicators is still close to that of the comparison scheme 4, and is better than other comparison schemes, indicating that the model can still maintain a good modeling ability of the real trajectory distribution when generating more possible trajectories, which reflects the adaptability of the model to multimodal distribution under high number predictions. 10 In terms of indicators, the present invention is also superior to all other methods, which further proves that the model not only has advantages in generating a small number of trajectories, but also can better balance the accuracy and diversity of trajectories when generating more diverse trajectories.

[0229] Fig.11 The module ablation experiment results of the present invention on the public dataset are given. The specific experimental comparison schemes are as follows: Experimental group 1 does not use scene map information and only uses vehicle trajectory as input; Experimental group 2 uses scene map information as input, but removes the lane sequence prediction task and only retains the vehicle trajectory prediction task; Experimental group 3 uses a set of MLPs to directly obtain the common feature latent variable Z from the concatenated features. S The mean and variance of the corresponding distribution are used to replace the PoE in the latent variable calculation. Comparing the present invention with Experimental Group 1, it can be seen that adding map information allows the model to perform better in ADE. K and FDE K There are obvious improvements in both indicators. Comparing the improvement ratios of the two indicators, we can also see that map information has a significant impact on FDE. KComparing the present invention with experimental group 2, it can be seen that the addition of the lane sequence prediction task improves the prediction accuracy of the algorithm, verifying the importance of the lane sequence prediction task. Comparing the present invention with experimental group 3, it can be seen that the method using PoE has a greater impact on the ADE 1 and FDE 1 The experimental results show that adding map information and lane sequence prediction tasks can improve model performance, among which map information contributes more to performance improvement; in addition, the use of the PoE method further enhances the model performance and has advantages in two indicators.

[0230] The above description of the embodiments is to facilitate the understanding and application of the present invention by those skilled in the art. It is obvious that those skilled in the art can easily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative work. Therefore, the present invention is not limited to the above embodiments. Improvements and modifications made by those skilled in the art to the present invention based on the disclosure of the present invention should be within the protection scope of the present invention.

Claims

1. A multi-mode vehicle trajectory prediction method integrating vehicle motion and lane information, comprising the following steps: (1) Obtain scene map data and vehicle historical observation data; (2) Preprocessing the data obtained in step (1) to obtain a vehicle observation data time series and a lane space tensor sequence; (3) Construct a multi-modal vehicle trajectory prediction model that integrates vehicle motion and lane information, including: The feature extraction network uses MLP and LSTM to extract the temporal motion features of each vehicle and the geometric and semantic features of each lane at t = 0 and t = F from the vehicle observation data time series and lane spatial tensor sequence; Multi-vehicle spatial relationship feature encoding network, for any vehicle i, first calculate the relative position of vehicle i and other vehicles based on the position information of vehicle i and other vehicles at time t=0 in the vehicle observation data time series, then use MLP to calculate the relationship features of vehicle i and other vehicles based on the node features of all vehicles and the motion features of vehicle i, and then use GRU to update the node features of all vehicles based on the relationship features; finally, determine the surrounding vehicles adjacent to vehicle i based on the relative position, and accumulate the node features of these surrounding vehicles to obtain the multi-vehicle relationship features of vehicle i; The lane global feature encoding network calculates the geometric and semantic features of all lanes at time t=0 through the attention mechanism to obtain the global features of the scene lanes; Latent variable encoding network, using MLP to calculate the vehicle personalized feature latent variable z according to the motion characteristics of each vehicle and the multi-vehicle relationship characteristics VP The prior and posterior distribution parameters of the lane are calculated by using MLP according to the geometric and semantic features of each lane at t = 0 and t = F as well as the global features of the scene lane. CP The prior and posterior distribution parameters are calculated by MLP and PoE based on the motion characteristics and multi-vehicle relationship characteristics of each vehicle, the geometric and semantic characteristics of each lane at t = 0 and t = F, and the global characteristics of the scene lane. S The prior and posterior distribution parameters of ; define the latent variable z according to these distribution parameters VP 、z CP 、z S The corresponding prior and posterior distributions are obtained, and then the prior and posterior distributions of the three latent variables are sampled K times respectively to obtain the corresponding sample matrix, where K is a natural number not less than 1; Vehicle trajectory multi-mode prediction network, for any vehicle i, according to z VP and z S The corresponding sample matrix uses MLP to restore the sample matrix of the latent variable of the vehicle i trajectory feature, and then the sample matrix is ​​combined with the motion features and multi-vehicle relationship features of vehicle i, the vehicle observation data at time t=0, and the global features of the scene lane as inputs. LSTM uses autoregression to predict the position of vehicle i from time t=1 to time t=F as the prediction result; The vehicle future surrounding lane sequence prediction network, for any lane j around any vehicle i at time t = F, according to z CP and z S The corresponding sample matrix uses MLP to restore the sample matrix of lane j’s feature latent variables, and then the sample matrix and the geometric and semantic features of lane j at time t=0 and the global features of the scene lane are used as input to calculate the geometric and semantic attribute sequence of lane j as the prediction result through the multi-head attention mechanism; (4) using the vehicle observation data time series and lane space tensor sequence in step (2) to train the multi-modal vehicle trajectory prediction model; (5) Use the trained multi-modal vehicle trajectory prediction model to predict the trajectory of the target vehicle.

2. The multi-mode vehicle trajectory prediction method according to claim 1, characterized in that: The scene map data includes geometric and semantic attribute sequences of all lanes in the scene, and the vehicle historical observation data includes movement information of all vehicles in the scene over a period of time in the past.

3. The multi-mode vehicle trajectory prediction method according to claim 1, characterized in that: In the step (2), the historical vehicle observation data is converted into a time series to obtain a vehicle observation data time series with a length of H+F, and the vehicle observation data time series is divided into two parts with time t=0 as the timeline: a historical vehicle observation data series and a future vehicle observation data series, wherein H is the length of the historical vehicle observation data series, and F is the length of the future vehicle observation data series; For the scene map data, it is converted into a tensor form to obtain two sets of lane space tensor sequences corresponding to time t=0 and time t=F, whose dimensions include the number of lanes, the number of spatial sampling points of the lanes, and the amount of information of the spatial sampling points.

4. The multi-mode vehicle trajectory prediction method according to claim 1, characterized in that: The following loss function is used in step (4): Train a multi-modal vehicle trajectory prediction model; in: and are the reconstruction losses for the vehicle trajectory prediction task and lane sequence prediction task, TC and DW are the total correlation and component dimension in the KL divergence regularization term, respectively. C , KLD , β1, and β2 are all weight coefficients.

5. The multi-mode vehicle trajectory prediction method according to claim 4, characterized in that: The reconstruction loss The expression is as follows: in: and They represent the kth prediction result and the true value of the position of the nth vehicle at time t, respectively. N is the number of vehicles in the scene, and || ||2 represents the L2 norm.

6. The multi-mode vehicle trajectory prediction method according to claim 4, characterized in that: The reconstruction loss The expression is as follows: in: and They respectively represent the kth prediction result and the true value of the lth spatial sampling point of the mth lane around the nth vehicle at time t=F regarding the geometric and semantic attribute sequence, M is the number of lanes around each vehicle, L is the number of spatial sampling points of the lanes, and || ||2 represents the L2 norm.

7. The multi-mode vehicle trajectory prediction method according to claim 4, characterized in that: The expressions of the total correlation TC and the component dimension DW are as follows: TC=KL[q(z VP ,z S )||q(z VP )q(z S )]+KL[q(z CP ,z S )||q(z CP )q(z S )] DW=KL[q(z VP )||p(z VP )]+KL[q(z CP )||p(z CP )]+2KL[q(z S )||p(z S )] Where: KL[A||B] represents the KL divergence of A and B, A and B are the variables of KL divergence, q(z CP ,z S ) represents z CP With z S The joint posterior probability density function, q(z VP ,z S ) represents z VP With z S The joint posterior probability density function, q(z CP ) and p(z CP ) represent z CP The posterior probability density function and the prior probability density function, q(z S ) and p(z S ) represent z S The posterior probability density function and the prior probability density function, q(z VP ) and p(z VP ) represent z VP The posterior probability density function and the prior probability density function of .

8. The multi-mode vehicle trajectory prediction method according to claim 4, characterized in that: The weight coefficient λ KLD The expression is as follows: Where: KLD (E) Weight coefficient λ in the Eth iteration of model training KLD The values ​​of λ0 and λ1 are respectively KLD The initial and final values ​​of the curve, κ is the factor that controls the rate of change of the curve, E c is the center point of the Sigmoid curve, and E is a natural number.

9. The multi-mode vehicle trajectory prediction method according to claim 1, characterized in that: In the step (5), the historical observation data of the target vehicle and the geometric and semantic attribute sequences of the lanes around the target vehicle are pre-processed and input into the trained multi-modal vehicle trajectory prediction model, so as to predict and output K groups of motion trajectories of the target vehicle in the future.

Citation Information

Patent Citations

  • Trajectory prediction method and system in automatic driving system

    CN114889638A

  • Future trajectory predictions in multi-actor environments for autonomous machine

    US12001958B2