Vehicle Driving Trajectory Prediction System and Method Based on Improved Graph Neural Network Model

By improving the graph neural network model and combining FPGCN and SocialLSTM models to process vehicle trajectory data, the problem of low processing efficiency of complex graph structure data in the prior art is solved, and more efficient and accurate vehicle trajectory prediction is achieved.

CN119580217BActive Publication Date: 2025-06-03CHANGCHUN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510143252.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-03
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The existing vehicle driving trajectory prediction technology is inefficient when processing complex graph structure data, making it difficult to effectively process complex graph structure data, resulting in large calculation overhead and poor detection accuracy.

Method used

The improved graph neural network model is adopted, combined with the FPGCN model and SocialLSTM model, and the historical trajectory diagram and current state diagram are processed through graph convolution layer, batch standardization layer, Dropout layer, differentiable pooling layer and feature fusion module to generate predicted trajectories.

Benefits of technology

It significantly improves the processing speed and accuracy of driving information characteristics, increases the model calculation efficiency, and enables the vehicle to make timely and correctly judges during driving, ensuring the personal safety of the driver.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580217B_ABST
    Figure CN119580217B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of traffic control systems, and relates to a vehicle driving trajectory prediction system and method based on an improved graph neural network model. The system includes an FPGCN model and a Social LSTM model; the FPGCN model includes a graph convolutional layer, a Dropout layer, a differentiable pooling layer, a feature fusion module, and a fully connected layer; graph-structured data obtains node feature vectors through the graph convolutional layer, the Dropout layer, and the differentiable pooling layer, and then the node features are processed by divide and conquer. Global features are generated through a dense attention mechanism and principal component analysis, and local features are generated through virtual padding and matrix calculation; after obtaining the global features and local features, they are fused using the feature fusion module, and then input into a fully connected layer to calculate the influence score of the historical trajectory graph on the current state graph; the preprocessed driving data and the data after the FPGCN model calculates the score are integrated and input into the Social LSTM model; while maintaining the model's expressive ability, the system significantly reduces the number of parameters and optimizes the calculation logic, increasing the prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of traffic control systems and relates to intelligent driving; specifically, it relates to a vehicle driving trajectory prediction system and method based on an improved graph neural network model. Background Art

[0002] With the progress of science and technology and the gradual improvement of traffic facilities, intelligent connected vehicles have started to develop rapidly, and now intelligent driving has become an important part of them. Through intelligent driving, vehicles can not only collect information around the vehicle for interaction, but also perform assisted driving, detect road conditions, and avoid dangers. When using intelligent driving, the safety of the driver must be ensured first, and efficient prediction of the vehicle driving trajectory is indispensable.

[0003] In the early vehicle trajectory prediction technology, methods based on the physical motion level were widely used, and trajectory prediction was carried out by extending the state of the current vehicle forward. This prediction method only has good effects within a relatively short time interval, but without considering the information of the surrounding vehicles and the state of the vehicle itself, it will have a negative impact on long-term prediction. Subsequently, the emerging machine learning technology provided various time series network models for trajectory prediction, such as the Recurrent Neural Network (RNN) or the Long Short-Term Memory (LSTM), which predict the driving trajectory by directly processing time series data; however, in actual situations, there are great problems in the time dependence of RNN and LSTM. If the time step is large, RNN will have gradient problems.

[0004] Trajectory prediction technology not only needs to consider spatio-temporal interaction information, but also needs to infer future road conditions through screening. Therefore, the interaction information processing model has become the key part to be improved in trajectory prediction technology. In the development process of previous trajectory prediction technologies, multiple papers have proved that using complex graph structure data as model input and information interaction is more interpretable for trajectory prediction, and it is easier to calculate feature relationships compared with other types of data to achieve the prediction of future trajectories. The experimental results are also more accurate compared with single-time series models. And Kawasaki et al. [Kawasaki A, Seki A. Multimodal trajectory predictions for urban environments using geometric relationships between a vehicle and lanes[C] / / 2020 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2020:9203-9209.] also combined lane information with KF filtering for trajectory prediction. However, when dealing with graph structure data, a suitable method needs to be found to handle the complex graph structure data. How to handle these complex relationships is a challenge. Conventional convolutional neural networks may collect redundant information when processing road vehicle data, which affects the model load, increases the computational overhead and may affect the detection accuracy. Summary of the Invention

[0005] In order to address the problems of high data complexity and low efficiency in current vehicle driving trajectory prediction technology, the present invention aims to provide a vehicle driving trajectory prediction system based on an improved graph neural network model. This system can not only efficiently process complex graph structure data transformed from vehicle information, but also improves the calculation strategy to be closer to some behaviors during vehicle driving. By combining the advantages of the deep feature divide-and-conquer structure and the attention mechanism, the processing speed and accuracy of driving information features are significantly improved, and at the same time, the model calculation efficiency is increased, enabling the vehicle to make correct judgments in a timely manner during driving and ensuring the personal safety of the driver.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A vehicle driving trajectory prediction system based on an improved graph neural network model. The prediction system is obtained by training an improved graph neural network model, and the improved graph neural network model includes an FPGCN model and a SocialLSTM model. Among them, the data input into the FPGCN model are two constructed graph-structured data, namely, a historical trajectory graph and a current state graph. The FPGCN model is used to calculate the influence score of the historical trajectory graph on the current state graph. After the preprocessed driving data is integrated with the data with the score calculated by the FPGCN model, it is input into the Social LSTM model, and the predicted trajectory is output after passing through a social pooling network.

[0008] Among them, the FPGCN model includes a graph convolutional layer, a batch normalization layer, a Dropout layer, a differentiable pooling layer, a feature fusion module, and a fully connected layer. The graph convolutional layer has three layers, and the scales of the three graph convolutional layers are 64, 32, and 16. After the graph-structured data passes through the graph convolution operation, a non-linear factor is introduced through an activation function, and then the input of each layer is normalized through the batch normalization layer. The constructed graph-structured data is aggregated through the three graph convolutional layers described above, so that each node considers not only its own feature information but also the relationship with neighbor nodes, and then it is input into the Dropout layer and the differentiable pooling layer. The Dropout layer randomly selects some nodes and forces them to be set to 0. The differentiable pooling layer calculates a cluster assignment matrix through the processed node feature matrix and the adjacency matrix of the graph, and aggregates the node features assigned to the same cluster through the cluster assignment matrix, thereby completing the pooling operation. After obtaining the node feature vectors through the graph convolutional layer, the Dropout layer, and the differentiable pooling layer, the node features are processed by divide-and-conquer, and global features are generated through a dense attention mechanism and principal component analysis, and local features are generated through virtual padding and matrix calculation. After obtaining the global features and local features, they are fused using the feature fusion module, and then input into a fully connected layer. The relationship score between the historical trajectory graph and the current state graph is obtained through the fully connected layer, and a numerical mapping of 0-1 is performed through the Sigmod function. The data with the score is integrated with the preprocessed driving data to form vehicle driving data with a score label, and then input into the Social LSTM model.

[0009] Preferably, as the present invention, the steps for training the improved graph neural network model are as follows:

[0010] Step S1. Select the vehicle dataset NGSIM as the vehicle driving trajectory dataset;

[0011] Step S2. Process the original dataset;

[0012] Step S3. After constructing the spatio-temporal graph-structured data, input the historical trajectory graph and the current state graph into the graph convolutional layer of the FPGCN model for processing to generate node feature vectors;

[0013] Step S4. The node features generated after the graph structure data undergoes graph convolution, Dropout, and differentiable pooling are subjected to divide-and-conquer processing, and global features are generated through a dense attention mechanism and PCA; local features are generated through virtual padding and matrix calculation;

[0014] Step S5. After obtaining the global features and local features, a feature fusion module is used for fusion, and then the result is input into a fully connected layer; after obtaining the output of the fully connected layer, the output is converted into a 0-1 distribution through the Sigmod function;

[0015] Step S6. The scored data is integrated with the preprocessed driving data to form vehicle driving data with score labels, which is then input into the Social LSTM model. After that, the predicted trajectory is output through the social pooling network; during the training process, the root mean square error loss function and the Adam optimizer are used, and the graph neural network model is continuously optimized and improved through Optuna.

[0016] As a further preference of the present invention, step S2 specifically includes the following steps:

[0017] Step S201. Perform wavelet denoising on the original data set, and then introduce a same-classification strategy. For vehicle driving data within a set time threshold, vehicle data with the same vehicle_ID and unchanged Lane_ID is labeled as LK; normal vehicle data with the same vehicle_ID but different Lane_IDs is labeled as LC, and secondary labeling is performed according to the positive or negative of the difference in Lane_ID before and after their trajectories. Positive values are labeled as TR, and negative values are labeled as TL;

[0018] Step S202. Detect outliers in the denoised and classified data set;

[0019] Step S203. Introduce a Savitzky-Golay smoothing window to segment the data on the normalized vehicle data;

[0020] Step S204. Convert complex data into graph structure data.

[0021] As a further preference of the present invention, in step S3, the three-layer graph convolutional layer aggregates information of the constructed graph structure data through a graph convolutional neural network, so that each node contains its own feature information. For each vehicle, at a fixed frame, its feature vector representation is:

[0022]

[0023] Wherein, H (l) is the feature matrix of the l layer, =U+I is the adjacency matrix with self-loops added, U is the adjacency matrix of the graph structure data, I is the identity matrix, with the same matrix size as U , a matrix with the main diagonal element value of 1 and other elements of 0, is 's degree matrix, that is, , represents 's i-th diagonal element, represents the element at the -th row and i -th column of the adjacent matrix j after adding self-loops, W (l) is the weight matrix of the l -th layer, σ() is the RELU activation function.

[0024] As a further preference of the present invention, before performing the dense attention mechanism in step S4, the feature vector output by the graph convolutional layer is shape-adjusted to [B, N, F], where B is the batch size, B = 1 when processing a single graph, N is the number of nodes, and F is the feature dimension;

[0025] After that, the average feature is calculated, and the calculation formula is:

[0026]

[0027] where, represents the average feature vector of the i -th graph, represents the feature vector of the i -th node of the j -th graph, represents the mask value of the i -th node of the j -th graph; after traversing and calculating, the average feature matrix is stacked; A sum ;

[0028] Then, the global transformation vector is calculated by calculating the feature mean of all nodes:

[0029]

[0030] where is the learnable weight matrix, is the average feature matrix, Global is the vector after global transformation, and tanh() is the hyperbolic tangent function;

[0031] Then, through the dot product operation of the eigenvector and the global transformation vector, the activation and weighted operations are calculated to obtain D ij :

[0032]

[0033] where, is the attention coefficient of the i th node in the j th graph, represents the eigenvector of the i th node in the j th graph, is the vector corresponding to the graph i in the global transformation vector, () is the sigmod activation function, which converts the dot product result into the attention weight in probability form, and calculates the attention coefficients of all nodes through iteration;

[0034] Finally, the node features are summed up for PCA processing;

[0035]

[0036] where, z i is the weighted aggregation eigenvector of the i th graph, represents the mask value of the i th node in the j th graph, is the attention coefficient of the i th node in the j th graph, represents the eigenvector of the i th node in the j th graph;

[0037] When performing principal component analysis, the number of principal components with the greatest influence is set to 3.

[0038] As a further preference of the present invention, the specific acquisition method of the local features in step S4 is:

[0039] When performing virtual padding processing on the eigenvector, fill the difference in the number of missing nodes, and set the filled elements to 0. The data dimension F 12 of the virtual padding is:

[0040]

[0041] where F 12 is the number of virtual padding, N 1 is the feature dimension of the eigenvector 1,N 2 is the feature dimension of the feature vector 2, represents taking the maximum value of the two dimensions, represents taking the minimum value of the two dimensions;

[0042] After virtual padding, an abstract feature matrix is obtained. The shape of the historical trajectory graph is (B1, N1, F), and the shape of the current state graph is (B2, N2, F); where B1 is the batch size of the historical trajectory graph, N1 is the number of nodes in the historical trajectory graph, B2 is the batch size of the current state graph, N2 is the number of nodes in the current state graph, and F is the feature dimension of each node embedding;

[0043] After that, the historical trajectory graph is multiplied by the transpose of the current state graph through matrix multiplication to obtain a feature relationship matrix with the shape of (N1, N2), and the obtained relationship matrix is normalized.

[0044] The present invention provides a vehicle driving trajectory prediction method based on an improved graph neural network model, and the method includes the following steps:

[0045] Step 1. Obtain historical trajectory data and current state data;

[0046] Step 2. Preprocess the data and convert it into graph-structured data;

[0047] Step 3. Input the historical trajectory and current state graph-structured data into the vehicle driving trajectory prediction system based on the above-mentioned improved graph neural network model, and use the FPGCN model to calculate the influence score of the historical trajectory graph on the current state graph;

[0048] Step 4. The preprocessed driving data and the data after the FPGCN model calculates the score are integrated and input into the SocialLSTM model, and the predicted trajectory is output after passing through the social pooling network.

[0049] Advantages and beneficial effects of the present invention:

[0050] (1) The improved graph neural network model provided by the present invention combines the core ideas of graph convolutional neural network and long short-term memory network with various optimization strategies, such as: Dense Attention Mechanism, Differentiable Pooling, social pooling, etc. These structures can significantly reduce the number of parameters and optimize the calculation logic while maintaining the model's expressive ability. This innovation is very suitable for realizing the trajectory prediction of target vehicles in complex traffic flows.

[0051] (2)The prediction system provided by the present invention performs a divide-and-conquer process on different-depth feature vectors. Each stage focuses on processing features of different granularities, outputs node features as global features and local features, and gradually extracts and integrates graph-structured information through a divide-and-conquer strategy, highlighting the main features for secondary data integration. This branch structure greatly improves the model's ability to perceive graph-structured features, enhances the feature acquisition effect, especially in complex target road conditions, increases the attention to surrounding vehicle information, enables it to adapt to driving data under different road conditions, and retains important local features without losing global features.

[0052] (3)The present invention selects and retains feature vectors with larger variances through PCA and reduces their dimensions. This design greatly reduces the risk of overfitting when the model processes global features. At the same time, when processing node features, important data for generating distribution maps are mapped through matrix multiplication, and feature fusion is achieved through the calculated local features and global features to supplement more detailed feature information.

[0053] (4)When extracting global features, the present invention uses an improved attention mechanism (dense attention mechanism) based on graph data. This improved mechanism can dynamically adjust the feature map and allocate different degrees of attention. This dynamic attention optimization strategy can enhance the context-aware interaction ability, enabling the model to improve the ability to process the feature relationships of complex vehicle graph-structured data, thereby optimizing the model performance and the processing efficiency of complex data, increasing the prediction accuracy, and reducing the model loss.

[0054] (5)The present invention adopts a pooling strategy of differentiable pooling, and automatically learns more efficient pooling operations through the backpropagation algorithm. After the combination of the dense attention mechanism and the differentiable pooling strategy (Diffpool strategy), the accuracy of feature extraction is enhanced, the context interaction perception ability is enhanced, the flexibility of the model is improved while maintaining a high feature expression ability, the number of parameters and the amount of calculation are greatly reduced (the computational complexity is reduced), and the model training efficiency and prediction speed are optimized.

[0055] (6)The present invention adds Batch Normalization and Dropout layers to the improved network. The adaptive hierarchical dimensionality reduction method enables the model to cluster when collecting complex graph-structured features, automatically reduces invalid redundant data and retains key information, and improves the computational efficiency and robustness of the model.

[0056] (7)The present invention adopts a Social LSTM model. Through social pooling and dynamic interaction modeling, the system can more efficiently converge and share through vehicle state information, thereby capturing the interaction rules between road data, finally plotting and outputting the coordinates of the output, and evaluating the model through comparison indicators.

[0057] (8) While maintaining a high prediction rate, the prediction system provided by the present invention significantly reduces the number of parameters and error values in processing complex graph-structured data. Experimental results show that the model performs excellently in the vehicle driving trajectory prediction task, can quickly calculate the surrounding information and make accurate predictions in a complex vehicle driving environment, and provides a new solution for achieving efficient and reliable vehicle trajectory prediction. Description of the Drawings

[0058] By referring to the following description in conjunction with the drawings, and with a more comprehensive understanding of the present invention, other objects and results of the present invention will become more apparent and easier to understand. In the drawings:

[0059] Figure 1 is the data processing flow chart of the improved graph neural network model (Hybrid FPGCN model) provided by the present invention;

[0060] Figure 2 is the RMSE index value of the Hybrid FPGCN model of the present invention and other models on NGSIM data. Detailed Embodiment

[0061] To enable those skilled in the art to better understand the technical solutions and their advantages of the present invention, the present application will be described in detail below with reference to the drawings, but it is not intended to limit the protection scope of the present invention.

[0062] As Figure 1 shown, this embodiment provides a vehicle driving trajectory prediction system based on an improved graph neural network model. The prediction system is obtained by training an improved graph neural network model (Hybrid FPGCN model), and the improved graph neural network model includes an FPGCN model and a Social LSTM model; wherein, the data input to the FPGCN model is two constructed graph-structured data, namely a historical trajectory graph and a current state graph (target graph). The FPGCN model is used to calculate the influence score of the historical trajectory graph on the current state graph. After the preprocessed driving data is integrated with the data whose scores are calculated by the FPGCN model, it is input into the Social LSTM model, and the predicted trajectory is output after passing through the social pooling network;

[0063] Among them, the FPGCN model includes a graph convolutional layer (GCN), a batch normalization layer, a Dropout layer, a differentiable pooling layer (diffPool), a feature fusion module, and a fully connected layer; the graph convolutional layer has three layers, and the scales of the three graph convolutional layers are: 64, 32, 16. After the graph structure data passes through the graph convolution operation, a non-linear factor is introduced through an activation function (such as the relu activation function) to alleviate the gradient vanishing problem, and then the input of each layer is normalized by the batch normalization layer, thereby accelerating the convergence of the model. The constructed graph structure data is aggregated through the three graph convolutional layers described above, so that each node considers not only its own feature information but also the relationship with neighbor nodes, and then is input into the Dropout layer and the differentiable pooling layer; the Dropout layer randomly selects some nodes and forces them to be set to 0, and the selected nodes no longer participate in the forward and backward propagation of the training process, thereby reducing the risk of overfitting in training, and random selection can improve the generalization ability of the model; the differentiable pooling layer (diffPool) calculates the cluster assignment matrix through the processed node feature matrix and the adjacency matrix of the graph, and aggregates the node features assigned to the same cluster through the cluster assignment matrix, thereby completing the pooling operation. Since the cluster assignment matrix is adaptive, the pooling strategy will also change with the change of the input data. This technology can adaptively change the pooling strategy to help the model learn the hierarchical data of the graph; after obtaining the node feature vectors through the graph convolutional layer, the Dropout layer, and the differentiable pooling layer, the node features are then processed by divide-and-conquer. The global features are generated through an improved attention mechanism (dense attention mechanism) based on graph data and principal component analysis, and the local features are generated through virtual padding and matrix calculation; after obtaining the global features and local features, they are fused using the feature fusion module, and then input into a fully connected layer. The relationship score between the historical trajectory graph and the current state graph is obtained through the fully connected layer, and a numerical mapping of 0-1 is performed through the Sigmod function; after the data of the obtained score is integrated with the preprocessed driving data into vehicle driving data with score labels, it is input into the Social LSTM model.

[0064] In this embodiment, the above constructed model is trained, and the specific method is as follows:

[0065] Step S1. Select the vehicle dataset NGSIM-US-101 as the vehicle driving trajectory dataset. This dataset is the vehicle trajectory data collected by the Next Generation Simulation project initiated by the US Federal Highway Administration, focusing on the vehicle driving conditions on roads such as US-101 in Los Angeles, California. It is an open dataset specifically for studying vehicle driving trajectories. Compared with traditional vehicle trajectory datasets (such as HighD, Argoverse 1.0, etc.), the NGSIM dataset pays more attention to the surrounding road information during vehicle driving, especially the information data of the ego vehicle and surrounding vehicles when the vehicle's behavior changes.

[0066] Step S2. Perform data processing on the original dataset. In the data processing stage, a series of methods are adopted to convert it into a form that the model can process. The specific processing methods are as follows:

[0067] Step S201. Perform wavelet denoising on the original dataset, and then introduce the same classification strategy to further improve the representativeness of the dataset.

[0068] Specifically, for vehicle driving data within a set time threshold, vehicle data with the same vehicle_ID and unchanged Lane_ID is labeled as LK (Lane_Keep); for normal vehicle data with the same vehicle_ID but different Lane_IDs, it is labeled as LC (Lane_Change), and secondary labeling is performed according to the positive and negative values of the difference in Lane_ID before and after their trajectories. A positive value is TR (Turn_Right), and a negative value is TL (Turn_Left), so as to classify various behavior data more finely. Through this same classification strategy, the trajectory data of different behaviors can be effectively distinguished, facilitating the model to further learn the data characteristics of this behavior during training.

[0069] In this embodiment, the quantity of the original data and each classified data is shown in Table 1;

[0070] Table 1 Quantity of original data and each classified data

[0071] Category label Quantity Original data 1048,576 Lane_Keep 286,496 Turn_Right 373,219 Turn_Left 213,608

[0072] Step S202. Perform outlier detection on the denoised and classified dataset;

[0073] Specifically, the code reads the data and normalizes the numerical features therein to ensure that all numerical values are on the same scale (between 0 and 1).

[0074] Step S203. A Savitzky-Golay smoothing window is introduced to segment the data on the normalized vehicle data. A low-order polynomial is applied to fit each data point in each segment, and then the values of the low-order polynomial are used to replace the original data points. This can reduce the impact of extreme data on the model calculation amount and have a negative impact on model training, and retain the local features of the data to the greatest extent; finally, the processed data is sorted by index.

[0075] Specifically, the Savitzky-Golay smoothing method is as follows:

[0076] Savitzky-Golay smoothing (abbreviated as S-G smoothing) is based on the least squares method, analyzes and retains the useful information in the data, and reduces the impact of extreme data on the model and experimental results. In the present invention, the smoothing window width is selected as 2m + 1, that is, the number of original data points n in this window (n = 2m + 1).

[0077] Assume that the original data points in the window can be fitted with a polynomial of degree k - 1, that is

[0078] ;

[0079] where: i = (-m, -m + 1,... 0, 1,... m - 1, m), and 1 such polynomial can be obtained for each of the n original data points in this window. n such polynomials form a system of k linear equations, and k fitting parameters a j (j = 0, 1, 2,... k - 1) need to be solved. Generally, the selected filtering window width z w should be greater than or at least equal to k. When z w = k, the algebraic method can be selected to solve the fitting result; if z w > k, the least squares method is selected to obtain the result.

[0080] ;

[0081] Simplified to:

[0082] ;

[0083] The least squares solution of A is:

[0084] ;

[0085] Then the filtered value of Y The calculation formula is as follows:

[0086] ;

[0087] where: , X is a matrix containing data indices for polynomial fitting. is a matrix of coefficient sets, determined solely by the X matrix. The matrix is a (2m + 1)×(2m + 1) matrix, A is a polynomial coefficient vector, E represents the error vector with a length of 2m + 1, reflecting the error situation during the smoothing process, Y represents the original input data, and the S-G smoothing fitting equation can be obtained according to the coefficient matrix. For the smoothed data, through visual comparison of the data gradient changes, the present invention finds that the smoothed experimental data is significantly more suitable for trajectory prediction than the original data.

[0088] Step S204. Convert complex data into graph-structured data;

[0089] Before the classified and processed data is input into the model, the complex vehicle data needs to be converted into a graph structure to facilitate the graph convolutional neural network in the model to extract features from it.

[0090] In this embodiment, it is planned to predict future trajectories through historical trajectories. The graph convolutional neural network needs to capture the important features of historical trajectories during certain behaviors. Therefore, before convolution, the data needs to be converted into a graph structure for pairwise processing, which are hereinafter referred to as I 1 (historical trajectory graph) and I 2 (current state graph) to construct a spatio-temporal relationship graph for the vehicle to be predicted. Assume that the total number of frames of a certain vehicle that appears in the training set is F n , and the experiment hopes to predict the trajectory after F n+1 through the model and verify it by comparison.

[0091] Specifically, data is extracted from the NGSIM dataset, such as vehicle ID (Vehicle_ID), timestamp (Total_Frames, Frame_ID), coordinates (Local_X, Local_Y), instantaneous speed (v_Vel), instantaneous acceleration (v_Acc), etc. to construct graph-structured data. Taking the vehicle with Vehicle_ID as a at Frame_ID as i as an example to construct a node f a :

[0092]

[0093]

[0094] Among them, f a represents the data of vehicle a's own node at Frame_ID as i, represents the relative coordinates of vehicle a at time i, are respectively the instantaneous speed and instantaneous acceleration of vehicle a at time i; Ga Represents the vehicle driving data that appears near vehicle a at time i. Assume that there are S vehicles around vehicle a at time i. n vehicles, denoted as S 1 , S 2 .. S i ..S n for differentiation. Represents the relative coordinates of vehicle S i . Respectively, are the instantaneous speed and instantaneous acceleration of vehicle S at time i. Taking i the f a node as the central node of the graph at this time, it is necessary to calculate the relative positions of vehicle a and the surrounding vehicles through coordinates to construct the edges of the undirected graph; taking vehicle S near vehicle a at time i as an 1 example, the edge value d between the two vehicle nodes is: i as follows:

[0095]

[0096] where and are the current relative coordinates of vehicle S. 1 .

[0097] In step S3, after constructing the spatio-temporal graph structure data, input the historical trajectory graph and the current state graph (target graph) into the graph convolutional layer (GCN) of the FPGCN model for processing to generate node feature vectors.

[0098] Traditional GCN only depends on the graph convolutional layer and is difficult to pay attention to the global context information. It only focuses on node relationships but ignores the global spatio-temporal feature relationships, especially in complex road conditions (such as peak hours) and cannot effectively process information. To solve this problem, in this embodiment, after obtaining the node feature vectors, the node features are processed by divide-and-conquer (the node feature vectors are processed in two paths) to obtain global features and local features.

[0099] In addition, due to the characteristics of GCN, convolutional operations exceeding three layers will cause gradient problems. Therefore, this application uses a three-layer graph convolutional network with layer scales of 64, 32, and 16; the three-layer graph convolutional network can greatly retain the results during the process of collecting features.

[0100] Specifically, through the graph convolutional neural network (with scales of 64, 32, 16), the constructed graph structure data is aggregated with information, so that each node contains its own feature information. For each vehicle (node), at a fixed frame, its feature vector can be expressed as:

[0101]

[0102] where H(l) is the feature matrix of the l layer, =U+I is the adjacency matrix with self-loops added, U is the adjacency matrix of the graph structure data, I is the identity matrix (a matrix with the same scale as U , with the main diagonal element values being 1 and other elements being 0), is 's degree matrix, that is , represents 's i-th diagonal element, represents the -th i row and j -th column element of the adjacent matrix W (l) is the weight matrix of the l layer, σ() is the RELU activation function.

[0103] Furthermore, due to the complexity of the graph structure data, the model load will inevitably increase in the experiment, affecting the computational efficiency. Therefore, before the node feature divide-and-conquer processing, a Dropout layer (with a set dropout probability of 0.3 in the experiment) and a differentiable pooling strategy (differentiable pooling layer) are added after the graph convolution layer. By the Dropout layer, the risk of overfitting is prevented, and the generalization ability of the model is enhanced. By the differentiable pooling strategy, the pooling method is dynamically adjusted with high flexibility to help the model learn the hierarchical data of the graph, further reducing the complexity of the graph while retaining key information.

[0104] Step S4. The node features generated after graph convolution, Dropout, and differentiable pooling of the graph structure data are subjected to divide-and-conquer processing, and global features are generated through a dense attention mechanism and PCA; local features are generated through virtual padding and matrix calculation;

[0105] Specifically, the method for obtaining global features is as follows:

[0106] After obtaining the node features, in order to enhance the connection between local features, enable context interaction, increase the efficiency of the model in processing complex data, and enhance the context awareness ability, in this embodiment, the attention mechanism is improved to a dense attention mechanism (DAM) to make it more adaptable to complex graph feature data. It considers the pairwise relationships between all factors in the input sequence when calculating the attention score, thereby generating a dense attention score matrix, which helps the model make more accurate predictions and significantly improves the performance of the model in processing complex data. The specific steps are:

[0107] (1) Increase the data dimension;

[0108] For the feature vector (with shape [N, F]) output by the graph convolutional layer, it is necessary to adjust the shape to [B, N, F] by adding a batch dimension (where B is the batch size, usually B = 1 when processing a single graph, N is the number of nodes, and F is the feature dimension). This operation can meet the need to process complex data and calculate the attention scores without changing the essential content of the data.

[0109] (2)Calculate the average feature:

[0110] In the calculation process of the dense attention mechanism, considering the features of itself and neighboring nodes, the average feature is calculated through an adjustable mask matrix Q (a matrix composed of 0 and 1, with the same dimension as the node feature dimension, indicating whether the corresponding node is valid):

[0111]

[0112] Among them, represents the average feature vector of the i th graph, represents the feature vector of the i th node in the j th graph, represents the mask value of the i th node in the j th graph; after traversing and calculating and stacking, the average feature matrix sum A sum can be obtained.

[0113] (3)Global transformation vector:

[0114] The global transformation vector can provide context information for the dense attention mechanism. In this way, when calculating the attention coefficients, each node feature will interact with the global transformation vector. The global transformation vector is calculated by computing the mean of the features of all nodes:

[0115]

[0116] Among them is a learnable weight matrix, is the average feature matrix, Global is the vector after global transformation, and tanh() is the hyperbolic tangent function, whose role is to introduce non-linear activation so that the model can learn complex data.

[0117] (4)Calculate the attention coefficients:

[0118] The attention coefficient enables the model to adaptively assign different weights to each node of the dense data, so as to better capture the feature information of the dense data. The coefficient is calculated through the dot product operation of the feature vector and the global transformation vector, and the activation and weighting operations. D ij :

[0119]

[0120] Among them, is the attention coefficient of the i th node in the j th graph, represents the feature vector of the i th node in the j th graph, is the vector corresponding to the graph i in the global transformation vector, () is the sigmod activation function, which converts the dot product result into the attention weight in probability form. Through iteration, the attention coefficients of all nodes can be calculated.

[0121] (5)Weighted summation:

[0122] The purpose of weighting the node features is to assign different weights to features of different importance to highlight the important node features; the purpose of summing the node features is to integrate all node information and represent it uniformly for subsequent PCA processing:

[0123]

[0124] Among them, z i is the weighted aggregated feature vector of the i th graph, represents the mask value of the i th node in the j th graph, is the attention coefficient of the i th node in the j th graph, represents the feature vector of the i th node in the j th graph; the feature vector generated by the graph convolution layer has stronger feature expressiveness after being processed by the improved attention mechanism (dense attention mechanism), thereby improving the model's processing ability for complex graph data.

[0125] After obtaining the data processed by the global feature, through principal component analysis (PCA), the clustered feature points are centered, decomposed, the principal components are selected, and the data is transformed to obtain the feature projection calculated from the global feature, completing the global feature processing step in the divide-and-conquer processing structure.

[0126] It should be noted that in this embodiment, the existing PCA is directly utilized. The key lies in parameter adjustment. The number of principal components retaining the greatest influence is set to 3 to reduce the model complexity and simultaneously reduce the risk of overfitting.

[0127] Specifically, the method for obtaining local features is as follows:

[0128] (1) After obtaining the node features, both the nodes and the side lengths are determined. However, to avoid affecting the experimental results due to differences in the data structures or dimensions of the two feature data in the later stage of the experiment, virtual padding processing needs to be performed on the feature vectors to fill the difference in the number of missing nodes. In the experiment, the value of the filled element is set to 0, which can keep the data dimensions consistent. Assuming the two feature dimensions are (N 1 , N 2 ), then the data dimension F 12 of the virtual padding is:

[0129]

[0130] where F 12 is the number of virtual padding, N 1 is the feature dimension of feature vector 1, N 2 is the feature dimension of feature vector 2, represents taking the maximum value of the two dimensions, represents taking the minimum value of the two dimensions.

[0131] (2) An abstract feature matrix is performed. The two features input into the model in this embodiment are the historical trajectory graph feature and the target graph feature. Taking the historical trajectory graph as an example, its shape and size are (B1, N1, F), where B1 is the batch size of the historical trajectory graph (indicating the number of graphs during batch processing), N1 is the number of nodes in the historical trajectory graph, and F is the feature dimension embedded in each node; similarly, the shape of the target graph is (B2, N2, F), B2 is the batch size of the target graph (current state graph), N2 is the number of nodes in the target graph (current state graph), and then batch processing is performed on the abstracted graph.

[0132] (3) Convert the abstract features processed in step 2 from the sparse format to the dense format;

[0133] Specifically, multiply the historical trajectory graph by the transpose of the target graph through matrix multiplication to obtain a feature relationship matrix with the shape of (N1, N2), and perform normalization processing on the obtained relationship matrix; at this time, the mapped values show the density of the node data feature relationship.

[0134] Step S5. After obtaining the global features and local features, use the feature fusion module to fuse them, and then input them into the fully connected layer;

[0135] Specifically, the globally feature after PCA mapping is fused with the locally feature after matrix transformation, and then input into a fully connected layer; the function of the fully connected layer is to map the fused features.

[0136] Let the feature vector be f, the weight matrix of the fully connected layer be W and the bias term be b, then the output of the fully connected layer can be expressed as:

[0137]

[0138] Among them: the output of the fully connected layer is Z, that is, the corresponding score; W is the weight matrix of N out ×C in , N out is the number of output features, and C in is the number of input features.

[0139] In this embodiment, after obtaining the output of the fully connected layer, the output is usually converted into a 0-1 distribution through the Sigmod function. The formula of the Sigmod function is:

[0140]

[0141] Among them: z is the input of the function, e is the base of the natural logarithm (approximately equal to 2.71828), is the output of the function, and its value range is (0, 1). The closer the value is to 0, the lower the degree of relationship between the historical trajectory map and the target map. The closer the value is to 1, the closer the relationship degree is. When facing similar road conditions, the probability of making a driving behavior with the same trajectory as the history is greater.

[0142] Step S6. Integrate the scored data with the preprocessed driving data into vehicle driving data with score labels, and then input it into the Social LSTM model. After passing through the social pooling network, the predicted trajectory is output.

[0143] It should be noted that in this embodiment, the differentiable pooling layer, the Social LSTM model and the social pooling network can all adopt existing network structures. The key of this application is to combine them with the divide-and-conquer structure. Therefore, the specific network structures and data processing processes of this application will not be introduced in detail.

[0144] During the training process of this embodiment, the root mean square error (RMSE) loss function and the Adam optimizer are used, and hyperparameters such as the learning rate, batch size, and number of training epochs are optimized through Optuna to improve the performance and accuracy of the model. In the testing phase, the FPGCN model is used to generate relationship scores to integrate the processed NGSIM vehicle data, completing the input preparation for the Social LSTM model to perform trajectory prediction; for example: to perform trajectory prediction on vehicle v at the i th frame, the trained model processes the historical trajectory map (frames 1 to i -1 frames as the historical trajectory map) and the target trajectory map ( i frame as the target trajectory map), outputs the coordinates for subsequent time, and evaluates the prediction effect and the RMSE value to ensure that the system can accurately and efficiently process vehicle driving data, providing strong technical support for efficient intelligent driving.

[0145] Most of the existing trajectory predictions use the attention mechanism and Transformer for encoding and decoding during feature acquisition, only using global features, so the root mean square error (RMSE) of some prediction values will be a bit higher. The prediction system provided in this embodiment obtains global features and local features through a divide-and-conquer structure. After the local features and global features are fused, they are added to the feature calculation. The global features are the main part, and the local features are used as information complementation after calculation, which is equivalent to considering road vehicle data more comprehensively than previous methods to achieve better prediction effects. To prove that the model provided in this embodiment is more suitable for vehicle trajectory prediction and has better prediction effects, this application compares it with other existing models, and the specific results are as Figure 2 shown.

[0146] It can be seen from Figure 2 that the Hybrid FPGCN model provided by the present invention performs more excellently in terms of the root mean square error (RMSE) index compared with other prediction models. The average value of the RMSE index in the first 30 frames (3s) is reduced by 0.31, 0.31, 0.25, 1.74, 0.27, 0.16, 0.69, 0.19 compared with the CS-LSTM, M-LSTM, TS-GAN, BRDP, WSIP, HMNET, S-GCN, MFP models, and the RMSE values in the time periods of 10, 20, and 30 frames are all lower than those of the above models. It can be seen that the present invention has more excellent capabilities in processing graph structure data and trajectory prediction. Embodiment

[0147] This embodiment provides a vehicle driving trajectory prediction method based on an improved graph neural network model, and the method includes the following steps:

[0148] Step 1. Obtain historical trajectory data and current state data;

[0149] Step 2. Preprocess the data and convert it into graph-structured data;

[0150] Step 3. Input the historical trajectory and current state graph-structured data into the vehicle driving trajectory prediction system based on the improved graph neural network model described in Embodiment 1, and use the FPGCN model to calculate the influence score of the historical trajectory graph on the current state graph;

[0151] Step 4. Integrate the preprocessed driving data with the data whose scores have been calculated by the FPGCN model, input them into the SocialLSTM model, and output the predicted trajectory after passing through the social pooling network.

[0152] The present invention also provides an electronic device, including: one or more processors, a memory; wherein, the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the vehicle driving trajectory prediction method based on the improved graph neural network model described in Embodiment 2.

[0153] The present invention also provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the vehicle driving trajectory prediction method based on the improved graph neural network model described in Embodiment 2.

[0154] Those skilled in the art can understand that all or part of the functions of the above various methods / modules can be implemented in a hardware manner or in a computer program manner. When all or part of the functions in the above embodiments are implemented in a computer program manner, the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory, random access memory, magnetic disk, optical disk, hard disk, etc. The above functions are implemented by a computer executing the program. For example, storing the program in the memory of the device, when the program in the memory is executed by the processor, the above all or part of the functions can be implemented.

[0155] In addition, when all or part of the functions in the above embodiments are implemented in a computer program manner, the program can also be stored in a storage medium such as a server, another computer, magnetic disk, optical disk, flash drive or mobile hard disk, downloaded or copied and saved to the memory of the local device, or the system of the local device is updated in version. When the program in the memory is executed by the processor, the above all or part of the functions in the above embodiments can be implemented.

[0156] The above uses specific examples to illustrate the present invention, which is only for helping to understand the present invention and is not intended to limit the present invention. For those skilled in the technical field to which the present invention pertains, several simple deductions, deformations or substitutions can also be made according to the idea of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A vehicle driving trajectory prediction system based on an improved graph neural network model, wherein the prediction system is obtained by training the improved graph neural network model, and is characterized in that: The improved graph neural network model includes an FPGCN model and a SocialLSTM model; wherein the data input to the FPGCN model are two constructed graph structured data, namely, a historical trajectory graph and a current state graph; the FPGCN model is used to calculate the impact score of the historical trajectory graph on the current state graph; the pre-processed driving data is integrated with the data whose scores are calculated by the FPGCN model and then input into the SocialLSTM model, and the predicted trajectory is output after the social pooling network; Among them, the FPGCN model includes a graph convolution layer, a batch normalization layer, a Dropout layer, a differentiable pooling layer, a feature fusion module, and a fully connected layer; the graph convolution layer has three layers, and the scales of the three graph convolution layers are: 64, 32, and 16 respectively. After the graph convolution operation, the graph structure data introduces nonlinear factors through the activation function, and then the batch normalization layer is used to standardize the input of each layer. The constructed graph structure data is aggregated through the three-layer graph convolution layer, so that each node considers its own feature information while also considering the relationship with the neighboring nodes, and then inputs the Dropout layer and the differentiable pooling layer; the Dropout layer randomly selects some nodes and forces them to be 0; the differentiable pooling layer compares the processed node feature matrix with the graph The cluster assignment matrix is ​​calculated based on the adjacency matrix, and the node features assigned to the same cluster are aggregated through the cluster assignment matrix to complete the pooling operation; the node feature vector is obtained through the graph convolution layer, the Dropout layer, and the differentiable pooling layer, and then the node features are divided and conquered, and the global features are generated through the dense attention mechanism and principal component analysis, and the local features are generated through virtual filling and matrix calculation; after obtaining the global features and local features, they are fused using the feature fusion module and then input into a fully connected layer, and the relationship score between the historical trajectory graph and the current state graph is obtained through the fully connected layer, and the Sigmod function is used for 0-1 numerical mapping; the scored data is integrated with the pre-processed driving data into vehicle driving data with score labels and then input into the Social LSTM model.

2. A vehicle driving trajectory prediction system based on an improved graph neural network model according to claim 1, characterized in that: The steps for training the improved graph neural network model are: Step S1. Select the vehicle dataset NGSIM as the vehicle driving trajectory dataset; Step S2. Processing the original data set; Step S3. After constructing the spatiotemporal graph structure data, the historical trajectory graph and the current state graph are input into the graph convolution layer of the FPGCN model to generate a node feature vector; Step S4. The node features generated by graph structure data after graph convolution, Dropout, and differentiable pooling are divided and conquered, and global features are generated through dense attention mechanism and PCA; local features are generated through virtual filling and matrix calculation; Step S5. After obtaining the global features and local features, the features are fused using the feature fusion module and then input into the fully connected layer; after obtaining the output of the fully connected layer, the output is converted into a 0-1 distribution through the Sigmod function; Step S6. The scored data is integrated with the pre-processed driving data into vehicle driving data with score labels, and then input into the SocialLSTM model, and then outputs the predicted trajectory after passing through the social pooling network; During the training process, the root mean square error loss function and Adam optimizer are used, and the graph neural network model is continuously optimized and improved through Optuna.

3. A vehicle driving trajectory prediction system based on an improved graph neural network model according to claim 2, characterized in that: Step S2 specifically includes the following steps: Step S201. Perform wavelet denoising on the original data set, and then introduce the same classification strategy. The vehicle data with the same vehicle_ID and unchanged Lane_ID in the vehicle driving data within the set time threshold is marked as LK; the normal vehicle data with the same vehicle_ID but different Lane_ID is marked as LC, and the positive and negative values ​​of the Lane_ID difference before and after the trajectory are marked again, with positive values ​​as TR and negative values ​​as TL; Step S202: Perform outlier detection on the denoised and classified data set, and normalize the numerical features therein to ensure that all values ​​are on the same scale; Step S203. Introduce a Savitzky-Golay smoothing window to perform data segmentation on the normalized vehicle data; Step S204: Convert complex data into graph structure data.

4. A vehicle driving trajectory prediction system based on an improved graph neural network model according to claim 2, characterized in that: In step S3, the graph convolution layer described in the third layer aggregates the constructed graph structure data through the graph convolution neural network, so that each node contains its own feature information. For each vehicle, its feature vector is represented as follows in a fixed frame: Among them, H (l) is the feature matrix of the lth layer, is the adjacency matrix with self-loops added, U is the adjacency matrix of the graph structure data, I is the identity matrix, the matrix size is the same as U, the main diagonal element value is 1, and the other elements are 0. yes The degree matrix of express The i-th diagonal element of Represents the neighbor matrix after adding the self-loop The element in the i-th row and j-th column of (l) is the weight matrix of the lth layer, and σ() is the RELU activation function.

5. A vehicle driving trajectory prediction system based on an improved graph neural network model according to claim 2, characterized in that: Step S4: Before performing the dense attention mechanism, the feature vector output by the graph convolution layer is reshaped to [B, N, F], where B is the batch size, B = 1 when processing a single graph, N is the number of nodes, and F is the feature dimension; Then calculate the average feature, the calculation formula is: Among them, A i represents the average eigenvector of the i-th graph, x ij represents the feature vector of the jth node of the i-th graph, Q ij Represents the mask value of the jth node of the i-th graph; through A i After traversing and calculating, the average feature matrix A is obtained by stacking sum ; Then calculate the global transformation vector Global by calculating the feature mean of all nodes: Global=tanh(W A ·A Sum ) Where W A ∈R F×F is a learnable weight matrix, A Sum ∈R B×F is the average feature matrix, Global is the vector after global transformation, and tanh() is the hyperbolic tangent function; Then, through the dot product operation of the feature vector and the global transformation vector, activation and weighting operations are calculated to obtain D ij : D ij =σ(x ij ·Global i ) Among them, D ij is the attention coefficient of the jth node in the i-th graph, xij represents the feature vector of the jth node in the i-th graph, and Global i is the vector corresponding to graph i in the global transformation vector, σ() is the sigmoid activation function, which converts the dot product result into the attention weight in the form of probability, and calculates the attention coefficient of all nodes through iteration; Finally, the node features are summed and used for PCA processing; Among them, z i is the weighted aggregate feature vector of the i-th graph, Q ij represents the mask value of the jth node in the i-th graph, D ij is the attention coefficient of the jth node in the i-th graph, x ij Represents the feature vector of the jth node of the i-th graph; In principal component analysis, the number of principal components with the greatest influence is set to 3.

6. A vehicle driving trajectory prediction system based on an improved graph neural network model according to claim 2, characterized in that: The specific method of obtaining the local features in step S4 is: When virtual filling is performed on the feature vector, the missing node number difference is filled, and the filled elements are set to 0. The virtual filling data dimension F 12 for: F 12 =max(N1,N2)-min(N1,N2) Among them, F 12 is the number of virtual fillings, N1 is the feature dimension of feature vector 1, N2 is the feature dimension of feature vector 2, max(N1, N2) means taking the maximum value of the two dimensions, and min(N1, N2) means taking the minimum value of the two dimensions; After virtual filling, the abstract feature matrix is ​​formed. The shape of the history trajectory graph is (B1, N1, F), and the shape of the current state graph is (B2, N2, F); where B1 is the batch size of the history trajectory graph, N1 is the number of nodes in the history trajectory graph, B2 is the batch size of the current state graph, N2 is the number of nodes in the current state graph, and F is the feature dimension embedded in each node; Then, the historical trajectory graph is multiplied by the transpose of the current state graph through matrix multiplication to obtain a feature relationship matrix with a shape of (N1, N2), and the obtained relationship matrix can be normalized.

7. A vehicle trajectory prediction method based on an improved graph neural network model, characterized in that: The method comprises the following steps: Step 1. Obtain historical trajectory data and current status data; Step 2. Preprocess the data and convert it into graph structure data; Step 3. Input the historical trajectory and current state graph structure data into the vehicle driving trajectory prediction system based on the improved graph neural network model according to any one of claims 1 to 6, and use the FPGCN model to calculate the impact score of the historical trajectory graph on the current state graph; Step 4. The preprocessed driving data is integrated with the data whose scores are calculated by the FPGCN model and then input into the SocialLSTM model, which outputs the predicted trajectory after passing through the social pooling network.

Citation Information

Patent Citations

  • Coal mine underground miner trajectory identification method based on dynamic space-time semantic joint embedding

    CN116524227A

  • Automatic trajectory prediction method based on graph spatial-temporal pyramid

    WO2024193334A1