A vehicle interaction perception trajectory prediction system and method based on graph convolution

CN118570763BActive Publication Date: 2026-10-09GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410618387.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2026-10-09
Estimated Expiration
2044-05-17

AI Technical Summary

Technical Problem

但是,这些研究大部分都是采用静态图结构,由于每个节点的相邻节点数量会随着时间而发生改变,因此无法更好的解释运动车辆间的动态空间交互

Benefits of technology

[0033] (1) A dynamic graph convolutional DG-GCN unit is proposed to extract multi-semantic vehicle dynamic interaction information from the feature matrix, and its reliability is verified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118570763B_ABST
    Figure CN118570763B_ABST
Patent Text Reader

Abstract

The application discloses a vehicle interaction perception trajectory prediction system and method based on graph convolution, the prediction system comprising an input preprocessing module, a DGAC module and a trajectory prediction module connected in sequence; vehicle data collected is subjected to the input preprocessing module to generate a feature matrix and graph topology data, is then subjected to a multilayer DGAC module to extract multisemantic graph space-time features and selectively pay attention to the difference between the spatial dimension, the time dimension and the channel dimension and key information by using an attention mechanism, and finally is subjected to the trajectory prediction module to generate a predicted trajectory of a vehicle in a surrounding environment for a future period of time. The application is also verified by experiments on two real data sets, and the experiments show that the model has stronger graph representation capability and is more suitable for the field of vehicle trajectory prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle interaction, and more specifically, to a method for predicting vehicle interaction perception trajectory based on graph convolution. Background Technology

[0002] Autonomous driving has increasingly come into focus in recent years. From its initial conception to the current widespread deployment of trial vehicles by major companies, it's clear that autonomous driving technology is gradually becoming a reality and a major trend in intelligent transportation. As a key research area in artificial intelligence, autonomous driving is receiving significant attention from both academia and industry. Currently, the perception, decision-making, and control modules of autonomous driving have made rapid progress; however, full deployment requires verification of its safety. To achieve safe and efficient driving, autonomous vehicles need to autonomously and reasonably perceive the movement of surrounding road users and accurately predict their future trajectories, thereby assisting the vehicle in making safe and immediate decisions.

[0003] The focus and objective of this invention is to propose a vehicle trajectory prediction method applicable to multiple scenarios, thereby improving the accuracy of trajectory prediction for surrounding traffic participants and ultimately enhancing the safety of autonomous vehicles.

[0004] The current status of research and technology both domestically and internationally is mainly shown below:

[0005] First, it fails to fully extract the spatiotemporal interaction information between moving vehicles. Current research primarily employs graph neural networks to extract this information, treating target vehicles as nodes and surrounding vehicles as adjacent nodes. Interactions between vehicles are represented by edges, and finally, graph neural networks are used to obtain the dependencies between each traffic participant, thereby modeling the interaction between the target vehicle and its surroundings. However, most of these studies use static graph structures. Since the number of adjacent nodes for each node changes over time, they cannot adequately explain the dynamic spatial interactions between moving vehicles.

[0006] Second, the impact of surrounding vehicles on trajectory interactions varies. Most existing methods employ attention mechanisms to simultaneously address the temporal and spatial differences between vehicles. However, most trajectory prediction models only consider the spatial (interaction information) and temporal (frame time information) differences between vehicles using attention mechanisms, without considering applying attention mechanisms at the channel dimension (position, speed, and other features of trajectory data) to selectively enhance the importance of different features. Compared to traditional physical model methods, graph neural network-based methods, when modeling complex interactions between vehicles, recognize that these interactions change over time. Therefore, they cannot be considered a static graph structure. Because static graph topologies are fixed, graph neural networks, with increasing layers, lack the ability to model multi-level semantic information across different layers, thus failing to fully represent the dynamic interactions between vehicles. Furthermore, the impact of surrounding vehicles on the target vehicle varies; how to handle these variations is a crucial issue for improving prediction accuracy. Summary of the Invention

[0007] To address the aforementioned problems, this invention first proposes a vehicle interaction perception trajectory prediction system based on graph convolution.

[0008] A vehicle interaction perception trajectory prediction system based on graph convolution includes an input preprocessing module, a DGAC module, and a trajectory prediction module connected in sequence.

[0009] The collected vehicle data is processed by a preprocessing module to generate a feature matrix and graph topology. Then, a multi-layer DGAC module is used to extract multi-semantic graph information and selectively focus on the differences between spatial, temporal, and channel dimensions and key information. Finally, the trajectory prediction module generates the predicted trajectories of all vehicles for a future time period.

[0010] The input data X to the input preprocessing module is all the observed vehicles at historical time step t. h Position and velocity in the middle, i.e.

[0011]

[0012] in,

[0013]

[0014] p (t) Let x represent the position coordinates (x, y) and velocity (u, v) of vehicle N within the range observed at time t.

[0015] The data Y output by the trajectory prediction module is the sum of all observed vehicles from time t. h+1 to t h +t fThe prediction of future speed, t f It is the range of future prediction time steps;

[0016]

[0017] in,

[0018]

[0019] The input preprocessing module includes creating a feature matrix and a graph topology, and the specific implementation process is as follows:

[0020] (1) Creating a feature matrix: Given a traffic scenario, assuming that in the past T h V vehicles were observed at time step, and the raw data was preprocessed into a three-dimensional tensor. Let C represent the coordinates (x, y) and velocities (u, v) of a vehicle in the scene at different time steps, and normalize all coordinates and velocities to the range (-1, 1);

[0021] (2) Creating the graph topology: For each observed time step t, an undirected graph G is constructed. t =(V t E t ), where node V t Represents the vehicle and edge E at time t. t This represents the spatial interaction between vehicles; where the set of nodes at time step t is defined as E, where the edge set is t Defined in time step t

[0022] At each time step t, spatial interactions are defined to occur only at distances. Between the two cars inside, T close The threshold is configurable, and it defines that vehicles must belong to adjacent lanes or the same lane, that is... Therefore, E at each time step t t Adjacency matrix Defined in the following way:

[0023]

[0024] Assume n is the number of vehicles observed at time step t, and n is always less than V, where V is the maximum number of vehicles in the scene within the manually set time step t. Therefore, given the trajectory of observed vehicles n≤V and time step T... h The present invention can obtain the adjacency matrix as described above. Finally, the feature matrix and adjacency matrix containing time information They are input together into the model (i.e., the DGAC module).

[0025] The DGAC module includes a dynamic graph convolutional DG-GCN unit, an enhanced channel attention mechanism (SCA) unit, and a temporal convolutional layer (TCN) unit.

[0026] The Dynamic Graph Convolutional (DG-GCN) unit groups feature information and then uses multiple different types of graph convolution operations to extract multi-semantic spatial information. Its input information is first processed by a 1×1 convolution into K groups of feature information, each group consisting of C / K channels. Then, each feature group is convolved with a corresponding A... i Perform graph convolution operation; finally, concatenate the K feature groups along the channel dimension, and then perform a 1×1 convolution operation on the merged feature information to generate an output with multi-level graph semantics.

[0027] The Enhanced Channel Attention (SCA) unit comprises three sub-units: Spatial Attention (SAM), Temporal Attention (TAM), and Channel Attention (CAM). The input feature map is sequentially fed into the three sub-units to generate attention maps along different dimensions. The attention maps are multiplied by the original feature map to enhance the corresponding features.

[0028] The temporal convolutional layer (TCN) unit is the new feature generated after passing through the DG-GCN and SCA units. The data is fed into a multi-scale temporal convolutional layer (TCN); the kernel size of the TCN is set to (1×3), (1×5), and (1×7) to force them to process the data along different time dimensions.

[0029] The trajectory prediction module employs a GRU-based encoder-decoder module to predict the future trajectories of all observed vehicles. Specifically, in the first decoding step, the encoder's hidden features and all vehicle velocity data from the last observation time step are fed into the decoder to predict vehicle trajectories. In subsequent decoding steps, the decoder uses its own hidden features and the predicted velocities of all objects from the previous time step as input for prediction. Then, all expected future trajectories are calculated using the predicted trajectories at time steps t. f The predicted speed is averaged, and then the average speed (Δx, Δy) is added to the vehicle's final historical position, thus successfully converting the result into (x, y) coordinates.

[0030] This invention utilizes time-series trajectory data, further processing it to obtain graph structure data and feature representation matrices that change over time. The processed new data is then fed into the DGAC-NET model for multi-semantic extraction of vehicle dynamic spatiotemporal interaction information. Finally, an encoding-decoding method is used to predict the future trajectories of surrounding vehicles and the target vehicle. Experiments on two real-world datasets validate the model, demonstrating its stronger graph representation capabilities and greater suitability for vehicle trajectory prediction.

[0031] A vehicle interaction perception trajectory prediction method based on graph convolution is applied to the system described above.

[0032] Compared with existing technologies, the beneficial effects of the present invention are as follows:

[0033] (1) A dynamic graph convolutional DG-GCN unit is proposed to extract multi-semantic vehicle dynamic interaction information from the feature matrix, and its reliability is verified.

[0034] (2) A lightweight channel attention enhancement mechanism (SCA) unit is proposed and embedded in the back of each graph convolutional layer. This unit can help the system model learn to selectively focus on important information of inter-vehicle interactions and important information of frames, thereby enhancing the expression of channel information (position, speed, etc.).

[0035] (3) The proposed method model DGAC-NET was validated in simple highway intersection scenarios and complex urban road scenarios, and compared with other prediction models, proving the effectiveness of the present invention.

[0036] Training and validation were performed on two public datasets (NGSIM dataset and ApolloScape trajectory dataset). Results show that the proposed method can predict more realistic vehicle trajectories in both simple highway intersection scenarios and complex urban road scenarios. Furthermore, in congested highway intersection scenarios (10 or more vehicles), the root mean square error (RMSE) of the predicted trajectory is reduced by 29% compared to the classic CS-LSTM model. In complex urban road scenarios, compared to the official baseline model (TrafficPredict), it improves accuracy by nearly 87% on weighted average displacement error (WSADE) and by nearly 91% on weighted final displacement error (WSFDE). In summary, the proposed DGAC module effectively enhances the model's prediction accuracy. Attached Figure Description

[0037] Figure 1 This is a diagram of the overall architecture of the trajectory prediction model.

[0038] Figure 2 This is a diagram of the DG-GCN architecture.

[0039] Figure 3 This is the SCA architecture diagram.

[0040] Figure 4 This is a diagram of the vehicle trajectory prediction module.

[0041] Figure 5 A visualization of trajectory predictions for the NGSIM dataset. Detailed Implementation

[0042] The accompanying drawings are for illustrative purposes only and should not be construed as limiting this patent. To better illustrate this embodiment, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions.

[0043] It will be understood by those skilled in the art that certain well-known structures and their descriptions may be omitted in the accompanying drawings. The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0044] This invention formulates the trajectory prediction problem as predicting the future trajectories of all objects in an observed scene based on their historical trajectories. Considering that predicting an object's velocity is easier than predicting its position, historical positions and velocities are input into the model of this invention, allowing the model to predict future velocities. Then, the final position prediction is obtained by summing the predicted future velocities and the last observed position.

[0045] The input X to the model of this invention is all observed vehicles at time t. h Position and velocity within a historical time step.

[0046]

[0047] in,

[0048]

[0049] This represents the position coordinates (x, y) and velocity (u, v) of N vehicles within the observed range at time t. The output Y represents the total number of observed vehicles from time t. h+1 to t h +t f The prediction of future speed, t f It is the range of future prediction time steps.

[0050]

[0051] in,

[0052]

[0053] The overall framework of this method model is as follows: Figure 1 As shown, this model consists of an input preprocessing module, a DGAC module, and a trajectory prediction module. Vehicle data is processed by the input preprocessing module to generate corresponding feature matrices and graph topology data. The DGAC module then extracts multi-semantic graph information and selectively focuses on the differences between spatial dimensions (inter-vehicle interaction information), temporal dimensions (information for each frame), and channel dimensions (position, speed features, etc.) and key information using attention. Finally, the trajectory prediction module generates predicted trajectories for all vehicles at a future time.

[0054] The specific implementation steps of the above system model are as follows:

[0055] I. Feature Matrix Creation and Graph Structure Construction

[0056] (1) Creation of the feature matrix

[0057] Given a traffic scenario, assuming in the past T h V vehicles were observed at time step. This invention preprocesses the raw data into a three-dimensional tensor. like Figure 1 As shown in the figure. In this embodiment, C is set to represent the coordinates (x, y) and velocities (u, v) of the vehicle in the scene at different time steps, and all coordinates and velocities are normalized to the range (-1, 1).

[0058] (2) Creation of graph topology

[0059] In traffic scenarios, vehicle motion is significantly influenced by the motion of surrounding vehicles. Therefore, this invention proposes that representing the interdependencies between vehicles as an undirected graph is effective. Specifically, for each observed time step t, an undirected graph G is constructed. t =(V t E t ), where node V t Represents the vehicle and edge E at time t. t This represents the spatial interaction between vehicles. The set of nodes at time step t is defined as... E, where the edge set is t Defined in time step t At each time step t, it is defined that spatial interactions will only occur within a distance D, where Between the two cars, T close The threshold is configurable, and this invention defines that vehicles must belong to adjacent lanes or the same lane, that is... Therefore, E at each time step t t Adjacency matrix Defined in the following way:

[0060]

[0061] Assume n is the number of vehicles observed at time step t, and n is always less than V, where V is the maximum number of vehicles in the scene within the manually set time step t. Therefore, given the trajectory of observed vehicles n≤V and time step T... h The present invention can obtain the adjacency matrix as described above. Finally, the feature matrix and adjacency matrix containing time information Input them into the model together.

[0062] II. Extracting Spatiotemporal Information with Multiple Semantics Between Vehicles

[0063] The DGAC module consists of a Dynamic Graph Convolutional (DG-GCN) unit, a Channel Enhancement Attention (SCA) unit, and a Temporal Convolutional (TCN) unit. The main function of the DGAC module is to aggregate multi-semantic spatiotemporal graph structural information using dynamic graphs, and selectively enhance inter-vehicle interaction information, temporal information, and channel (position and velocity) information using the attention mechanism. Finally, the enhanced multi-semantic spatial information feature matrix X is... s The information is then fed into the next module, the Temporal Convolutional Network (TCN) unit, where multi-scale temporal information is extracted.

[0064] Dynamic graph convolution DG-GCN unit

[0065] This invention designs a Dynamic Graph Convolutional (DG-GCN) unit, which groups feature information and then uses multiple different types of graph convolution operations to extract multi-semantic spatial information. In DG-GCN, the input is first processed into K groups of feature information (each feature group consists of C / K channels) through a 1×1 convolution. Then, each feature group is coupled with the corresponding A... i Perform graph convolution operations. Finally, concatenate the K feature groups along the channel dimension, and then perform a 1×1 convolution operation on the merged feature information to generate an output with multi-level graph semantics. For example... Figure 2 As shown:

[0066] Where A i It consists of two learnable parts, first OA i It is constructed from the original data based on Formula 5. Secondly, this invention learns the data-related dynamic term DA. i and CA i The difference between these two dynamic terms is that DA i It is unrelated to the channel, while CA i It is related to the channel. Figure 2 As can be seen, these two dynamic terms are obtained by performing average pooling along the time dimension, given the input feature X.a and X b Then, based on different operations, the dynamic term DA is obtained. i and CA i .

[0067] Where DA i The definition is as follows:

[0068]

[0069] In the formula: DA i This is a dynamic term related to the non-channel information of the input features. X is the input feature matrix, and AvgPool t (·) represents average pooling along the time dimension, g represents a convolutional layer with a 1×1 kernel, T represents the matrix transpose operation, and Softmax(·) is the Softmax operation.

[0070] Where CA i The definition is as follows:

[0071] CA i =Tanh(Pairwise(X) a -X b (7)

[0072] Where: CA i For the dynamic terms related to the channels of the input features, Pairwise(X) a -X b ) represents calculating the pairwise distance between two matrices, and Tanh(·) represents the Tanh activation function.

[0073] Where A i It consists of three parts: the adjacency matrix OA of the graph topology. i and two dynamic terms DA i and CA i Where α and β are learnable parameters, as shown in the following formula:

[0074] A i =OA i +α×DA i +β×CA i (8)

[0075] In the formula: A i This is the dynamically expanded adjacency matrix, where OA i It is the topological adjacency matrix of the data-generated graph, and α and β are learnable model parameters.

[0076] Enhanced Channel Attention Mechanism SCA Unit

[0077] It comprises three sub-modules: Spatial Attention Unit (SAM), Temporal Attention Unit (TAM), and Channel Attention Unit (CAM). The multi-semantic feature map representation output from the previous layer's DG-GCN unit is sequentially fed into these three sub-modules, generating attention scores along different dimensions that capture spatial, temporal, and channel dependencies. These attention maps are then multiplied with the original feature maps to enhance the corresponding features. A residual connection is added to all attention modules to stabilize training.

[0078] Spatial Attention Subunit (SAM): This module helps the upper-level modules obtain a multi-semantic spatial graph representation X. G The spatial dependencies between key vehicle interactions are captured, and the calculation formula is shown below:

[0079] M s =σ(g s (AvgPool t (X G ))) (9)

[0080] X s =M S X G +X G (10)

[0081] In the formula: input The graph representation generated for the upper-level unit DG-GCN, which encompasses multiple semantic spatial information, It is the spatial attention score, AvgPool t This indicates that the feature matrix is ​​averaged along the time dimension. s This indicates a one-dimensional convolution in the spatial dimension. σ represents the sigmoid activation function. X s To capture the graph representation of spatial dependencies.

[0082] Temporal Attention Subunit (TAM): This module helps identify key information about the vehicle's position within a frame, thus successfully capturing temporal dependencies in the data. The calculation formula is as follows:

[0083] M t =σ(g t (AvgPool s (X s (11)

[0084] X st =M t X s +X s (12)

[0085] In the formula: This represents the time attention score. (AvgPool) sThis indicates that the feature matrix is ​​averaged along the spatial dimension. t This indicates that a one-dimensional convolution was performed in the time dimension. X st To capture the graph representation of spatiotemporal dependencies.

[0086] Channel Attention (CAM) subunit: This module helps the model enhance channel (vehicle position, speed) features based on spatial and temporal differences in the input samples. The calculation formula is as follows:

[0087] M t =σ(g c2 (δ(g c1 (AvgPool st (X st ))+g c1 MaxPool st (X st (13)

[0088] X stc =M c X st +X st (14)

[0089] In the formula: It is the channel attention score, AvgPool st MaxPool st These represent average pooling and max pooling along the spatial and temporal dimensions, respectively. c1 and g c2 These are two linear layers that operate along the channel dimension. δ represents the ReLU activation function. This means that the spatial and temporal differences of the input samples are captured, thereby enhancing the representation of the channel (vehicle position, speed) feature map.

[0090] TCN temporal convolutional layer unit

[0091] This invention generates new features after passing through DG-GCN and SCA attention mechanisms. The data is input into the temporal convolutional units. This invention sets the kernel size in the temporal convolutional units to (1×3), (1×5), and (1×7) to force them to perform multi-scale data extraction along the temporal dimension, thus extracting richer temporal information dependencies. By adding appropriate padding and stride, it is ensured that each layer has an output feature map of the expected size. Its dimensions are consistent with the input features.

[0092] III. Trajectory Prediction Module

[0093] This invention employs a GRU-based encoder-decoder module to predict the future trajectories of all observed vehicles. For example... Figure 4 As shown, the graph representation feature matrix of the last layer is fed into the encoder-decoder module to obtain the final future trajectory. In the first decoding step, the encoder's hidden features and all vehicle velocity data from the last observation time step are fed into the decoder to predict vehicle speeds. In subsequent decoding steps, the decoder uses its own hidden features and the predicted velocities of all objects from the previous time step as input for prediction. Then, all the expected future trajectories at time steps t are calculated. f The predicted speed is averaged, and then the average speed (Δx, Δy) is added to the vehicle's final historical position, successfully converting the result into (x, y) coordinates. Since the vehicle does not move at a constant speed, this invention adds residual connections between them (…). Figure 4 The dashed lines in the diagram are used to simulate the changes in the predicted velocity.

[0094] IV. Experimental Verification

[0095] 1. Dataset and Measurement Methods

[0096] This invention evaluates the proposed scheme using two existing trajectory datasets: NGSIM's I-80 (reference: J. Colyar and J. Halkias, "Us highway 80 dataset," Federal Highway Administration (FHWA), Tech.Rep. FHWA-HRT-07-030, 2007.) and US-101 (reference: "Us highway 101 dataset," Federal Highway Administration (FHWA), Tech.Rep. FHWA-HRT-07-030, 2007.) and the ApolloScape trajectory dataset (reference: Y. Ma, X. Zhu, S. Zhang, R. Yang, W. Wang, and D. Manocha, "Trafficpredict: Trajectory prediction for heterogeneous traffic-agents," arXiv preprint arXiv:1811.02146, 2018.).

[0097] The NGSIM datasets were captured at a frequency of 10 Hz over 45 minutes, divided into 15-minute intervals for light, moderate, and congested traffic conditions. This invention follows the approach of Deo et al. (references: N. Deo and MM Trivedi, “Multi-modal trajectory prediction of surrounding vehicles with maneuver based LSTMs,” in Proc. IEEE Intell. Vehicles Symp. (IV), Jun. 2018. and Deo N, Trivedi MM. Convolutional social pooling for vehicle trajectory prediction [C] / / Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. 2018.) to divide these two datasets into training and test sets. Each trajectory was segmented into 8-second segments, with the first 3 seconds used to observe historical trajectories and the remaining 5 seconds used for predicting real trajectories. For fair comparison, this invention downsampled each segment by a factor of 2, similar to Deo et al., resulting in 5 frames per second. The dataset was split into training, validation, and test sets in a ratio of 0.7:0.1:0.2.

[0098] The ApolloScape trajectory dataset was collected by operating the Apollo data collection vehicle in urban areas during peak hours (Reference: X. Huang, P. Wang, X. Cheng, D. Zhou, Q. Geng, and R. Yang, “The apolloscape open dataset for autonomous driving and its application,” 2018.). Traffic data was collected, including camera-based images and LiDAR-based point clouds, and object trajectories were calculated using object detection and tracking algorithms. In total, the trajectory dataset includes 53 minutes of training sequences and 50 minutes of test sequences, captured at a rate of 2 frames per second. Object ID, object type, object location, object size, and orientation angle are provided.

[0099] This invention employs two evaluation metrics, with the root mean square error (RMSE) used on the NGSIM dataset to evaluate model performance. The model used for construction is then trained and validated against a comparison model.

[0100]

[0101] In the formula: n is the sample size, and m is the number of vehicles observed in the scene. and y ij These represent the predicted value of the future trajectory and the actual value of the actual trajectory, respectively. Here, i = 1, 2, ..., m represents different vehicle nodes, and j = 1, 2, ..., T represents the prediction time frame.

[0102] In the ApolloScape trajectory dataset, this invention, following official requirements, uses Weighted Average Displacement Error (WSADE) and Weighted Final Displacement Error (WSFDE) to evaluate the model's performance. The Average Displacement Error (ADE) is the average Euclidean distance between all predicted and actual positions within the prediction time, and the Final Displacement Error (FDE) is the average Euclidean distance between the final predicted position and the corresponding actual position. The formulas are as follows:

[0103]

[0104] In the formula: m is the number of vehicles observed in each driving scenario, and T is the length of the prediction window. and y ij These are the predicted value of the future trajectory and the actual value of the actual trajectory, respectively, where i = 1, 2, ..., m represents different vehicle nodes, and j = 1, 2, ..., T represents the prediction time frame.

[0105]

[0106] In the formula: and y iT These are the predicted value of the future trajectory and the actual value of the actual trajectory, respectively, for the last frame.

[0107] WSADE=D v ×ADE v +D p ×ADE p +D b ×ADE b (18)

[0108] WSFDE=D v ×FDE v +D p ×FDE p +D b ×FDE b (19)

[0109] In the formula: where D v D p D bRelated to the reciprocals of the average speeds of vehicles, pedestrians, and bicycles in the dataset, the official values ​​for these three variables are set to 0.20, 0.58, and 0.22, respectively, with the subscript v representing vehicles, p representing pedestrians, and p representing bicycles or electric bicycles.

[0110] 2. Experimental Details

[0111] The experimental model was implemented using PyTorch 1.9.0. The model had 6 DGAC layers with a hidden size of 64, and a 2-layer GRU for trajectory prediction. The batch size was 128 samples, and the model was trained for 5 epochs. An Adam optimizer with a learning rate of 0.001 was used. The experimental hardware environment consisted of a Linux operating system, a 12vCPU Intel(R) Xeon(R) Silver 4214R CPU @ 2.40GHz, 90GB of memory, and an RTX 3080 Ti (12GB) for accelerating the neural network training process.

[0112] 3. Ablation test

[0113] This invention also conducts an ablation study on the proposed DG-GCN and the lightweight SCA attention mechanism. The baseline uses the GRIP++ model (Li, X.; Ying, X.; Chuah, MCGRIP++: Enhanced Graph-Based Interaction-Aware Trajectory Prediction for Autonomous Driving. arXiv2020, arXiv:2007.), which uses the graph construction method proposed in this invention for modeling. The remaining graph operation (GC) layers are set to 3 layers, the hidden layers are set to 64, and the prediction modules all use the GRU SeqtoSeq framework with consistent parameters. The ablation analysis results are shown in Table 1.

[0114] This invention sets the parameters of the DGAC module to be consistent with the Baseline. Specifically, DGC-Baseline replaces the graph convolution GCN module in the graph operation layer with the DG-GCN module proposed in this invention; GAC-Baseline introduces the SCA attention mechanism after the original graph operation layer GCN; and DGAC-Baseline adds the DG-GCN module proposed in this invention to replace the original graph convolution GCN module and adds a lightweight SCA attention mechanism.

[0115] Table 1 Ablation experiments of different modules

[0116]

[0117] Table 1 shows a good exploration of the performance of the proposed DGAC module. It can be seen that using different modules individually provides only a slight improvement, but combining the two modules significantly improves the model's performance in trajectory prediction. This indicates that combining the two modules can greatly enhance the graph representation capability, thereby improving prediction accuracy.

[0118] 4. Comparison and results of different models in the NGSIM dataset.

[0119] In the NGSIM dataset, the comparison models are as follows: V-LSTM is a single-input-single-output model that does not consider interaction with surrounding vehicles; CS-LSTM (reference: N. Deo and MM Trivedi, “Multi-modal trajectory prediction of surrounding vehicles with maneuver based LSTMs,” in Proc. IEEE Intell. Vehicles Symp. (IV), Jun. 2018.), a deep learning model that considers interaction between vehicles; MATF (reference: T. Zhao et al., “Multi-agent tensor fusion for contextual trajectory prediction,” in Proc. IEEE / CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2019.), STA (reference: LIN L, LI WZ, BI HK, et al. Vehicle trajectory prediction using LSTMs with spatial-temporal attention mechanisms[J]. IEEE Intelligent Transportation Systems Magazine, 2021.), SIT (reference: LI XL, XIAJ, CHEN XY, et al. SIT: a spatial interaction-aware transformer-based model for freeway trajectory prediction[J]. ISPRS International Journal of Geo-Information, 2022.), is a state-of-the-art method for trajectory prediction incorporating an attention mechanism. GRIP++ (Li, X.; Ying, X.; Chuah, MCGRIP++: Enhanced Graph-Based Interaction-Aware Trajectory Prediction for Autonomous Driving. arXiv 2020, arXiv:2007.) is a model that uses graph neural networks for trajectory prediction. This model predicts the trajectory of a single vehicle within a scene. In this embodiment, this model is reused to predict the trajectory of all vehicles. The comparison results are shown in Table 2.

[0120] Table 2 Comparison of prediction performance of different models on the NGSIM dataset

[0121]

[0122] As shown in Table 2, compared with other models, the DGAC-NET model proposed in this invention performs similarly to other models in long-term prediction and has a lower RMSE evaluation index in the short to medium time. Furthermore, in congested scenarios (with 10 or more surrounding vehicles), compared with the CS-LSTM model, the RMSE evaluation index of the predicted vehicle trajectory is reduced by 29% over a 5-second time period. This further illustrates that the higher the complexity of spatial interactions between vehicles in road scenarios, the stronger the graph representation capability of DGAC-NET, making it more suitable for vehicle trajectory prediction in complex scenarios.

[0123] The vehicle trajectory visualization effect of the NGSIM dataset is as follows: Figure 5 As shown, this figure represents the vehicle trajectory distribution in the current frame (3 seconds later in the historical timeframe). The left side of the figure shows the prediction scenario when there are fewer than 10 surrounding vehicles, and the right side shows the prediction scenario when there are more than 10 surrounding vehicles. Blue squares represent the target vehicle, black squares represent surrounding vehicles interacting with the target vehicle (less than 50 feet), and white squares represent vehicles not within the target vehicle's field of vision. The black line represents the historical trajectory from the previous three seconds, the blue line represents the actual trajectory position for the next 5 seconds, and the yellow dashed line represents the predicted future trajectory position. This invention shows that when there are no surrounding vehicles, the prediction will have a slight error, but when there are surrounding vehicles interacting, the vehicle trajectory prediction effect is better. Often, when there are no surrounding vehicles, the vehicle has a higher degree of freedom, which leads to a decrease in prediction accuracy. However, when there are more surrounding vehicles, the target vehicle has more constraints, making its trajectory prediction more accurate. When the cumulative number of vehicles in the scene exceeds 10, this invention refers to the scene as a congestion scene. As can be seen from Table 2, the higher the complexity of spatial interaction between vehicles in a road scene, the smaller the overall error in predicting the future trajectory of the vehicle. This further proves that the model has a stronger ability to extract spatial interaction information in complex environments.

[0124] 5. Comparison and Results of Different Models in the ApolloScape Trajectory Dataset

[0125] In the ApolloScape dataset, the following models are compared: TrafficPredict (Reference: Ma Y, Zhu X, Zhang S, et al. Trafficpredict: Trajectory prediction for heterogeneous traffic-agents[C] / / Proceedings of the AAAI conference on artificialintelligence.2019,33(01):6120-6127.), this model uses a hierarchical LSTM model for trajectory prediction, and is also the benchmark model for the ApolloScape dataset; StarNet (Reference: Zhu Y, Qian D, Ren D, et al. Starnet: Pedestrian trajectory prediction using deep neural network in startopology[C] / / 2019 IEEE / RSJ International Conference on Intelligent Robots and Systems(IROS).IEEE,2019:8075-8080.), this model uses a star topology to model the interaction between pedestrians and uses an LSTM network to predict the future trajectory of pedestrians; Transformer (Giuliari F, Hasan I, Cristani M, et al. Transformer networks for trajectory forecasting[C] / / 202025th international conference on pattern recognition(ICPR).IEEE,2021:10335-10342.), This model proposes to use the original Transformer to predict pedestrian trajectories, modeling each person individually without considering any complex interactions between people or between people and the scene. The comparison results are shown in Table 3.

[0126] Table 3 compares the model proposed in this invention with the official benchmark model TrafficPredict on the ApolloScape trajectory dataset. The proposed model is more accurate than the benchmark model, achieving an accuracy improvement of nearly 86%. Furthermore, if the data contains five features (position x-value, position y-value, object length, object width, and heading), adding new features—length and width—allows the model to better identify object types, resulting in an accuracy improvement of nearly 5% compared to the GRIP++ model. As shown in the table, this model significantly reduces the average displacement error (ADE) and final displacement error (FDE) in predicting vehicles, pedestrians, bicycles, and electric vehicles. Therefore, it can be seen that the module proposed in this invention can fully extract dynamic spatiotemporal information, thereby enhancing the model's trajectory prediction capabilities.

[0127] Table 3: Results in the ApolloScape trajectory dataset

[0128]

[0129] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A vehicle interaction perception trajectory prediction system based on graph convolution, characterized in that, It includes an input preprocessing module, a DGAC module, and a trajectory prediction module connected in sequence; The collected vehicle data is input into the preprocessing module to create a feature matrix and graph topology. Then, it is processed by a multi-layer DGAC module to extract multi-semantic graph information and selectively focus on the differences between spatial, temporal, and channel dimensions and key information using attention. Finally, it is processed by the trajectory prediction module to generate the predicted trajectory of the vehicle in the surrounding environment at a certain time in the future. The DGAC module includes a dynamic graph convolutional DG-GCN unit, an enhanced channel attention mechanism (SCA) unit, and a temporal convolutional layer (TCN) unit. The Dynamic Graph Convolutional (DG-GCN) unit groups feature information and then uses multiple different types of graph convolution operations to extract multi-semantic spatial information. Its input information first goes through a Convolution processing The feature information of the group, each feature group consists of Channel composition; then each feature group is associated with the corresponding Perform graph convolution operations; finally, Each feature group is concatenated along the channel dimension, and the merged feature information is then processed... Convolution operations generate outputs with multi-level graph semantics; The Enhanced Channel Attention (SCA) unit comprises three sub-units: Spatial Attention (SAM), Temporal Attention (TAM), and Channel Attention (CAM). The input feature map is sequentially fed into the three sub-units to generate attention maps along different dimensions. The attention maps are multiplied by the original feature map to enhance the corresponding features. The temporal convolutional layer (TCN) unit is the new feature generated after passing through the DG-GCN and SCA units. The data is fed into a multi-scale temporal convolutional layer (TCN); the kernel size of the TCN is set to (1×3), (1×5), and (1×7) to force them to process the data along different time dimensions.

2. The vehicle interaction perception trajectory prediction system based on graph convolution according to claim 1, characterized in that, The data input by the input preprocessing module It is all observed vehicles at historical time steps Position and velocity in the middle, i.e. (1) in, (2) Indicates in Within the range of the time observed Vehicle position coordinates and speed ; The trajectory prediction module outputs data It is all observed vehicles from time arrive The prediction of future speed, among which It is the range of future prediction time steps; (3) in, (4)。 3. The vehicle interaction perception trajectory prediction system based on graph convolution according to claim 2, characterized in that, The input preprocessing module includes creating a feature matrix and a graph topology, and the specific implementation process is as follows: Creating a feature matrix: Given a traffic scenario, assuming in the past... Observed at time step The vehicle data is preprocessed into a three-dimensional tensor. Let C represent the coordinates of the vehicle within the scene at different time steps. and speed Features, and normalize all coordinates and velocities to Scope; Constructing a graph topology: for each observed time step , an undirected graph is constructed , wherein nodes represent vehicles at time and edges represent spatial interactions between vehicles; wherein the node set at a time step is defined as , wherein the edge set of is defined in the time step as ; At each time step It defines that spatial interaction will only occur at a distance. Between the two cars inside, The threshold is configurable, and it defines that vehicles must belong to adjacent lanes or the same lane, that is... ; Therefore, at each time step of Adjacency matrix Defined in the following way: (5) Assumption It is a time step The number of currently observed vehicles, and It must be less than , It is a manually set time step The maximum number of vehicles in the interior scene; therefore, given the observed vehicles... trajectory and time step To obtain the adjacency matrix as described above. Finally, the feature matrix and adjacency matrix containing time information Input them into the model together.

4. The vehicle interaction perception trajectory prediction system based on graph convolution according to claim 3, characterized in that, The trajectory prediction module employs a GRU-based encoder-decoder module to predict the future trajectories of all observed vehicles. Specifically, in the first decoding step, the encoder's hidden features and all vehicle velocity data from the last observation time step are fed into the decoder to predict vehicle trajectories. In subsequent decoding steps, the decoder uses its own hidden features and the predicted velocities of all objects from the previous time step as input for prediction. Then, all the expected future trajectories are calculated. The results are averaged, and then the average speed is... Including the vehicle's final historical location, the result was successfully converted. coordinate.

5. A method for predicting vehicle interaction perception trajectory based on graph convolution, characterized in that, It is a method applied to the system described in any one of claims 1-4.

Citation Information

Patent Citations

  • Vehicle track prediction method based on dynamic interaction graph convolution

    CN114802296A

  • Multi-modal space-time model for accurate motion prediction based on visual fusion

    CN117315603A