Future interactive perception vehicle trajectory prediction method based on graph converter

The future interactive perception vehicle trajectory prediction method based on the graph converter combines future and historical interaction models, integrates perception safety indicators and map information, solves the problem of future trajectory prediction deviation in existing methods, achieves accurate and robust trajectory prediction, and reduces the risk of vehicle collision.

CN120716754APending Publication Date: 2025-09-30KUNMING UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511160910.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing interaction-based vehicle trajectory prediction methods mainly focus on the interaction behaviors in historical trajectories, which leads to mismatch between future trajectories and historical trajectories, large prediction deviations, and inability to accurately capture future interaction behaviors between vehicles.

Method used

A future interaction perception vehicle trajectory prediction method based on graph converter is adopted. Through the vehicle state embedding module, dual-view interaction module, map information integration module and feature fusion module, the future interaction model and the historical interaction model are combined, and the attention weight is adjusted using the perception safety index to capture the future interaction features between vehicles and integrate map information for feature fusion.

Benefits of technology

The accuracy and stability of trajectory prediction have been significantly improved, and the future trajectory of the vehicle can be accurately predicted, the risk of collision is reduced, and the model's social information perception ability is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120716754A_ABST
    Figure CN120716754A_ABST
Patent Text Reader

Abstract

The invention discloses a future interactive perception vehicle trajectory prediction method based on a graph converter, and relates to the technical field of vehicle trajectory prediction, and the method comprises a vehicle state embedding module which is used for extracting the hidden state of an original trajectory as a vehicle state vector; the double-view interaction module comprises a future interaction model, implicitly represents the trend of future interaction behaviors by utilizing the safety sensed by a human driver, and is also provided with a parallel historical interaction model; the map information integration module is used for extracting map features related to trajectory prediction through a gating selection mechanism; and the feature fusion module takes the multi-dimensional features of the target vehicle as input, carries out feature fusion through an edge offset multi-head attention mechanism, and further generates a prediction trajectory of the target vehicle through a linear layer. Therefore, by adopting the future interactive perception vehicle trajectory prediction method based on the graph converter, the future interactive behavior between the vehicles can be effectively captured by utilizing the perception safety index, and stable and accurate trajectory prediction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle trajectory prediction, and in particular to a future interactive perception vehicle trajectory prediction method based on a graph converter. Background Art

[0002] With the rapid development of communication technology, the deep integration of physical objects and the internet has driven ubiquitous information exchange. This paradigm shift, known as the Internet of Things (IoT), has profoundly impacted diverse sectors, including residential environments, industrial production, and transportation systems. The Internet of Vehicles (IoV), a specialized branch of IoT technology, connects vehicles to in-vehicle networks, enabling real-time sharing of driving status and road traffic conditions. Within this framework, trajectory prediction becomes a critical task, laying the foundation for downstream applications such as decision support and motion planning. By accurately predicting vehicle trajectories, connected cars can proactively adjust driving strategies, make informed decisions, and reduce collision risks, significantly improving road safety in hybrid autonomous driving scenarios.

[0003] In recent years, a large number of research results have emerged in the field of vehicle trajectory prediction. Early research primarily focused on modeling the dynamics of individual vehicles, emphasizing the inherent characteristics of vehicle mechanics. These physics-based methods aim to capture the evolving patterns of historical trajectories. However, such methods often ignore the constraints imposed by road conditions and surrounding vehicles. To address these limitations, mobility-based methods have been introduced to enhance traditional physics-based models by incorporating road geometry and the influence of neighboring vehicles. Despite some progress, mobility-based methods still fail to accurately capture the complex interactions between vehicles. With the emergence of deep learning as a transformative paradigm for pattern recognition and data analysis, interaction-based methods have become a promising research direction in trajectory prediction, capable of modeling complex inter-vehicle interactions. Some studies represent connected vehicles as graph structures and utilize graph neural networks (GNNs) to model these interactions, thereby capturing deep interaction features. Other studies have expanded on this foundation, considering not only inter-vehicle interactions but also interactions with map elements, employing Transformer networks or GNNs to expand the model's receptive domain and extract richer high-level features for trajectory prediction. In summary, the interaction-based approach has significantly superior prediction performance than the physics-based and maneuver-based approaches.

[0004] However, existing interaction-based methods primarily focus on modeling interactions within historical trajectories. This limited representation restricts the timeframe of predictions, leading to mismatches between future and historical trajectories and prediction bias. Therefore, a future interaction modeling solution is urgently needed that can implicitly infer and capture future interactions between vehicles, improving the reliability and accuracy of trajectory predictions. Summary of the Invention

[0005] The purpose of the present invention is to provide a future interaction-aware vehicle trajectory prediction method based on a graph converter, which can simultaneously model individual dynamics, inter-vehicle interactions, and map information, and use the safety perceived by human drivers to implicitly characterize the trends of future interactive behaviors, thereby achieving accurate, robust, and social information-aware trajectory prediction.

[0006] To achieve the above objectives, the present invention provides a future interactive perception vehicle trajectory prediction method based on a graph converter, comprising a vehicle state embedding module, a dual-view interaction module, a map information integration module, and a feature fusion module;

[0007] The vehicle state embedding module represents the original trajectory of the vehicle as a multivariate time series. The embedding layer converts the original trajectory of the vehicle into a feature sequence and uses the LSTM network to extract the hidden state of the future time interval as the vehicle state vector.

[0008] The dual-view interaction module includes a parallel future interaction model and a historical interaction model:

[0009] The future interaction model integrates collision time, exposure time collision time, and time-integrated collision time into future interaction indicators, which are used to adjust the attention weights of the target vehicle and surrounding vehicles. It also includes a guided graph transformer with two message propagation mechanisms. It integrates future interaction indicators into the attention calculation process, thereby converting potential future interactions into explicit node features to obtain the future interaction features of the target vehicle.

[0010] The historical interaction model, which takes the vehicle state vector as input and extracts the historical interaction features of the target vehicle through a guided graph transformer that does not include future interaction indicators;

[0011] The map information integration module extracts map features related to trajectory prediction through a gated selection mechanism based on the channel attention operator;

[0012] The feature fusion module combines the node features, future interaction features, historical interaction features, and map features of the target vehicle into a sequence, and performs feature fusion through an edge-biased multi-head attention mechanism, and then generates the predicted trajectory of the target vehicle through a linear layer.

[0013] Furthermore, in the future interaction model, two message propagation mechanisms include from surrounding vehicles to target vehicle and from target vehicle to surrounding vehicles.

[0014] Furthermore, during the message propagation from surrounding vehicles to the target vehicle, the node characteristics of the target vehicle are updated as follows:

[0015]

[0016] in,

[0017]

[0018] Where, the superscript l represents the number of layers of the guided graph transformer, represents the node features of the target vehicle v0 at the lth layer, represents the set of surrounding vehicles of the target vehicle υ0, represents the adjacent vehicles υ on the lth layer m Node features, W t represents the learnable parameter, α m,0 represents the adjacent vehicle v m The attention weight of the target vehicle υ0; a t Obtained by the feedforward neural network, ⊙ represents the Hadamard product, Φ t Represents the LeakyReLU activation function.

[0019] Furthermore, during the message propagation from the target vehicle to the surrounding vehicles, the node features of the surrounding vehicles are updated as follows:

[0020]

[0021] in,

[0022]

[0023] Where, α 0,m Indicates the target vehicle υ0's relationship with the adjacent vehicle υ m The attention weight, a s It is obtained by the feedforward neural network with message propagation mechanism.

[0024] Furthermore, the gated selection mechanism based on the channel attention operator includes calculating the attention score:

[0025]

[0026] Where β is the attention score, W m represents the learnable parameters, f channel represents flattened convolution, Φ map represents the softmax layer;

[0027] Update the flattened convolution:

[0028] f′ channel =Φ u (f channel ⊙β+f channel );

[0029] Where f′ channel is the updated flattened convolution, Φ u Represents an expand operation.

[0030] Furthermore, the edge-biased multi-head attention mechanism includes updating the multi-dimensional features of the target vehicle:

[0031]

[0032]

[0033] Where, γ i,j represents the attention coefficient of the jth feature of the i-th attention head, represents the features of the i-th attention head after update, c j represents the jth feature in the sequence, W Q 、W K 、W V Represent the linear transformation matrices of query, key, and value respectively, They represent the edge bias for distinguishing the connection relationship between different features, d e represents the hidden layer dimension, γ i,j represents the attention coefficient between the i-th and j-th features of the target vehicle, represents the updated i-th feature of the target vehicle.

[0034] Therefore, the present invention adopts the above-mentioned future interactive perception vehicle trajectory prediction method based on the graph converter, which has the following technical effects:

[0035] (1) The present invention sets up a dual-view interaction module. While extracting the historical interaction relationship between the target vehicle and surrounding vehicles, it adjusts the attention weight of the interaction graph by introducing a safety perception indicator, indirectly inferring the future interaction behavior, and captures the high-order correlation between state vectors through a guided graph transformer, significantly improving the accuracy and stability of trajectory prediction.

[0036] (2) The present invention integrates map information into the model and adopts a gated selection mechanism to obtain driving scene context information. At the same time, it adopts a feature fusion module and a side-biased multi-head attention mechanism to integrate feature information from different perspectives, thereby enhancing the model's ability to identify key connections between features, thereby achieving robust and accurate trajectory prediction.

[0037] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 Schematic diagram of advantages of future interaction modeling in an embodiment of a future interaction perception vehicle trajectory prediction method based on a graph converter;

[0039] Figure 22. Schematic diagram of the overall framework of BaTF for vehicle trajectory prediction in an embodiment of a future interactive perception vehicle trajectory prediction method based on a graph converter;

[0040] Figure 3 2. Schematic diagram of two-dimensional TTC calculation in an embodiment of a future interactive perception vehicle trajectory prediction method based on a graph converter;

[0041] Figure 4 is a schematic diagram of a guided graph converter in an embodiment of a future interactive perception vehicle trajectory prediction method based on a graph converter;

[0042] Figure 5 2. A schematic diagram of a gated selection mechanism based on a channel attention operator in an embodiment of a future interactive perception vehicle trajectory prediction method based on a graph converter;

[0043] Figure 6 Schematic diagram of an edge-biased multi-head attention mechanism in an embodiment of a future interactive perception vehicle trajectory prediction method based on a graph converter;

[0044] Figure 7 is the prediction accuracy of different models on the HighD dataset in an embodiment of the future interactive perception vehicle trajectory prediction method based on the graph converter, where (a) is the ADE indicator and (b) is the FDE indicator;

[0045] Figure 8 is the prediction accuracy of different models on the INTERACTION dataset in an embodiment of the future interactive perception vehicle trajectory prediction method based on the graph converter, where (a) is the ADE index and (b) is the FDE index;

[0046] Figure 9 Figure 1 is the result of ablation experiments on different datasets in an embodiment of a future interactive perception vehicle trajectory prediction method based on a graph converter, where (a) is the ablation result on the HighD dataset, and (b) is the ablation result on the INTERACTION dataset.

[0047] Figure 10 The effect of different hyperparameters on trajectory prediction accuracy in the embodiment of the future interactive perception vehicle trajectory prediction method based on the graph converter, where (a) is the effect of the number of BGT layers L, and (b) is the number of convolution kernels n. c The impact of (c) is the number of attention heads h d the impact of;

[0048] Figure 11 : This is a comparison chart of the BaTF predicted trajectory, HyperMTP predicted trajectory, and actual trajectory in different scenarios in an embodiment of the future interactive perception vehicle trajectory prediction method based on the graph converter, where (a) is a highway, (b) is a ramp, and (c) is a roundabout. DETAILED DESCRIPTION

[0049] The present invention can be explained in more detail by the following examples. The purpose of disclosing the present invention is to protect all changes and improvements within the scope of the present invention. The present invention is not limited to the following examples.

[0050] Example 1

[0051] like Figure 1 As shown, under traditional interaction methods, two vehicles traveling in different directions are predicted to maintain their original driving state, resulting in trajectory distortion and even potential collisions in real-world traffic scenarios. To address this limitation, this invention captures future vehicle interactions through the perspective of perceived safety—the current safety level reflects the psychological state that influences the driver's future behavior, which in turn affects the way vehicles interact with each other. A dual-view interaction module approach is designed to integrate historical and future vehicle interaction information.

[0052] like Figure 2 As shown in the figure, the present invention provides a future interactive perception vehicle trajectory prediction method based on a graph converter, referred to as BaTF, which includes four parts: vehicle state embedding module, dual-view interaction module, map information integration module, and feature fusion module, as follows:

[0053] 1. Vehicle status embedding module:

[0054] The original trajectory of vehicle i is represented as a multivariate time series In a given time interval t′, each multivariate data point is defined as are the horizontal and vertical coordinates respectively, l i 、ω i are the length and width of vehicle i, respectively. In order to initialize these variables into learnable features, this embodiment uses an embedding layer:

[0055]

[0056] Where W emb 、b emb are the learnable weight matrix and bias vector, Φ emb is a nonlinear activation function. After this transformation, the original trajectory of the vehicle is converted into a feature sequence

[0057] To capture individual dynamic characteristics, this example uses an LSTM network to model the temporal dependencies in the sequence:

[0058]

[0059] Where, Φ LSTM represents the LSTM network, θv Represents a learnable parameter. The feature sequence obtained by the embedding layer from the retrospective time window T As the input of the LSTM network, the hidden state v at the near future time interval t is extracted i , as the vehicle state vector required by subsequent modules.

[0060] 2. Dual-view interaction module:

[0061] Unlike existing prediction models that focus solely on historical vehicle interaction trajectories, this embodiment integrates future interaction information, providing a more comprehensive and predictive approach to trajectory prediction. Specifically, the interaction graph is defined as a unified directed graph, where the target vehicle node connects to all surrounding vehicle nodes. Secondly, because the relationships in the graph optimally reflect different perspectives of vehicle interactions, this embodiment directly uses the same vehicle state vector as input to initialize the node features of both future and historical interaction graphs, differing only in the setting of the respective relationship weights.

[0062] 1. Future interaction model:

[0063] To capture the future interactions between the target vehicle and surrounding vehicles, this embodiment draws on the concept of perceived safety by human drivers. In the field of traffic safety, perceived safety is often characterized by three risk metrics: time to collision (TTC), exposure time to collision (TET), and time-integrated time to collision (TIT). These metrics describe the one-way interaction between any two vehicles, as follows:

[0064] (1) Time to Collision (TTC), which indicates the remaining time until a collision occurs if the two vehicles continue to move forward along their current driving trajectories. This indicator serves as an immediate risk assessment tool and has an early warning function. Since traditional TTC is only applicable to one-dimensional longitudinal driving scenarios, this embodiment extends it to two-dimensional space to adapt to more complex driving scenarios and demonstrates the rationality of the proposed two-dimensional TTC measurement method. The two-dimensional TTC calculation formula for the target vehicle i and the adjacent vehicle j within the time interval t is:

[0065]

[0066] in,

[0067]

[0068] Where, represents the relative distance between the target vehicle i and the adjacent vehicle j, express The rate of change, Represent the two-dimensional coordinates and speeds of the target vehicle i and the adjacent vehicle j respectively.

[0069] like Figure 3 As shown, the relative distance between the target vehicle i and the adjacent vehicle j is The relative speed between the two vehicles is represented by the yellow vector. On the other hand, we need to calculate the relative position unit vector and relative speed The inner product between the two vectors is greater than 90°, but the angle β between the two vectors is greater than 90°, that is, cosβ<0. To address this problem, this embodiment uses a negative sign in the two-dimensional TTC calculation formula.

[0070] (2) Time to collision (TET), which measures the duration that the TTC remains within the safety threshold over a specified period of time:

[0071]

[0072] Where, τ sc Indicates the minimum time step for measuring the TTC value without changing, which is set to 0.1s here; δ i (t k ) is a binary variable:

[0073]

[0074] Where, The threshold of the safety TTC is set to 2.5s in this embodiment.

[0075] (3) Time-integrated Time to Collision (TIT): Since TET treats all TTC values ​​within the threshold equally, this indicator cannot distinguish different risk levels. To address this limitation, TIT provides a more detailed representation of safety levels by combining TTC curves:

[0076]

[0077] The higher the values ​​of TTC, TET and TIT, the greater the potential collision risk exposure, which leads to a decrease in perceived safety. In this embodiment, these three perceived safety indicators are integrated into a comprehensive future interaction indicator To adapt to the changes in the weights of different relationships in the future interaction graph.

[0078] In addition, this embodiment also proposes a guided graph transformer (BGT) to transform potential future interactions into explicit node features, thereby promoting information dissemination and extracting potential interaction features. Figure 4 As shown, the blue and green nodes represent the target vehicle and surrounding vehicles, respectively; the guided graph transformer contains two message propagation mechanisms (represented by lines of different colors and types): from surrounding vehicles to the target vehicle and from the target vehicle to surrounding vehicles.

[0079] For the message propagation from surrounding vehicles to target vehicle, assume that the target vehicle is v0 and the surrounding vehicles are Then the node feature v0 of the target vehicle can be updated as follows:

[0080]

[0081] Where, the superscript l represents the number of layers of the guided graph transformer, represents the adjacent vehicle v on the lth layer m Node features, W t represents the learnable parameter, α m,0 represents the adjacent vehicle v m Attention weight for target vehicle v0. In order to integrate the perceived safety index and effectively model the future correlation between the target vehicle and surrounding vehicles, this embodiment incorporates the future interaction index into the attention calculation process:

[0082]

[0083] Where a t Obtained by the feedforward neural network, ⊙ represents the Hadamard product, Φ t Represents the LeakyReLU activation function.

[0084] For message propagation from the target vehicle to surrounding vehicles, the node characteristics of the surrounding vehicles are updated as follows:

[0085]

[0086] Where, α 0,m Indicates the target vehicle v0's response to the adjacent vehicle v m The attention weight, a s is obtained by a feedforward neural network customized for message propagation. In other embodiments, a s with a t The same feed-forward neural network can be used to obtain

[0087] Through the iterative message propagation mechanism, the node features of the target vehicle v0 output by the last layer of the guided graph transformer are used as the low-dimensional future interaction representation of the target vehicle v0.

[0088] 2. Historical interaction model:

[0089] Given the importance of historical vehicle states in trajectory prediction, this embodiment uses parallel branches of the historical interaction model and the future interaction model. When constructing the historical interaction model, the logic of the future interaction model is followed, with the difference that the input of the guided graph transformer is composed of the hidden states generated by the vehicle state embedding module. Specifically, this embodiment does not include the perceived safety index e jThe general guided graph transformer of BGT , then the historical interaction model is expressed as:

[0090]

[0091] Where, Represent the historical characteristics of target vehicles in the lth and l-1th layers respectively; Represent the historical features of the surrounding vehicles in the lth and l-1th layers respectively. In the first layer, the vehicle state is embedded in the extracted initial feature sequence

[0092] In summary, the future interaction features of the target vehicle are obtained through the dual-view interaction module and historical interaction features L is the number of layers of the guided graph transformer.

[0093] 3. Map information integration module:

[0094] Given the significant impact of map elements on the future driving trajectory of the target vehicle, this embodiment integrates map information into the model to improve trajectory prediction accuracy. Although convolutional neural networks (CNNs) are often used to extract map features, they often produce non-negligible noise interference, leading to learning bias and thus reducing prediction accuracy. To this end, this embodiment proposes a gated selection mechanism based on a channel attention operator, which can adaptively highlight the driving scene features most relevant to trajectory prediction. Figure 5 As shown in Figure 2, the calculation formula for the attention score is:

[0095]

[0096] Where W m represents the learnable parameters, f channel represents flattened convolution, Φ map Represents the softmax layer.

[0097] Then, update the flattened convolutional layer as follows:

[0098] f′ channel =Φ u (f channel ⊙β+f channel );

[0099] Among them, f′ channel is the updated flattened convolution, Φ u Represents an expand operation.

[0100] Finally, the learned map features are used express.

[0101] 4. Feature fusion module:

[0102] Integrate the outputs of the above three modules into a sequence Feature fusion is performed through the edge-biased multi-head attention mechanism to effectively capture the correlation between these features, thereby improving the trajectory prediction effect.

[0103] like Figure 6 As shown, the multi-dimensional features of the target vehicle v0 are updated as follows:

[0104]

[0105] Where, γ i,j represents the attention coefficient of the jth feature of the i-th attention head, represents the features of the i-th attention head after update, c j represents the jth feature of the sequence, W Q 、W K 、W V Represent the linear transformation matrices of query, key, and value respectively, They represent the edge bias for distinguishing the connection relationship between different features, d e represents the hidden layer dimension.

[0106] Then, the updated four features are concatenated into a fused feature vector Used for trajectory prediction.

[0107] Fusion feature vector Generate the predicted trajectory of the target vehicle through the linear layer In model training, this embodiment uses mean square error (MSE) as the loss function, supplemented by a regularization term to alleviate the overfitting problem:

[0108]

[0109] Where, represents the true trajectory of the target vehicle, and λ is the regularization term.

[0110] Example 2

[0111] To evaluate the effectiveness of our method, BaTF, we conducted experiments on two mainstream real-world trajectory datasets: the HighD dataset and the INTERACTION dataset. To ensure fair comparison with other trajectory prediction models, we strictly adhered to the pre-defined data partitioning schemes provided by each dataset. During the experiment, the trajectory data was segmented into 10-second segments, with the first 6 seconds serving as model input and the remaining 4 seconds used for prediction calculations.

[0112] In addition, this embodiment evaluates BaTF by comparing it with a series of cutting-edge benchmark models: LaneGCN introduces a multi-view expansion mechanism based on the graph convolution algorithm to accurately capture the complex associations in the lane topology; TNT, as a target-driven prediction model, can not only encode the interaction relationship between vehicles, but also integrate the interaction between vehicles and map elements; HiVT adopts a translation-invariant driving scene representation method to capture the interaction between vehicles from a global and local perspective; HDGT constructs a graph transformation framework to realize the encoding of various interaction semantics in heterogeneous driving graphs; HEAT adopts an edge-enhanced attention mechanism to model individual dynamic behaviors and interaction relationships through heterogeneous graphs. The HCAGCN extends HEAT by incorporating interaction dynamics into spatiotemporal dynamic graphs for encoding. HyperMTP uses hyper-relational graphs to simulate high-order interactions between vehicles, pedestrians, and map elements. FFINet incorporates future feedback into interaction modeling, enabling this method to capture potential future interactions through cross-temporal aggregation. TP-EGT develops a collision-aware graph transformer to capture complex social interactions between traffic agents. I2T integrates long-term dependencies between intention cues and future motion. EPHGT introduces an edge-enhanced heterogeneous graph transformer to implement priority-based feature aggregation to model interactions. Furthermore, two commonly used evaluation metrics were used in the experiments: average displacement error (ADE), which evaluates the accuracy of trajectory prediction from a holistic perspective by calculating the average Euclidean distance between the predicted trajectory and the true trajectory; and final displacement error (FDE), which focuses on the prediction error of the final time interval, which is particularly important for downstream tasks such as connected vehicle motion planning.

[0113] In the specific implementation process, BaTF was implemented using the PyTorch framework. The training process consisted of 300 cycles, and an early stopping strategy was adopted. The Adam optimizer was used with an initial learning rate of 0.001 and a 20% decay every 50 cycles. The three key hyperparameters - the number of BGT layers (L), the number of convolution kernels in the map information fusion module (n c ) and the number of attention heads h in the feature fusion module d They are set to 4, 20, and 4 respectively. It is worth noting that no data augmentation techniques are used in the experiment.

[0114] like Figure 7 and Figure 8As shown in the figure, all models perform better on the HighD dataset than on the INTERACTION dataset. This difference is mainly due to the significantly higher complexity of the traffic scenes in the INTERACTION dataset, which makes accurate trajectory prediction more challenging. Moreover, the models consistently outperform the FDE metric on the ADE metric, indicating that while predicting the final position of the vehicle is crucial, achieving accurate prediction is more difficult. Notably, the performance of all models drops off precipitously as the prediction time horizon increases. This highlights the importance of examining predictions over longer time frames, which is crucial for downstream connected vehicle tasks such as decision-making and motion planning to ensure traffic safety.

[0115] In addition, when comparing different models, it was found that LaneGCN and TNT performed worse than other baseline models. This is mainly due to their limited ability to represent interactions between vehicles, which restricts these models. HDGT and HiVT have similar network structures for capturing interactive relationships, but HDGT performs better. This performance difference is mainly due to HDGT's excellent modeling ability of map information - it significantly improves the efficiency of interactive message passing by integrating map elements into the global graph. HCAGCN, as an extended version of HEAT, shows better performance than HEAT due to its more powerful interaction dynamics capture ability. FFNet is the only baseline model that explicitly considers future vehicle interactions. However, its performance lags behind HyperMTP. This may be because its shallow model architecture fails to fully exploit the input data to support future interactions, and its information interaction method is limited to a short-range field of view. TPEGT, I2T, and EPHGT are three multimodal trajectory prediction methods that integrate driving knowledge and aim to enhance the representation ability of vehicle-to-vehicle interactions.

[0116] However, their performance is still insufficient compared to the BaTF of the present invention. This demonstrates that this embodiment, by introducing the concept of "perceptual driving safety" to characterize future interactions and employing a graph transformer architecture, is able to capture deep connections between vehicle states, effectively transforming driving knowledge into contextual information for future driving behavior, significantly improving trajectory prediction accuracy. Furthermore, this embodiment also found that the BaTF model exhibited excellent prediction stability across different prediction timeframes, significantly outperforming all baseline models. This improves prediction accuracy while enhancing prediction robustness.

[0117] On the other hand, this example also compares the inference time of BaTF and other baseline models, as shown in Table 1.

[0118] Table 1 Inference time of different models on two datasets

[0119]

[0120]

[0121] As can be seen, the BaTF model not only achieves prediction times comparable to the current state-of-the-art baseline methods, but also outperforms the best-performing HyperMTP baseline model in inference speed. This advantage highlights the high efficiency of BaTF, enabling real-time trajectory prediction. This efficiency is due to the careful selection of model components and the optimization of hyperparameters—a scientific configuration that effectively reduces redundant computation.

[0122] This example also compares the predicted trajectory of BaTF with the best performing baseline model HyperMTP, and combines it with real trajectory data of safety-critical scenarios for visualization analysis. Figure 11 As shown, in three typical traffic scenarios (freeway, ramp, and roundabout), the trajectories predicted by HyperMTP often closely match the motion trajectories of two vehicles, easily leading to potential collisions. In contrast, the BaTF model generates trajectories that are closer to the actual driver, effectively avoiding these accidents. This demonstrates that modeling future vehicle interactions not only improves prediction accuracy but also enhances trajectory realism, ultimately enabling intelligent trajectory prediction with social awareness capabilities.

[0123] Example 3

[0124] To evaluate the effectiveness of key components of BaTF, this example conducts ablation studies by creating multiple model variants. These variants are designed to exclude or modify specific components of BaTF, thereby quantifying the technical effects of these components.

[0125] (1) To evaluate the effectiveness of future interaction modeling, a variant model without future interaction function (w / oFI) was designed, which removes the future interaction process from the dual-view interaction module. Figure 9 As shown in Figure 3, compared to the full BaTF model, this variant performs significantly worse and even lags behind the best baseline model, HyperMTP. These results highlight the important role of future interaction modeling in improving prediction accuracy and further verify the necessity of incorporating learned future interaction features into the trajectory prediction framework.

[0126] (2) To evaluate the effectiveness of BGT, a variant model without BGT was designed by replacing BGT in the dual-view interaction module with a regular GAT. The performance of this variant is still inferior to the full BaTF model, which highlights the advantages of BGT in capturing interactions and improving trajectory prediction accuracy. However, compared with the best-performing baseline model HyperMTP, this variant still outperforms, fully demonstrating the inherent superiority and robustness of the BaTF architecture - even with adjustments to specific components, it can still maintain excellent performance.

[0127] (3) To evaluate the effectiveness of map information, a variant model without map information was designed, completely excluding the map information integration module. The results show that map information significantly improves the prediction performance. This improvement is intuitive because the ability to perceive the surrounding driving environment plays a key role in accurate trajectory prediction.

[0128] (4) To evaluate the effectiveness of edge bias, a variant model without EB is designed by removing the edge bias of the multi-head attention mechanism in the feature fusion module. The results show that the introduction of edge bias can effectively integrate feature information from different perspectives and enhance the model's ability to identify key connections between features, thereby achieving robust and accurate trajectory prediction.

[0129] Example 4

[0130] This paper also explores the impact of three key hyperparameters in BaTF: the number of BGT layers (L), the number of convolution kernels (n c ) and the number of attention heads (h d ).

[0131] like Figure 10 As shown in (a), as the horizontal axis L increases, the model performance gradually improves, reaching a peak at L = 3, and then tends to stabilize. To balance performance and computational efficiency, this embodiment sets the L value of BaTF to 4.

[0132] like Figure 10 As shown in (b), the number of convolution kernels (n c ) has a similar effect to the hyperparameter L, increasing n c The prediction performance can be improved, and the selection of the optimal value requires a balance between efficiency and performance.

[0133] like Figure 10 As shown in (c), increasing the number of attention heads (h d ) can continuously improve the prediction performance, where h d =4 is the best. But when h d When it exceeds 4, the performance starts to degrade, which is due to overfitting of the model.

[0134] Therefore, the present invention adopts the above-mentioned future interaction perception vehicle trajectory prediction method based on graph converter, which significantly improves the credibility and accuracy of trajectory prediction by implicitly inferring and capturing future interaction behaviors between vehicles, making the predicted trajectory closer to the actual situation, thereby effectively reducing the risk of collision.

[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A future interactive perception vehicle trajectory prediction method based on a graph converter, characterized by: It includes vehicle status embedding module, dual-view interaction module, map information integration module, and feature fusion module; The vehicle state embedding module represents the original trajectory of the vehicle as a multivariate time series. The embedding layer converts the original trajectory of the vehicle into a feature sequence and uses the LSTM network to extract the hidden state of the future time interval as the vehicle state vector. The dual-view interaction module includes a parallel future interaction model and a historical interaction model: The future interaction model integrates collision time, exposure time collision time, and time-integrated collision time into future interaction indicators, which are used to adjust the attention weights of the target vehicle and surrounding vehicles. It also includes a guided graph transformer with two message propagation mechanisms. It integrates future interaction indicators into the attention calculation process, thereby converting potential future interactions into explicit node features to obtain the future interaction features of the target vehicle. The historical interaction model, which takes the vehicle state vector as input and extracts the historical interaction features of the target vehicle through a guided graph transformer that does not include future interaction indicators; The map information integration module extracts map features related to trajectory prediction through a gated selection mechanism based on the channel attention operator; The feature fusion module combines the node features, future interaction features, historical interaction features, and map features of the target vehicle into a sequence, and performs feature fusion through an edge-biased multi-head attention mechanism, and then generates the predicted trajectory of the target vehicle through a linear layer.

2. The future interactive perception vehicle trajectory prediction method based on graph converter according to claim 1 is characterized in that In the future interaction model, two message propagation mechanisms include from surrounding vehicles to target vehicle and from target vehicle to surrounding vehicles.

3. The future interactive perception vehicle trajectory prediction method based on graph converter according to claim 2 is characterized in that During the message propagation from surrounding vehicles to the target vehicle, the node characteristics of the target vehicle are updated as follows: in, Where, the superscript l represents the number of layers of the guided graph transformer, represents the node features of the target vehicle v0 at the lth layer, represents the set of surrounding vehicles of the target vehicle υ0, represents the adjacent vehicle v on the lth layer m Node features, W t represents the learnable parameter, α m,0 represents the adjacent vehicle v m The attention weight to the target vehicle v0; a t Obtained by the feedforward neural network, ⊙ represents the Hadamard product, Φ t Represents the LeakyReLU activation function.

4. The future interactive perception vehicle trajectory prediction method based on graph converter according to claim 2 is characterized in that During the message propagation from the target vehicle to the surrounding vehicles, the node features of the surrounding vehicles are updated as follows: in, Where, α 0,m Indicates the target vehicle v0's response to the adjacent vehicle v m The attention weight, a s It is obtained by the feedforward neural network with message propagation mechanism.

5. The future interactive perception vehicle trajectory prediction method based on graph converter according to claim 1 is characterized in that The gated selection mechanism based on the channel attention operator includes the calculation of the attention score: Where β is the attention score, W m represents the learnable parameters, f channel represents flattened convolution, Φ map represents the softmax layer; Update the flattened convolution: f′ channel =Φ u (f channel ⊙β+f channel ); Where f′ channel is the updated flattened convolution, Φ u Represents an expand operation.

6. The future interactive perception vehicle trajectory prediction method based on graph converter according to claim 1 is characterized in that The edge-biased multi-head attention mechanism involves updating the multi-dimensional features of the target vehicle: Where, γ i,j represents the attention coefficient of the jth feature of the i-th attention head, represents the features of the i-th attention head after update, c j represents the jth feature in the sequence, W Q 、W K 、W V Represent the linear transformation matrices of query, key, and value respectively, They represent the edge bias for distinguishing the connection relationship between different features, d e represents the hidden layer dimension, γ i,j represents the attention coefficient between the i-th and j-th features of the target vehicle, represents the updated i-th feature of the target vehicle.