Vehicle multi-modal trajectory prediction method based on HGT network
By adopting a multimodal trajectory prediction method based on HGT network in autonomous driving technology, the complexity and uncertainty of vehicle future trajectory prediction are solved, and more accurate and stable trajectory prediction is achieved, which improves the safety and efficiency of autonomous driving vehicles.
Patent Information
- Application Number
- CN202510098781.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-30
AI Technical Summary
In autonomous driving technology, the prediction of future trajectories of vehicles faces complexity and uncertainty, especially because driving intentions cannot be directly observed and complex interactive behaviors between vehicles, making it difficult for the prior art to generate accurate and stable trajectory predictions.
Using the vehicle multimodal trajectory prediction method based on HGT network, the local space fusion module, environment perception module, global interaction module and trajectory decoding module are constructed to capture the dynamic interaction relationship, environmental characteristics and global information between vehicles, and generate multimodal future trajectory prediction.
It improves the accuracy and stability of vehicle trajectory prediction, can generate accurate future trajectories under changing road conditions, and enhances the safety and efficiency of autonomous driving vehicles.
Smart Images

Figure CN120067573A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of trajectory prediction, and particularly relates to a vehicle multi-modal trajectory prediction method based on the HGT network. Background Art
[0002] With the rapid development of society, driving a car has become the preferred mode of travel for modern the public, which has greatly promoted the flow of people and the exchange of goods. However, at the same time, the sharp increase in the number of vehicles has also led to a series of unprecedented severe challenges such as traffic congestion, frequent accidents, and environmental pollution. In this predicament, autonomous driving technology, as a frontier field of scientific and technological innovation, is gradually changing from a concept to reality, providing a new direction for solving traffic and environmental problems.
[0003] In the core field of autonomous driving technology, accurately predicting the future trajectory of a vehicle is the key cornerstone to ensure the safe and efficient operation of autonomous vehicles. However, this task faces numerous challenges. The primary problem is that driving intention is not directly observable, which makes the future trajectory of the vehicle exhibit complex, variable, and multi-modal distribution characteristics in space. To address this challenge, the prediction process not only requires a meticulous analysis of the static elements of the road environment, such as road morphology, traffic signals, construction areas, etc., but also needs to capture and deeply analyze dynamic information in real time, such as the driving intentions and subtle changes in speed fluctuations of surrounding vehicles. Even more intractable is that the complex interaction behaviors between vehicles, such as avoidance, following, and overtaking, greatly increase the uncertainty and complexity of trajectory prediction, posing higher requirements for the accuracy and robustness of the prediction model.
[0004] In summary, in order to predict the future behavior of a vehicle, a vehicle trajectory prediction method that can improve the accuracy of trajectory prediction needs to be proposed. Summary of the Invention
[0005] Object of the Invention: The present invention provides a vehicle multi-modal trajectory prediction method based on the HGT network, which can perform multi-modal trajectory prediction on the surrounding vehicles of autonomous vehicles in the face of changing road conditions and intricate interactions between vehicles, and generate accurate and stable future trajectories.
[0006] Technical Solution: The vehicle multi-modal trajectory prediction method based on the HGT network described in the present invention includes the following steps:
[0007] (1) Split the Argoverse trajectory prediction dataset into a training set and a validation set, then perform preprocessing, and save the trajectory data and lane line data in the local space;
[0008] (2) Construct an HGT network, where the HGT network includes a local spatial fusion module, an environmental perception module, a global interaction module, and a trajectory decoding module; the local spatial fusion module encodes vehicle interaction relationships, mines local interaction patterns, and simultaneously extracts temporal features to capture its dynamic changes; then it fuses the vehicle temporal features and interacts the output features with the environmental features, enabling the vehicle to perceive environmental changes and adjust its behavior; the environmental perception module captures the impact of the surrounding environmental features on the vehicle's future behavior; the global interaction module further enhances the vehicle's perception ability of the impact of surrounding vehicles from a global perspective; the trajectory decoding module decodes the vehicle features using a multi-layer perceptron to generate a predicted trajectory and calculates a score to evaluate the quality of the predicted trajectory;
[0009] (3) Feed the preprocessed training data into the HGT network for training, and use the AdamW optimizer to perform backpropagation to update the weights to optimize the HGT network;
[0010] (4) Feed the test data into the trained HGT network for testing, and use visualization techniques to display the results predicted in the test phase.
[0011] Further, the preprocessing in step (1) is to segment the data according to time frames, save the position information of all vehicles within the time frame; use the saved trajectory data of the first 20 frames to predict the trajectory positions of the next 30 frames; and use the position information data of the 19th frame trajectory to obtain the lane line information at the segmentation point moment using the API and save it.
[0012] Further, the local spatial fusion module in step (2) includes a spatial interaction module, a temporal extraction module, and a temporal fusion module.
[0013] Further, the spatial interaction module builds a graph network and constructs a message passing module and an update module; maps each vehicle data in space to a node in the graph one by one, passes messages between nodes through the message passing module, and uses the update module to complete the update of the node features;
[0014] The message calculation in the message passing process of the message passing module is as follows:
[0015]
[0016] where Q is the feature of the central node, K and V are the features of the connected nodes, d k is the dimension of the feature K, softmax is the activation function, dropout is the regularization operation, and message i is the message feature to be passed to the central node i;
[0017] The message update process in the update module is as follows:
[0018]
[0019] Among them, messag represents the message set generated by each connected node in the message passing module, gate is a gating unit, and MLP represents a linear layer. is the original feature of node i at time t. represents the updated feature of node i.
[0020] Furthermore, the implementation process of the timing extraction module is as follows:
[0021] First, the features with continuous time dimensions are dimensionally increased through a multi-layer perceptron to enhance the model's ability to store information; before extracting the timing features, the position information of the data features is recorded by setting position encoding; the attention mechanism is used to capture the key features in the continuous trajectory, and these key information are fused into the original trajectory through residual connection; the feed-forward neural network is used to further mine more useful features from the original data to enrich and improve the feature representation.
[0022] Furthermore, the timing fusion module captures the information between time frames and effectively fuses the features of each vehicle in each time frame. The implementation process is as follows:
[0023] First, a learnable [SUMMARY] frame is added to the original frame sequence, and the dimension of this frame is the same as that of the remaining frames; then, position encoding is added to the features of each time frame to retain the feature information contained in the input order; finally, the data processed above is sent into the Encoder-F structure, which focuses on the mutual relationship between frames to achieve the fusion of the features of each frame.
[0024] Furthermore, the implementation process of the environment perception module in step (2) is as follows:
[0025] Extract the key environmental features from historical moments, interact the features of each vehicle with the environmental features to effectively integrate the historical environmental features; generate a comprehensive feature representation that fuses the historical environmental information.
[0026] The historical environmental features are represented as:
[0027] S = {lane k,i , lane k,j , cross k , head k , control k} (3)
[0028] Among them, lane k,i and lane k,j represent the positions of the kth center line,, cross kIndicates whether the k-th lane intersects, head k Records the orientation of the k-th lane (going straight, turning left, turning right), control k Indicates whether the k-th lane is under traffic control;
[0029] In the environmental perception module, the lane lines are represented as:
[0030]
[0031] The edges between the vehicle and the lane lines are represented as:
[0032] c i,j =(a i , l j ) (5)
[0033] The edge attribute features are represented as:
[0034] f j ={cross j , head j , control j}} (6)
[0035] Among them, a i represents vehicle i, l j represents lane line j, c i,j represents the connection edge between vehicle i and lane line j; f j represents the attribute of the connection edge between vehicle i and lane line j;
[0036] In order to fuse the lane line information into each vehicle feature, a topological graph containing lane lines and vehicles is constructed. During the message passing process, a cross-attention mechanism is adopted to capture the impact of the lane line environmental information on the vehicle, and a gated unit is used to control the update of the features. The specific fusion process is as follows:
[0037]
[0038] Among them, gate represents the gated unit, MLP represents the linear layer, and concat(·,·) represents the concatenation operation. al i represents vehicle i after fusing the environmental features.
[0039] Furthermore, the implementation process of the global interaction module in step (2) is as follows:
[0040] Global information is obtained through local information interaction, and the global information is used to act on the local features in reverse. A new graph network is constructed with vehicles as nodes. The message passing function selects the attention mechanism to expand the range of message passing in the graph network and broaden the receptive field of the model, enabling the model to understand and process data from a more macroscopic perspective.
[0041] Furthermore, the trajectory decoding module in step (2) is a dual-branch structure, where one branch uses two layers of MLP to perform decoding operations to obtain prediction results of multiple future positions, and the operation is as follows:
[0042]
[0043] Among them, Fglobal is the feature output by the global interaction module, scale is a hyperparameter, indicating the uncertainty of the trajectory, MLP is a multi-layer perceptron, The result value of the predicted trajectory for the i-th agent;
[0044] The other branch obtains the trajectory score through MLP as follows:
[0045]
[0046] in, is the score corresponding to the trajectory predicted by the i-th agent.
[0047] Beneficial effects: Compared with the prior art, the present invention has the following beneficial effects: the present invention constructs a local space fusion module. For vehicle data segmented by time frames, the spatial interaction module enables the model to capture the local dynamic interaction relationship between vehicles. For trajectory data segmented by vehicles, a time series extraction module is designed to enhance the model's perception of time series relationships. Finally, the space and time series relationships are integrated; the present invention designs an environment perception module to capture the potential impact of surrounding environment features on vehicle behavior patterns through the interaction between vehicle and environmental features; the present invention proposes a global interaction module, which first obtains global feature information through the local features of the interacting vehicles; then, the extracted global features are reacted to the local features to improve the vehicle's perception of the global feature information; these modules work together to deepen the model's learning ability of historical trajectories and environmental features, and generate accurate and stable future trajectories. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a schematic diagram of the HGT network structure proposed by the present invention;
[0049] Figure 2 It is a schematic diagram of the GCN structure proposed in the present invention;
[0050] Figure 3 It is a schematic diagram of message transmission in the present invention;
[0051] Figure 4 It is a schematic diagram of the structure of the timing extraction module proposed by the present invention;
[0052] Figure 5 It is a schematic diagram of the structure of the timing fusion module proposed in the present invention;
[0053] Figure 6 It is a schematic diagram of the global interaction module structure proposed by the present invention;
[0054] Figure 7 It is a model prediction result diagram in a turning environment;
[0055] Figure 8 It is a model prediction result diagram in a straight - running environment;
[0056] Figure 9 It is a prediction result diagram using the present invention. Specific implementation manner
[0057] The following further elaborates on the specific technical solutions of the present invention in conjunction with the accompanying drawings.
[0058] The present invention proposes a vehicle multi - modal trajectory prediction method based on the HGT (Hierarchical graph and transformer network) network. First, the Argoverse trajectory prediction dataset is segmented to divide the training set and the validation set, and then pre - processed to save the trajectory data and lane line data in the local space; then, an HGT network as shown in Figure 1 is constructed. The HGT network includes four core modules: local space fusion, environment perception, global interaction, and trajectory decoding, which jointly constitute an efficient trajectory prediction system; immediately afterwards, the pre - processed training data is fed into the HGT network for training, and the AdamW optimizer is used for backpropagation to update the weights to optimize the HGT network; the test data is fed into the trained HGT network for testing, and the visualization technology is used to display the prediction results in the test phase.
[0059] As shown in Figure 1 the HGT network includes four core modules: local space fusion, environment perception, global interaction, and trajectory decoding, which jointly constitute an efficient trajectory prediction system. During training, the data is first pre - processed and segmented by time frames. The vehicle interaction relationship is encoded in the local space to mine the local interaction pattern, and at the same time, the time features are extracted to capture its dynamic changes; then, the temporal fusion module further fuses the vehicle time features and interacts the output features with the environmental features, enabling the vehicle to perceive environmental changes and adjust its behavior; then, the global interaction module further enhances the vehicle's perception ability of the influence of surrounding vehicles from a global perspective; finally, a multi - layer perceptron is used to decode the vehicle features to generate the predicted trajectory, and a score is calculated to evaluate the quality of the predicted trajectory.
[0060] When designing the spatial interaction module, the primary task is to construct a graph network, as shown in Figure 2As shown in the figure. Each vehicle in space is mapped one by one to a node in the graph. To comprehensively consider the mutual influence between each vehicle and its surrounding connected vehicles, a node message passing module is designed, as Figure 3 shown. The core of this module is to utilize the self-attention mechanism, which can accurately capture and quantify the complex interaction relationships between vehicles. Finally, the node features are updated with the help of an update module.
[0061] In the spatial interaction module, the message calculation in the message passing process is as follows:
[0062]
[0063] where Q is the feature of the central node, K and V are the features of the connected nodes, d k is the dimension of the feature K, softmax is the activation function, dropout is the regularization operation, and message i is the message feature to be passed to the central node i.
[0064] The message update process in the spatial interaction module is as follows:
[0065]
[0066] where messag represents the set of messages generated by each adjacent node in the message passing module, gate is a gating unit, MLP represents a linear layer, is the original feature of node i at time t, represents the updated feature of node i.
[0067] After the interaction, the vehicle has the ability to perceive surrounding vehicles in the local space. However, since the data is segmented by time frames, the local interaction shows a discrete state in the time dimension. To add continuous temporal relationships in the local fusion module, a temporal extraction module is designed in this paper, as Figure 4 shown. For the temporal extraction module, the input data is segmented by vehicle and is continuous in the time dimension. Specifically, the network first performs dimensionality increase processing on the features that are continuous in the time dimension through a multi-layer perceptron to enhance the model's ability to store information. Given the sequential nature of the features due to the continuous nature of the data in the time dimension, position encoding is set to record the position information of the data features before extracting the temporal features. On this basis, the attention mechanism is used to capture the key features in the continuous trajectory, and these key information are fused into the original trajectory through residual connections. Finally, a feed-forward neural network is used to further mine more useful features from the original data to enrich and improve the feature representation.
[0068] The temporal fusion module is as Figure 5As shown, the workflow is as follows: First, a learnable [SUMMARY] frame is added to the original frame sequence, and the dimension of this frame is consistent with the rest of the frames; then, positional encoding is added to the features of each time frame to retain the feature information contained in the input order; finally, the data processed above is sent into the Encoder-F structure, which focuses on the relationship between frames, thereby achieving the fusion of the features of each frame. This design enables the local feature fusion module to make more full use of the information in the time dimension and provide richer and more representative feature representations for subsequent processing.
[0069] To capture the potential impact of environmental features on the vehicle's driving trajectory, first, key environmental features are extracted from historical moments, and these features can reflect the state of the environment in which the vehicle was located at past time steps; then, an environmental perception module is designed to interact the features of each vehicle with the environmental features to effectively integrate the historical environmental features; finally, after being processed by this module, a comprehensive feature representation that fuses historical environmental information can be generated.
[0070] The historical environmental features are represented as:
[0071] S = {lane k,i ,lane k,j ,cross k ,head k ,control k} (3)
[0072] Among them, lane k,i and lane k,j represent the position of the k-th center line,, cross k represents whether the k-th lane intersects, head k records the orientation of the k-th lane (going straight, turning left, turning right), control k represents whether the k-th lane is under traffic control.
[0073] In the environmental perception module, the lane line can be represented as:
[0074]
[0075] The edge between the vehicle and the lane line can be represented as:
[0076] c i,j =(a i ,l j ) (5)
[0077] Its edge attribute features are represented as:
[0078] f j ={cross j, head j , control j} (6)
[0079] Among them, a i represents vehicle i, l j represents lane line j, c i,j represents the connection edge between vehicle i and lane line j. f j represents the attribute of the connection edge between vehicle i and lane line j.
[0080] To fuse lane line information into each vehicle feature, a topological graph containing lane lines and vehicles is constructed. During the message passing process, a cross-attention mechanism is adopted to capture the impact of the environmental information of the lane line on the vehicle. And a gated unit is used to control the update of the features. The specific fusion process is as follows:
[0081]
[0082] Among them, gate represents the gated unit, MLP represents the linear layer, indicating the concat(·, ·) splicing operation, al i represents vehicle i after fusing the environmental features.
[0083] Encoding from only the perspective of a single vehicle has limitations and ignores global feature information, thus having a certain impact on the performance and accuracy of the model. To effectively solve this problem, a global interaction module is constructed, as Figure 6 shown. By integrating global information through this module, the receptive field of the model is effectively broadened, enabling the model to understand and process data from a more macroscopic perspective.
[0084] Inside the global interaction module, a new graph network is constructed with vehicles as nodes. The message passing function of this network still selects the attention mechanism. Through multiple uses of message passing for message update, the range of message passing in the graph network is further expanded, effectively increasing the receptive field of the vehicle. In this process, the self-attention mechanism plays a crucial role. It ensures that each vehicle can conduct comprehensive information interaction with other vehicles, effectively promoting the effective communication between local features. Moreover, the self-attention mechanism can guide the network to accurately identify and highly value those key factors that have a significant impact on global features, thereby improving the model's ability to capture and utilize global information.
[0085] The trajectory decoding module is designed as a two-branch structure. One branch uses two-layer MLP for decoding operations to obtain prediction results of multiple future positions. The operations are as follows:
[0086]
[0087] Among them, Fglobal is the feature output by the global interaction module, scale is a hyperparameter representing the uncertainty of the trajectory, and MLP is a multi-layer perceptron. is the result value of the predicted trajectory of the i-th agent.
[0088] Similarly, another branch obtains the trajectory score through the MLP as follows:
[0089]
[0090] Among them, is the score corresponding to the predicted trajectory of the i-th agent.
[0091] In this embodiment, first, the Argoverse motion prediction dataset is obtained from the official website and partitioned and feature constructed. Specifically, the historical state of the moving object in the first two seconds is used as the input feature, and the next three seconds are used as the true label value of the predicted position. The motion position of the vehicle at a future moment is predicted, and data samples and labels are constructed and partitioned according to 8:2. The data is segmented according to time frames, and the position information of all vehicles within the time frame is saved. The position information data (segmentation point) of the 19th frame trajectory is used to obtain the lane line information at the segmentation point moment using the API and saved. The data is stored in the Argoverse / datasets folder, which includes two sub-folders, namely train and val. The preprocessed vehicle trajectory data is input into HGT for training, with an initial learning rate of 0.0005, the AdamW optimizer is used, batch is 16, epochs is 80, and model training starts. Then the trained HGT model is tested on the validation set, and finally the test results are visually displayed using visualization techniques.
[0092] The specific visualization effect is as Figures 7 to 9 shown, including different scenarios, such as: going straight, turning, etc. To prove that the model can have a multi-modal nature, finally, the visualization technique is used to output all 6 predicted trajectories. Figures 7 to 9 In, the green dashed line represents the historical 20-frame trajectory, i.e., the input data, the red solid line represents the true predicted trajectory, i.e., the true label predicted by the model, and the blue solid line represents the trajectory predicted by the HGT model. The x-axis in the figure represents the time frame, and the prediction task uses the first 20 frames to predict the next 30 frames. The y-axis represents the y value of the trajectory coordinates. From Figures 7 to 9 the fitting effect, it can be seen that for different environmental conditions, the HGT model designed by the present invention can well predict the future trajectory. Regarding the uncertainty of the future trajectory, i.e., multi-modality, visualizing the 6 possible trajectories obtained by prediction shows that the HGT model has the ability of multi-modal prediction.
[0093] It should be noted that the above is only used to illustrate the technical solution of the present invention rather than to limit it. Those of ordinary skill in the art should understand that the technical solution of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solution of the present invention, and all of them should be covered within the scope of the claims of the present invention.
Claims
1. A vehicle multimodal trajectory prediction method based on HGT network, characterized in that: The following steps are involved: (1) Segment the Argoverse trajectory prediction dataset into a training set and a validation set, and then perform preprocessing to save the trajectory data and lane line data in the local space; (2) Constructing an HGT network, the HGT network includes a local space fusion module, an environment perception module, a global interaction module and a trajectory decoding module; the local space fusion module encodes the vehicle interaction relationship, mines the local interaction mode, and extracts the time feature to capture its dynamic change; then the vehicle time feature is fused, and the output feature is interacted with the environment feature, so that the vehicle perceives the environmental change and adjusts the behavior; the environment perception module captures the influence of the surrounding environment features on the future behavior of the vehicle; the global interaction module further enhances the vehicle's perception of the influence of the surrounding vehicles from a global perspective; the trajectory decoding module uses a multi-layer perceptron to decode the vehicle features to generate a predicted trajectory, and calculates the score to evaluate the quality of the predicted trajectory; (3) The preprocessed training data is sent to the HGT network for training, and the HGT network is optimized by back-propagation updating the weights using the AdamW optimizer; (4) The test data is sent to the trained HGT network for testing, and the prediction results of the test phase are displayed using visualization technology.
2. A vehicle multimodal trajectory prediction method based on HGT network according to claim 1, characterized in that: The preprocessing in step (1) is to segment the data according to the time frame, save the position information of all vehicles in the time frame; use the saved trajectory data of the first 20 frames to predict the trajectory position of the last 30 frames; and use the position information data of the 19th frame trajectory to obtain the lane line information at the segmentation point using the API and save it.
3. The vehicle multimodal trajectory prediction method based on HGT network according to claim 1 is characterized in that: The local space fusion module in step (2) includes a space interaction module, a time series extraction module and a time series fusion module.
4. The vehicle multimodal trajectory prediction method based on HGT network according to claim 3 is characterized in that: The spatial interaction module builds a graph network, and constructs a message passing module and an update module; each vehicle data in the space is mapped one by one to a node in the graph, the message passing module is used to pass messages between nodes, and the update module is used to complete the update of node features; Message passing module message passing process message calculation is as follows: Among them, Q is the feature of the central node, K and V are the features of the connected nodes, and d k is the dimension of feature K, softmax is the activation function, dropout is the regularization operation, and message i is the message feature to be delivered to the central node i; The message update process in the update module is as follows: Among them, messag represents the message set generated by each connected node in the message passing module, gate is a gating unit, MLP represents the linear layer, is the original feature of node i at time t, Represents the updated features of node i.
5. The vehicle multimodal trajectory prediction method based on HGT network according to claim 3 is characterized in that: The timing extraction module implementation process is as follows: Firstly, the continuous features in the time dimension are processed by multi-layer perceptron to enhance the model's ability to store information. Before extracting the time series features, the position information of the data features is recorded by setting position encoding. The attention mechanism is used to capture the key features in the continuous trajectory, and these key information are integrated into the original trajectory through residual connection. The feedforward neural network is used to further mine more useful features from the original data to enrich and improve the feature representation.
6. The vehicle multimodal trajectory prediction method based on HGT network according to claim 3 is characterized in that: The time series fusion module captures the information between time frames and effectively fuses the features of each vehicle in each time frame. The implementation process is as follows: First, a learnable [SUMMARY] frame is added to the original frame sequence, and the dimension of this frame is consistent with the rest of the frames. Then, position encoding is added to the features of each time frame to retain the feature information contained in the input order. Finally, the processed data is sent to the Encoder-F structure, which focuses on the relationship between frames, thereby realizing the fusion of the features of each frame.
7. The vehicle multimodal trajectory prediction method based on HGT network according to claim 1 is characterized in that: The implementation process of the environment perception module in step (2) is as follows: Extract key environmental features from historical moments, interact the features of each vehicle with the environmental features, and effectively integrate the historical environmental features; generate a comprehensive feature representation that integrates historical environmental information; The historical environmental characteristics are expressed as: S={lane k,i ,lane k,j ,cross k ,head k ,control k } (3) Among them, lane k,i and lane k,j represents the position of the kth center line, cross k Indicates whether the kth lane intersects, head k Record the direction of the kth lane (straight ahead, left turn, right turn), control k Indicates whether the kth lane is under traffic control; The lane line in the environment perception module is represented as: The edge between the vehicle and the lane line is represented as: c i,j =(a i ,l j ) (5) The edge attribute feature is expressed as: f j ={cross j ,head j ,corttrol j } (6) Among them, a i represents vehicle i, l j represents lane line j, c i,j represents the edge connecting vehicle i and lane line j; f j Represents the attributes of the edge connecting vehicle i and lane line j; In order to integrate lane information into each vehicle feature, a topological map containing lanes and vehicles is constructed. The cross-attention mechanism is used in the message transmission process to capture the impact of lane information on the vehicle, and the gate unit is used to control the update of features. The specific fusion process is as follows: Among them, gate represents the gated unit, MLP represents the linear layer, and concat(·,·) represents the concatenation operation. i Represents vehicle i after integrating environmental features.
8. The vehicle multimodal trajectory prediction method based on HGT network according to claim 1 is characterized in that: The global interaction module implementation process in step (2) is as follows: Global information is obtained through local information interaction, and the global information is fed back to local features. A new graph network is constructed with vehicles as nodes. The message passing function uses the attention mechanism to expand the scope of message passing in the graph network and broaden the receptive field of the model, so that the model can understand and process data from a more macro perspective.
9. The vehicle multimodal trajectory prediction method based on HGT network according to claim 1 is characterized in that: The trajectory decoding module in step (2) is a dual-branch structure, where one branch uses two layers of MLP to perform decoding operations to obtain prediction results of multiple future positions. The operation is as follows: Among them, Fglobal is the feature output by the global interaction module, scale is a hyperparameter, indicating the uncertainty of the trajectory, MLP is a multi-layer perceptron, The result value of the predicted trajectory for the i-th agent; The other branch obtains the trajectory score through MLP as follows: in, is the score corresponding to the trajectory predicted by the i-th agent.
Citation Information
Cited By
Pedestrian trajectory prediction method based on multi-clue transformation network
CN120726085A
Passenger train delay propagation prediction method based on HGT-GRU model
CN121616443A