Vehicle trajectory prediction method based on physical social soft attention Transformer

By adopting a physical social soft attention Transformer structure in vehicle trajectory prediction, combining physical scenarios, social interactions and local-global attention mechanisms, the problem that existing methods are difficult to model vehicle social interactions and time dependencies is solved, and higher prediction accuracy and multi-dimensional feature extraction capabilities are achieved.

CN120145000AInactive Publication Date: 2025-06-13NANTONG UNIV +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510200219.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN120145000A_ABST
    Figure CN120145000A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle trajectory prediction method based on physical social soft attention Transformer, and the method comprises the steps: constructing a scene attention sharing mechanism for the condition of physical environment information loss in a conventional method, and enhancing the influence of a physical environment on attention; aiming at the condition of lacking complex vehicle behavior information, constructing a social soft attention module to learn a vehicle track intention from the vehicle interaction information; constructing a Transform network of a local-global attention mechanism to capture global and local scale features in allusion to the situation that local features of the vehicle are missing; inputting the trajectory coordinates of the vehicle and the physical scene features into a Transform network to obtain high-level spatial-temporal features under multiple factors; and then, a physical social soft attention Transform network is constructed. According to the method, through learning of complex vehicle behavior information and a physical scene, the prediction precision of a vehicle track prediction task in the complex scene is enhanced, the ability of the model to extract multi-dimensional spatial-temporal characteristics of the track is enhanced, and the prediction performance of track prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision and artificial intelligence, and particularly relates to a deep learning method for vehicle trajectory prediction, especially a vehicle trajectory prediction method based on physical social soft attention and local-global Transformer network. Background Art

[0002] With the development of intelligent transportation and autonomous driving technologies, trajectory prediction in autonomous driving technology is a research hotspot and has important research value. Trajectory prediction for vehicles plays a key role in ensuring safety and decision-making planning. Existing vehicle trajectory prediction methods mainly rely on traditional recurrent neural networks (such as LSTM) or time series modeling based on convolutional neural networks (CNN). However, a single network structure is difficult to effectively model the social interaction behavior and time dependence relationship between vehicles, resulting in insufficient prediction accuracy.

[0003] The Transformer network has achieved remarkable results in natural language processing and vision tasks, but when directly applied to vehicle trajectory prediction, it faces challenges in modeling spatio-temporal dependence and social interaction relationships. Therefore, a prediction method that can combine the advantages of the Transformer network and effectively capture the social relationship and spatio-temporal features between vehicles is needed. Summary of the Invention

[0004] Object of the Invention: Aiming at the above problems, the present invention aims to provide a vehicle trajectory prediction method based on a physical social soft attention Transformer structure, which can effectively model the social relationship and spatio-temporal features between vehicles and improve the prediction accuracy. This method constructs a scene attention sharing mechanism to enhance the influence of the physical environment on attention for the lack of physical environment information in traditional prediction methods; constructs a social soft attention module to learn vehicle trajectory intentions from vehicle interaction information for the lack of complex vehicle behavior information; constructs a Transformer network with a local-global attention mechanism to capture global and local scale features for the lack of vehicle local features; inputs the trajectory coordinates and physical scene features of the vehicle into the Transformer network to obtain high-level spatio-temporal features under multiple factors; and then constructs a physical social soft attention Transformer network. The present invention enhances the prediction accuracy in the vehicle trajectory prediction task in a complex scene by learning complex vehicle behavior information and physical scenes, enhances the model's ability to extract multi-dimensional spatio-temporal features of the trajectory, and is beneficial to improving the prediction performance of trajectory prediction.

[0005] Technical Solution: A vehicle trajectory prediction method based on physical social soft attention Transformer.

[0006] It includes the following steps:

[0007] Step 1) Collect vehicle data at special traffic nodes such as crossroads, university campuses, and zebra crossings and transmit it to the server. Extract the vehicle trajectory observation sequence within the target area from the original data. The data form of the sequence includes the two-dimensional plane coordinates of the vehicle and the corresponding time steps. Preprocess the data sequence to reduce anomalies and missing values. Convert it into a vehicle feature matrix according to the vehicle ID, vehicle coordinates, and time nodes, and split it into a training set and a test set;

[0008] Step 2) Construct a physical scene attention sharing module to enhance the influence of the physical environment on attention, forming an environmental feature matrix and enhancing the influence of the physical environment on attention. This module includes a physical scene attention sharing mechanism;

[0009] Step 3) Construct a social soft attention module to calculate the social attention factor between vehicles and enhance the model's prediction ability for complex vehicle behavior information. This module includes the F ssa function;

[0010] Step 4) Construct a Transformer network with a local-global attention mechanism to capture global and local scale features. This module is based on a global branch of the multi-head attention mechanism and a local branch composed of a convolutional layer and group normalization to capture time-dependent features;

[0011] Step 5) Design a spatial feature extraction module based on the physical social soft attention Transformer. This network is composed of a physical scene attention sharing module, a social attention module, and a local-global attention mechanism Transformer module. Calculate the social relationship between vehicles, extract and learn global and local features to generate future trajectory prediction results and predict spatio-temporal feature representations. Use the training set to train the physical social soft attention Transformer network, output the predicted coordinate information including time steps, and use the test set to test the accuracy of the model.

[0012] Furthermore, in the above Step 1), collect vehicle data distributed at crossroads, zebra crossings, and university campuses. The data form includes the ID of the vehicle, two-dimensional plane longitude and latitude coordinates, and time stamps. There are N vehicles in a period of time [1, T pred , and the position of each vehicle i at each time node t is represented by the time two-dimensional coordinates (x i t , y i t ), where t ∈ {1, 2, 3, …, T pred}, and i ∈ {1, 2, 3, …, N}. Convert it into a vehicle feature matrix according to the vehicle ID, vehicle coordinates, and time nodes, and split it into a training set and a test set. Use the graph G t =(Vt , A t ) is used to describe the graph constructed from vehicle data, where represents the vehicle node at time t, The initial value is the observed vehicle longitude and latitude coordinates is the edge node of graph G t and is calculated through function F ssa .

[0013] Further, in step 2), a physical scene attention sharing module is constructed to enhance the influence of the physical environment on attention, forming an environmental feature matrix to obtain a physical environment attention factor and enhancing the influence of the physical environment on attention. This module includes a physical scene attention sharing mechanism. The scene image matrix preprocessed by VGG19 is input into the network, and aggregated with the scenes in the dataset to obtain the physical scene weight V p , where I scene represents the pictures in the dataset, and W vgg19 is the pre-trained network weight.

[0014] V p = VGG19(I scene ; W vgg19 )

[0015] Thus, the scene attention factor is further obtained W att contains the parameters in the scene attention, is the position of vehicle i at time t

[0016]

[0017] Thus where φ(·) is the embedding layer, and W e is the weight of the embedding layer, and SceneAtt is the scene attention function.

[0018] Further, in step 3), a social soft attention module is constructed to calculate the social attention factor between vehicles and enhance the model's prediction ability for complex vehicle behavior information. This module includes F ssa function. represents the velocity vector of the node at time t, that is represents the velocity vector of the node at time t, that is

[0019]

[0020] Among them, is the above matrix A tAfter regularization, a new adjacency matrix is obtained; α and β respectively represent and the included angle between, and θ is the hyperparameter of the vehicle node self-attention. Thus, a new graph

[0021] Furthermore, in step 4), a Transformer network that constructs a local-global attention mechanism captures global and local scale features. This module is based on a global branch of the multi-head attention mechanism, a convolutional layer, and a local branch composed of group normalization to capture time-dependent features. In the global branch part, the multi-head self-attention mechanism is adopted, and the 8-head self-attention mechanism is used to shorten the distance between long-distance features. In the local branch part, 2 parallel convolutional layers and group normalization are used to extract local context. Each convolutional layer captures different types of local information, and group normalization divides the channels into multiple groups, and each group is normalized to maintain the local statistical characteristics within the group.

[0022]

[0023] Among them, MSA is the processing process of the multi-head attention mechanism, which inputs the encoder to reshape the graph shape and shorten the distance between long-distance dependent features. The convolutional kernel sizes of the convolutional layers are 3x3 and 1x1 respectively, and finally the features extracted from the global and local branches are summed.

[0024] Furthermore, in step 5), a spatial feature extraction module based on the physical-social soft attention Transformer is constructed. This network consists of a physical scene attention sharing module, a social soft attention module, and a local-global attention mechanism Transformer module. The scene attention sharing module is obtained from step 2, the social soft attention module is obtained from step 3, and the local-global attention mechanism Transformer module is obtained from step 4. The specific steps are as follows:

[0025] Step 5-1: Initialize the network structure, network weights, batch size, step size, and maximum number of iterations, optimizer, and loss function;

[0026] Step 5-2: Determine the structure of the scene attention sharing module to obtain vehicle feature data that fuses the physical scene;

[0027] Step 5-3: Load the data in step 5-2 into the social soft attention module, update the vehicle interaction data, and obtain complex multi-level vehicle feature data;

[0028] Step 5-4: Load the vehicle data in Step 5-3 into the local-global attention mechanism Transformer module in the physical social soft attention Transformer network, and train the physical social attention Transformer network using the training set;

[0029] Step 5-5: Determine the evaluation function and output the test results of the model using the test set.

[0030] Beneficial effects: The vehicle trajectory prediction of the present invention constructs a vehicle trajectory prediction method based on a physical social soft attention Transformer structure for the multi-dimensional complex data brought by complex vehicle interactions and information interference of the physical environment itself. A physical scene attention mechanism is constructed to learn environmental information from a specific physical scene, combined with a Transformer network of a social soft attention mechanism and a local-global attention mechanism, effectively modeling the social relationship and time dependence between vehicles, further extracting multi-dimensional features, ensuring the continuity and efficiency of feature expression, and improving the accuracy of prediction results.

[0031] The present invention aims to provide a vehicle trajectory prediction method based on a physical social soft attention Transformer structure, which effectively models the social relationship and spatio-temporal features between vehicles and improves the prediction accuracy. A social soft attention module is constructed to learn the vehicle trajectory intention from vehicle interaction information for the lack of complex vehicle behavior information in traditional prediction methods; a physical scene attention sharing mechanism is constructed to enhance the influence of the physical environment on attention for the lack of physical environment information; a Transformer network with a local-global attention mechanism is constructed to capture global and local scale features for the lack of local vehicle features; high-level spatio-temporal features under multiple factors are obtained by inputting the trajectory coordinates of the vehicle and physical scene features into the Transformer network; and then a physical social soft attention Transformer network is constructed. The present invention enhances the prediction accuracy when predicting vehicle trajectories in complex scenarios by learning complex vehicle behavior information and physical scenes, enhances the ability of the model to extract multi-dimensional spatio-temporal features of trajectories, is applicable to complex dynamic scenarios, and is beneficial to improving the prediction performance of trajectory prediction. Description of the Drawings

[0032] Figure 1 It is a schematic diagram of the steps of a vehicle trajectory prediction method based on a physical social soft attention Transformer of the present invention;

[0033] Figure 2 It is a flowchart of a vehicle trajectory prediction method based on a physical social soft attention Transformer of the present invention;

[0034] Figure 3 Structural diagram of the social soft attention module for a vehicle trajectory prediction method based on physical social soft attention Transformer of the method of the present invention;

[0035] Figure 4 Structural diagram of the local-global Transformer network for a vehicle trajectory prediction method based on physical social soft attention Transformer of the method of the present invention.

[0036] Figure 5 Structural diagram of the multi-head attention mechanism for a vehicle trajectory prediction method based on physical social soft attention Transformer of the method of the present invention.

[0037] Figure 6 Training iteration diagram for a vehicle trajectory prediction method based on physical social soft attention Transformer of the present invention;

[0038] Figure 7 Comparison diagram of the real data and predicted data of the test set for a vehicle trajectory prediction method based on physical social soft attention Transformer of the present invention. Detailed implementation manners

[0039] The technical method of the present invention will be further described in detail below with reference to the accompanying drawings of the specification.

[0040] As Figure 1-2 shown, a vehicle trajectory prediction method based on physical social soft attention Transformer. The method includes the following steps:

[0041] Step 1) Collect vehicle data at special traffic nodes such as crossroads, university campuses, and zebra crossings and transmit them to the server. Extract the vehicle trajectory observation sequence in the target area from the original data. The data form of the sequence includes the two-dimensional plane coordinates of the vehicle and the corresponding time step. Preprocess the data sequence to reduce anomalies and missing values. Convert it into a vehicle feature matrix according to the vehicle ID, vehicle coordinates, and time node, and split it into a training set and a test set.

[0042] In the said step 1), collect vehicle data distributed at crossroads, zebra crossings, and university campuses. The data form includes the ID of the vehicle, two-dimensional plane longitude and latitude coordinates, and time stamps. There are N vehicles in a period of time [1, T pred , and the position of each vehicle i at each time node t is represented by the time two-dimensional coordinates (x i t , y i t ), where t ∈ {1, 2, 3,..., T pred} where \(i\in\{1,2,3,\ldots,N\}\). Convert the vehicle ID, vehicle coordinates, and time node into a vehicle feature matrix, and split it into a training set and a test set. Use the graph \(G\) t \(=(V\) t , \(A\) t ) to describe the graph constructed from vehicle data, where represents the vehicle node at time \(t\), The initial value of is the observed vehicle longitude and latitude coordinates is the edge node of graph \(G\) t and is calculated through the function \(F\) ssa .

[0043] Step 2) Construct a physical scene attention sharing module to enhance the influence of the physical environment on attention, form an environmental feature matrix, and enhance the influence of the physical environment on attention. This module includes a physical scene attention sharing mechanism, as Figure 2 shown;

[0044] In the above-mentioned step 2), construct a physical scene attention sharing module to enhance the influence of the physical environment on attention, form an environmental feature matrix to obtain a physical environment attention factor, and enhance the influence of the physical environment on attention. This module includes a physical scene attention sharing mechanism. Input the preprocessed scene image matrix of VGG19 into the network, and aggregate it with the scenes in the dataset to obtain the physical scene weight \(V\) p , where \(I\) scene represents the pictures in the dataset, and \(W\) vgg19 is the pre-trained network weight.

[0045] \(V\) p \(= VGG19(I\) scene ; \(W\) vgg19 )

[0046] Further obtain the scene attention factor \(W\) att contains the parameters in the scene attention, is the position of vehicle \(i\) at time \(t\)

[0047]

[0048] Thus where \(\varphi(\cdot)\) is the embedding layer, and \(W\) e is the weight of the embedding layer.

[0049] Step 3) Construct a social soft attention module to calculate the social attention factor between vehicles and enhance the model's prediction ability for complex vehicle behavior information. This module includes the \(F\) ssa function;

[0050] In step 3), a social soft attention module is constructed to calculate the social attention factor between vehicles, enhancing the model's prediction ability for complex vehicle behavior information. This module includes the F ssa function. denotes the speed vector of the node at time t, i.e., denotes the speed vector of the node at time t, i.e.,

[0051]

[0052] where is the above matrix A t After regularization processing, a new adjacency matrix is obtained. Thus, a new graph

[0053] Step 4) constructs a Transformer network with a local-global attention mechanism to capture global and local scale features. This module is based on a global branch of the multi-head attention mechanism, a convolutional layer, and a local branch composed of group normalization to capture time-dependent features, such as Figure 4 shown;

[0054] In step 4), a Transformer network with a local-global attention mechanism is constructed to capture global and local scale features. This module is based on the multi-head attention mechanism to capture time-dependent features. In the global branch part, the multi-head self-attention mechanism is adopted, as Figure 5 shown, using an 8-head self-attention mechanism to shorten the distance between long-distance features. In the local branch part, 2 parallel convolutional layers and group normalization are used to extract local context. Each convolutional layer captures different types of local information, and group normalization divides the channels into multiple groups, and each group is normalized to maintain the local statistical characteristics within the group.

[0055]

[0056] where MSA is the processing process of the multi-head attention mechanism, reshaping the input encoder to the graph shape to shorten the distance between long-distance dependent features. The convolutional kernel sizes of the convolutional layers are 3x3 and 1x1 respectively, and finally the features extracted from the global and local branches are summed.

[0057] Step 5) Design a spatial feature extraction module based on a physical social soft attention Transformer. This network consists of a social attention module, a physical scene attention sharing module, and a local-global attention mechanism Transformer module. Calculate the social relationships between vehicles, extract and learn global and local features to generate future trajectory prediction results and predict spatio-temporal feature representations. Use the training set to train the social soft attention Transformer network, output the predicted coordinate information including time steps, and use the test set to test the accuracy of the model.

[0058] In the above step 5), construct a spatial feature extraction module based on a physical social soft attention Transformer. This network consists of a physical scene attention sharing module, a social soft attention module, and a local-global attention mechanism Transformer module. The physical scene attention sharing module is obtained from step 2, the social soft attention module is obtained from step 3, and the local-global attention mechanism Transformer module is obtained from step 4. The specific steps are as follows:

[0059] Step 5-1: Initialize the network structure, network weights, batch size, step size, maximum number of iterations, optimizer, and loss function;

[0060] Step 5-2: Determine the structure of the scene attention sharing module to obtain vehicle feature data fused with the physical scene;

[0061] Step 5-3: Load the data in step 5-2 into the social soft attention module, update the vehicle interaction data, and obtain complex multi-level vehicle feature data;

[0062] Step 5-4: Load the vehicle data in step 5-3 into the local-global attention mechanism Transformer module in the physical social soft attention Transformer network, and use the training set to train the physical social soft attention Transformer network, with the number of iterations as Figure 6 shown;

[0063] Step 5-5: Determine the evaluation function and use the test set to output the test results of the model, with the fitting situation as Figure 7 shown.

[0064] The vehicle trajectory prediction method of the present invention constructs a vehicle trajectory prediction method based on a physical social soft attention Transformer structure for the multi-dimensional complex data brought by the complex interaction of vehicles and the information interference of the physical environment itself. In view of the lack of complex vehicle behavior information in traditional prediction methods, a social soft attention module is constructed to learn the vehicle trajectory intention from vehicle interaction information; in view of the lack of physical environment information, a physical scene attention sharing mechanism is constructed to enhance the influence of the physical environment on attention; in view of the lack of vehicle local features, a Transformer network with a local-global attention mechanism is constructed to capture global and local scale features; the trajectory coordinates of the vehicle and the physical scene features are input into the Transformer network to obtain high-level spatio-temporal features under multiple factors; furthermore, a physical social soft attention Transformer network is constructed. By learning complex vehicle behavior information and physical scenes, the present invention enhances the prediction accuracy in the vehicle trajectory prediction task in complex scenes, enhances the model's ability to extract multi-dimensional spatio-temporal features of trajectories, is applicable to complex dynamic scenes, and is conducive to improving the prediction performance of trajectory prediction.

[0065] The above is only a preferred implementation mode of the present invention under the vehicle trajectory data set. The protection scope of the present invention is not limited to the above implementation mode. Any equivalent modifications and other modified changes made by those of ordinary skill in the art according to the content disclosed by the present invention shall be included in the protection scope recorded in the claims.

Claims

1. A vehicle trajectory prediction method based on physical social soft attention Transformer, characterized in that: The following steps are involved: Step 1: Construct a vehicle trajectory observation feature matrix and visual scene features based on vehicle trajectory data; the vehicle trajectory data includes vehicle ID, two-dimensional plane coordinates and timestamp, and the data is converted into a vehicle feature matrix according to vehicle ID, coordinates and time nodes, and divided into a training set and a test set; Step 2: Construct a physical scene attention sharing module, preprocess the scene image matrix through VGG19 and aggregate it with the vehicle feature matrix to generate the environment feature matrix and scene attention factor to enhance the impact of the physical environment on attention; Step 3: Build a social soft attention module to calculate the social attention factor based on the inter-vehicle speed vector, generate a vehicle interaction relationship diagram, and enhance the model's ability to predict complex vehicle behavior information; Step 4: Construct a Transformer network with a local-global attention mechanism to capture global and local time-dependent features through the global branch of the multi-head self-attention mechanism and the local branch composed of convolutional layers and group normalization; Step 5: Build a physical social soft attention Transformer network based on the modules of step 2, step 3 and step 4, train the network with the training set and optimize the loss function, and save the optimal weights; input the data to be predicted into the trained model, and output the vehicle trajectory prediction results for the future time step.

2. According to claim 1, a vehicle trajectory prediction method based on physical social soft attention Transformer is characterized in that: In step 1, the construction of the vehicle trajectory observation feature matrix includes: Extract the coordinates (x i t ,y i t ) represents, where t∈{1,2,3,…,T pred }, i∈{1,2,3,…,N}, T pred represents time, N represents the number of vehicles; the coordinates and timestamps are integrated into graph structure data according to the vehicle ID, the nodes in the graph represent the vehicle positions, and the edge nodes are generated by calculating the social soft attention function.

3. According to claim 1, a vehicle trajectory prediction method based on physical social soft attention Transformer is characterized in that: The implementation of the physical scene attention sharing module in step 2 includes: Aggregate the preprocessed scene image matrix with the vehicle feature matrix to generate the physical visual scene feature V p , the specific formula is: V p =VGG19(I scene ;W vgg19 ) Among them, I scene represents the scene image, W vgg19 are the pre-trained network weights; Calculating scene attention factor through embedding layer The specific formula is: Among them, SceneAtt is the scene attention function, W att Contains the parameters in scene attention, is the position of vehicle i at time t 4. The vehicle trajectory prediction method based on physical social soft attention Transformer according to claim 1 is characterized in that: The calculation of the social soft attention factor in step 3 includes: Among them, F ssa represents the social soft attention function, represents the time t The velocity vector of the node is represents the time t The velocity vector of the node is is the matrix A t After regularization, a new adjacency matrix is ​​obtained; α and β represent and , θ is the hyperparameter of the vehicle node self-attention; thus, the new graph is obtained 5. A vehicle trajectory prediction method based on physical social soft attention Transformer according to claim 4, characterized in that: The Transformer network of the local-global attention mechanism in step 4 includes: the global branch adopts an 8-head self-attention mechanism to shorten the dependency distance between long-distance features; the local branch adopts parallel convolution layers of 3x3 and 1x1 convolution kernels and group normalization to extract local context features; the features of the global and local branches are fused by summing, and the formula is: Among them, MSA is a multi-head attention mechanism.

6. The vehicle trajectory prediction method based on physical social soft attention Transformer according to claim 1 is characterized in that: The loss function in step 5 is a negative log-likelihood function based on Gaussian distribution, and the calculation formula is: Where: L is the loss function, is the actual position of the vehicle, It represents the probability of predicting trajectory points under the condition of given Gaussian distribution parameters; the parameters of the model are updated through back propagation and gradient descent to optimize the loss function.

7. The vehicle trajectory prediction method based on physical social soft attention Transformer according to claim 1 is characterized in that: The step 5 also includes: using a spatiotemporal convolutional network to further train the fused traffic flow feature matrix to improve the expression capability of spatiotemporal features.

8. The vehicle trajectory prediction method based on physical social soft attention Transformer according to claim 5 is characterized in that: The calculation formula of the attention weight A in the multi-head self-attention mechanism is: Where: Q is the query vector, which represents the location characteristics of the current vehicle; K is the key vector, which represents the location characteristics of other vehicles; d k is the scaling factor of the vector dimension; V is the value vector representing the vehicle's position data.

9. The vehicle trajectory prediction method according to claim 1, characterized in that: During the training process of the physical social soft attention Transformer network, the Adam optimizer is used to update parameters, and the maximum number of iterations is dynamically adjusted by the training curve.

10. The vehicle trajectory prediction method according to claim 1, characterized in that: The method is applicable to vehicle trajectory prediction in intersections, zebra crossings and campus scenes, and the output result is the future T pred A 2D sequence of coordinates of the time steps.

Citation Information

Cited By

  • Automatic driving vehicle motion prediction method based on scene understanding enhancement

    CN120339990A

  • Intelligent weight prediction method for dynamic scene

    CN120805990A

  • An intelligent weight prediction method for dynamic scenes

    CN120805990B