Map-free vehicle trajectory prediction method based on virtual lane and heterogeneous social graph network

By generating virtual lanes and building a heterogeneous social graph network, the dependence problem on high-precision maps in the existing methods is solved, and efficient and accurate vehicle trajectory prediction is achieved under map-free conditions, which is suitable for intelligent traffic and autonomous driving.

CN120492895AInactive Publication Date: 2025-08-15NANTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510585990.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing trajectory prediction methods rely on high-precision maps, resulting in high collection and maintenance costs, lag in complex or dynamically changing road environments, making it difficult to fully tap road and lane structure information in vehicle historical trajectory data, affecting prediction accuracy and applicability.

Method used

The virtual lane center line is generated through a density-based hierarchical clustering algorithm, and a two-way long and short-term memory network is used to encode lane features, and a heterogeneous graph is constructed to transmit information and feature fusion, combining the multi-head self-attention module to capture vehicle interaction characteristics and predict future trajectories.

Benefits of technology

Under the condition of no high-precision map, the accuracy and robustness of vehicle trajectory prediction are significantly improved, and the adaptability and prediction accuracy of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492895A_ABST
    Figure CN120492895A_ABST
Patent Text Reader

Abstract

The invention discloses a map-free vehicle trajectory prediction method based on a virtual lane and a heterogeneous social graph network. The method comprises the following steps: firstly, acquiring historical track data of a vehicle, and encoding time sequence characteristics through a single-layer LSTM; secondly, generating a virtual lane center line by adopting an HDBSCAN algorithm in a clustering manner, and coding lane features by utilizing bidirectional LSTM; then, constructing a heterogeneous graph containing vehicle nodes and virtual lane nodes, and realizing information transmission and feature updating through a graph convolutional network; a multi-head self-attention module is introduced to capture vehicle interaction characteristics; and finally fusing lane information to predict a future trajectory. According to the method, the dependence of a traditional method on a high-precision map is broken through, the trajectory prediction precision and robustness under the map-free condition are remarkably improved by dynamically generating the virtual lane and deeply fusing the road structure information, and the method is suitable for the fields of intelligent traffic and automatic driving. The method has the characteristics of high calculation efficiency and strong adaptability, and shows excellent performance in complex traffic scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning and intelligent transportation, and specifically provides a map-free vehicle trajectory prediction method based on virtual lanes and heterogeneous social graph networks. Background Art

[0002] With the rapid development of autonomous driving technology and intelligent transportation systems, vehicle trajectory prediction, as a key technology for achieving autonomous driving safety and improving traffic management efficiency, has become a hot topic of research both domestically and internationally. Existing trajectory prediction methods mostly rely on high-precision map information. However, HD maps are expensive to acquire and maintain, and in complex or dynamically changing real-world road environments, HD map information may be delayed or missing. Therefore, map-free trajectory prediction methods offer broader application prospects.

[0003] In recent years, the widespread deployment of numerous vehicle sensors and connected vehicle devices has enabled the collection of massive amounts of historical vehicle trajectory data. This data not only contains vehicle displacement information but also implicitly contains features such as road structure and vehicle travel direction. How to fully mine and utilize this implicit information to improve trajectory prediction accuracy is a key research direction. While existing methods have to some extent exploited information about social interactions between vehicles, they are still insufficient in mining and utilizing the implicit road structure, and struggle to fully account for the multiple factors that influence vehicle motion, including road orientation, lane constraints, and travel direction.

[0004] Therefore, how to fully explore and utilize the road and lane structures implicit in vehicle historical trajectory data in the absence of high-precision map information, and thereby improve the accuracy and applicability of vehicle trajectory prediction models, has become a key technical issue that needs to be urgently addressed in the field of intelligent transportation and autonomous driving. Summary of the Invention

[0005] Purpose of the invention: In response to the above problems, the purpose of the present invention is to provide a map-free vehicle trajectory prediction method based on virtual lanes and heterogeneous social graph networks. This method uses historical vehicle trajectory data and automatically generates virtual lane lines that reflect the actual road structure through density-based hierarchical clustering algorithm (HDBSCAN) clustering. Lane direction vector information is introduced when aggregating the representative trajectories of vehicles in the cluster, and a bidirectional long short-term memory network (BiLSTM) is used to encode the lane centerline and direction information; at the same time, the encoded lane features are used as structured nodes and vehicle nodes to jointly construct a heterogeneous graph, and a message passing mechanism is used to achieve a deep fusion of lane information and vehicle motion characteristics. Furthermore, virtual lane supplementary information is introduced in the input encoding and output decoding stages, thereby significantly improving the accuracy and robustness of vehicle trajectory prediction in the absence of high-precision maps.

[0006] Technical Solution: The specific steps of the map-free vehicle trajectory prediction method based on virtual lanes and heterogeneous social graph networks include:

[0007] Step 1) Obtain historical trajectory data of multiple vehicles from the traffic scene, convert the absolute coordinates into relative coordinates with the target vehicle as the reference, and extract the temporal dynamic characteristics of vehicle motion;

[0008] Step 2) Cluster the vehicle historical trajectories using a density-based hierarchical clustering algorithm (HDBSCAN) to generate virtual lane centerlines and extract the corresponding direction information. The lane centerlines and direction vectors are encoded using a bidirectional long short-term memory network (BiLSTM) to obtain virtual lane features.

[0009] Step 3) Construct a heterogeneous graph convolutional network, including vehicle nodes and virtual lane nodes, and define vehicle-vehicle edges and vehicle-lane edges. Use the graph neural network to transfer information and update features between nodes.

[0010] Step 4) Use the multi-head self-attention module to further capture the dependencies and interactions between vehicles and improve the expressiveness of the fusion features;

[0011] Step 5) In the output decoder, lane prior information is integrated and the fused features are decoded to predict the future trajectory of the vehicle. End-to-end training and multimodal strategies are used to optimize the model performance and evaluate the prediction accuracy.

[0012] Furthermore, in step 1), in the traffic scene, the trajectory information of each vehicle is recorded in the form of a time series, specifically represented as a two-dimensional coordinate

[0013]

[0014] in represents the trajectory position of vehicle i at time t, and Represent the horizontal and vertical coordinates of vehicle i at time t, t = -T h +1,…,0 represents the historical time step. In order to adapt to model training, the absolute coordinates are first converted to relative coordinates with the target vehicle as the reference, and the data is normalized by coordinate transformation.

[0015] Next, a discrete displacement sequence is generated for each vehicle and a vector definition is performed:

[0016]

[0017] in represents the displacement change of vehicle i at time t relative to the previous time t-1, and Represent the trajectory position of vehicle i at time t and time t-1 respectively. At the same time, a binary flag is introduced Right now Indicates the validity of the data at that time, and concatenates the two to form the input vector for each time step Finally, the input sequence of each vehicle is input into a single-layer long short-term memory network (LSTM) for encoding to obtain the hidden state

[0018]

[0019] in, is the hidden state of vehicle i at time t, w ecn represents the weight parameter, b ecn represents the bias parameter, Serves as initial features for subsequent vehicle modeling.

[0020] Furthermore, in step 2), the vehicle historical trajectory data is clustered to generate a virtual lane centerline and extract lane direction information. A bidirectional long short-term memory network is then used to encode the lane centerline and generate a lane feature representation. The specific steps are as follows:

[0021] Step 2-1: Cluster the historical trajectory data of the vehicle. Assume that the historical trajectory of each vehicle i is

[0022]

[0023] in represents the two-dimensional position of the vehicle at time t, T h The vehicle clusters for each virtual lane are obtained by using the density-based hierarchical clustering algorithm (HDBSCAN) to group the vehicles. Where k is the number of virtual lanes. For each cluster C k , calculate the average value of the positions of all vehicles in the cluster at time t = 0 as the representative point of the center line of the virtual lane corresponding to the cluster:

[0024]

[0025] in is the two-dimensional coordinate of vehicle i at time t = 0, |c K | is cluster c K The number of vehicles in.

[0026] Step 2-2: In order to describe the driving trend of the lane, the present invention extracts the direction information from the center line of the virtual lane. k Discretize into a series of points Where T Lis the number of discrete points on the lane centerline, Represents the position coordinates of the kth virtual lane at the i-th discrete point. In order to capture the driving direction of the lane, the direction vector between adjacent points is calculated:

[0027]

[0028] in Represents the direction vector of the lane centerline at time t. To simplify the model, the global direction vector d is obtained by calculating the difference between the first and last points of the lane. k :

[0029]

[0030] This represents the global direction of travel for the lane.

[0031] Step 2-3: Generate lane feature vectors based on the location and direction information of the virtual lane. The feature vector of each virtual lane k is It can be obtained by concatenating the lane position information and direction information and mapping them. Specifically, the feature vector generation process is:

[0032]

[0033] Among them, l k is the lane location information, d k is the direction vector, w lane is the learnable weight matrix, b lane is the bias vector. This gives the feature representation of each virtual lane. as input to subsequent models.

[0034] Step 2-4: To better capture the temporal characteristics of the lane centerline, encode the discrete point sequence of each virtual lane. First, the direction information of the lane's discrete point sequence is input into a bidirectional long short-term memory network (Bi-LSTM) for encoding. The bidirectional long short-term memory network can simultaneously capture the forward and backward dependencies of the lane centerline, thereby more comprehensively representing the lane characteristics. and the corresponding direction vector Construct the input vector:

[0035]

[0036] in, represents the enhanced feature vector of the kth virtual lane at time t. Then, these concatenated input vector sequences are input into the bidirectional long short-term memory network for encoding. The forward and reverse hidden states of the bidirectional long short-term memory network are:

[0037]

[0038] in, represents the forward hidden state of the kth virtual lane at time t, represents the reverse hidden state of the kth virtual lane at time t. After concatenating the forward and reverse hidden states, the lane feature representation is obtained:

[0039]

[0040] where l' k is the final feature vector of the lane, t * The aggregation moment of the feature.

[0041] Furthermore, in step 3), in this step, the vehicle nodes encoded by the vehicle and the virtual lane nodes encoded by the lane are constructed into a heterogeneous graph, and information transfer between the nodes is realized by defining vehicle-vehicle edges and vehicle-lane edges. The specific steps are as follows:

[0042] Step 3-1: Node construction of heterogeneous graph, including vehicle nodes and lane nodes. For vehicle nodes, the feature vector after LSTM encoding is used As the initial feature of the vehicle node. The node feature of each vehicle i is

[0043]

[0044] in Represents the initial node feature vector of vehicle i in the heterogeneous graph. For lane nodes, the virtual lane feature vector l' after bidirectional LSTM encoding k As the initial feature of the lane node. The node feature of each virtual lane k is

[0045]

[0046] in represents the initial node feature vector of the kth virtual lane in the heterogeneous graph, and d is the dimension of the lane feature.

[0047] Step 3-2: In order to establish the relationship between nodes in the heterogeneous graph, two types of edges are defined: vehicle-vehicle edges. Specifically, if i and j are the numbers of two vehicles, the edge characteristics between vehicles are

[0048]

[0049] in and They represent the positions of vehicles i and j at time t = 0. This edge feature represents the spatial relative position between vehicles i and j and can reflect the social interaction relationship between the vehicles.

[0050] Vehicle-lane edge: This edge represents the relationship between a vehicle and a virtual lane. The edge feature is defined as the distance from the vehicle to the lane centerline. For vehicle i and virtual lane k, the edge feature is

[0051]

[0052] where l k is the representative position of the virtual lane k (the median point of the lane centerline). This edge feature represents the spatial relationship between the vehicle and the lane, capturing the constraints of the vehicle traveling on the road.

[0053] Step 3-3: Model the relationships between nodes through a message passing mechanism. Specifically, the message passing process includes interactions between vehicle nodes and between vehicle nodes and lane nodes. This invention uses graph convolution operations to update node features, thereby enabling message passing between nodes. For vehicle node i, the update formula is:

[0054]

[0055] in is the input feature of the vehicle-vehicle edge, is the input feature of the vehicle-lane edge, W and b are learnable parameters, σ is the Sigmoid activation function for gating, and g is the ReLU activation function for nonlinear transformation. After multiple layers of graph convolution, the features of each node are integrated with information from its neighboring nodes to generate a richer feature representation.

[0056] Furthermore, in step 4), in this step, the vehicle's motion features are deeply integrated with the structural features of the virtual lane, and the multi-head self-attention module is used to further capture the dependencies and interaction characteristics between vehicles, thereby improving the expressiveness of the integrated features. The specific steps are as follows:

[0057] Step 4-1: Introduction of attention mechanism. In order to enable the model to learn the importance of information between different nodes, the present invention introduces a distance-based attention mechanism. Under this mechanism, the distance between vehicle node i and each virtual lane is As the basis for calculating the attention weight, the smaller the distance between the lanes, the greater the impact on the vehicle. First, calculate the distance between the vehicle node i and the lane node k:

[0058]

[0059] in represents the minimum distance between vehicle i and virtual lane k, is the position of vehicle i at time t=0, is the position information of discrete points in the virtual lane k.

[0060] Then, the attention weight α is calculated based on the distance between the vehicle and the lane i,k , the weight determines the degree of attention paid to the information of vehicle node i and lane node k:

[0061]

[0062] Where β is the influence of the adjustment distance on the weight, and k is the total number of virtual lanes.

[0063] Step 4-2: Information aggregation and fusion. After obtaining the attention weights between the vehicle and the lane, perform weighted aggregation on the features of the lane nodes. Specifically, the features of the vehicle node i and the features of all adjacent lane nodes are weighted summed to obtain the aggregated lane feature L i :

[0064]

[0065] where l' k is the feature of lane node k, is the attention weight, which indicates the degree of attention paid by vehicle node i to lane node k. For each vehicle node i, its final feature representation will be composed of the vehicle’s own features and the aggregated lane features L i After splicing, they are fused through linear transformation to obtain the final enhanced feature representation:

[0066]

[0067] Where W φ and b φ are the weights and biases of the linear transformation, and || represents the feature concatenation operation.

[0068] Step 4-3: By integrating the fusion features of the vehicle node The feature matrix is formed and these features are weighted and aggregated through the self-attention module. The query, key, and value between vehicle node i and other vehicle nodes j are calculated. Specifically, the query matrix, key matrix, and value matrix of each attention head h are:

[0069]

[0070] in The learned weight matrix is responsible for converting the input features Converted into query, key and value representations. Next, calculate the attention score of each vehicle node i to other vehicle nodes j and normalize it through the softmax function:

[0071]

[0072] in represents the attention weight of vehicle i to neighbor vehicle j in the h-th attention head, represents the set of nodes adjacent to node i, It is the dot product calculation of query and key, indicating the similarity between nodes. k is the dimension of the key vector. Then, the attention weights calculated are Pair matrix V h Perform weighted aggregation to obtain the output features of each node i under the attention head h:

[0073]

[0074] Where d is the dimension of the feature, used as a normalization factor.

[0075] Finally, the outputs of all heads are concatenated into a long vector and linearly mapped to obtain the final node feature representation:

[0076]

[0077] Among them, W o and b o are the weights and biases of the linear transformation.

[0078] Step 4-4: Fusion of vehicle features and aggregated lane features through MLP weighted summation.

[0079] Calculate the weights of attention features and lane features:

[0080] α att,i =softmax(W att tanh(W att1 A i +b att1 ))

[0081] α lane,i =softmax(W att tanh(W lane1 L i +b lane1 ))

[0082] Where W att , W att1 , W lane1 , b att1 , b lane1 are learnable parameters.

[0083] Weighted summation generates fusion features:

[0084]

[0085] The weighted sum feature is input into the two-layer MLP for nonlinear transformation:

[0086]

[0087] in is a learnable parameter, output Finally, the features are fused.

[0088] Furthermore, in step 5), in the output decoder module, the features of the vehicle node are processed by multi-head self-attention, and the final fusion feature f is input into the parallel decoder to output trajectories of different modes. The future trajectory of vehicle i can be expressed as the decoder output o i Add the vehicle's current position at t = 0

[0089]

[0090] in, is the predicted position of vehicle i at time l, y=1,…,T f represents the future prediction time step.

[0091] Beneficial effects: The trajectory prediction method of the present invention addresses the problem that existing map-free vehicle trajectory prediction methods are difficult to fully explore the implicit structure of roads and lane constraints in the absence of high-precision map information. The present invention proposes a map-free vehicle trajectory prediction based on dynamic generation of virtual lanes and heterogeneous social graph convolutional networks. The density-based hierarchical clustering algorithm is used to generate virtual lane centerlines and extract direction information, and the bidirectional long short-term memory network is used to encode lane features, thereby overcoming the limitations of traditional methods in road feature extraction. At the same time, by constructing a heterogeneous graph containing vehicle nodes and virtual lane nodes, and using a message passing mechanism to deeply fuse vehicle motion features and lane structure information, the model's adaptability and prediction accuracy to complex traffic scenarios are significantly improved. In addition, a multi-head self-attention module is introduced to capture the interaction characteristics between vehicles, and lane prior information is again integrated in the decoding stage to further enhance the robustness and accuracy of trajectory prediction. Ultimately, this method achieves efficient and accurate vehicle trajectory prediction in the absence of high-precision maps, providing important support for the practical application of intelligent transportation and autonomous driving technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] Figure 1 This is a schematic diagram of the steps of a map-free vehicle trajectory prediction method based on dynamic virtual lane generation and heterogeneous social graph convolutional network of the present invention;

[0093] Figure 2This is a flow chart of a map-free vehicle trajectory prediction method based on dynamic virtual lane generation and heterogeneous social graph convolutional network of the present invention;

[0094] Figure 3 This is a structural diagram of a heterogeneous graph module for a map-free vehicle trajectory prediction method based on dynamic virtual lane generation and heterogeneous social graph convolutional network of the present invention;

[0095] Figure 4 This is a multi-head attention module structure diagram of a map-free vehicle trajectory prediction method based on dynamic virtual lane generation and heterogeneous social graph convolutional network of the present invention;

[0096] Figure 5 This is a diagram showing the overall structure of a map-free vehicle trajectory prediction method model based on dynamic virtual lane generation and heterogeneous social graph convolutional networks.

[0097] Figure 6 This is a training iteration graph for a map-free vehicle trajectory prediction method based on dynamic virtual lane generation and heterogeneous social graph convolutional network of the present invention;

[0098] Figure 7 This is a comparison chart of the predicted and actual trajectories of a map-free vehicle trajectory prediction method based on dynamic virtual lane generation and heterogeneous social graph convolutional networks in the present invention. DETAILED DESCRIPTION

[0099] The technical method of the present invention will be further described in detail below with reference to the accompanying drawings.

[0100] like Figure 1-2 As shown in FIG, a map-free vehicle trajectory prediction method based on dynamic virtual lane generation and heterogeneous social graph convolutional network includes the following steps:

[0101] Step 1) Obtain historical trajectory data of multiple vehicles from the traffic scene, convert the absolute coordinates into relative coordinates with the target vehicle as the reference, and extract the temporal dynamic characteristics of vehicle motion;

[0102] In the step 1), in the traffic scene, the trajectory information of each vehicle is recorded in the form of a time series, specifically represented as a two-dimensional coordinate

[0103]

[0104] in represents the trajectory position of vehicle i at time t, and Represent the horizontal and vertical coordinates of vehicle i at time t, t = -T h+1,…,0 represents the historical time step. To adapt to model training, the absolute coordinates are first converted to relative coordinates with the target vehicle as the reference, and the coordinate transformation normalization is performed on the data. Next, a discrete displacement sequence is generated for each vehicle and a vector definition is performed:

[0105]

[0106] in represents the displacement change of vehicle i at time t relative to the previous time t-1, and Represent the trajectory position of vehicle i at time t and time t-1 respectively. At the same time, a binary flag is introduced Right now Indicates the validity of the data at that time, and concatenates the two to form the input vector for each time step Finally, the input sequence of each vehicle is input into a single-layer long short-term memory network (LSTM) for encoding to obtain the hidden state

[0107]

[0108] in, is the hidden state of vehicle i at time t, w ecn represents the weight parameter, b ecn represents the bias parameter, As the initial feature of the subsequent vehicle modeling, f LSTM It is the forward operation of the LSTM network.

[0109] Step 2) Cluster the vehicle historical trajectories using a density-based hierarchical clustering algorithm (HDBSCAN) to generate virtual lane centerlines and extract the corresponding direction information. The lane centerlines and direction vectors are encoded using a bidirectional long short-term memory (LSTM) network to obtain virtual lane features.

[0110] In step 2), the vehicle's historical trajectory data is clustered to generate a virtual lane centerline and extract lane direction information. A bidirectional long short-term memory network is then used to encode the lane centerline and generate a lane feature representation. The specific steps are as follows:

[0111] Step 2-1: Cluster the historical trajectory data of the vehicle. Assume that the historical trajectory of each vehicle i is

[0112]

[0113] in represents the two-dimensional position of the vehicle at time t, T hThe vehicle clusters for each virtual lane are obtained by using the density-based hierarchical clustering algorithm (HDBSCAN) to group the vehicles. Where k is the number of virtual lanes. For each cluster C k , calculate the average value of the positions of all vehicles in the cluster at time t = 0 as the representative point of the center line of the virtual lane corresponding to the cluster:

[0114]

[0115] in is the two-dimensional coordinate of vehicle i at time t = 0, |c K | is cluster c K The number of vehicles in.

[0116] Step 2-2: In order to describe the driving trend of the lane, the present invention extracts the direction information from the center line of the virtual lane. k Discretize into a series of points Where T L is the number of discrete points on the lane centerline, Represents the position coordinates of the kth virtual lane at the i-th discrete point. In order to capture the driving direction of the lane, the direction vector between adjacent points is calculated:

[0117]

[0118] in Represents the direction vector of the lane centerline at time t. To simplify the model, the global direction vector d is obtained by calculating the difference between the first and last points of the lane. k :

[0119]

[0120] This represents the global direction of travel for the lane.

[0121] Step 2-3: Generate lane feature vectors based on the location and direction information of the virtual lanes. The feature vector of each virtual lane k is It can be obtained by concatenating the lane position information and direction information and mapping them. Specifically, the feature vector generation process is:

[0122]

[0123] Among them, l k is the lane location information, d k is the direction vector, w lane is the learnable weight matrix, b lane is the bias vector. This gives the feature representation of each virtual lane. as input to subsequent models.

[0124] Step 2-4: To better capture the temporal characteristics of the lane centerline, we encode the discrete point sequence of each virtual lane. First, the direction information of the lane's discrete point sequence is input into a bidirectional long short-term memory network (Bi-LSTM) for encoding. The bidirectional long short-term memory network can simultaneously capture the forward and backward dependencies of the lane centerline, thereby more comprehensively representing the lane characteristics. and the corresponding direction vector Construct the input vector:

[0125]

[0126] in, represents the enhanced feature vector of the kth virtual lane at time t. Then, these concatenated input vector sequences are input into the bidirectional long short-term memory network for encoding. The forward and reverse hidden states of the bidirectional long short-term memory network are:

[0127]

[0128] in, represents the forward hidden state of the kth virtual lane at time t, represents the reverse hidden state of the kth virtual lane at time t. After concatenating the forward and reverse hidden states, the lane feature representation is obtained:

[0129]

[0130] where l' k is the final feature vector of the lane, t * The aggregation moment of the feature.

[0131] Step 3) Construct a heterogeneous graph convolutional network, including vehicle nodes and virtual lane nodes, and define vehicle-vehicle edges and vehicle-lane edges. Use the graph neural network to transfer information and update features between nodes.

[0132] In step 3), in this step, the vehicle nodes encoded by the vehicle and the virtual lane nodes encoded by the lane are constructed into a heterogeneous graph, and information transfer between the nodes is achieved by defining vehicle-vehicle edges and vehicle-lane edges. The specific steps are as follows:

[0133] Step 3-1: Node construction of heterogeneous graph, including vehicle nodes and lane nodes. For vehicle nodes, the feature vector after LSTM encoding is used As the initial feature of the vehicle node. The node feature of each vehicle i is

[0134]

[0135] in Represents the initial node feature vector of vehicle i in the heterogeneous graph. For lane nodes, the virtual lane feature vector l' after bidirectional LSTM encoding k As the initial feature of the lane node. The node feature of each virtual lane k is

[0136]

[0137] in represents the initial node feature vector of the kth virtual lane in the heterogeneous graph, and d is the dimension of the lane feature.

[0138] Step 3-2: In order to establish the relationship between nodes in the heterogeneous graph, two types of edges are defined: vehicle-vehicle edges. Specifically, if i and j are the numbers of two vehicles, the edge characteristics between vehicles are

[0139]

[0140] in and They represent the positions of vehicles i and j at time t = 0. This edge feature represents the spatial relative position between vehicles i and j and can reflect the social interaction relationship between the vehicles.

[0141] Vehicle-lane edge: This edge represents the relationship between a vehicle and a virtual lane. The edge feature is defined as the distance from the vehicle to the lane centerline. For vehicle i and virtual lane k, the edge feature is

[0142]

[0143] where l k is the representative position of the virtual lane k (the median point of the lane centerline). This edge feature represents the spatial relationship between the vehicle and the lane, capturing the constraints of the vehicle traveling on the road.

[0144] Step 3-3: Model the relationships between nodes through a message passing mechanism. Specifically, the message passing process includes interactions between vehicle nodes and between vehicle nodes and lane nodes. This invention uses graph convolution operations to update node features, thereby enabling message passing between nodes. For vehicle node i, the update formula is:

[0145]

[0146] in is the input feature of the vehicle-vehicle edge, is the input feature of the vehicle-lane edge, W and b are learnable parameters, σ is the Sigmoid activation function for gating, and h is the ReLU activation function for nonlinear transformation. After multiple layers of graph convolution, the features of each node are integrated with information from its neighboring nodes to generate a richer feature representation.

[0147] Step 4) Use the multi-head self-attention module to further capture the dependencies and interactions between vehicles and improve the expressiveness of the fusion features;

[0148] In step 4), the vehicle's motion features are deeply integrated with the structural features of the virtual lane. The multi-head self-attention module is used to further capture the dependencies and interaction characteristics between vehicles, thereby improving the expressiveness of the integrated features. The specific steps are as follows:

[0149] Step 4-1: Introduction of attention mechanism. In order to enable the model to learn the importance of information between different nodes, the present invention introduces a distance-based attention mechanism. Under this mechanism, the distance between vehicle node i and each virtual lane is As the basis for calculating the attention weight, the smaller the distance between the lanes, the greater the impact on the vehicle. First, calculate the distance between the vehicle node i and the lane node k:

[0150]

[0151] in represents the minimum distance between vehicle i and virtual lane k, is the position of vehicle i at time t=0, is the position information of discrete points in the virtual lane k.

[0152] Then, the attention weight α is calculated based on the distance between the vehicle and the lane i,k , the weight determines the degree of attention paid to the information of vehicle node i and lane node k:

[0153]

[0154] Where β is the influence of the adjustment distance on the weight, and k is the total number of virtual lanes.

[0155] Step 4-2: Information aggregation and fusion. After obtaining the attention weights between the vehicle and the lane, perform weighted aggregation on the features of the lane nodes. Specifically, the features of the vehicle node i and the features of all adjacent lane nodes are weighted summed to obtain the aggregated lane feature L i :

[0156]

[0157] where l' kis the feature of lane node k, is the attention weight, which indicates the degree of attention paid by vehicle node i to lane node k. For each vehicle node i, its final feature representation will be composed of the vehicle’s own features and the aggregated lane features L i After splicing, they are fused through linear transformation to obtain the final enhanced feature representation:

[0158]

[0159] Where W φ and b φ are the weights and biases of the linear transformation, and || represents the feature concatenation operation.

[0160] Step 4-3: By integrating the fusion features of the vehicle node The feature matrix is formed and these features are weighted and aggregated through the self-attention module. For vehicle node i and other vehicle nodes j, the query, key, and value between them are calculated. Specifically, the query matrix, key matrix, and value matrix of each attention head h are:

[0161]

[0162] in The learned weight matrix is responsible for converting the input features Converted into query, key and value representations. Next, calculate the attention score of each vehicle node i to other vehicle nodes j and normalize it through the softmax function:

[0163]

[0164] in represents the attention weight of vehicle i to neighbor vehicle j in the h-th attention head, represents the set of nodes adjacent to node i, It is the dot product calculation of query and key, indicating the similarity between nodes. k is the dimension of the key vector. Then, the attention weights calculated are Pair matrix V h Perform weighted aggregation to obtain the output features of each node i under the attention head h:

[0165]

[0166] Where d is the dimension of the feature, used as a normalization factor.

[0167] Finally, the outputs of all heads are concatenated into a long vector and linearly mapped to obtain the final node feature representation:

[0168]

[0169] Among them, W o and b o are the weights and biases of the linear transformation.

[0170] Step 4-4: Fuse vehicle features and aggregated lane features through MLP and weighted summation.

[0171] Calculate the weights of attention features and lane features:

[0172] α att,i =softmax(W att tanh(W att1 A i +b att1 ))

[0173] α lane,i =softmax(W att tanh(W lane1 L i +b lane1 ))

[0174] Where W att , W att1 , W lane1 , b att1 , b lane1 are learnable parameters.

[0175] Weighted summation generates fusion features:

[0176]

[0177] The weighted sum feature is input into the two-layer MLP for nonlinear transformation:

[0178]

[0179] in is a learnable parameter, output Finally, the features are fused.

[0180] Step 5) In the output decoder, lane prior information is integrated and the integrated features are decoded to predict the future trajectory of the vehicle. End-to-end training and multimodal strategies are used to optimize the model performance and evaluate the prediction accuracy.

[0181] In step 5), in the output decoder module, the features of the vehicle node are processed by multi-head self-attention to obtain the final fusion feature f iThe input is fed into the parallel decoder to output trajectories of different modes. The future trajectory of vehicle i can be represented as the decoder output o i Add the vehicle's current position at t = 0

[0182]

[0183] in, is the predicted position of vehicle i at time l, t=1,…,T f represents the future prediction time step.

[0184] In response to the above problems, the present invention proposes a map-free vehicle trajectory prediction method based on dynamic generation of virtual lanes and heterogeneous social graph convolutional networks. This method uses historical vehicle trajectory data and automatically generates virtual lane lines that reflect the actual road structure through density-based hierarchical clustering algorithm (HDBSCAN) clustering. Lane direction vector information is introduced when aggregating the representative trajectories of vehicles in the cluster, and a bidirectional long short-term memory network (LSTM) is used to encode the lane centerline and direction information. At the same time, the encoded lane features are used as structured nodes and vehicle nodes to jointly construct a heterogeneous graph, and a message passing mechanism is used to achieve deep fusion of lane information and vehicle motion characteristics. Furthermore, virtual lane supplementary information is introduced in the input encoding and output decoding stages, thereby significantly improving the accuracy and robustness of vehicle trajectory prediction in the absence of high-precision maps.

[0185] The above description is merely an embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A map-free vehicle trajectory prediction method based on virtual lanes and heterogeneous social graph networks, characterized by: The following steps are involved: Step 1: Obtain historical vehicle trajectory data, convert absolute coordinates into relative coordinates, and use a single-layer long short-term memory network to encode the temporal characteristics of vehicle motion to obtain the initial vehicle features; Step 2: Based on the initial features of the vehicle, a density-based hierarchical clustering algorithm is used to cluster the historical trajectories and generate a virtual lane centerline l k And extract the direction information, and obtain the lane features through bidirectional long short-term memory network encoding; Step 3: Using the initial vehicle features as vehicle node features and the lane features as virtual lane node features, a heterogeneous graph containing vehicle-vehicle edges and vehicle-lane edges is constructed. Inter-node information transfer and feature updates are performed through a graph convolutional network to obtain updated vehicle node features. Step 4: Based on the updated vehicle node features and lane features, a multi-head self-attention module is introduced to capture the interaction characteristics between vehicles, and feature fusion is performed with lane prior information to obtain enhanced features; Step 5: Input the enhanced features into the decoder and integrate the lane information to predict the future trajectory of the vehicle.

2. The method according to claim 1, characterized in that The step 1 specifically includes: Step 1-1: Obtain historical trajectory data of multiple vehicles within a preset time window. The trajectory of each vehicle is represented as a two-dimensional coordinate sequence. in represents the trajectory position of vehicle i at time t, and Represent the horizontal and vertical coordinates of vehicle i at time t, t = -T h +1,…,0 represents the historical time step, T h is the length of the historical time window; Step 1-2: Calculate the relative displacement of each vehicle in adjacent time steps and Represent the trajectory position of vehicle i at time t and time t-1 respectively, and introduce the validity flag Constructing the input vector Step 1-3: Input sequence Input the single-layer long short-term memory network for encoding, and finally obtain the initial feature vector of the vehicle 3. The method according to claim 1, characterized in that The step 2 specifically includes: Step 2-1: The positions of N vehicles at time t=0 Perform density clustering to obtain K clusters c K ; Calculate the virtual lane centerline for each cluster in is the two-dimensional coordinate of vehicle i at time t = 0, |c K | is cluster c K the number of vehicles in the Step 2-2: Set each virtual lane k Discretize into multiple points and calculate the direction vector of adjacent points t=1,…,T L -1, where Represents the position coordinates of the k-th virtual lane at the t-th discrete point; Calculate the global direction vector in Represents the two-dimensional coordinates of the last point of the k-th simulated lane centerline in the discrete point sequence, Indicates the displacement from the starting point to the end point of the virtual lane centerline, represents the Euclidean distance between the first and last points; Step 2-3: Position the lane and direction information After splicing, the data is input into the bidirectional long short-term memory network for encoding; Concatenate the final hidden states of the forward and reverse LSTM networks to obtain the lane feature vector where l' k is the final feature vector of the lane, t * is the aggregation moment of the feature, w lane is the weight matrix, b lane is the bias vector, represents the forward hidden state of the kth virtual lane at time t, represents the reverse hidden state of the kth virtual lane at time t.

4. The method according to claim 3, characterized in that In step 3, the node feature of each vehicle i is: Lane node features are: Where d is the dimension of lane features; The edge features between vehicles are: where τ i,0 and τ j,0 Represent the positions of vehicles i and j at time t = 0 respectively; for vehicle i and virtual lane k, the edge feature is where l k is the representative position of the virtual lane k, that is, the median point of the lane centerline.

5. The method according to claim 4, characterized in that In step 3, the relationship between nodes is modeled through the information transmission mechanism. For vehicle node i, the update formula is defined as: in is the input feature of the vehicle-vehicle edge, is the input feature of the vehicle-lane edge, W and b are learnable parameters, σ is the Sigmoid activation function for gating, and g is the ReLU activation function for nonlinear transformation.

6. The method according to claim 1, characterized in that The step 4 specifically includes: Step 4-1: Calculate the minimum distance between vehicle i and lane node k in represents the minimum distance between vehicle i and virtual lane k, is the position of vehicle i at time t=0, is the position information of discrete points in virtual lane k; Calculate attention weight based on minimum distance Where β is the influence of the adjustment distance on the weight, and k is the total number of virtual lanes; Step 4-2: Aggregate lane features based on attention weights where l' k is the feature of lane node k, is the attention weight, which indicates the degree of attention paid by vehicle node i to lane node k; Fusion of vehicle features and lane feature L i , get the fusion features Where W φ and b φ are the weights and biases of the linear transformation, || represents the feature concatenation operation; Step 4-3: By integrating the fusion features of the vehicle node Form a feature matrix and perform weighted aggregation on these features through the self-attention module to obtain the output feature of each node i under the attention head h where Q h , K h 、V h They are the query matrix, key matrix and value matrix of the attention head h, d k is the dimension of the key vector; The outputs of all heads are concatenated into a long vector and linearly mapped to obtain the final node feature representation. Among them, W o and b o are the weights and biases of the linear transformation; Step 4-4: A i and L i Weighted summation generates fusion features Will Input two layers of MLP for nonlinear transformation and output the final fusion feature f i .

7. The method according to claim 1, characterized in that The step 5 specifically includes: The final fusion feature f i Input multi-layer perceptron decoder; Predict the trajectory of multiple steps into the future. The future trajectory of vehicle i is represented by the decoder output o i Add the vehicle's current position at t = 0 A multimodal output strategy is adopted to predict multiple possible trajectories and their probability distribution.

8. The method according to any one of claims 1 to 7, characterized in that The spacing shape of the center lines of the virtual lanes is a stripe, a dot matrix or a custom geometric pattern.

9. The method according to any one of claims 1 to 7, characterized in that The method is applicable to map-free vehicle trajectory prediction in augmented reality and autonomous driving scenarios.

Citation Information

Cited By

  • Physical system trajectory prediction method and system based on isotropic graph neural network

    CN121389682A