Track prediction method based on space-time interaction feature fusion

By constructing a traffic scene expression model and using a multi-scale spatiotemporal graph convolution model, the spatial and temporal interaction characteristics of traffic participants are extracted, and the problem that intelligent driving system is difficult to accurately predict the future behavior of surrounding traffic participants in complex traffic scenarios is solved, and more efficient trajectory prediction is achieved.

CN119928890AActive Publication Date: 2025-05-06SAIC GM WULING AUTOMOBILE CO LTD

Patent Information

Application Number
CN202510056758.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-06
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

The existing intelligent driving system is difficult to accurately predict the future behavior of surrounding traffic participants in complex traffic scenarios, resulting in low trajectory prediction accuracy.

Method used

By constructing a traffic scene expression model, the spatiotemporal interaction characteristics of traffic participants are extracted, and trajectory prediction is performed using a multi-scale spatiotemporal graph convolution model and a trajectory prediction model of the encoder-decoder structure.

Benefits of technology

Effectively capture the spatial correlation of traffic participants' behavior on different time scales, improve the accuracy of trajectory prediction, and provide more accurate information on movement of surrounding obstacles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119928890A_ABST
    Figure CN119928890A_ABST
Patent Text Reader

Abstract

The invention provides a trajectory prediction method based on space-time interaction feature fusion. The trajectory prediction method comprises the following steps: collecting traffic participant information on a road in real time; constructing a traffic scene expression model based on the traffic participant information; constructing a multi-scale space-time diagram convolution model to enable the multi-scale space-time diagram convolution model to extract space-time interaction features of the traffic participants based on the traffic scene expression model; and constructing a trajectory prediction model to enable the trajectory prediction model to perform trajectory prediction based on the space-time interaction characteristics, and obtaining a prediction trajectory. According to the trajectory prediction method based on space-time interaction feature fusion provided by the invention, the traffic scene expression is constructed and the space-time interaction features of the traffic participants are extracted, so that the finally obtained vehicle prediction trajectory can capture the spatial correlation of the behaviors of the traffic participants on different time scales; and multi-scale space-time fusion interaction features are extracted, so that the precision of trajectory prediction can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent driving technology, and in particular to a trajectory prediction method based on spatiotemporal interactive feature fusion. Background Art

[0002] Accurately predicting the future trajectory of surrounding traffic participants is crucial to the intelligent driving system. It is a prerequisite for intelligent driving cars to conduct reasonable and efficient behavioral interactions and collision detection with other participants. In complex traffic scenarios, the future behavior of surrounding traffic participants is highly uncertain, and it is difficult for intelligent driving systems to characterize the highly uncertain future behavior of surrounding traffic participants with a deterministic prediction trajectory. Therefore, each decision cycle of the intelligent driving system not only needs to perform collision risk detection based on the most likely behavior of surrounding traffic participants, but also needs to comprehensively examine the probability distribution of all possible future behaviors.

[0003] In order to comprehensively and truthfully describe the highly uncertain future behaviors of surrounding traffic participants, some recent technical studies have achieved multimodal trajectory prediction of surrounding traffic participants by considering the road topology. However, due to the small number of trajectory modes predicted, it still cannot comprehensively describe the highly uncertain future behaviors of surrounding traffic participants. At the same time, many methods lack consideration of vehicle kinematics, and the generated trajectories sometimes do not conform to the vehicle kinematic constraints, making it difficult to provide accurate surrounding obstacle motion information for subsequent decision-making and planning links.

[0004] There are also many prediction methods that have begun to improve trajectory prediction accuracy by modeling the behavioral interactions between traffic participants. However, existing trajectory prediction methods do not adequately model the spatial and temporal correlations of behavioral interactions, resulting in low trajectory prediction accuracy. Most of these trajectory prediction methods only extract behavioral interactions from spatial correlations, and rarely pay attention to the temporal correlations of interactions. Although other works model both spatial and temporal correlations at the same time, they do so separately, ignoring the capture of more spatial correlations at different time scales. Summary of the invention

[0005] The present invention aims to provide a trajectory prediction method based on the fusion of spatiotemporal interaction features to solve the above technical problems. By constructing a traffic scene expression and extracting the spatiotemporal interaction features of traffic participants, the spatial correlation of the behaviors of traffic participants at different time scales can be captured, and the accuracy of trajectory prediction can be effectively improved.

[0006] In order to solve the above technical problems, the present invention provides a trajectory prediction method based on spatiotemporal interactive feature fusion, comprising the following steps:

[0007] Collect information of traffic participants on the road in real time;

[0008] Construct a traffic scene expression model based on traffic participant information;

[0009] Construct a multi-scale spatiotemporal graph convolution model to extract the spatiotemporal interaction features of traffic participants based on the traffic scene expression model;

[0010] Construct a trajectory prediction model to enable the trajectory prediction model to perform trajectory prediction based on spatiotemporal interaction features and obtain a predicted trajectory.

[0011] The above scheme constructs traffic scene expressions and extracts the spatiotemporal interaction characteristics of traffic participants, so that the final vehicle prediction trajectory can capture the spatial correlation of traffic participants' behaviors at different time scales, extracts multi-scale spatiotemporal fusion interaction features, and can effectively improve the accuracy of trajectory prediction.

[0012] Furthermore, the traffic scene expression model includes a Euclidean distance adjacency matrix, a potential collision adjacency matrix and an adaptive adjacency matrix; the traffic scene expression model is constructed based on traffic participant information, including:

[0013] Based on the traffic participant information, the Euclidean distance adjacency matrix is ​​constructed, which is specifically expressed as:

[0014]

[0015] In the formula, is the Euclidean distance adjacency matrix An element in , which represents the adjacency relationship between traffic participants i and j at time t; represents the distance between traffic participants i and j at time t; D close Represents the distance threshold between neighbor participants; 1 means that at time t, traffic participant i and traffic participant j are neighbor participants; 0 means that at time t, traffic participant i and traffic participant j are not neighbor participants.

[0016] It should be noted that once a traffic participant and a vehicle are identified as neighbors, the behavior of the traffic participant will be considered to have an impact on the vehicle trajectory in the subsequent prediction stage. close The behavior of traffic participants outside the vehicle will still have an impact on the vehicle. Therefore, in order to solve the problem that the fixed distance threshold of the Euclidean distance adjacency matrix cannot be adjusted dynamically, this solution obtains the current speed and current direction of the traffic participant based on the traffic participant information, obtains several directed line segments, and constructs the future trajectory model of the traffic participant based on the directed line segments, which is specifically expressed as:

[0017]

[0018] Where, Lpred Indicates the length of the directed line segment; u cur is the current speed of the traffic participant; t pred is the prediction time domain; θ represents the direction of the future trajectory; θ cur The current direction of the traffic participant;

[0019] The potential collision adjacency matrix is ​​constructed based on the future trajectory model of traffic participants, which is specifically expressed as:

[0020]

[0021] In the formula, is the potential collision adjacency matrix An element in represents the potential collision relationship between traffic participants i and j at time t. In this calculation, traffic participant i is the vehicle itself; AB is the future trajectory of the vehicle calculated based on the future trajectory model of the traffic participant; CD is the future trajectory of the current traffic participant calculated based on the future trajectory model of the traffic participant; if there is an intersection (x, y), it means that the future trajectory of the vehicle overlaps with the future trajectory of the current traffic participant, and there is a potential collision relationship. At this time, Assign a value of 1.

[0022] It should be noted that the potential collision adjacency matrix A simplified collision risk assessment algorithm is used to find traffic participants that have mutual impact. Because the future trajectory is still unknown in the trajectory prediction stage, this scheme uses a directed line segment generated according to the current speed and current direction of the participant as the future trajectory for collision detection.

[0023] On the basis of the above-mentioned construction of the Euclidean distance adjacency matrix and the potential collision adjacency matrix, in order to continue to explore the hidden behavioral interaction relationship, an adaptive adjacency matrix can be constructed based on the traffic participant information, which is specifically expressed as:

[0024]

[0025] In the formula, is the latent adaptive adjacency matrix An element in , which represents the adaptive relationship between traffic participants i and j at time t; E i Indicates that traffic participant i is taken as the source node, E j Indicates that traffic participant j is taken as the target node; ReLU is the activation function, which is used to eliminate weak connections and increase the nonlinearity of the network layer; SoftMax is the normalization function, which is used to normalize the adaptive adjacency matrix; the adaptive adjacency matrix It is used to capture the dependency between the source node and the target node embedding, that is, to represent the interaction between traffic participants i and j.

[0026] It should be noted that node embedding is to map a node to a multi-dimensional vector and express it with this multi-dimensional vector, or more commonly, to replace this node as input to the subsequent neural network layer.

[0027] Furthermore, the multi-scale spatiotemporal graph convolution model is constructed so that the multi-scale spatiotemporal graph convolution model extracts the spatiotemporal interaction features of traffic participants based on the traffic scene expression model, including:

[0028] Construct a multi-scale spatiotemporal graph convolutional model consisting of multiple expanded TCN-GCN spatiotemporal layers; the GCN in each spatiotemporal layer is used to capture the interactive relationships between traffic participants at different time scales, and the expanded TCN in each spatiotemporal layer is used to capture the movement trends of traffic participants over time;

[0029] The multi-scale spatiotemporal graph convolution model is used to extract the spatiotemporal interaction features of traffic participants based on the traffic scene expression model;

[0030] Among them: Dilated TCN stands for Dilated Temporal Convolutional Network; GCN stands for Graph Convolutional Network.

[0031] The expanded TCN constructed by the above scheme achieves the best balance between prediction performance and computational efficiency, providing feasibility for the algorithm to be deployed on low-computing domain control chips; and constructs a way to represent the traffic scene expression model through three adjacency matrices: Euclidean distance adjacency matrix, potential collision adjacency matrix and adaptive adjacency matrix, which is conducive to fully modeling prior knowledge and thus optimizing the performance of the GCN graph convolution layer.

[0032] Furthermore, in the multi-scale spatiotemporal graph convolution model, for any dilated TCN, the dilated convolution operation F(t) at time t is specifically expressed as:

[0033]

[0034] Wherein: Q represents a one-dimensional input sequence; f represents a filter; a represents the element index in the filter; k1 represents the size of the filter; d represents the dilation factor, which is used to control the distance between the elements of the convolution kernel; in the multi-scale spatiotemporal graph convolution model, the d value of the odd-numbered dilated TCN-GCN spatiotemporal layer is 1, and the d value of the even-numbered dilated TCN-GCN spatiotemporal layer is 2.

[0035] Furthermore, in the multi-scale spatiotemporal graph convolution model, for any GCN, its output spatiotemporal interaction feature Z is specifically expressed as:

[0036]

[0037] In the formula, k2 represents the diffusion step; A power series representing the Euclidean distance adjacency matrix; A power series representing the potential collision adjacency matrix; represents the power series of the adaptive adjacency matrix; X is the output result of the expanded TCN; is the learnable parameter matrix in the convolution operation; Indicates the last observation timestamp t obs The Euclidean distance adjacency matrix of ; Indicates the last observation timestamp t obs The potential collision adjacency matrix of ; Indicates the last observation timestamp t obs The Euclidean distance adjacency matrix, the potential collision adjacency matrix and the adaptive adjacency matrix constitute a traffic scene expression model.

[0038] Furthermore, the constructing of the trajectory prediction model so as to make the trajectory prediction model perform trajectory prediction based on the spatiotemporal interaction features and obtain the predicted trajectory includes:

[0039] Constructing a trajectory prediction model of an encoder-decoder structure; wherein both the encoder and the decoder have LSTM neural networks inside;

[0040] The spatiotemporal interaction features are input into the encoder, and after learning and processing by the LSTM neural network, the context feature information is obtained;

[0041] The context feature information is input into the decoder, processed and learned by the LSTM neural network, and combined with the current position of the vehicle to obtain several sequence points;

[0042] Perform trajectory prediction based on several sequence points to obtain the predicted trajectory.

[0043] Furthermore, the constructing of the trajectory prediction model so as to make the trajectory prediction model perform trajectory prediction based on the spatiotemporal interaction features and obtain the predicted trajectory includes:

[0044] Construct a trajectory prediction model, wherein the trajectory prediction model includes a target lane prediction sub-model, a trajectory cluster prediction sub-model and a trajectory evaluation output sub-model; wherein:

[0045] The target lane prediction submodel is used to obtain the vehicle target lane based on pre-acquired vehicle historical trajectory data and lane data;

[0046] The trajectory cluster prediction submodel is used to obtain the vehicle initial trajectory cluster based on the pre-acquired vehicle historical trajectory data and the vehicle target lane;

[0047] The trajectory evaluation output sub-model is used to construct a vehicle trajectory cluster based on the vehicle initial trajectory cluster and evaluate the vehicle trajectory cluster based on the spatiotemporal interaction characteristics, calculate the probability distribution of the vehicle trajectory cluster, and obtain the predicted trajectory.

[0048] Furthermore, the target lane prediction sub-model is used to obtain the vehicle target lane based on the pre-acquired vehicle historical trajectory data and lane data, including:

[0049] Based on the pre-acquired vehicle historical trajectory data and lane data, the relative position of the vehicle and the candidate lane is obtained;

[0050] Based on the cross-attention mechanism and the relative position of the vehicle and the candidate lane, the relevance of the center line of the candidate lane to the vehicle's historical trajectory is calculated;

[0051] The target lane of the vehicle is obtained based on the correlation of the center line of the candidate lane with the historical trajectory of the vehicle.

[0052] In the above scheme, if all lanes around the vehicle are directly sampled without consideration and the corresponding vehicle trajectory clusters are generated, the number of vehicle trajectory clusters will be too large, increasing the complexity of the prediction model. Therefore, in order to reduce the sampling space of the prediction trajectory cluster, this scheme uses the cross-attention mechanism to first select the target lane where the vehicle is most likely to travel, and then generate a vehicle trajectory cluster based on this lane.

[0053] Furthermore, obtaining the target lane of the vehicle mainly depends on the relative position relationship between the vehicle's historical trajectory and the centerline of the lanes around it, that is, the relative position of the vehicle and the candidate lane. In order to characterize the above relative position relationship, the Frenet coordinates s and d of the vehicle's historical trajectory on the centerline of each candidate lane are first calculated. Among them, s represents the longitudinal distance along the centerline, and d represents the lateral distance in the direction perpendicular to the centerline. For curved lanes, the Frenet coordinate system can more clearly and directly describe the relative position relationship between the vehicle and the lane than the Cartesian coordinate system. After using the Frenet coordinates to describe the relative position relationship between the vehicle and the candidate lane, this scheme can use the cross-attention mechanism to process the relationship between the sequence of the vehicle's historical trajectory and the sequence of the centerline path points of the candidate lane. Among them, the cross-attention mechanism and the self-attention mechanism both belong to the attention network, which models and captures the dependency of the input sequence through the (Q, K, V) triple tuple. According to the different input sources of the three vectors Q, K, and V, in this scheme, the candidate lane centerline can be used as the query vector Q of the cross-attention mechanism, and the vehicle historical trajectory can be used as the key vector K and the value vector V, thereby constructing the self-attention mechanism and the cross-attention mechanism.

[0054] Furthermore, in the self-attention mechanism, Q, K, and V are the same feature input vector X multiplied by three different weight matrices W. Q ,W K ,W V The three generated vectors of the same dimension, i.e. Q = XW Q ,K=XW K ,V=XW V , the weight matrix can be obtained through neural network training. The essence of the self-attention mechanism is to calculate the mutual influence between the internal elements of the input vector X, which can be implemented by using the Transformer network. The query vector Q in the cross-attention mechanism is composed of the lane center line X A K and V are converted from the vehicle's historical trajectory X. B Transformed, it is mainly used to model the relationship between two different sequences, that is, Q = X A W Q ,K=X B W K ,V=X B W V .

[0055] Therefore, in the target lane prediction sub-model, the lane centerline X A As the query vector Q of the cross attention mechanism, the vehicle history trajectory X B As the key vector K and the value vector V, the two are used together to perform cross-attention operation to obtain the relevance of the lane centerline to the historical trajectory. The specific steps are as follows:

[0056] 1) Feature Mapping: Q = X A W Q ,K=X B W K ,V=X B W V , X A is the lane centerline, X B is the vehicle history trajectory;

[0057] 2) Similarity calculation: Through the dot product operation of Q and K, the similarity score between the query vector Q and the key vector K is calculated. These similarity scores represent the similarity between X A Each element in B The relationship between all elements in can be expressed as:

[0058]

[0059] Among them, S is the similarity score matrix, is a scaling factor used to prevent the dot product value from being too large;

[0060] 3) Attention weight calculation: The similarity score is normalized by the Softmax function to obtain the attention weight, which reflects the X A Each element in B The importance of all elements in can be expressed as:

[0061] W a A→B =Softmax(S A→B )

[0062] Among them, W a is the attention weight;

[0063] 4) Weighted summation: The attention weights are weighted summed with the value vector V to obtain the output representation after cross-attention adjustment, which can be expressed as:

[0064] O A→B =W a A→B V B

[0065] In the formula, O A→B The dimension of O is consistent with the dimension of the query vector Q. Through the above attention mechanism, O A→B Each element in the query vector Q contains the association information of the key vector K and the value vector V, which models the lane centerline sequence X. A and the vehicle history trajectory sequence X B The mutual relationship between them.

[0066] Furthermore, after uniformly encoding the lane centerline information and the vehicle's historical trajectory to obtain their mutual connection, all candidate lanes can be constructed as an [n, c]-dimensional vector: each candidate lane is represented by a c-dimensional vector, and the number of lanes is increased sequentially to n lanes to form an [n, c]-dimensional lane array. Among them, n is the number of candidate lanes, and c is the characteristic dimension of the lane. Then, the Transformer self-attention network is used to learn the connection and difference between the candidate lanes, and the multi-layer perceptron MLP and Softmax are used to output the probability distribution of the candidate lanes. Finally, considering the possible prediction error, this scheme strikes a balance between sampling efficiency and prediction accuracy, and selects the top three lanes with the largest probability as the vehicle target lane prediction output.

[0067] Furthermore, the trajectory cluster prediction sub-model is used to obtain the vehicle initial trajectory cluster based on the pre-acquired vehicle historical trajectory data and the vehicle target lane, including:

[0068] Construct the Frenet coordinate system of the lane centerline of the vehicle target lane;

[0069] In the Frenet coordinate system of the lane centerline, the vehicle target lane is spatially sampled based on the local path planning method to obtain the first path polynomial curve;

[0070] Convert the first path polynomial curve into a Cartesian coordinate system to obtain a second path polynomial curve;

[0071] The curvature of the second path polynomial curve is detected, and the path polynomial curve that meets the preset requirements is screened to obtain the initial vehicle trajectory cluster.

[0072] It should be noted that in order to generate a predicted trajectory cluster that conforms to the road topology and vehicle kinematics, the idea of ​​local path planning is applied to perform spatial sampling of the target lane, and a polynomial curve is used to express the initial vehicle trajectory cluster, which may include the following steps:

[0073] 1) Sampling is performed in the Frenet coordinate system of the lane centerline, where s represents the longitudinal distance along the centerline and d represents the lateral distance perpendicular to the centerline. The target lane centerline is fitted as a function of the parameter s through a cubic spline curve:

[0074]

[0075] In the formula, is the spline curve coefficient, x, y is the value of the center line point in the Cartesian coordinate system. The parameterized center line can facilitate the subsequent restoration of the trajectory cluster to the Cartesian coordinate system through s interpolation.

[0076] 2) The longitudinal distance s and lateral distance d of the vehicle initial trajectory cluster in Frenet coordinates are constructed as the fourth-order and fifth-order polynomials of the prediction time t respectively:

[0077]

[0078] In the formula, for the quartic polynomial, the initial state can be combined and terminal state Five equations are established for solution. The above states correspond to the initial longitudinal distance, initial longitudinal velocity, initial longitudinal acceleration, terminal longitudinal velocity, and terminal longitudinal acceleration. Similar to the lateral distance d, the fifth-order polynomial coefficients are obtained. The initial states are then and terminal state Five equations are created and solved.

[0079] 3) Sampling different terminal states S t and terminal state D tBy combining the above, a series of first path polynomial curves in Frenet coordinates can be generated, and then the center line parameterized in step 1) is interpolated and restored to the Cartesian coordinate system to obtain the second path polynomial curve. Among them, the terminal speed sampling range is set as follows:

[0080]

[0081] In the formula, is the current speed, which is obtained by the historical trajectory point difference; t pred is the prediction time domain; a max,brake and a max,acc The maximum deceleration and maximum acceleration are respectively. Then the terminal speed is sampled at a resolution of 1km / h, the terminal lateral distance is sampled at a resolution of 0.5m, and the driving area of ​​the vehicle along the target lane is sampled through longitudinal and lateral combinations, and the second path polynomial curve is generated.

[0082] 4) Perform curvature detection on the second path polynomial curve to eliminate unreasonable candidate trajectories with too large curvature. The trajectory cluster curvature calculation can use the triangle circumscribed circle curvature method, and finally select the path polynomial curve that meets the preset requirements according to the calculation results to obtain the vehicle initial trajectory cluster.

[0083] Furthermore, the trajectory evaluation output sub-model is used to construct a vehicle trajectory cluster based on the vehicle initial trajectory cluster and evaluate the vehicle trajectory cluster based on the spatiotemporal interaction characteristics, calculate the probability distribution of the vehicle trajectory cluster, and obtain the predicted trajectory, including:

[0084] The vehicle trajectory cluster and the vehicle historical trajectory are uniformly represented for heterogeneous information to construct the vehicle trajectory cluster;

[0085] The vehicle trajectory clusters and the temporal and spatial interaction features are integrated and encoded through the cross-attention mechanism, and the preset Transformer self-attention model is used to capture the temporal motion characteristics of the vehicle trajectory clusters. Finally, the probability distribution of the vehicle trajectory clusters is output through the MLP multi-layer perceptron and the Softmax layer.

[0086] The predicted trajectory is obtained based on the probability distribution of vehicle trajectory clusters.

[0087] It should be noted that on structured roads, human drivers will first determine which lane the surrounding vehicles are driving in, and then continue to determine their relative position relationship with the vehicle on this basis. Therefore, lane guidance information is crucial for trajectory prediction. However, how to encode multi-source heterogeneous information such as lane lines and historical trajectories of adjacent vehicles and learn their mutual relationship with neural networks is a major difficulty. Because lane lines are a series of waypoints, unlike the historical trajectory information of adjacent vehicles, the waypoint sequence does not imply speed and acceleration characteristics. In addition, in general, there is not only one lane around the target vehicle, which greatly hinders the network model from learning the relationship between the vehicle's historical trajectory and the correct lane during training.

[0088] In order to solve the above technical problems, this solution splices the vehicle's historical trajectory with each trajectory in the predicted vehicle initial trajectory cluster to construct a complete vehicle trajectory cluster from the starting point of the historical trajectory to the end point of the predicted trajectory. The predicted vehicle trajectory cluster itself is a motion trajectory containing speed and acceleration information. At the same time, when generating the predicted vehicle trajectory cluster, the road topology is also considered, reflecting the changing trend of lane direction and position. Therefore, by combining the historical trajectory with the vehicle initial trajectory cluster, the unified representation of static map information and dynamic vehicle motion trajectory heterogeneous information is achieved.

[0089] The Transformer's self-attention mechanism can be further used to learn continuous motion features from historical vehicle trajectories to predict vehicle trajectory clusters. The complete vehicle trajectory cluster constructed above is the x, y coordinates in the Cartesian coordinate system. In the scenario guided by the lane centerline, converting the x, y coordinates to the Frenet coordinates s and d based on the lane centerline can obtain a more direct and essential relative position relationship, thereby accelerating the training of the neural network. However, due to special scenarios such as curves, the distance s of the xy trajectory projected along the lane centerline is not linearly related to the vehicle speed. Therefore, the trajectory feature vector designed in this scheme uses the more essential vehicle speed v and the distance d in the direction perpendicular to the lane centerline to characterize the vehicle's forward and backward movement trend.

[0090] Furthermore, after the road topology information and vehicle motion trajectory information are uniformly represented, the predicted vehicle trajectory cluster that is complete in time series includes the vehicle's historical trajectory and predicted candidate trajectory. Therefore, by inputting each complete trajectory in time series, the motion characteristics of the trajectory points before and after the trajectory sequence can be encoded through the self-attention mechanism, and the best candidate trajectory that best matches the previous and next motion trends can be found. Therefore, the trajectory evaluation output sub-model can be constructed to evaluate the vehicle's initial trajectory cluster to obtain the probability distribution and predicted trajectory of the final vehicle trajectory cluster.

[0091] The above scheme can predict the future trajectory of the vehicle itself and surrounding vehicles. It takes into account the road topology and vehicle kinematic constraints, and finally obtains a multimodal predicted vehicle trajectory cluster and its probability distribution that fully characterizes the uncertain behavior of the vehicle. It can provide more accurate surrounding obstacle motion information for subsequent decision-making and planning links, and obtain the final predicted trajectory. The generated vehicle trajectory cluster revolves around the vehicle's maximum probability predicted trajectory, covering almost all possible vehicle movement intentions, and generates corresponding probabilities for each trajectory in the predicted trajectory cluster, which has the advantage of accurate long-term multimodal trajectory prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] Figure 1 A schematic flow chart of a trajectory prediction method based on spatiotemporal interactive feature fusion provided by an embodiment of the present invention;

[0093] Figure 2 An example of a time-space diagram provided for an embodiment of the present invention;

[0094] Figure 3 A multi-modal trajectory prediction flow chart considering road topology and vehicle kinematic constraints provided by one embodiment of the present invention;

[0095] Figure 4 A schematic diagram of an encoder-decoder structure of an LSTM neural network provided in one embodiment of the present invention;

[0096] Figure 5 A schematic diagram of a target lane prediction sub-model provided by an embodiment of the present invention;

[0097] Figure 6 A schematic diagram of a vehicle initial trajectory cluster provided by an embodiment of the present invention;

[0098] Figure 7 A schematic diagram of a triangle circumscribed circle curvature method provided by an embodiment of the present invention;

[0099] Figure 8 A schematic diagram of a method for splicing a historical trajectory with an initial vehicle trajectory provided by an embodiment of the present invention;

[0100] Fig. 9 A schematic diagram of a trajectory assessment output sub-model architecture provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0101] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0102] See also Figure 1 , this embodiment provides a trajectory prediction method based on spatiotemporal interactive feature fusion, comprising the following steps:

[0103] S1: Real-time collection of traffic participant information on the road;

[0104] In this embodiment, traffic participant information includes category information such as vehicles, bicycles, pedestrians, and other basic information, where category information is a dimension of the state quantity. When solving the adjacency matrix later, different adjacency matrices are not explicitly and specifically generated according to different categories, but their influence is generally and uniformly considered or learned indirectly in the potential collision risk and adaptive adjacency matrix.

[0105] S2: Construct a traffic scene expression model based on traffic participant information;

[0106] S3: Construct a multi-scale spatiotemporal graph convolution model to extract the spatiotemporal interaction features of traffic participants based on the traffic scene expression model;

[0107] S4: Construct a trajectory prediction model so that the trajectory prediction model performs trajectory prediction based on the spatiotemporal interaction features to obtain a predicted trajectory.

[0108] This embodiment constructs a traffic scene expression and extracts the spatiotemporal interaction features of traffic participants, so that the final vehicle prediction trajectory can capture the spatial correlation of traffic participants' behaviors at different time scales, extracts multi-scale spatiotemporal fusion interaction features, and can effectively improve the accuracy of trajectory prediction.

[0109] It should be noted that the traffic scene expression model represents the traffic scene as a space-time graph in graph theory, where the nodes in the graph represent traffic participants, and the edges in the graph represent the interactive relationship between two traffic participants. Then, this space-time graph can be mathematically represented as an adjacency matrix. The following embodiments consider the three aspects of Euclidean distance, potential collision and adaptability, and construct three adjacency matrices: Euclidean distance adjacency matrix, potential collision adjacency matrix and adaptive adjacency matrix. The three adjacency matrices are combined to jointly express the true relationship between the participants in the traffic scene.

[0110] In one embodiment, the traffic scene expression model includes a Euclidean distance adjacency matrix, a potential collision adjacency matrix, and an adaptive adjacency matrix. Figure 2 , in order to model the dynamic interaction relationship between traffic participants, Figure 2 The traffic scene in a is represented as Figure 2 b and Figure 2 c shows the space-time graph G t = {V t ,E t}. Among them, G t Represents the traffic behavior impact relationship diagram at time t with traffic participants as nodes. The nodes in the time-space graph represent traffic participants, and the edges in the graph represent the interaction relationship between two traffic participants. The node set V t Each node in E corresponds to a traffic participant. t is a set of edges, representing the relationship between two participants at time t. Then the adjacency matrix can be used to mathematically represent the spatiotemporal graph structure, specifically:

[0111] The traffic scene expression model based on traffic participant information is constructed, including: constructing a Euclidean distance adjacency matrix based on traffic participant information, for details, see Figure 2 b, Figure 2 b depicts the Euclidean distance adjacency matrix scene expression method, which uses distance to characterize and abstract traffic scenes. Each solid circle represents a traffic participant node. The Euclidean distance adjacency matrix can be expressed as:

[0112]

[0113] In the formula, is the Euclidean distance adjacency matrix An element in , which represents the adjacency relationship between traffic participants i and j at time t; represents the distance between traffic participants i and j at time t; D close Represents the distance threshold between neighbor participants; 1 means that at time t, traffic participant i and traffic participant j are neighbor participants; 0 means that at time t, traffic participant i and traffic participant j are not neighbor participants.

[0114] It should be noted that once a traffic participant and a vehicle are identified as neighboring participants, the behavior of the traffic participant will be considered to have an impact on the vehicle trajectory in the subsequent prediction stage. Figure 2 As shown in b, from the perspective of the red target vehicle, the radius is D closeAll traffic participants within the blue circle are its neighbors. However, in real traffic, the behavior of traffic participants outside the blue circle will still have an impact on the red target vehicle. For example, in the figure, two pedestrians are crossing the road in front of the red target vehicle. For a human driver, the red target vehicle will generally begin to judge the risk of collision with the two pedestrians at this time, and eventually slow down to avoid the closer pedestrian. However, since both pedestrians are outside the blue circle, this pedestrian with a significant impact is not taken into account by the Euclidean distance adjacency matrix.

[0115] Therefore, in order to solve the problem that the fixed distance threshold of the Euclidean distance adjacency matrix cannot be adjusted dynamically, this embodiment proposes an adjacency matrix that takes into account the potential collision risk: Used to consider the possibility of future collisions. Adjacency matrix A simplified collision risk assessment algorithm is used to find traffic participants that have an impact on each other. Figure 2 c continues to use the pedestrian crossing the road as an example to illustrate the core idea of ​​this collision assessment algorithm. Because the future trajectory is still unknown in the trajectory prediction stage, this embodiment uses a directed line segment generated according to the participant's current speed and direction as the future trajectory for collision detection. Therefore, directed line segments AB, CD, and EF represent Figure 2 a and Figure 2 The future trajectories of the vehicle and two pedestrians in b are calculated and generated using the constant speed kinematic model, which is specifically expressed as:

[0116]

[0117] Where, L pred Indicates the length of the directed line segment; u cur is the current speed of the traffic participant; t pred is the prediction time domain; θ represents the direction of the future trajectory; θ cur is the current direction of the traffic participant; it is simplified to always be consistent with the current direction. This model provides a simple and fast method to assess the risk of collision, that is, if the line segments AB and CD intersect at one point, it is considered that there is a risk of collision between the vehicle and the pedestrian. Obviously, when a collision may occur between two participants, it is necessary to model their behavioral interactions, because the possibility of a collision will force them to change their original motion patterns. Therefore, a potential collision adjacency matrix is ​​constructed based on the future trajectory model of traffic participants, which is specifically expressed as:

[0118]

[0119] In the formula, is the potential collision adjacency matrix An element in represents the potential collision relationship between traffic participants i and j at time t. In this calculation, traffic participant i is the vehicle itself; AB is the future trajectory of the vehicle calculated based on the future trajectory model of the traffic participant; CD is the future trajectory of the current traffic participant calculated based on the future trajectory model of the traffic participant; if there is an intersection (x, y), it means that the future trajectory of the vehicle overlaps with the future trajectory of the current traffic participant, and there is a potential collision relationship. At this time, Assign a value of 1.

[0120] Based on the above construction of the Euclidean distance adjacency matrix and the potential collision adjacency matrix, in order to continue to explore the hidden behavioral interaction relationship, this embodiment constructs an adaptive adjacency matrix To model the subtle spatial dependencies and interactions between traffic participants. Adaptive adjacency matrix The main idea is to embed the source node into E i and the target node embedding E j Multiply to construct a dynamic and learnable adjacency matrix. Specifically, an adaptive adjacency matrix can be constructed based on traffic participant information, which is specifically expressed as:

[0121]

[0122] In the formula, is the latent adaptive adjacency matrix An element in , which represents the adaptive relationship between traffic participants i and j at time t; E i Indicates that traffic participant i is taken as the source node, E j Indicates that traffic participant j is taken as the target node; ReLU is the activation function, which is used to eliminate weak connections and increase the nonlinearity of the network layer; SoftMax is the normalization function, which is used to normalize the adaptive adjacency matrix; the adaptive adjacency matrix It is used to capture the dependency between the source node and the target node embedding, that is, to represent the interaction between traffic participants i and j.

[0123] It should be noted that node embedding is to map a node to a multidimensional vector and express it with this multidimensional vector, or more generally, to replace the node input into the subsequent neural network layer. Both node embeddings are randomly initialized to a set of learnable parameters.

[0124] It should be noted that, based on the traffic scene expression model, the spatial correlation of traffic participants' behaviors at different time scales can be captured by stacking multiple expanded TCN-GCN spatiotemporal layers, so as to realize multi-scale spatiotemporal fusion interaction feature extraction. Thanks to the receptive fields of different scales, the GCN of each spatiotemporal layer can capture the interactive relationships between participants at different time scales. More specifically, at the bottom layer, GCN encodes short-term behavioral interaction information, while at the top layer, GCN understands the interactive relationship from a longer-term time level. However, as the number of stacked layers increases, the number of model parameters and the computational burden increase significantly. In order to solve this problem, when stacking multiple spatiotemporal layers, this embodiment adopts an expanded TCN temporal convolutional network, improves it and achieves a balance between prediction accuracy and computational efficiency. Specifically:

[0125] The multi-scale spatiotemporal graph convolution model is constructed so that the multi-scale spatiotemporal graph convolution model extracts the spatiotemporal interaction features of traffic participants based on the traffic scene expression model, including:

[0126] Construct a multi-scale spatiotemporal graph convolutional model consisting of multiple expanded TCN-GCN spatiotemporal layers; the GCN in each spatiotemporal layer is used to capture the interactive relationships between traffic participants at different time scales, and the expanded TCN in each spatiotemporal layer is used to capture the movement trends of traffic participants over time;

[0127] The multi-scale spatiotemporal graph convolution model is used to extract the spatiotemporal interaction features of traffic participants based on the traffic scene expression model;

[0128] Among them: Dilated TCN stands for Dilated Temporal Convolutional Network; GCN stands for Graph Convolutional Network.

[0129] The expanded TCN constructed in this embodiment achieves the best balance between prediction performance and computational efficiency, providing feasibility for the algorithm to be deployed on low-computing domain control chips; and constructs a method of representing the traffic scene expression model through three adjacency matrices: Euclidean distance adjacency matrix, potential collision adjacency matrix and adaptive adjacency matrix, which is conducive to fully modeling prior knowledge and thus optimizing the performance of the GCN graph convolution layer.

[0130] In one embodiment, in the multi-scale spatiotemporal graph convolution model, for any dilated TCN, the dilated convolution operation F(t) at time t is specifically expressed as:

[0131]

[0132] Wherein: Q represents a one-dimensional input sequence; f represents a filter; a represents the element index in the filter; k1 represents the size of the filter; d represents the dilation factor, which is used to control the distance between the elements of the convolution kernel; in the multi-scale spatiotemporal graph convolution model, the d value of the odd-numbered dilated TCN-GCN spatiotemporal layer is 1, and the d value of the even-numbered dilated TCN-GCN spatiotemporal layer is 2.

[0133] In this embodiment, in order to extract the correlation of behavioral interactions between traffic participants in the time dimension, the present invention uses a time convolution TCN network to process time series information. Compared with the recurrent neural network RNN ​​(RNN), the TCN network is a method based on a convolutional neural network, which has the advantages of parallel computing and gradient stability, and can extract feature information on a longer time scale. The present invention captures the movement trend of traffic participants over time by stacking dilated time convolution layers with different receptive fields as a TCN model.

[0134] Furthermore, if Figure 3 As shown, this embodiment shows three different ways to implement dilated convolution. Figure 3 The stacked dilation TCN method shown in a is to design a dilation factor sequence of 1, 2; 1, 2; 1, 2; 1, 2, and then apply it to the parameter dilation of each layer of convolution function Conv1d (called G-WaveNet TCN in this embodiment). If this method is directly applied to trajectory prediction, the time consumption is still very high. On the contrary, Figure 3 The exponential expansion time convolution shown in b greatly speeds up the calculation speed, but excessively increases the loss of local information, resulting in a decrease in prediction accuracy. Figure 3 c explores a new implementation method of stacked dilated TCN layers, that is, assigning 1 or 2 to the stride parameter of the Conv1d function according to the odd and even layers to implement the dilated TCN layer.

[0135] In one embodiment, in order to extract the mutual influence of the behaviors of spatially adjacent traffic participants, this embodiment further uses a graph convolution network to extract the dependencies of each node on the adjacency matrix based on the traffic scene expressions of three adjacency matrices: Euclidean distance, potential collision, and adaptive. Graph convolution has outstanding advantages in extracting information about surrounding traffic participants due to its ability to aggregate and transform adjacency information of nodes. Through the design of the adjacency matrix, prior knowledge can be integrated into the prediction to improve the prediction accuracy. After the Chebyshev spectral filter can be simplified by first-order approximation, standard convolution can be used to process graph structure data, and any GCN output spatiotemporal interaction feature Z in the multi-scale spatiotemporal graph convolution model is specifically expressed as:

[0136]

[0137] In the formula, k2 represents the diffusion step; A power series representing the Euclidean distance adjacency matrix; A power series representing the potential collision adjacency matrix; represents the power series of the adaptive adjacency matrix; X is the output result of the expanded TCN; is the learnable parameter matrix in the convolution operation; Indicates the last observation timestamp t obs The Euclidean distance adjacency matrix of ; Indicates the last observation timestamp t obs The potential collision adjacency matrix of ; Indicates the last observation timestamp t obs The Euclidean distance adjacency matrix, the potential collision adjacency matrix and the adaptive adjacency matrix constitute a traffic scene expression model.

[0138] It should be noted that compared with the standard convolution operation, the essence of graph convolution is still the convolution operation implemented by the two-dimensional convolution layer Conv2D. The key difference lies in the processing of the input signal, which is multiplied by Item, which absorbs the structural information of the graph well, so that the expert prior knowledge can be integrated into the prediction model through the adjacency matrix design. So far, a GCN layer and an expanded TCN layer together constitute a time-space layer. Then, by stacking multiple TCN-GCN spatiotemporal layers, the spatial correlation of traffic participants' behaviors at different time scales is captured, and multi-scale spatiotemporal fusion interaction feature extraction is realized. Based on this, this embodiment can capture long-term and short-term behavioral interaction patterns at the same time. For example, in a scene where pedestrians cross the lane ahead, the target vehicle will generally make decisions on a more macroscopic time scale and will not give way to pedestrians because of the temporary stop of pedestrians when crossing the road. GCN can capture this long-term interaction in combination with the top TCN layer, and in the following vehicle scenario, the rear vehicle may maintain a constant relative speed and distance with the front vehicle for a short period of time. The present invention can also learn this short-term behavioral interaction through the cooperation of the bottom TCN layer and the corresponding GCN layer.

[0139] In one embodiment, constructing a trajectory prediction model so as to make the trajectory prediction model perform trajectory prediction based on the spatiotemporal interaction features and obtain a predicted trajectory includes:

[0140] Constructing a trajectory prediction model of an encoder-decoder structure; wherein both the encoder and the decoder have LSTM neural networks inside;

[0141] The spatiotemporal interaction features are input into the encoder, and after learning and processing by the LSTM neural network, the context feature information is obtained;

[0142] The context feature information is input into the decoder, processed and learned by the LSTM neural network, and combined with the current position of the vehicle to obtain several sequence points;

[0143] Perform trajectory prediction based on several sequence points to obtain the predicted trajectory.

[0144] In this embodiment, an encoder-decoder structure based on an LSTM neural network can output a trajectory point sequence to obtain a predicted trajectory. Figure 4 As shown in Figure 1, the encoder-decoder structure consists of two main parts: encoder and decoder. The core network layers inside the encoder and decoder are both two-layer LSTM neural networks. First, the spatiotemporal interaction features are passed into the encoder, which can be converted into context information after being processed by the encoder and saved as h t ,have:

[0145] h t =LSTM(Z t )

[0146] In the formula, Z t Represents the spatiotemporal interaction characteristics; the encoding process inside the encoder can be described by the following formula:

[0147] i t =sigmoid(W i Z t +U i h t-1 +b i )

[0148] f t =sigmoid(W f Z t +U f h t-1 +b f )

[0149] o t =sigmoid(W o Z t +U o h t-1 +b o )

[0150] a t =tanh(W a Z t +U a h t-1 +b a )

[0151] c t =f t c t-1 +i t a t

[0152] h t =o t tanh(c t )

[0153] Among them, i t ,f t ,o t They are the input gate, forget gate, and output gate of the LSTM network, respectively. i ,W f ,W o is the linear transformation weight of the corresponding network layer, sigmoid is the nonlinear activation function, c t The current state information or memory extracted by the LSTM network.

[0154] Next, the context information is passed into the decoder. After the LSTM network in the decoder processes and learns, it is combined with the current position [x0, y0] of the target vehicle i to iteratively generate a set of point sequences to obtain the predicted trajectory according to the following formula:

[0155]

[0156] Finally, the predicted trajectory is obtained based on several sequence points.

[0157] In one embodiment, constructing a trajectory prediction model so as to make the trajectory prediction model perform trajectory prediction based on the spatiotemporal interaction features and obtain a predicted trajectory includes:

[0158] Construct a trajectory prediction model, wherein the trajectory prediction model includes a target lane prediction sub-model, a trajectory cluster prediction sub-model and a trajectory evaluation output sub-model; wherein:

[0159] The target lane prediction submodel is used to obtain the vehicle target lane based on pre-acquired vehicle historical trajectory data and lane data;

[0160] The trajectory cluster prediction submodel is used to obtain the vehicle initial trajectory cluster based on the pre-acquired vehicle historical trajectory data and the vehicle target lane;

[0161] The trajectory evaluation output sub-model is used to construct a vehicle trajectory cluster based on the vehicle initial trajectory cluster and evaluate the vehicle trajectory cluster based on the spatiotemporal interaction characteristics, calculate the probability distribution of the vehicle trajectory cluster, and obtain the predicted trajectory.

[0162] The trajectory prediction model constructed in this embodiment can generate a vehicle initial trajectory cluster that meets the road topology and vehicle kinematic constraints. First, the relative relationship between the vehicle's historical trajectory and the vehicle's target lane is used, and then the target lane is spatially sampled to generate a vehicle initial trajectory cluster that meets the road topology. Finally, the probability of each trajectory in the vehicle's initial trajectory cluster being a predicted trajectory is evaluated through the trajectory evaluation output sub-model. The most likely future trajectory can be obtained through its probability distribution. At the same time, the probability distribution also expresses the uncertainty of the behavior of the adjacent vehicle, which is a prerequisite for decision-making planning to deal with the uncertain behavior of the adjacent vehicle.

[0163] In one embodiment, the target lane prediction sub-model is used to obtain the vehicle target lane based on pre-acquired vehicle historical trajectory data and lane data, including:

[0164] Based on the pre-acquired vehicle historical trajectory data and lane data, the relative position of the vehicle and the candidate lane is obtained;

[0165] Based on the cross-attention mechanism and the relative position of the vehicle and the candidate lane, the relevance of the center line of the candidate lane to the vehicle's historical trajectory is calculated;

[0166] The target lane of the vehicle is obtained based on the correlation of the center line of the candidate lane with the historical trajectory of the vehicle.

[0167] In this embodiment, if all lanes around the vehicle are directly sampled without consideration and the corresponding vehicle trajectory clusters are generated, the number of vehicle trajectory clusters will be too large, increasing the complexity of the prediction model. Therefore, in order to reduce the sampling space of the prediction trajectory cluster, this embodiment uses the cross attention mechanism to build a Figure 5 The target lane prediction sub-model shown first selects the target lane where the vehicle is most likely to travel, and then generates a vehicle trajectory cluster based on this lane.

[0168] In one embodiment, the acquisition of the target lane of the vehicle mainly relies on the relative position relationship between the historical trajectory of the vehicle and the center line of the lanes around it, that is, the relative position of the vehicle and the candidate lane. In order to characterize the above relative position relationship, the Frenet coordinates s and d of the historical trajectory of the vehicle on the center line of each candidate lane in the adjacent lane set are first calculated. Among them, s represents the longitudinal distance along the center line, and d represents the lateral distance in the direction perpendicular to the center line. For curved lanes, the Frenet coordinate system can more clearly and directly describe the relative position relationship between the vehicle and the lane than the Cartesian coordinate system. After using the Frenet coordinates to describe the relative position relationship between the vehicle and the candidate lane, this embodiment can use the cross-attention mechanism (Cross-Attention) to process the relationship between the sequence of the historical trajectory of the vehicle and the sequence of the path points of the center line of the candidate lane. Among them, the cross-attention mechanism and the self-attention mechanism (Self-Attention) both belong to the attention network, which models and captures the dependency relationship of the input sequence through the (Q, K, V) triple tuple. According to the different input sources of the three vectors Q, K, and V, in this embodiment, the center line of the candidate lane can be used as the query vector Q of the cross-attention mechanism, and the vehicle historical trajectory can be used as the key vector K and the value vector V, thereby constructing a self-attention mechanism and a cross-attention mechanism.

[0169] In one embodiment, Q, K, and V in the self-attention mechanism are respectively multiplied by three different weight matrices W for the same feature input vector X Q ,W K ,W V The three generated vectors of the same dimension, i.e. Q = XW Q ,K=XW K ,V=XW V , the weight matrix can be obtained through neural network training. The essence of the self-attention mechanism is to calculate the mutual influence between the internal elements of the input vector X, which can be implemented by using the Transformer network. The query vector Q in the cross-attention mechanism is composed of the lane center line X A K and V are converted from the vehicle's historical trajectory X. B Transformed, it is mainly used to model the relationship between two different sequences, that is, Q = X A W Q ,K=X B W K ,V=X B W V .

[0170] Therefore, in the target lane prediction sub-model, the lane centerline X A As the query vector Q of the cross attention mechanism, the vehicle history trajectory X BAs the key vector K and the value vector V, the two are used together to perform cross-attention operation to obtain the relevance of the lane centerline to the historical trajectory. The specific steps are as follows:

[0171] 1) Feature Mapping: Q = X A W Q ,K=X B W K ,V=X B W V , X A is the lane centerline, X B is the vehicle history trajectory;

[0172] 2) Similarity calculation: Through the dot product operation of Q and K, the similarity score between the query vector Q and the key vector K is calculated. These similarity scores represent the similarity between X A Each element in B The relationship between all elements in can be expressed as:

[0173]

[0174] Among them, S is the similarity score matrix, is a scaling factor used to prevent the dot product value from being too large;

[0175] 3) Attention weight calculation: The similarity score is normalized by the Softmax function to obtain the attention weight, which reflects the X A Each element in B The importance of all elements in can be expressed as:

[0176] W a A→B =Softmax(S A→B )

[0177] Among them, W a is the attention weight;

[0178] 4) Weighted summation: The attention weights are weighted summed with the value vector V to obtain the output representation after cross-attention adjustment, which can be expressed as:

[0179] O A→B =W a A→B V B

[0180] In the formula, O A→B The dimension of O is consistent with the dimension of the query vector Q. Through the above attention mechanism, O A→B Each element in the query vector Q contains the association information of the key vector K and the value vector V, which models the lane centerline sequence X. Aand the vehicle history trajectory sequence X B The mutual relationship between them.

[0181] In one embodiment, after uniformly encoding the lane centerline information and the vehicle historical trajectory to obtain their mutual connection, all candidate lanes can be constructed as an [n, c]-dimensional vector: each candidate lane is represented by a c-dimensional vector, and sequentially increased to n lanes to form an [n, c]-dimensional lane array. Among them, n is the number of candidate lanes, and c is the characteristic dimension of the lane. Next, the Transformer self-attention network is used to learn the connection and difference between the candidate lanes, and the multi-layer perceptron MLP and Softmax are used to output the probability distribution of the candidate lanes. Finally, considering the possible prediction error, this embodiment strikes a balance between sampling efficiency and prediction accuracy, and selects the top three lanes with the largest probability as the vehicle target lane prediction output.

[0182] In one embodiment, the trajectory cluster prediction sub-model is used to obtain the vehicle initial trajectory cluster based on the pre-acquired vehicle historical trajectory data and the vehicle target lane, including:

[0183] Construct the Frenet coordinate system of the lane centerline of the vehicle target lane;

[0184] In the Frenet coordinate system of the lane centerline, the vehicle target lane is spatially sampled based on the local path planning method to obtain the first path polynomial curve;

[0185] Convert the first path polynomial curve into a Cartesian coordinate system to obtain a second path polynomial curve;

[0186] The curvature of the second path polynomial curve is detected, and the path polynomial curve that meets the preset requirements is screened to obtain the initial vehicle trajectory cluster.

[0187] It should be noted that in order to generate a predicted trajectory cluster that conforms to the road topology and vehicle kinematics, the idea of ​​local path planning is applied to perform spatial sampling of the target lane, and a polynomial curve is used to express the initial vehicle trajectory cluster, which may include the following steps:

[0188] 1) Sampling is performed in the Frenet coordinate system of the lane centerline, where s represents the longitudinal distance along the centerline and d represents the lateral distance perpendicular to the centerline. The target lane centerline is fitted as a function of the parameter s through a cubic spline curve:

[0189]

[0190] In the formula, is the spline curve coefficient, x, y is the value of the center line point in the Cartesian coordinate system. The parameterized center line can facilitate the subsequent restoration of the trajectory cluster to the Cartesian coordinate system through s interpolation.

[0191] 2) The longitudinal distance s and lateral distance d of the vehicle initial trajectory cluster in Frenet coordinates are constructed as the fourth-order and fifth-order polynomials of the prediction time t respectively:

[0192]

[0193] In the formula, for the quartic polynomial, the initial state can be combined and terminal state Five equations are established for solution. The above states correspond to the initial longitudinal distance, initial longitudinal velocity, initial longitudinal acceleration, terminal longitudinal velocity, and terminal longitudinal acceleration. Similar to the lateral distance d, the fifth-order polynomial coefficients are obtained. The initial states are then and terminal state Five equations are created and solved.

[0194] 3) Sampling different terminal states S t and terminal state D t By combining the first path polynomial curves in a series of Frenet coordinates, the second path polynomial curves are obtained by interpolating the center line parameterized in step 1) and restoring to the Cartesian coordinate system. In this embodiment, the terminal speed sampling range is set as:

[0195]

[0196] In the formula, is the current speed, which is obtained by the historical trajectory point difference; t pred is the prediction time domain; a max,brake and a max,acc The maximum deceleration and maximum acceleration are respectively. Then the terminal speed is sampled at a resolution of 1km / h, the terminal lateral distance is sampled at a resolution of 0.5m, and the driving area of ​​the vehicle along the target lane is sampled through longitudinal and lateral combinations, and the second path polynomial curve is generated.

[0197] 4) Perform curvature detection on the second path polynomial curve to eliminate unreasonable candidate trajectories with too large curvature. The trajectory cluster curvature calculation can use the triangle circumscribed circle curvature method, and finally select the path polynomial curve that meets the preset requirements according to the calculation results to obtain the vehicle initial trajectory cluster.

[0198] For details, please refer to Figure 6 As shown, the light green trajectory cluster is a feasible trajectory that meets the curvature constraint, and the blue trajectory cluster is an infeasible trajectory. The curvature calculation of the triangle circumscribed circle curvature method is performed as follows:

[0199] See also Figure 7 First, assume that A, B, and C are three continuous discrete points on the predicted trajectory, and a, b, and c are their opposite sides. In ΔABC, the cosine theorem yields:

[0200]

[0201] Then, the curvature k can be obtained from the law of sines:

[0202]

[0203] In one embodiment, the trajectory evaluation output sub-model is used to construct a vehicle trajectory cluster based on the vehicle initial trajectory cluster and evaluate the vehicle trajectory cluster based on the spatiotemporal interaction characteristics, calculate the probability distribution of the vehicle trajectory cluster, and obtain the predicted trajectory, including:

[0204] The vehicle initial trajectory cluster and the vehicle historical trajectory are uniformly represented for heterogeneous information to construct the vehicle trajectory cluster;

[0205] The vehicle trajectory clusters and the temporal and spatial interaction features are integrated and encoded through the cross-attention mechanism, and the preset Transformer self-attention model is used to capture the temporal motion characteristics of the vehicle trajectory clusters. Finally, the probability distribution of the vehicle trajectory clusters is output through the MLP multi-layer perceptron and the Softmax layer.

[0206] The predicted trajectory is obtained based on the probability distribution of vehicle trajectory clusters.

[0207] It should be noted that on structured roads, human drivers will first determine which lane the surrounding vehicles are driving in, and then continue to determine their relative position relationship with the vehicle on this basis. Therefore, lane guidance information is crucial for trajectory prediction. However, how to encode multi-source heterogeneous information such as lane lines and historical trajectories of adjacent vehicles and learn their mutual relationship with neural networks is a major difficulty. Because lane lines are a series of waypoints, unlike the historical trajectory information of adjacent vehicles, the waypoint sequence does not imply speed and acceleration characteristics. In addition, in general, there is not only one lane around the target vehicle, which greatly hinders the network model from learning the relationship between the vehicle's historical trajectory and the correct lane during training.

[0208] In order to solve the above technical problems, this embodiment splices the vehicle historical trajectory with each trajectory in the predicted vehicle initial trajectory cluster to construct a complete vehicle trajectory cluster from the starting point of the historical trajectory to the end point of the predicted trajectory. The predicted vehicle trajectory cluster itself is a motion trajectory containing speed and acceleration information. At the same time, when generating the predicted vehicle trajectory cluster, the road topology is also considered, reflecting the changing trend of lane direction and position. Therefore, the historical trajectory and the vehicle initial trajectory are spliced ​​according to Figure 8By splicing and reorganizing, the unified representation of static map information and dynamic vehicle motion trajectory heterogeneous information is achieved.

[0209] The Transformer's self-attention mechanism can be further used to learn continuous motion features from historical vehicle trajectories to predicted vehicle trajectory clusters. The complete vehicle trajectory cluster constructed above is the x, y coordinates in the Cartesian coordinate system. In a scenario guided by the lane centerline, converting the x, y coordinates to Frenet coordinates s and d based on the lane centerline can obtain a more direct and essential relative position relationship, thereby accelerating the training of the neural network. However, due to special scenarios such as curves, the distance s of the xy trajectory projected along the lane centerline is not linearly related to the vehicle speed. Therefore, the trajectory feature vector designed in this embodiment uses the more essential vehicle speed v and the distance d in the direction perpendicular to the lane centerline to characterize the vehicle's forward and backward movement trend. According to Figure 8 Splicing is performed to directly splice each candidate prediction trajectory behind the historical trajectory. The dimension of the historical trajectory is [3,5], where 3 represents the feature vector [x,y,v]; the dimension of the prediction trajectory is [3,15]; the dimension of the complete trajectory after splicing is [3,20]. After the splicing is completed, the feature vector [x,y,v] will be further converted into a feature vector [v,d], and the dimension of the spliced ​​trajectory is [2,20].

[0210] Furthermore, after the road topology information and vehicle motion trajectory information are uniformly represented, the predicted vehicle trajectory cluster that is complete in time series includes the vehicle's historical trajectory and predicted candidate trajectory. Therefore, by inputting each complete trajectory in time series, the motion characteristics of the trajectory points before and after the trajectory sequence can be encoded through the self-attention mechanism, and the best candidate trajectory that best matches the previous and next motion trends can be found. Therefore, the trajectory evaluation output sub-model can be constructed to evaluate the vehicle's initial trajectory cluster to obtain the probability distribution and predicted trajectory of the final vehicle trajectory cluster.

[0211] Further, according to the above idea, the trajectory evaluation output sub-model can be constructed as follows Fig. 9The evaluation model shown in the figure specifically includes three parts: behavior interaction feature fusion encoding, prediction trajectory cluster self-attention feature extraction and probability distribution decoding. Among them: first, after extracting the multi-scale spatiotemporal interaction features of traffic participants, the predicted vehicle trajectory cluster and the spatiotemporal interaction features are integrated and encoded through Cross-Attention, so that the vehicle trajectory cluster can pay attention to the behavior interaction information of traffic participants. Here, the predicted trajectory cluster is used as the query Q vector of the Cross-Attention module, and the spatiotemporal interaction features are used as the key vector K and value vector V of Cross-Attention. The output of the Cross-Attention module maintains the same dimension as the predicted trajectory cluster. Then, the standard Transformer self-attention network is used to capture the motion features of the trajectory cluster in time series, including the feature [v, d], that is, the vehicle speed v and the distance d in the direction perpendicular to the center line of the lane. In addition, the rate of change of acceleration and distance d can be indirectly learned. Finally, the probability distribution evaluation of the predicted trajectory cluster is modeled as a classification problem. The probability distribution of the predicted trajectory cluster is output through the MLP multi-layer perceptron and the Softmax layer, and finally a multimodal predicted trajectory cluster and its probability distribution that fully characterizes the uncertain behavior of the predicted vehicle are obtained. Among them: Softmax normalized exponential function is a commonly used classifier in deep neural networks, and its output is the probability corresponding to different trajectories.

[0212] This embodiment can predict the future trajectory of the vehicle itself and surrounding vehicles, taking into account the road topology and vehicle kinematic constraints, and finally obtains a multimodal predicted vehicle trajectory cluster and its probability distribution that fully characterizes the uncertain behavior of the vehicle, which can provide more accurate surrounding obstacle motion information for subsequent decision-making and planning links, and obtain the final predicted trajectory. The generated vehicle trajectory cluster revolves around the vehicle's maximum probability predicted trajectory, covering almost all possible vehicle movement intentions, and generates corresponding probabilities for each trajectory in the predicted trajectory cluster, which has the advantage of accurate long-term multimodal trajectory prediction.

[0213] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A trajectory prediction method based on spatiotemporal interactive feature fusion, characterized in that: The following steps are involved: Collect information of traffic participants on the road in real time; Construct a traffic scene expression model based on traffic participant information; Construct a multi-scale spatiotemporal graph convolution model to extract the spatiotemporal interaction features of traffic participants based on the traffic scene expression model; Construct a trajectory prediction model to enable the trajectory prediction model to perform trajectory prediction based on spatiotemporal interaction features and obtain a predicted trajectory.

2. The trajectory prediction method based on spatiotemporal interactive feature fusion according to claim 1 is characterized in that: The traffic scene expression model includes a Euclidean distance adjacency matrix, a potential collision adjacency matrix and an adaptive adjacency matrix; the traffic scene expression model is constructed based on traffic participant information, including: Based on the traffic participant information, the Euclidean distance adjacency matrix is ​​constructed, which is specifically expressed as: In the formula, is the Euclidean distance adjacency matrix An element in , which represents the adjacency relationship between traffic participants i and j at time t; represents the distance between traffic participants i and j at time t; D close Represents the distance threshold between neighbor participants; 1 means that at time t, traffic participant i and traffic participant j are neighbors of each other; 0 means that at time t, traffic participant i and traffic participant j are not neighbors; Based on the traffic participant information, the current speed and current direction of the traffic participant are obtained, a number of directed line segments are obtained, and a future trajectory model of the traffic participant is constructed based on the directed line segments, which is specifically expressed as follows: Where, L pred Indicates the length of the directed line segment; u cur is the current speed of the traffic participant; t pred is the prediction time domain; θ represents the direction of the future trajectory; θ cur The current direction of the traffic participant; The potential collision adjacency matrix is ​​constructed based on the future trajectory model of traffic participants, which is specifically expressed as: In the formula, is the potential collision adjacency matrix An element in represents the potential collision relationship between traffic participants i and j at time t. In this calculation, traffic participant i is the vehicle itself; AB is the future trajectory of the vehicle calculated based on the future trajectory model of the traffic participant; CD is the future trajectory of the current traffic participant calculated based on the future trajectory model of the traffic participant; if there is an intersection (x, y), it means that the future trajectory of the vehicle overlaps with the future trajectory of the current traffic participant, and there is a potential collision relationship. At this time, Assign a value of 1; Based on the information of traffic participants, an adaptive adjacency matrix is ​​constructed, which is specifically expressed as: In the formula, is the latent adaptive adjacency matrix An element in , which represents the adaptive relationship between traffic participants i and j at time t; E i Indicates that traffic participant i is taken as the source node, E j Indicates that traffic participant j is taken as the target node; ReLU is the activation function; SoftMax is the normalization function; adaptive adjacency matrix It is used to capture the dependency between the source node and the target node embedding, that is, to represent the interaction between traffic participants i and j.

3. The trajectory prediction method based on spatiotemporal interactive feature fusion according to claim 1 is characterized in that: The multi-scale spatiotemporal graph convolution model is constructed so that the multi-scale spatiotemporal graph convolution model extracts the spatiotemporal interaction features of traffic participants based on the traffic scene expression model, including: Construct a multi-scale spatiotemporal graph convolutional model consisting of multiple expanded TCN-GCN spatiotemporal layers; the GCN in each spatiotemporal layer is used to capture the interactive relationships between traffic participants at different time scales, and the expanded TCN in each spatiotemporal layer is used to capture the movement trends of traffic participants over time; The multi-scale spatiotemporal graph convolution model is used to extract the spatiotemporal interaction features of traffic participants based on the traffic scene expression model; Among them: Dilated TCN stands for Dilated Temporal Convolutional Network; GCN stands for Graph Convolutional Network.

4. The trajectory prediction method based on spatiotemporal interactive feature fusion according to claim 3 is characterized in that: In the multi-scale spatiotemporal graph convolution model, for any dilated TCN, the dilated convolution operation F(t) at time t is specifically expressed as: Wherein: Q represents a one-dimensional input sequence; f represents a filter; a represents the element index in the filter; k1 represents the size of the filter; d represents the dilation factor, which is used to control the distance between the elements of the convolution kernel; in the multi-scale spatiotemporal graph convolution model, the d value of the odd-numbered dilated TCN-GCN spatiotemporal layer is 1, and the d value of the even-numbered dilated TCN-GCN spatiotemporal layer is 2.

5. The trajectory prediction method based on spatiotemporal interactive feature fusion according to claim 3 is characterized in that: In the multi-scale spatiotemporal graph convolution model, for any GCN, its output spatiotemporal interaction feature Z is specifically expressed as: In the formula, k2 represents the diffusion step; A power series representing the Euclidean distance adjacency matrix; A power series representing the potential collision adjacency matrix; represents the power series of the adaptive adjacency matrix; X is the output result of the expanded TCN; is the learnable parameter matrix in the convolution operation; Indicates the last observation timestamp t obs The Euclidean distance adjacency matrix of ; Indicates the last observation timestamp t obs The potential collision adjacency matrix of ; Indicates the last observation timestamp t obs The Euclidean distance adjacency matrix, the potential collision adjacency matrix and the adaptive adjacency matrix constitute a traffic scene expression model.

6. A trajectory prediction method based on spatiotemporal interactive feature fusion according to any one of claims 1 to 5, characterized in that: The constructing of the trajectory prediction model so as to make the trajectory prediction model perform trajectory prediction based on the spatiotemporal interaction features and obtain the predicted trajectory includes: Constructing a trajectory prediction model of an encoder-decoder structure; wherein both the encoder and the decoder have LSTM neural networks inside; The spatiotemporal interaction features are input into the encoder, and after learning and processing by the LSTM neural network, the context feature information is obtained; The context feature information is input into the decoder, processed and learned by the LSTM neural network, and combined with the current position of the vehicle to obtain several sequence points; Perform trajectory prediction based on several sequence points to obtain the predicted trajectory.

7. A trajectory prediction method based on spatiotemporal interactive feature fusion according to any one of claims 1 to 5, characterized in that: The constructing of the trajectory prediction model so as to make the trajectory prediction model perform trajectory prediction based on the spatiotemporal interaction features and obtain the predicted trajectory includes: Construct a trajectory prediction model, wherein the trajectory prediction model includes a target lane prediction sub-model, a trajectory cluster prediction sub-model and a trajectory evaluation output sub-model; wherein: The target lane prediction submodel is used to obtain the vehicle target lane based on pre-acquired vehicle historical trajectory data and lane data; The trajectory cluster prediction submodel is used to obtain the vehicle initial trajectory cluster based on the pre-acquired vehicle historical trajectory data and the vehicle target lane; The trajectory evaluation output sub-model is used to construct a vehicle trajectory cluster based on the vehicle initial trajectory cluster and evaluate the vehicle trajectory cluster based on the spatiotemporal interaction characteristics, calculate the probability distribution of the vehicle trajectory cluster, and obtain the predicted trajectory.

8. The trajectory prediction method based on spatiotemporal interactive feature fusion according to claim 7, characterized in that: The target lane prediction submodel is used to obtain the vehicle target lane based on the pre-acquired vehicle historical trajectory data and lane data, including: Based on the pre-acquired vehicle historical trajectory data and lane data, the relative position of the vehicle and the candidate lane is obtained; Based on the cross-attention mechanism and the relative position of the vehicle and the candidate lane, the relevance of the center line of the candidate lane to the vehicle's historical trajectory is calculated; The target lane of the vehicle is obtained based on the correlation of the center line of the candidate lane with the historical trajectory of the vehicle.

9. The trajectory prediction method based on spatiotemporal interactive feature fusion according to claim 7, characterized in that: The trajectory cluster prediction sub-model is used to obtain the vehicle initial trajectory cluster based on the pre-acquired vehicle historical trajectory data and the vehicle target lane, including: Construct the Frenet coordinate system of the lane centerline of the vehicle target lane; In the Frenet coordinate system of the lane centerline, the vehicle target lane is spatially sampled based on the local path planning method to obtain the first path polynomial curve; Convert the first path polynomial curve into a Cartesian coordinate system to obtain a second path polynomial curve; The curvature of the second path polynomial curve is detected, and the path polynomial curve that meets the preset requirements is screened to obtain the initial vehicle trajectory cluster.

10. The trajectory prediction method based on spatiotemporal interactive feature fusion according to claim 7, characterized in that: The trajectory evaluation output sub-model is used to construct a vehicle trajectory cluster based on the vehicle initial trajectory cluster and evaluate the vehicle trajectory cluster based on the spatiotemporal interaction characteristics, calculate the probability distribution of the vehicle trajectory cluster, and obtain the predicted trajectory, including: The vehicle initial trajectory cluster and the vehicle historical trajectory are uniformly represented for heterogeneous information to construct the vehicle trajectory cluster; The vehicle trajectory clusters and the temporal and spatial interaction features are integrated and encoded through the cross-attention mechanism, and the preset Transformer self-attention model is used to capture the temporal motion characteristics of the vehicle trajectory clusters. Finally, the probability distribution of the vehicle trajectory clusters is output through the MLP multi-layer perceptron and the Softmax layer. The predicted trajectory is obtained based on the probability distribution of vehicle trajectory clusters.

Citation Information

Patent Citations

  • Multi-modal space-time model for accurate motion prediction based on visual fusion

    CN117315603A

  • Trajectory prediction method considering dynamic interaction between vehicles

    CN118553108A

Cited By

  • Model training method and trajectory prediction method based on multi-dimensional feature fusion

    CN120561884A

  • Robot trajectory prediction method based on time-frequency wavelet transform and graph network

    CN121048642A

  • Trajectory prediction method based on spatio-temporal interaction feature fusion

    WO2026153158A1