A vehicle trajectory prediction method for roundabout scenarios

CN122830718APending Publication Date: 2026-09-29YANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610762697.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

第一类为循环神经网络,该方法能够实现单车轨迹的预测,但是车辆之间交互较差,难以应用于多车辆的复杂交通场景

Benefits of technology

[0052]本发明提供一种面向环岛场景的车辆轨迹预测方法,采用基于Transformer的并行解码机制,在单次前向传播中生成未来多个时间步的解码特征,提高推理效率,减少自回归结构中的误差累积问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122830718A_ABST
    Figure CN122830718A_ABST
Patent Text Reader

Abstract

The application discloses a kind of roundabouts scene-oriented vehicle trajectory prediction method, belong to automatic driving and trajectory prediction technical field, method includes obtaining roundabout vehicle historical trajectory data, vehicle trajectory is filtered and sliding time window is divided, based on the position of the last frame effective vehicle of historical trajectory is decentralized, and adjacent matrix is constructed, extract the speed, acceleration, heading angle and effectiveness mask of vehicle under local rectangular coordinate system as input state;Using space-time graph convolutional encoder extracts historical space-time interaction features;Using learnable time query vector and Transformer decoder generates future time series features, recovers future multi-modal trajectory by acceleration update, velocity and displacement kinematics recursion.The application can predict the future trajectory of the vehicle by observing the historical trajectory of the vehicle, accurately and reasonably predict the trajectory of the roundabout vehicle, and provide protection for the safe driving of the vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle trajectory prediction technology, specifically to a vehicle trajectory prediction method for roundabout scenarios. Background Technology

[0002] With the development of autonomous driving technology, autonomous vehicles are gradually becoming a part of daily traffic. For autonomous vehicles, the ability to accurately predict the future trajectories of other vehicles in the vicinity can effectively ensure their own driving safety and avoid traffic accidents.

[0003] Existing vehicle trajectory prediction methods can be mainly divided into three categories. The first category is recurrent neural networks (RNNs), which can predict the trajectory of a single vehicle, but suffer from poor interaction between vehicles, making them difficult to apply to complex traffic scenarios with multiple vehicles. The second category is graph neural networks (Graph Neural Networks), which characterize the spatial relationships between vehicles, but in the sequential processing time dimension, spatiotemporal features interfere with each other, and their long-term memory is weak, easily forgetting information from earlier periods. The third category is Transformers, which achieve unified processing of spatial interaction and temporal dependencies, but often employs autoregressive decoding, limiting the efficiency of real-time prediction, and most models lack physical constraints, especially in special traffic scenarios such as roundabouts.

[0004] The existing three types of vehicle trajectory prediction technologies have the following four shortcomings, which are particularly prominent in roundabout scenarios and seriously restrict the safety and reliability of autonomous driving systems.

[0005] First, the methods suffer from poor scenario adaptability. Existing mainstream trajectory prediction methods are mostly suitable for scenarios such as straight roads like highways and urban roads, or simple intersections, but are unsuitable for roundabout scenarios. The unique circular geometry of roundabouts causes continuous changes in curvature as vehicles travel along the roundabout, and current methods lack modeling and analysis for these continuous curvature changes. Directly using existing straight-road-based models to predict trajectories within roundabouts easily leads to the prediction of the vehicle's future position as a straight line, resulting in a significant deviation between the predicted trajectory and the actual roundabout, and even issues such as crossing road boundaries. This lack of scenario adaptability significantly reduces the prediction accuracy of existing methods in roundabout scenarios, failing to meet the high-precision trajectory prediction requirements of autonomous driving systems.

[0006] Second, vehicle intent modeling is overly simplistic. Existing methods model driving intent in a simplistic way, such as simple three-class classification (left turn, straight, right turn) or five-class classification (left lane change, right lane change, left turn, right turn, straight). However, roundabout scenarios have multiple entrances and exits (typically 3-5), and vehicle intent includes not only whether to exit the roundabout but also which exit to choose. Therefore, existing simple classification modeling methods cannot cover the continuous, high-dimensional, multi-entry and multi-exit decision space in roundabout scenarios. Furthermore, existing methods are mostly single-modal outputs, i.e., predicting only the most probable trajectory, which easily leads to the omission of key behavioral modalities, increases the probability of decision-making errors, and ignores the intent uncertainty in roundabout scenarios.

[0007] Third, there is a lack of physical constraints. Most existing purely data-driven deep learning methods directly output future coordinates, lacking explicit physical constraints. Specifically, this manifests as a lack of mathematical consistency between acceleration, velocity, and position, potentially leading to abrupt velocity changes, discontinuous acceleration, and other problems that render the predicted trajectory impossible to execute in reality. This lack of physical plausibility not only affects prediction accuracy but may also mislead autonomous driving systems into making dangerous decisions, creating safety hazards.

[0008] Fourth, inference efficiency is insufficient and error accumulation occurs. Most methods employ autoregressive decoding, which generates future trajectory points incrementally, with each prediction depending on the output of the previous moment. This mechanism causes inference time to increase with the prediction duration; for example, predicting a 5-second trajectory (25 frames) requires 25 forward propagations. Furthermore, autoregressive decoding suffers from error accumulation; earlier prediction errors propagate to later moments, leading to a significant decrease in long-term prediction accuracy. Summary of the Invention

[0009] The purpose of this invention is to overcome the shortcomings of the prior art and provide a vehicle trajectory prediction method for roundabout scenarios. By modeling the spatiotemporal interaction relationship of vehicles in roundabout scenarios, performing multimodal prediction of future target points, and combining the Transformer parallel decoding mechanism with physical kinematic constraints, the method can accurately predict the future trajectory of vehicles in roundabout scenarios.

[0010] To achieve the above objectives, the present invention is implemented using the following technical solution:

[0011] This invention provides a vehicle trajectory prediction method for roundabout scenarios, comprising:

[0012] Data preprocessing involves acquiring historical vehicle trajectory data for a roundabout scenario, filtering the historical trajectory data by vehicle type, and constructing samples. Based on these samples, an adjacency matrix is ​​constructed through decentralized processing. Based on these samples, the vehicle's velocity, acceleration, heading angle, and validity mask in a local Cartesian coordinate system are extracted to construct a multidimensional input feature vector.

[0013] The multidimensional input feature vector and adjacency matrix are input into the spatiotemporal graph convolutional encoder to extract the historical spatiotemporal interaction features of the target vehicle and its neighboring vehicles;

[0014] A learnable time query vector corresponding to the future prediction length is constructed, and combined with the historical spatiotemporal interaction features, future temporal features are generated through a Transformer decoder; the acceleration increment and position residual at each moment and in each mode are predicted based on the future temporal features, and the future multimodal trajectory is recursively recovered based on the acceleration increment and position residual;

[0015] Output the predicted future multimodal trajectory of the target vehicle in a roundabout scenario;

[0016] The loss function is obtained by selecting the mode with the smallest average displacement error, and the parameters of the spatiotemporal graph convolutional encoder and Transformer decoder are optimized based on the loss function.

[0017] Vehicle trajectory prediction is performed using a spatiotemporal graph convolutional encoder and a Transformer decoder with optimized parameters.

[0018] Furthermore, the sample construction method involves dividing the continuous trajectory of historical trajectory data using a sliding time window approach, with each time window serving as a sample.

[0019] Furthermore, the decentralized processing specifically includes: taking the average position of all valid vehicles in the last frame of the historical trajectory as the current scene center, and normalizing the position coordinates of each vehicle at each moment within the time window.

[0020] Furthermore, the adjacency matrix is ​​constructed as follows: the interaction relationship between vehicles is determined based on the Euclidean distance between each vehicle in the last frame of the historical trajectory. When the distance between two vehicles is less than the preset interaction radius, it is determined that there is an interaction relationship between them. The adjacency matrix is ​​then normalized to obtain the adjacency matrix used for graph convolution calculation.

[0021] Furthermore, the spatiotemporal graph convolutional encoder includes one input normalization layer and two spatiotemporal graph convolutional blocks, wherein the spatiotemporal graph convolutional blocks include spatial graph convolution and temporal convolution.

[0022] Furthermore, learnable edge importance coefficients are introduced into the spatial graph convolution, which are multiplied element-wise with the adjacency matrix to scale the connection strength of different spatial partitions.

[0023] Further, the construction of a learnable time query vector corresponding to the future prediction length, combined with the historical spatiotemporal interaction features, generates future temporal features through a Transformer decoder; based on the future temporal features, the acceleration increment and position residual are predicted for each moment and each mode, and the future multimodal trajectory is recursively recovered based on the acceleration increment and position residual, including:

[0024] Define the vehicle's velocity and acceleration in the last frame of the historical trajectory as follows: , , No. The mode in the th ... The acceleration increment output at each future moment is The positional residual is The predicted time step is Future multimodal displacement increments are derived through acceleration updates and two-way kinematics. :

[0025] ;

[0026] ;

[0027] ;

[0028] in, , , They represent the first The mode in the th ... The acceleration, velocity, and displacement increments at a future moment. This represents the hyperbolic tangent activation function, used to constrain the position residual within a finite range; by adjusting the displacement increment... The system performs cumulative calculations and combines the position of the last frame of the historical trajectory to reconstruct the future absolute trajectory.

[0029] Furthermore, the average displacement error is calculated as follows:

[0030] Define vehicle The The modality in the future The predicted displacement increment at each time point is The actual displacement increment is In the scene The average displacement error corresponding to each mode for:

[0031] ;

[0032] in, for Time vehicle The validity code is used to determine the vehicle. In the Whether it is valid at this moment, when When the vehicle is valid, This indicates that the vehicle is invalid. For the final number of effective vehicles, Indicates the historical observation time step.

[0033] Furthermore, the step of selecting the mode with the smallest average displacement error for loss function calculation includes:

[0034] The mode with the smallest average displacement error is selected as the optimal mode, and the target point loss, trajectory regression loss, direction consistency loss and endpoint consistency loss are calculated.

[0035] A loss function is constructed based on the target point loss, trajectory regression loss, direction consistency loss, and endpoint consistency loss.

[0036] The loss function The formula is expressed as follows:

[0037] ;

[0038] In the formula, , , and Indicates weight, Indicates the loss at the target point. Indicates the loss of endpoint consistency. Indicates the trajectory regression loss. This represents the loss of directional consistency.

[0039] Furthermore, the target point loss The calculation formula is as follows:

[0040] ;

[0041] in, Represents the smoothing L1 loss function. Indicates a validity mask. Indicates according to vehicle The position of the endpoint of the true future trajectory relative to the last frame of the historical trajectory;

[0042] Endpoint consistency loss The calculation formula is as follows:

[0043] ;

[0044] in, Indicates vehicle At the last moment Validity coding, Indicates vehicle The future cumulative endpoint displacement is obtained through physical recursion in the optimal mode. The predicted target point corresponding to the optimal mode is: ;

[0045] Trajectory Regression Loss The calculation formula is as follows:

[0046] ;

[0047] in, This represents the vehicle corresponding to the optimal mode. The future displacement increment sequence, Indicates vehicle The true future displacement increment sequence;

[0048] Directional consistency loss The calculation formula is as follows:

[0049] ;

[0050] in, This represents the cosine similarity.

[0051] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0052] This invention provides a vehicle trajectory prediction method for roundabout scenarios. It adopts a Transformer-based parallel decoding mechanism to generate decoding features for multiple future time steps in a single forward propagation, thereby improving inference efficiency and reducing the error accumulation problem in autoregressive structures.

[0053] This invention also employs physical kinematics recursion to recover future trajectories, ensuring that the predicted trajectories are executable while maintaining smoothness and meeting kinematic constraints that conform to actual driving characteristics.

[0054] This invention also improves the accuracy of trajectory endpoints, the rationality of trajectory directions, and the overall executability through the synergistic optimization of target point loss, trajectory regression loss, direction consistency loss, and endpoint consistency loss. Attached Figure Description

[0055] Figure 1 This is a flowchart illustrating a vehicle trajectory prediction method for a roundabout scenario according to an embodiment of the present invention.

[0056] Figure 2 for Figure 1 Data processing flowchart in the middle;

[0057] Figure 3 for Figure 1 A flowchart of spatiotemporal graph convolutional coding in [the context of the original text].

[0058] Figure 4 for Figure 1 Flowchart of multimodal target point prediction in [the context of the original text].

[0059] Figure 5 for Figure 1 A flowchart of the parallel decoding of physical constraints in the process;

[0060] Figure 6 To display the predicted trajectory Figure 1 ;

[0061] Figure 7 To display the predicted trajectory Figure 2 ;

[0062] Figure 8 To display the predicted trajectory Figure 3 ;

[0063] Figure 9 To display the predicted trajectory Figure 4 . Detailed Implementation

[0064] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.

[0065] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are used only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0066] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0067] Example 1

[0068] Please see Figure 1 This embodiment introduces a vehicle trajectory prediction method for roundabout scenarios, including the following steps:

[0069] S100. Data preprocessing: Obtain historical trajectory data of vehicles in the roundabout scene; filter the historical trajectory data by vehicle type and construct a sample set; perform decentralization processing on the samples to construct an adjacency matrix; extract the vehicle's speed, acceleration, heading angle and validity mask in the local Cartesian coordinate system based on the samples, and construct a multi-dimensional input feature vector.

[0070] In this embodiment, the vehicle trajectory data used comes from the publicly available RounD dataset, the first large-scale vehicle trajectory dataset specifically designed for unsignal-controlled roundabout scenarios. Released by RWTH Aachen University in 2021, this dataset contains the complete motion trajectories of over 110,000 vehicles, with most of the data originating from a relatively complex roundabout.

[0071] Specifically, please refer to Figure 2 The data preprocessing process is as follows:

[0072] The raw data is stored in the form of trajectories, and each trajectory record contains information such as the vehicle's spatial position, speed, acceleration, and vehicle size. When using this raw data, the trajectories of non-motorized vehicles must first be filtered out, and only the data of cars, vans, trucks, and trailers should be retained.

[0073] To construct the trajectory prediction task, a sliding time window method is used to segment continuous trajectories, with each time window serving as a sample. Only vehicles that exist continuously within the entire time window are retained as valid targets, thus ensuring that valid vehicles have both complete historical and future trajectories. A sample set is constructed based on the segmented samples.

[0074] To improve the model's generalization ability to roundabouts in different geographical locations and to facilitate the construction of a model of a vehicle driving on a curve with dynamic constraints, in this embodiment, the 13-dimensional information of the original data in the sample set is refined and transformed into a 6-dimensional spatiotemporal dynamic feature vector.

[0075] Finally obtained the vehicle exist 6-dimensional input feature vector at time step The formula is expressed as follows:

[0076] ;

[0077] in, and These represent the vehicle's speed in the x and y directions, respectively. and These represent the vehicle's acceleration in the x and y directions, respectively. Indicates the vehicle's heading angle. The validity mask represents whether the vehicle is valid at a given moment. Compared to traditional trajectory prediction methods that only use position coordinates, this invention introduces additional dynamic features such as x-axis and y-axis acceleration, enabling the vehicle graph structure to not only reflect spatial proximity relationships but also express differences in vehicle motion states, thereby improving modeling capabilities in complex interactive scenarios.

[0078] Although the Round dataset provides vehicle positions in a local Cartesian coordinate system, the overall distribution of vehicles across different time windows still exhibits translational bias. If the original position coordinates are used directly for training, the model may become overly reliant on the scene's positional distribution, thus weakening its generalization ability. Therefore, in this embodiment, the vehicle positions within each time window are scene-centered, converting the trajectory coordinates into motion relative to the local center of the scene. This eliminates interference from absolute positions, allowing the model to focus on the relative motion of the vehicles, enhancing its generalization ability while avoiding the risk of overfitting.

[0079] Specifically, based on the sample set, the coordinates of the last frame of all vehicle historical trajectories are extracted from the original data to calculate the scene center. Since the last frame of the historical trajectory is closest to the starting point of the future predicted trajectory, it can more accurately reflect the real-time interaction relationships and vehicle spatial distribution within the roundabout. Define a scene with N valid vehicles, where the... The last frame of the vehicle's historical trajectory is The time, corresponding to the vehicle at that time. The position coordinates are The formula for calculating the local center of the current traffic scenario is as follows:

[0080] ;

[0081] The local centers of all vehicles in the current scene are calculated and denoted as... ;in This indicates the total number of vehicles in the current scenario.

[0082] Next, the position of each vehicle in all frames within the time window is translated to obtain the normalized local coordinates:

[0083] ;

[0084] in, Indicates vehicle At any moment Location coordinates, This represents the normalized position coordinates.

[0085] Furthermore, considering that vehicle acceleration, size, and heading angle have inherent physical meaning and do not depend on a global coordinate system, only the spatial position needs to be decentralized, while other dynamic and geometric features, such as acceleration and heading angle, remain unchanged. This decentralized approach effectively eliminates the influence of different roundabout geometries and different road coordinate systems on model training, enabling the model to have stronger cross-scene generalization capabilities.

[0086] In trajectory prediction models based on graph neural networks, the interactions between vehicles are often modeled using graph structures. Vehicles within each time slice are represented as nodes in the graph, and the interaction strength between vehicles is described using an adjacency matrix, thus characterizing the complex vehicle interaction behaviors in a roundabout scenario. and vehicles Euclidean distance The formula is expressed as follows:

[0087] ;

[0088] The Euclidean distance is calculated using the coordinate data before decentralization. The distance matrix between each vehicle can be calculated based on the following formula:

[0089] ;

[0090] This embodiment uses a distance threshold-based method to construct adjacency relationships. When the distance between vehicles is large, the interaction between them can be considered weak, or even negligible. When the distance between two vehicles is less than a preset interaction radius... If they are in a relationship, then they are considered to have an interaction relationship; otherwise, they are considered not to have an interaction relationship. This indicates whether there is an interaction relationship between vehicles.

[0091] ;

[0092] Finally, we obtain an adjacency matrix that can represent the vehicle interaction relationships. To ensure the stability of graph convolution training, the obtained adjacency matrix is... Perform symmetric normalization:

[0093] ;

[0094] in, This represents the normalized adjacency matrix. The purpose of this is to add a connection to each node, pointing to itself, so as to avoid ignoring the characteristics of the node itself. It is a degree matrix used to represent the number of neighbors of each node.

[0095] S200. After completing trajectory data processing, this invention employs a spatiotemporal graph convolutional network (ST-GCN) to construct a spatiotemporal graph convolutional encoder to jointly encode the spatiotemporal features of the target vehicle and its neighboring vehicles, thereby capturing the complex dynamic interaction information between vehicles in a roundabout scenario. The spatiotemporal graph convolutional encoder consists of a single input normalization layer and two layers of spatiotemporal graph convolutional blocks (spatial graph convolution and temporal convolution). The core task of the spatiotemporal graph convolutional network is to decouple and model the spatial interactivity and temporal evolution of vehicles in a roundabout scenario. The encoder input tensor is in the form of... ,in This represents the number of independent samples (samples segmented by a time window) processed simultaneously during training. Representing feature dimension, Indicates the historical observation time step, This indicates the number of vehicle nodes contained in each frame. The spatiotemporal graph convolutional encoder used in this embodiment introduces learnable edge importance coefficients in the spatial convolution path and adaptively scales the connection strength of different spatial partitions by multiplying them element-wise with the normalized adjacency matrix. The process of the spatiotemporal graph convolutional encoder is as follows: Figure 3 As shown.

[0096] The core logic of spatial convolution is to enable the target vehicle to perceive the dynamic state of surrounding vehicles by weighted summation of the features of neighboring nodes. The spatial graph convolutional layer is responsible for feature aggregation based on the topological connections between vehicles at each observation time step. First, the system receives feature tensors from the input layer. Secondly, the adjacency matrix after symmetric normalization is used. Determine the interaction relationships between nodes. Spatial graph convolution uses a learnable transformation matrix. The formula for extracting linear correlations between features is as follows:

[0097] ;

[0098] in, This represents the normalized adjacency matrix corresponding to the f-th spatial partition. This is the feature mapping matrix corresponding to the spatial partition. Spatial partitioning refers to splitting the interactions between vehicles at the same time into multiple sub-adjacency matrices according to different connection types, with each sub-adjacency matrix corresponding to a spatial relationship pattern. The ReLU function is used as a non-linear activation function to give the model the ability to handle complex, non-linear driving logic.

[0099] In roundabout scenarios, the contributions of vehicle interactions under different spatial relationship patterns are not entirely consistent. The innovation of this application lies in improving the spatial graph convolutional layer. By introducing learnable edge importance coefficients into the spatial convolutional path, the adjacency matrix corresponding to different spatial partitions is adaptively scaled, enhancing the model's ability to express key interaction relationships. The specific improvement method is as follows:

[0100] Let the edge importance coefficient tensor corresponding to the l-th layer be . By introducing learnable edge importance coefficients The model can automatically adjust the contribution weights of different spatial partitions or adjacency patterns. Element-wise multiplication with the normalized adjacency matrix is ​​used to scale the connectivity strength across different spatial partitions. Compared to a traditional fixed adjacency matrix, The introduction of this factor allows the model to break free from the constraint of a single physical distance threshold. The spatial convolution formula after adding the edge importance coefficient is:

[0101] ;

[0102] Representing the tensor after spatial convolution, the temporal convolutional layer is a crucial module immediately following spatial convolution. By leveraging the sliding aggregation of one-dimensional convolutional kernels along the time axis, this layer achieves semantic dimensionality enhancement from discrete observation points to continuous motion trends, effectively filtering random sampling noise from the sensor. More importantly, it can capture the trends of the vehicle's physical inertial constraints and driving intentions through temporal weight allocation. The tensor after spatial convolution... As input, the temporal convolution process of the l-th layer can be represented as:

[0103] ;

[0104] in, This represents a one-dimensional convolution operation performed along the time dimension. This represents the output features of temporal convolution.

[0105] The above formula can be used The temporal convolution process unfolds as follows:

[0106] ;

[0107] in, Represents the set of all valid vehicles. Indicates time, The temporal convolution kernel size represents the temporal domain length that the model can capture; Indicates the first Layer-time convolution kernel at the 1st Learnable parameters at each location; This is a local temporal offset within the convolutional kernel, ensuring that every historical moment within the window can participate in the synthesis of the current feature; This is the center alignment offset, used to ensure that the output tensor after convolution remains spatially aligned with the input tensor on the time axis.

[0108] Temporal convolution, by aggregating dynamic states such as velocity, acceleration, and heading angle between adjacent historical moments, can extract more stable representations of motion trends, thereby enhancing the model's ability to represent complex behavioral patterns such as yielding, merging, detouring, and driving out in roundabout scenarios.

[0109] To improve the training stability of the spatiotemporal graph convolutional encoder and alleviate the gradient decay problem during deep feature extraction, this invention sets a fusion path between the input features and the convolutional output features within each spatiotemporal graph convolutional block. Thus, the output of the l-th layer graph convolutional block represents the historical spatiotemporal features. It can be expressed by the formula as follows:

[0110] ;

[0111] in, For input features The residual reference features are obtained after dimension alignment. When the number of input and output channels is the same, the residual reference features can be directly given by the input features; when the number of input and output channels is inconsistent, a linear mapping or dimension alignment operation is performed on the input features to ensure tensor dimension matching during residual fusion. Through this feature fusion mechanism, the model can extract higher-level spatiotemporal semantic representations layer by layer while preserving the original dynamic information, thereby improving the convergence of network training and the feature representation ability in complex interaction scenarios.

[0112] S300, Multimodal Target Point Prediction:

[0113] In a roundabout, the future direction of a vehicle is usually not uniquely determined. After entering the roundabout, a vehicle can either continue to circle around or exit from different exits, thus exhibiting significant multimodal characteristics in its future destination. To characterize this uncertainty, this invention includes a multimodal target point prediction module for simultaneously predicting... The coordinates of each candidate target point represent a potential driving intention. The target points guide the predicted trajectory, providing spatiotemporal constraints for subsequent physical kinematic reconstruction, ensuring the logical completeness and physical coherence of the long-term predicted trajectory, resulting in a smooth and continuous predicted trajectory. The multimodal target point prediction process is as follows: Figure 4 As shown.

[0114] Target point coordinate regression using contextual features of the last frame of the historical trajectory As input, contextual features are fused with changes in vehicle motion state and interactions between vehicles in historical trajectories. The multimodal target point prediction process can be represented as follows:

[0115] ;

[0116] in, Indicates the predicted A set of target points Indicates the first The two-dimensional relative displacement coordinates of each candidate target point relative to its current position. This represents a target point regression network consisting of two fully connected layers and a ReLU nonlinear activation function. The model selects the output features corresponding to the last frame of the historical trajectory in the encoded sequence as context features, and maps these context features to multiple two-dimensional candidate target points via the target point regression network. Through this process, this invention can simultaneously predict multiple future candidate target points based on the same historical spatiotemporal features, thereby achieving multimodal representation of vehicle driving intentions in roundabout scenarios.

[0117] In this embodiment, the predicted multimodal target points are not directly used as input to the decoder for trajectory feature generation at each time step. Instead, they are constrained by target point regression loss and trajectory endpoint consistency, forming a collaborative optimization relationship with the parallel trajectory decoding module. Through this mechanism, the model can improve the rationality and stability of the predicted trajectory at the endpoint position while maintaining the expressive power of the multimodal endpoint.

[0118] S400, Physical Constraint Parallel Decoding:

[0119] For long-term trajectory prediction tasks, traditional cyclic decoding methods typically employ a frame-by-frame recursive strategy, i.e., the first frame... The prediction of time depends on the first The output at each time point not only causes the early errors to accumulate and propagate to later time points, but also causes the inference overhead to increase linearly with the prediction duration, which is not conducive to real-time prediction in island scenarios.

[0120] Therefore, this embodiment employs a parallel decoding mechanism based on Transformer Decoder to generate the complete prediction trajectory in one step, reducing error accumulation and improving prediction stability. First, the system constructs a learnable time query vector corresponding to the future prediction length. .in, Indicates the future number The time query vector is a learnable time query vector for a specific future time. It guides the decoder to extract the most relevant information for that time from the historical spatiotemporal features output by the encoder. By combining the time query vector with the historical spatiotemporal features output by the encoder, the decoder can simultaneously generate decoding features for multiple future time moments in a single forward propagation. This mechanism avoids the serial decoding of traditional autoregressive structures, which uses the prediction result of the previous time moment as the input of the next time moment. This not only improves prediction efficiency but also reduces the impact of error accumulation on the prediction result in long-term prediction. Secondly, in the decoding part, this invention improves upon the traditional autoregressive decoding method by introducing a physically recursive decoding method. By progressively updating the acceleration, velocity, and displacement increments, the continuity and executability of the predicted trajectory are ensured. This physically recursive decoding method allows the model to avoid trajectory divergence in long-term prediction, helping to reduce prediction trajectory errors. The process of the physically constrained parallel decoding mechanism is as follows: Figure 5 As shown.

[0121] Define the historical spatiotemporal characteristics of the ST-GCN encoder output as follows: The velocity and acceleration corresponding to the last frame of the historical trajectory are denoted as follows: and , and They are all two-dimensional vectors. , For the future... At any given moment, a learnable time query vector is introduced. The parallel decoding process can then be represented as:

[0122] ;

[0123] in, This represents the future temporal characteristics of the Transformer decoder output.

[0124] Based on the future temporal characteristics output by the decoder, the system further regresses the acceleration increment and position residual at each time step and in each mode. Define the... The mode in the th ... The acceleration increment output at each future moment is The positional residual is The future multimodal displacement increment is The state update process under physical constraints is as follows:

[0125] ;

[0126] ;

[0127] ;

[0128] in, To predict the time step, This represents the hyperbolic tangent activation function, used to limit the position residual to a finite range, while maintaining physical recursive stability and locally compensating for the deviation between the ideal second-order kinematic model and the real trajectory.

[0129] Based on the displacement increment at each time point This can be further accumulated to obtain the future relative trajectory. : ;

[0130] Combined with the coordinates of the last frame of the historical trajectory This will give you the absolute trajectory of the future. :

[0131] ;

[0132] In this embodiment, the method of predicting acceleration increments and position residuals, and then recovering the complete trajectory through velocity and displacement kinematics recursion, ensures that the trajectory is smooth and conforms to physical constraints, avoiding unreasonable situations such as sudden changes in velocity or direction, and ensuring the executability of the predicted trajectory.

[0133] S500, Training Loss Function

[0134] This embodiment constructs a collaborative optimization loss function system consisting of trajectory regression loss, target point loss, direction consistency loss, and endpoint consistency loss. This system is used to simultaneously constrain the local motion details of the future trajectory, the final target point position, and the overall physical rationality. Within a multimodal prediction framework, this loss system guides the model to learn a future trajectory representation that conforms to kinematic laws and possesses endpoint discrimination capabilities.

[0135] In multimodal trajectory prediction tasks, the model will simultaneously output... There are several candidate future trajectories. Since the actual future trajectory corresponds to only one of the potential behavioral patterns, it is necessary to first consider... The optimal mode, which best matches the actual trajectory, is selected from the candidate modes. The vehicle is defined. The The modality in the future The predicted displacement increment at each time point is The actual displacement increment is So in the scene, the first... The average displacement error corresponding to each mode for:

[0136] ;

[0137] in, for Time vehicle The validity code is used to determine the vehicle. In the Whether it is valid at this moment, when When the vehicle is valid, This indicates that the vehicle is invalid. Let this be the final number of effective vehicles. The system selects the mode with the smallest average displacement error as the optimal mode for the current sample. Let the optimal mode be the [number of]th [mode]. With several modes, the target point for optimal mode prediction is obtained as follows: .

[0138] In the proposed loss function, the multimodal target point loss and endpoint loss have a significant impact on the long-term prediction capability of the trajectory. Specifically, the target point loss is used to calculate and constrain the deviation between the target point given by the multimodal target point prediction module and the true endpoint, making the predicted target point closer to the true endpoint. Building on this, the endpoint consistency loss further establishes the connection between the predicted trajectory and the target point. By constraining the consistency between the predicted trajectory endpoint and the predicted target point, the predicted trajectory can gradually converge to a reasonable endpoint position. These two losses work together to ensure the stability and rationality of the model's long-term trajectory prediction capability.

[0139] According to the vehicle The position of the endpoint of the true future trajectory relative to the last frame of the historical trajectory is The target point loss was calculated. for:

[0140] ;

[0141] in, This represents the smoothing L1 loss function. When the prediction error is small, it adopts a quadratic form to ensure a smooth and stable optimization process; when the prediction error is large, it adopts a linear form to reduce the excessive influence of abnormal biases on the training process. Therefore, this loss function can improve the robustness of model training while ensuring regression accuracy. The target point loss is used to constrain the predicted target point to be close to the true future endpoint position, thereby enhancing the model's ability to distinguish between roundabout exit selection and endpoint position.

[0142] To further ensure the consistency between multimodal target point prediction and physical recursive trajectory, this invention constructs an endpoint consistency loss. (Definition of vehicle...) The future cumulative endpoint displacement obtained through physical recursion in the optimal mode is: The prediction target point corresponding to the optimal mode is Then the endpoint consistency loss for:

[0143] ;

[0144] The endpoint consistency loss is used to ensure that the trajectory endpoint obtained by parallel decoding is consistent with the target point prediction result.

[0145] Next, the trajectory regression loss is constructed, and the vehicle corresponding to the optimal mode is defined. The future displacement increment sequence is ,vehicle The true future displacement increment sequence is Trajectory regression loss for:

[0146] ;

[0147] Trajectory regression loss is used to constrain the predicted trajectory to be as close as possible to the true trajectory over the entire future time range, ensuring not only the accuracy of the endpoint but also the continuity and local accuracy of the trajectory process.

[0148] In roundabout prediction, simply requiring proximity is insufficient to guarantee a reasonable trajectory direction. Therefore, this embodiment further introduces a direction consistency loss to constrain the predicted displacement increment direction to remain consistent with the actual displacement increment direction. Direction Consistency Loss for:

[0149] ;

[0150] in, Let cosine similarity be denoted as . Direction consistency loss is used to constrain future motion directions, reducing the likelihood of abrupt changes in direction in the predicted trajectory within a roundabout scenario. Let weights be added to the four losses mentioned above. , , and This represents their proportion in the total loss. Then the final loss function... The formula is expressed as:

[0151] ;

[0152] By using the loss function designed to comprehensively consider local motion details, the final target point position, and overall physical rationality, the model can automatically select the prediction mode that is closest to the real behavior among multimodal candidate trajectories, while simultaneously taking into account local trajectory accuracy, endpoint position accuracy, directional rationality, and endpoint consistency, thereby improving the accuracy and physical feasibility of vehicle trajectory prediction in roundabout scenarios.

[0153] In this embodiment, the loss function is used to optimize the parameters in the spatiotemporal graph convolutional encoder, edge importance coefficient, multimodal target point prediction module, and Transformer decoder. The parameter set in this model is set as follows:

[0154] ;

[0155] in, The learnable parameters of the spatiotemporal graph convolutional encoder are represented. This represents the importance coefficient of learnable edges. This represents the learnable parameters of the multimodal target point prediction module. This represents the learnable parameters of the Transformer decoder, including the parameters of its internal trajectory prediction output.

[0156] Under this model, the Adam optimizer is used to adaptively optimize the model parameters. This process can be represented as follows:

[0157] ;

[0158] in, For model learning rate, and The first Second and third The parameter set for each model iteration The gradient of the total loss function with respect to the current model parameters represents the rate of change of the total loss as the parameters change. This indicates the adaptive parameter optimization direction obtained by the Adam optimizer based on the current gradient and historical gradient information. Through this optimization process, the model parameters will gradually optimize in the direction of reducing the total loss function.

[0159] After optimizing the model parameters in this round, the optimized parameters can be used for the next round of model training and vehicle trajectory prediction.

[0160] The technical solution of this embodiment will be further illustrated by a specific example below:

[0161] The experimental verification of this invention uses the RounD dataset for testing. The RounD dataset contains 23 CSV files; files 1-14 are used for model training, files 15-17 for model validation, and files 18-23 for model testing. The model reads the historical trajectory of the past 2 seconds and predicts the likely trajectory of the target vehicle (Node 0) in the next 5 seconds. A trajectory. Then from... The prediction mode with the smallest Final Displacement Error (FDE) is selected from the trajectories. The final displacement error of this mode is taken as the minimum final displacement error, denoted as minFDE. The average displacement error corresponding to this mode is calculated and denoted as minADE. A total of 30,303 test samples were processed in this experiment. The test results of this invention are compared with those of the multimodal GRIP model, and the results are shown in Tables 1 and 2.

[0162] Table 1. minADE of this method and the multimodal GRIP model

[0163]

[0164] As shown in Table 1, with the increase of prediction time, the average displacement error of the predicted optimal trajectory gradually increases from 0.1148 meters to 0.9334 meters. This invention outperforms the multimodal GRIP model in minADE from 1 second to 5 seconds, and the overall trajectory is closer to the true trajectory.

[0165] Table 2. minFDE of this method and the multimodal GRIP model

[0166]

[0167] According to the data in Table 2, the endpoint displacement error of the predicted optimal trajectory increased from 0.1379 meters to 2.2653 meters. Similarly, the present invention outperforms the multimodal GRIP model in minFDE from 1 second to 5 seconds, and the prediction of the trajectory endpoint is more accurate.

[0168] Figure 6 , Figure 7 , Figure 8 and Figure 9 To illustrate some of the predicted trajectories, the black lines in the figure represent the historical trajectories of the target vehicle, the colored dashed lines represent the predicted multimodal future trajectories, the green solid lines represent the actual future trajectories of the target vehicle, the colored crosses mark the multimodal target points, and the red dots represent the actual target points. Figure 6 , Figure 7, Figure 8 and Figure 9 These represent the prediction results for test samples 5300, 5500, 14600, and 16600, respectively.

[0169] Example 2: This example provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described in Example 1.

[0170] Example 3: This example provides a computer device, including:

[0171] Memory, used to store computer programs / instructions;

[0172] A processor for executing the computer program / instructions to implement the steps of any of the methods described in Embodiment 1.

[0173] Example 4: This example provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the method described in any one of Examples 1.

[0174] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

[0175] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0176] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0177] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0178] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure and not to limit its protection scope. Although this disclosure has been described in detail with reference to the above embodiments, those skilled in the art should understand that after reading this disclosure, they can still make various changes, modifications or equivalent substitutions to the specific implementation of the invention, but these changes, modifications or equivalent substitutions are all within the protection scope of the pending claims.

Claims

1. A vehicle trajectory prediction method for roundabout scenarios, characterized in that, include: Data preprocessing: Obtain historical trajectory data of vehicles in the roundabout scene; filter vehicle types from the historical trajectory data and construct samples. Based on the samples, an adjacency matrix is ​​constructed through decentralized processing; Based on the sample, the vehicle's velocity, acceleration, heading angle, and validity mask in the local Cartesian coordinate system are extracted to construct a multi-dimensional input feature vector; The multidimensional input feature vector and adjacency matrix are input into the spatiotemporal graph convolutional encoder to extract the historical spatiotemporal interaction features of the target vehicle and its neighboring vehicles; A learnable time query vector corresponding to the future prediction length is constructed, and combined with the historical spatiotemporal interaction features, future temporal features are generated through a Transformer decoder; the acceleration increment and position residual at each moment and in each mode are predicted based on the future temporal features, and the future multimodal trajectory is recursively recovered based on the acceleration increment and position residual; The loss function is obtained by selecting the mode with the smallest average displacement error, and the parameters of the spatiotemporal graph convolutional encoder and Transformer decoder are optimized based on the loss function. Vehicle trajectory prediction is performed using a spatiotemporal graph convolutional encoder and a Transformer decoder with optimized parameters.

2. The vehicle trajectory prediction method for roundabout scenarios according to claim 1, characterized in that, The sample is constructed by dividing the continuous trajectory of historical trajectory data using a sliding time window method, with each time window serving as a sample.

3. The vehicle trajectory prediction method for roundabout scenarios according to claim 1, characterized in that, The decentralized processing specifically includes: taking the average position of all valid vehicles in the last frame of the historical trajectory as the center of the current scene, and normalizing the position coordinates of each vehicle at each moment within the time window.

4. The vehicle trajectory prediction method for roundabout scenarios according to claim 3, characterized in that, The adjacency matrix is ​​constructed as follows: the interaction relationship between vehicles is determined based on the Euclidean distance between each vehicle in the last frame of the historical trajectory. When the distance between two vehicles is less than the preset interaction radius, it is determined that there is an interaction relationship between them. The adjacency matrix is ​​then normalized to obtain the adjacency matrix used for graph convolution calculation.

5. The vehicle trajectory prediction method for roundabout scenarios according to claim 1, characterized in that, The spatiotemporal graph convolutional encoder includes one input normalization layer and two spatiotemporal graph convolutional blocks, wherein the spatiotemporal graph convolutional blocks include spatial graph convolution and temporal convolution.

6. The vehicle trajectory prediction method for roundabout scenarios according to claim 5, characterized in that, In the convolution part of the spatial graph, learnable edge importance coefficients are introduced, which are multiplied element-wise with the adjacency matrix to scale the connection strength of different spatial partitions.

7. The vehicle trajectory prediction method for roundabout scenarios according to claim 1, characterized in that, The process involves constructing a learnable time query vector corresponding to the future prediction length, combining it with the historical spatiotemporal interaction features, and generating future temporal features through a Transformer decoder. Based on these future temporal features, the process predicts the acceleration increment and position residual for each moment and each mode, and recursively recovers the future multimodal trajectory based on the acceleration increment and position residual. This includes: Define the vehicle's velocity and acceleration in the last frame of the historical trajectory as follows: , , No. The mode in the th ... The acceleration increment output at each future moment is The positional residual is The predicted time step is Future multimodal displacement increments are derived through acceleration updates and two-way kinematics. : ; ; ; in, , , They represent the first The mode in the th ... The acceleration, velocity, and displacement increments at a future moment. This represents the hyperbolic tangent activation function, used to constrain the position residual within a finite range; by adjusting the displacement increment... The system performs cumulative calculations and combines the position of the last frame of the historical trajectory to reconstruct the future absolute trajectory.

8. The vehicle trajectory prediction method for roundabout scenarios according to claim 1, characterized in that, The average displacement error is calculated as follows: Define vehicle The The modality in the future The predicted displacement increment at each time point is The actual displacement increment is In the scene The average displacement error corresponding to each mode for: ; in, for Time vehicle The validity code is used to determine the vehicle. In the Whether a time is valid, when When the vehicle is valid, This indicates that the vehicle is invalid. For the final number of effective vehicles, Indicates the historical observation time step.

9. The vehicle trajectory prediction method for roundabout scenarios according to claim 8, characterized in that, The step of selecting the mode with the smallest average displacement error for loss function calculation includes: The mode with the smallest average displacement error is selected as the optimal mode, and the target point loss, trajectory regression loss, direction consistency loss and endpoint consistency loss are calculated. A loss function is constructed based on the target point loss, trajectory regression loss, direction consistency loss, and endpoint consistency loss. The loss function The formula is expressed as follows: ; In the formula, , , and Indicates weight, Indicates the loss at the target point. Indicates the loss of endpoint consistency. Indicates the trajectory regression loss. This represents the loss of directional consistency.

10. The vehicle trajectory prediction method for roundabout scenarios according to claim 9, characterized in that, The target point loss The calculation formula is as follows: ; in, Represents the smoothing L1 loss function. Indicates a validity mask. Indicates according to vehicle The position of the endpoint of the true future trajectory relative to the last frame of the historical trajectory; Endpoint consistency loss The calculation formula is as follows: ; in, Indicates vehicle At the last moment Validity coding, Indicates vehicle The future cumulative endpoint displacement is obtained through physical recursion in the optimal mode. The predicted target point corresponding to the optimal mode is: ; Trajectory Regression Loss The calculation formula is as follows: ; in, This represents the vehicle corresponding to the optimal mode. The future displacement increment sequence, Indicates vehicle The true future displacement increment sequence; Directional consistency loss The calculation formula is as follows: ; in, This represents the cosine similarity.