Method and system for predicting track of vulnerable road user based on motion query enhancement
Generate extended static queries through normalization and clustering, and optimize decoder connections using cross-connection dynamic query, the problem of motion intention query coverage limitations and dynamic query connections in the prior art is solved, and more accurate prediction of the future trajectory of vulnerable road users is achieved.
Patent Information
- Application Number
- CN202510297766.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The existing motion intention query has limitations when covering the motion mode, and cannot effectively represent the motion intention of the disadvantaged road users. The dynamic query is not closely connected in the cascade decoder, resulting in inaccurate prediction results.
By normalizing the current position and direction of the moving sample, the clustering center is used as a static query, and the observation sequence of the moving sample is added to the static query to generate an extended static query. Use cross-connect dynamic queries to establish connections between shallow and deep decoders, and combine cross-attention module to aggregate trajectory features and map features to optimize dynamic queries.
Enhanced multimodal intent representation of queries, can more accurately predict the future trajectory of vulnerable road users, cover more diverse motion patterns and intentions, and improve the effectiveness of dynamic queries.
Smart Images

Figure CN120216607A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of autonomous driving, and particularly relates to a method and system for predicting the trajectories of vulnerable road users based on enhanced motion queries. Background Art
[0002] Influenced by historical and environmental information, trajectory prediction predicts future trajectories by inferring the motion intentions of traffic participants, which is crucial for the safety and comfort of autonomous driving systems. However, due to the uncertainty of the intentions of vulnerable road users and the complexity of the traffic environment, the motion intentions of vulnerable road users essentially exhibit diversity, that is, there are multiple reasonable future trajectories for a given historical trajectory, which makes the trajectory prediction task extremely challenging.
[0003] The development of query design in Transformer has inspired researchers to start using queries for multimodal trajectory prediction. One type of method is to represent the motion multimodality of traffic scenarios through a set of motion intention queries. However, the existing motion intention queries have two limitations. First, the coverage of static queries for motion patterns is limited by the effective range of motion samples. The existing static queries are within the boundary samples of motion patterns, resulting in the inability of static queries to effectively cover these motion patterns and thus unable to comprehensively represent the motion intentions of vulnerable road users. In addition, dynamic queries are gradually optimized in the cascaded decoder. Since the decoder layers are connected sequentially, each decoder only optimizes the target based on the prediction of the previous decoder, resulting in weak connections between two relatively distant decoders. In addition, this sequential connection will cause the dynamic queries from the previous decoder, whether effective or not, to be propagated to the next decoder. Therefore, deeper decoders may be affected by cascading errors and generate worse prediction results. To solve these problems, the present invention proposes a method for predicting the trajectories of vulnerable road users based on enhanced motion queries. Summary of the Invention
[0004] The object of the present invention is to provide a method for predicting the trajectories of vulnerable road users based on enhanced motion queries that can enhance the multimodal intention representation of queries, so as to achieve accurate prediction of the future trajectories of vulnerable road users.
[0005] To achieve the above object, on the one hand, the present invention provides a method for predicting the trajectories of vulnerable road users based on enhanced motion queries, including the following steps:
[0006] Align the current position and direction of each motion sample through normalization, cluster the end points of the normalized motion samples, use the obtained cluster centers as the generated static queries, and add the observation sequences of the motion samples to the generation process of the static queries to obtain extended static queries;
[0007] Based on the cascaded Transformer decoding layer, a cross-connection dynamic query is used to establish a connection between the dynamic queries output by the shallow decoder and the deep decoder. Combining the cross-attention module to aggregate and enhance the trajectory feature O and the map feature C to obtain the updated query feature for each layer, and an optimized dynamic query is obtained.
[0008] Taking the optimized dynamic query as the input of the motion decoder module, in the motion decoder module, an MLP layer is used for each decoding layer to predict the probability and parameters of each Gaussian component using the Gaussian mixture model at each time step, and the predicted center of the extracted Gaussian component is used to obtain the predicted trajectory.
[0009] Integrate the predicted trajectories of multiple homogeneous motion decoder modules, adopt an integration strategy for similar motion patterns, based on the predicted trajectories of different motion decoder modules and the scores corresponding to each predicted trajectory, calculate the distances between all the predicted trajectories using non-maximum suppression, and select those with distances less than the set threshold as the prediction results.
[0010] Furthermore, the current position and direction of each motion sample are aligned through normalization. Clustering the endpoints of the normalized motion samples includes:
[0011] For a given motion sample The current position and direction of each motion sample are aligned through the following process:
[0012]
[0013]
[0014] where and are the position and velocity of the VRU at time t respectively; T is the length of the observation sequence; R(θ T ) is the rotation matrix with the rotation angle θ T , that is, the angle between the position at time T and the coordinate axis. Through the above transformation, the current positions of all motion samples are aligned to the origin, and their motion directions are aligned to a relatively restricted space. The distance from the target of the motion sample to the origin is called the effective range d, and during the normalization process, d = T p . Using the clustering algorithm to cluster the endpoints of the normalized motion samples, and taking the obtained cluster centers as the generated static queries K is the number of static queries.
[0015] Furthermore, the observation sequence of the motion sample is added to the generation process of the static query, and the extended static query is obtained, including:
[0016] Add the observed sequence of motion samples to the generation process of the static query, and extend the effective range d from T p to T+T p ; Translate the starting position of the observed sequence of all motion samples to the origin, rotate it around the set motion direction, and given a motion sample its normalization process is as follows:
[0017]
[0018] where is the starting point of the observed sequence.
[0019] Furthermore, obtaining the enhanced trajectory feature O includes the following steps: Based on the trajectory data and map data after vectorization processing, extract the VRU trajectory feature and VRU map feature, splice and perform cross-attention calculation to obtain the map feature C, and combine the VRU trajectory feature, map feature and trajectory history feature to extract the interaction feature of the future trajectory; Pass the VRU trajectory feature and the interaction feature of the future trajectory through feature splicing and three MLP layers to obtain the enhanced trajectory feature O.
[0020] Furthermore, based on the trajectory data and map data after vectorization processing, extracting the VRU trajectory feature and VRU map feature and performing splicing and cross-attention calculation to obtain the map feature C includes:
[0021] Adopt the historical trajectory encoding based on Transformer, use the polyline encoder to encode each polyline into an input token feature of the Transformer encoder, which is used to aggregate the features of each polyline into the VRU trajectory feature, adopt the map encoding based on Transformer similar to the trajectory encoding, encode the road map, which is used to aggregate the features of each polyline into the VRU map feature, and the obtained VRU map feature Cp and VRU trajectory feature O p After splicing, input it into the Transformer encoder for cross-attention calculation to obtain the map feature C.
[0022] Furthermore, based on the cascaded Transformer decoding layer, use the cross-connection dynamic query to establish a connection between the dynamic queries output by the shallow decoder and the deep decoder, and combine the cross-attention module to aggregate the enhanced trajectory feature O and the map feature C to obtain the updated query feature of each layer, and the optimized dynamic query includes:
[0023] Update the dynamic query output by each decoder according to the following formula:
[0024]
[0025] where D l(q, k, v) is the l-th decoder layer with cross-attention mechanism in the Transformer, where 0 ≤ l ≤ L - 1; F c is the feature of the VRU to be predicted; is the interaction feature between the VRU to be predicted, neighboring VRUs, and map polylines; N is the total number of interaction objects, is the k-th dynamic query in the l-th decoder layer, and the dynamic query is initialized by the static query;
[0026] The cross-connection dynamic query is implemented by pooling connection or dense connection.
[0027] Furthermore, the pooling connection is specifically: pooling all the dynamic queries of intermediate layers to the last decoder layer. For the final layer L, its expression is:
[0028]
[0029] The dense connection is expressed by connecting the dynamic queries of each layer and all previous layers, as:
[0030]
[0031] D l+1 is the (l + 1)-th decoder layer with cross-attention mechanism in the Transformer, w p is the balance coefficient, is the k-th dynamic query in the p-th decoder layer.
[0032] Furthermore, the predicted trajectories of multiple homogeneous motion decoder modules are integrated, adopting an integration strategy for similar motion patterns. Based on the predicted trajectories of different motion decoder modules and the scores corresponding to each predicted trajectory, it includes:
[0033] Given a set of models {M1, M2, …, M J}, each model gives K predicted trajectories and corresponding scores; the predicted trajectory and the corresponding score are the k-th predicted trajectory of model M j , where T p is the length of the future trajectory, k ∈ {1, 2, …, K}, j ∈ {1, 2, …, J}. Assume that the k-th predicted trajectory of each model has a similar motion pattern, and its corresponding predicted score is For confidence integration, calculate the average value of the confidence of the k-th trajectory given by all J models as the confidence of the k-th predicted trajectory after integration;
[0034] For trajectory integration: By calculating the end point of the k-th predicted trajectory given by J models the average trajectory end point is obtained Take the one that is closest to the end point as the final integrated target point, while the corresponding predicted trajectory is selected as the k-th predicted trajectory for integration. For K predictions, K trajectories are obtained and the confidence levels corresponding to the K trajectories
[0035] Furthermore, during the training of the motion decoder module:
[0036] Use an MLP layer for each decoding layer to predict the probability p of each Gaussian component of the Gaussian mixture model at each time step k and the parameters (μ x , σ x , μ y , σ y , ρ). Then, use the predicted centers of the extracted Gaussian components to obtain the predicted trajectory. The dynamic queries of each decoder layer are optimized with the ground truth through a hard assignment strategy. The cross-entropy loss is used to maximize the probability of the selected positive Gaussian components. The final loss is the weighted sum of all Gaussian regression losses of each decoder layer, and the weights of each decoder layer are equal.
[0037] On the other hand, the present invention provides a vulnerable road user trajectory prediction system based on motion query enhancement, including a static query acquisition module, a dynamic query acquisition module, and a prediction module;
[0038] The static query acquisition module is used to align the current position and direction of each motion sample through normalization, cluster the end points of the normalized motion samples, use the obtained cluster centers as the generated static queries, and add the observation sequences of the motion samples to the generation process of the static queries to obtain extended static queries;
[0039] The dynamic query acquisition module is based on a cascaded Transformer decoding layer, uses cross-connection dynamic queries to establish connections between the dynamic queries output by the shallow decoder and the deep decoder, combines the cross-attention module to aggregate and enhance the trajectory feature O and the map feature C to obtain the updated query feature of each layer, and obtains optimized dynamic queries;
[0040] The prediction module is used to take the optimized dynamic query as the input of the motion decoder module. In the motion decoder module, an MLP layer is used for each decoding layer to predict the probability and parameters of each Gaussian component of the Gaussian mixture model at each time step, and the predicted center of the Gaussian component is used to obtain the predicted trajectory. The predicted trajectories of multiple homogeneous motion decoder modules are integrated. An integration strategy of similar motion patterns is adopted. Based on the predicted trajectories of different motion decoder modules and the scores corresponding to each predicted trajectory, the distance between all predicted trajectories is calculated using non-maximum suppression, and those with a distance less than the set threshold are selected as the prediction results.
[0041] Compared with the prior art, the present invention has at least the following beneficial effects: The present invention extends the effective range of motion samples from their future trajectory sequences to the complete trajectory sequences. As the time range increases, richer motion behaviors are introduced into the clustering samples, enabling the generated static queries to cover more diverse motion patterns and intentions. The cross dynamic query introduces cross connections across decoders to replace direct connections, which is beneficial for the optimization and update of dynamic queries and improves the effectiveness of the final dynamic query. Such processing can more effectively extract features that can reflect the motion trends of vulnerable road users, making the prediction results more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of existing static queries and the expanded static query range;
[0043] Figure 2 It is the basic framework of the model of the method for predicting the trajectories of vulnerable road users based on motion query enhancement of the present invention;
[0044] Figure 3 It is a high-definition map marker obtained after vectorizing the trajectory and the map;
[0045] Figure 4 It is a schematic diagram of the Transformer decoder network structure based on motion query;
[0046] Figure 5 It is different connection methods for dynamic query optimization;
[0047] Figure 6 It is the model integration method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] First, the following four concepts are described:
[0050] 1) Motion sample: Used to generate static queries. Each motion sample is a sequence of trajectories, including an observation sequence and a future sequence. The end point of the observation sequence is the motion state at the current moment, and the end point of the future sequence is called the intention. The motion state includes position, speed, and angle. A motion sample is the trajectory of a specific VRU in a specific traffic scenario.
[0051] 2) Traffic scenario: Describes the motion state of VRUs and map information. The motion state includes the positions and speeds of the predicted VRU, neighboring VRUs, and the host vehicle. The map information includes the topological structure of the road and traffic signs. Usually, the future trajectory of the target VRU needs to be predicted based on the states of other VRUs and the map.
[0052] 3) Static query: Multiple intentions predefined on motion samples, which are independent of the training and inference of the model. Represented by a set of vectors, such as points or sequences.
[0053] 4) Dynamic query: Multiple predicted intentions updated in the decoding modules of multiple cascaded decoders after being initialized by static queries. Each decoder layer generates a set of dynamic queries, represented by a set of vectors. Refer to Figure 1 。
[0054] Example 1, as Figure 2 shown, the method for predicting the trajectory of vulnerable road users based on motion query enhancement provided by this application has the following specific implementation steps:
[0055] The first step is to vectorize the trajectory data and map data. For the prediction target, its trajectory is a time-related directed spline curve. All these elements can be approximated as a sequence of vectors: for the map data, select a starting point and direction, uniformly sample key points from the spline curve at the same spatial distance, and sequentially connect adjacent key points into vectors; for the trajectory, key points can be sampled starting from t = 0 at a fixed time interval (such as 0.1 seconds) and connected into vectors. At a sufficiently small spatial or time interval, the generated polyline can approximate the original map and trajectory. The vectorization process of the data performs a one-to-one mapping between continuous trajectories, map annotations, and vector sets, and the coordinate points can form a graph representation on the vector set for encoding using graph neural networks. The map is divided into multiple polylines P using vectorization j 。
[0056] Each vector v j belonging to the polyline P i is regarded as a node in the graph, and the node features are given by the following formula: where, and are the coordinates of the vector starting point and the vector ending point, and d itself can be represented as 2D coordinates (x i , y i ) or 3D coordinates (x i , y i , z i ); a i corresponds to attribute features, such as object type, timestamp of the trajectory, or road feature type or speed limit of the lane; j is the integer ID of the polyline P j , indicating that v i ∈ P j . To make the input node features invariant to the position of the prediction target, the coordinates of all vectors are normalized so that the input nodes are centered on the position at the last observation moment of the prediction target.
[0057] Second, extract and fuse the trajectory features and map features. After the trajectory data and map data are vectorized, first, use the Transformer-based historical trajectory encoding to represent the historical states of N traffic participants as O in ∈ R N×T×K , where T is the number of historical frames and K is the number of state information (e.g., position, heading angle, and speed). For trajectories with fewer than T frames, zeros are filled at the positions of the missing frames; then, use a simple polyline encoder to encode each polyline into a feature of an input token of the Transformer encoder: O p = Φ(MLP(O in ))), where MLP(·) is a multi-layer perceptron network and Φ is the max pooling operation used to aggregate the features of each polyline into the VRU trajectory feature O p ∈ R N×D , where the feature dimension is D; immediately afterwards, use the Transformer-based map encoding similar to the trajectory encoding to represent the road map as where M is the number of map polylines, n is the number of points in each polyline, and C m is the number of attributes of each point (e.g., position and road type). Similarly, its encoding method is: C p = Φ(MLP(C in ))), where MLP(·) is a multi-layer perceptron network and Φ is the max pooling operation used to aggregate the features of each polyline into the VRU map feature C p ∈ R M×D , where the feature dimension is D. After the obtained VRU map feature Cp and the VRU trajectory feature O p are concatenated together, they are sent into the Transformer encoder for cross-attention calculation to obtain the map feature C ∈ R M×DFinally, interactive modeling based on Transformer self-attention is adopted. Specifically, the j-th layer Transformer attention module can be expressed as:
[0058]
[0059] where MultiHeadAttn(·,·,·) is the multi-head attention layer, G0 = [O p , C p ∈ N (N+M)×D represents the concatenated feature composed of the VRU trajectory feature and the VRU map feature, and κ(·) represents the k nearest polylines of each query polyline found using the K-nearest neighbor algorithm. PE represents the sinusoidal positional encoding of the input tokens, and the latest position is used for encoding for each VRU.
[0060] After obtaining the above VRU trajectory feature O p ∈ R N×D and the map feature C ∈ R M×D , combined with the trajectory history feature, a regression head is used to densely predict the future trajectories and speeds of all VRUs, and further extract the interaction features of the future trajectories. That is: S 1:Tp = MLP(O p ), where S i ∈ R N×4 contains the future positions and speeds of each VRU at time step i, and T p is the number of future frames to be predicted. The same polyline encoder as O p is used to encode the future trajectory S 1:T of the VRU to encode the future state of the VRU as a feature O f ∈ R N×D , and then the enhanced trajectory feature O is obtained through feature concatenation and three MLP layers, O = MLP([O p , O f ), and the above enhanced trajectory feature O and the map feature C are used as the feature input F of the cascaded Transformer decoding layer.
[0061] As an example, the trajectory encoder includes a three-layer MLP with a feature dimension of 256. The input map information includes the position, direction, and type of each polyline point. Trajectory enhancement is performed in the trajectory encoder. Each map polyline contains up to 20 points, which is approximately 10 meters in the WOMD dataset. M = 768 nearest map polylines around the target position to be predicted are selected. The map encoder includes a five-layer MLP. Since the number of map polylines is much larger than the number of targets, the present invention uses a smaller feature dimension of 64. For the convenience of feature interaction and fusion, the trajectory encoder and the map encoder finally map the feature dimensions to 256 dimensions respectively through another linear layer. In the local self-attention of the encoder, the number of neighbors is set to 16. For the dense future prediction module, a three-layer MLP with a feature dimension of 512 is used to predict the future positions and velocities of all targets.
[0062] Through the cross-attention mechanism, the trajectory features and map features of the VRU are fused together and feature enhancement is performed to obtain the enhanced trajectory feature O, which provides additional future context information for the decoder network and helps the model predict a more scene-compliant future trajectory. Based on the vectorized data and the results of feature extraction and fusion, the generation and optimization of static queries and dynamic queries are carried out.
[0063] Reference Figure 3 , in the third step, the current positions and directions of each motion sample are aligned through a normalization process, and the endpoints of these normalized motion samples are clustered. The clustering centers obtained by clustering are used as the generated static queries. For a given motion sample The current positions and directions of each motion sample are aligned through a normalization process, and the process can be expressed as:
[0064]
[0065]
[0066] where and are the position and velocity of the VRU at time t respectively; T is the length of the observation sequence; R(θ T ) is the rotation matrix with a rotation angle of θ T , that is, the angle between the position at time T and the coordinate axis. Through the above transformation, the current positions of all motion samples are aligned to the origin, and the motion directions of the motion samples are aligned to a relatively restricted space. The distance from the target of the motion sample to the origin is called the effective range d, and under this normalization process, d = T p . Subsequently, a clustering algorithm (such as k-Means) is used to cluster the endpoints of these normalized motion samples, and the obtained clustering centers are used as the generated static queries Among them, K is the number of static queries.
[0067] The training and test traffic scenarios are also aligned through the above normalization method. Given a traffic scenario S = {C, I, M}, the normalization process can be expressed as:
[0068]
[0069] Among them, represents the motion state of the VRU to be predicted; is a set containing N neighboring VRUs; represents the motion state of the nth neighboring VRU; represents the map; is the position of the jth polyline among J map polylines at time t; after the traffic scenario is normalized, the static queries obtained by clustering can be directly applied to this normalized traffic scenario to represent the diverse future motion patterns of static queries.
[0070] Static queries can be applied to all traffic scenarios to cover the possible motion areas of the VRU to be predicted and represent diverse motion patterns. However, the limited effective range of motion samples causes the clustering centers to deviate from the marginal motion patterns, such as boundary samples. Since the VRU to be predicted in the training or test traffic scenario is normalized in the same way as the motion sample sampling, therefore, it is also difficult for static queries to cover these motion patterns of the VRU to be predicted. Use extended static queries to extend the coverage range of static queries to provide more diverse motion patterns. Add the observation sequence of motion samples to the generation process of static queries, and extend the effective range d of static queries from T p to T + T p . Translate the starting positions of all motion sample observation sequences to the origin, and then rotate them around the set motion direction. Given a motion sample Its normalization process is as follows:
[0071]
[0072] Among them, is the starting point of the observation sequence. In this way, the aligned motion sample distribution is more dispersed and can cover more diverse motion patterns.
[0073] Fourthly, use cross-connection dynamic queries to establish connections between the dynamic queries output by the shallow decoder and the deep decoder, and then coordinate the query optimization between the decoding layers. The static queries are updated sequentially in the cascaded decoder module. Each decoder learns the residuals of the previous decoder, and the output of each decoder is a dynamic query.
[0074] The Transformer-based motion query decoder module adopted by the present invention is as shown in Figure 4 . Its core is an N-layer stacked Transformer decoder, which continuously optimizes future trajectory prediction through iterative updates of motion queries. The decoder module includes: a static query component, a dynamic query update component, and a multi-modal trajectory prediction head.
[0075] The static query component reduces the uncertainty of future trajectories by leveraging different intention queries for different motion patterns. Specifically, first, κ representative intention points, i.e., static query points I ∈ R κ×2 , are generated by applying the k-means clustering algorithm to the end points of the real trajectories, where each intention point represents a motion pattern considering the motion direction and speed. Then, each static intention query is modeled as an intention feature through positional encoding as follows:
[0076] Q I = MLP(PE(I))
[0077] where PE(·) is the sine positional encoding, and Q I ∈ R κ×D .
[0078] The dynamic query update component aims to supplement global intention localization by iteratively collecting fine-grained trajectory features, thereby continuously fine-tuning the trajectory. Each dynamic query is also the positional embedding of a spatial point, which is initialized as the corresponding static intention point but is dynamically updated according to the predicted trajectory in each decoder layer. Specifically, given the predicted future trajectory in the j-th decoder layer, the dynamic search query update for the (j + 1)-th decoder layer is as follows:
[0079]
[0080] In each decoder layer, the static query is used to propagate information between different motion intentions. Specifically, the static query is used as the positional embedding of the self-attention module as follows:
[0081]
[0082] where C j-1 ∈ R κ×D is the query feature from the (j - 1)-th decoder layer, C 0 is initialized to zero, and is the updated query feature. Next, using the dynamic query as the query position feature of cross-attention, local features related to trajectory prediction are extracted from the scene context. Specifically, the query feature and the position feature are used as the query and the key respectively to decouple their contributions to the attention weights. Cross-attention modules are respectively used to aggregate the VRU trajectory feature O and the map feature C:
[0083]
[0084] where [·,·] represents feature concatenation; ∪ represents aggregating the trajectory feature and the map feature as the key and value of cross-attention; α(C) is an operation representing using the dynamic query as the query position feature of cross-attention to extract local features related to trajectory prediction from the scene context, which is used to collect local map features around L predicted trajectories for motion fine-tuning. Finally, the obtained C j ∈R κ×D is the updated query content feature for each motion query pair in the j-th layer.
[0085] Append a multi-modal trajectory prediction head to the output C j of each decoding layer to generate future trajectories. The Gaussian Mixture Model (GMM) is used to represent the distribution of the predicted trajectories at each time step. Specifically, for each future time step i ∈ 1, …, T, the probability p and the parameters (μ x , μ y , σ x , σ y , ρ) of each Gaussian component are predicted as follows:
[0086]
[0087] where, includes K Gaussian components N 1:K (μ x , σ x ; μ y , σ y ; ρ), and its probability distribution is p 1:K . The predicted distribution of the target VRU position at time step i can be expressed as:
[0088]
[0089] where, is the probability that the target appears at the spatial position o ∈ R 2 . The predicted trajectory can be generated by simply extracting the predicted centers of the Gaussian components.
[0090] The above is the specific network structure of each layer of the decoder, and the dynamic queries output by each decoder are updated according to the following formula:
[0091]
[0092] where D l (q, k, v) is the l-th decoder layer with cross-attention mechanism in the Transformer, where 0 ≤ l ≤ L - 1; F c is the feature of the VRU to be predicted; is the fused feature of the enhanced trajectory feature O and the map feature C; N is the total number of interaction objects, is the k-th dynamic query in the l-th decoder layer. In addition, the dynamic query is initialized by the static query;
[0093] On this basis, cross-connected dynamic queries are used to establish a connection between shallow dynamic queries and deep dynamic queries, thereby coordinating the query optimization between decoding layers. Referring to Figure 5 , the cross-connected dynamic query can be implemented in two ways. The first way is the pooling connection, which pools all the dynamic queries of the intermediate layers to the last decoder layer. For the final layer L, its expression is:
[0094]
[0095] where w p is the balance coefficient, is the k-th dynamic query in the p-th decoder layer.
[0096] The second way is the dense connection, which connects the dynamic queries of each layer and all previous layers, that is, the pooling connection is used in each decoder layer. The expression of the dense connection can be expressed as:
[0097]
[0098] The Transformer based on motion query enhancement (Query Enhanced Motion Transformer, QE-MTR) proposed by the present invention stacks 6 layers of Transformer decoders, and needs to predict the future motions of pedestrians, bicycles and vehicles, which are two types of VRUs. The motion patterns of traffic participants in each category are different. Therefore, a set of static queries of 64 points are generated for each category of traffic participants respectively. In the cross dynamic query, the balance coefficient w p is set to 0.1. The prediction head in each decoder layer adopts a three-layer MLP with an intermediate feature dimension of 512, and the model weights are not shared between different decoding layers.
[0099] Step 5: Perform model training and inference. During training, combine the static queries obtained by extending the static query strategy with the context features and input them into the motion decoder module of cross-dynamic queries. Specifically, use an MLP layer for each decoding layer to predict the probability p of each Gaussian component of the Gaussian Mixture Model (GMM) at each time step. k and the parameters (μ x ; σ x ; μ y ; σ y ; ρ), and finally obtain the predicted trajectory by using the predicted centers of the extracted Gaussian components.
[0100] As an example, the dynamic queries of each decoder layer are optimized with the ground truth through a hard assignment strategy as follows:
[0101]
[0102] where represents the index of the dynamic query closest to the ground truth end point; GT is the future trajectory; L l is the Gaussian regression loss, which is used to maximize the likelihood of the true position of the VRU at a certain time step ; (μ x , σ x ; μ y , σ y ; ρ) is the selected positive Gaussian component of the predicted trajectory for optimization; is the predicted probability of this selected positive Gaussian component. In the above formula, the cross-entropy loss is used to maximize the probability of the selected positive Gaussian component; the final loss is the weighted sum of all Gaussian regression losses of each decoder layer, and the weights are equal.
[0103] Step 6: After training multiple homogeneous models through the above steps, the model is the trained motion decoder module. Refer to Figure 6 , and perform prediction result integration on multiple homogeneous models. Given the model set {M1, M2, …, M J}, each model gives K predicted trajectories and the score corresponding to each predicted trajectory The predicted trajectory and the corresponding score R1 is the k-th predicted trajectory of model M j , where T p is the length of the future trajectory, k ∈ {1, 2, …, K}, j ∈ {1, 2, …, J}. Assume that the k-th predicted trajectory of each model has a similar motion pattern, and its corresponding predicted score is For confidence integration, calculate the average of the confidence levels of the k-th trajectory given by all J models. As the confidence level of the k-th predicted trajectory after integration, the calculation is as follows:
[0104]
[0105] Similarly, for trajectory integration, first calculate the end points of the k-th predicted trajectory given by J models to obtain the average trajectory end point
[0106]
[0107] Select the one closest to as the final integrated target point. And the corresponding predicted trajectory is selected as the k-th predicted trajectory for integration. For K predictions, K trajectories and their corresponding confidence levels can be obtained. Finally, to generate 6 future trajectories for evaluation, use Non-maximum Suppression (NMS) to calculate the distances between the end points of the predicted K = 64 trajectories, and select the top 6 predictions, where the distance threshold is set to 2.5 meters.
[0108] Example 2, based on the concept of the method of the present invention, also provides a vulnerable road user trajectory prediction system based on motion query enhancement, including a static query acquisition module, a dynamic query acquisition module, and a prediction module;
[0109] The static query acquisition module is used to align the current position and direction of each motion sample through normalization, cluster the end points of the normalized motion samples, use the obtained cluster centers as the generated static queries, and add the observation sequences of the motion samples to the generation process of the static queries to obtain extended static queries;
[0110] The dynamic query acquisition module is based on a cascaded Transformer decoding layer, uses cross-connection dynamic queries to establish connections between the dynamic queries output by the shallow decoder and the deep decoder, combines a cross-attention module to aggregate and enhance the trajectory feature O and the map feature C to obtain the updated query feature for each layer, and obtains optimized dynamic queries;
[0111] The prediction module is used to take the optimized dynamic query as the input of the motion decoder module. In the motion decoder module, an MLP layer is used for each decoding layer to predict the probability and parameters of each Gaussian component using a Gaussian mixture model at each time step, and the predicted center of the Gaussian component is used to obtain the predicted trajectory. The predicted trajectories of multiple homogeneous motion decoder modules are integrated, and an integration strategy for similar motion patterns is adopted. Based on the predicted trajectories of different motion decoder modules and the scores corresponding to each predicted trajectory, the distances between all predicted trajectories are calculated using non-maximum suppression, and those with distances less than the set threshold are selected as the prediction results.
[0112] To evaluate the proposed method for predicting the trajectories of vulnerable road users based on motion query enhancement, official evaluation tools are used on the WOMD v1.2.0 dataset to calculate evaluation metrics, including Soft mAP, mean average precision (mAP), minimum average displacement error (minADE), minimum final displacement error (minFDE), miss rate (MR), and overlap rate (OR). Among them, Soft mAP is the most important ranking metric, and MR is the second most important ranking metric. At the same time, for the sake of comparison, while evaluating the accuracy of the prediction method of this application, various state-of-the-art methods are also selected for comparison, including laneGCN, PnPNet, LeapNet, HDGT, HPTR, LLM-MTR, MGTr+, Control MTR, GTR, and MTR++. The specific details of the models participating in the prediction accuracy evaluation are shown in Table 1:
[0113] Table 1
[0114]
[0115] Table 1 shows the comparison of the metrics between the method of this application and other methods in the 2023 Waymo Dynamic Prediction Challenge. The results of the method of this application are excellent in the comprehensive evaluation metric soft mAP and exceed MTR++. This fully demonstrates the effectiveness of the method. In addition, the method of this application also achieved competitive results in other evaluation metrics. By integrating 5 homogeneous models, the indicators of this method have been effectively improved.
[0116] Table 2
[0117]
[0118] Table 2 compares the method of this application with the latest methods on WOMD v1.1.0, and it is superior to other methods in terms of mAP. Through the DSQ and BDQ designed in this chapter, the method of this application significantly exceeds MTR in all evaluation metrics.
[0119] Table 3
[0120]
[0121] Table 3 evaluates the effectiveness of DSQ and BDQ in the method of this application. Both DSQ and BDQ can improve the performance of mAP, while the complete QE-MTR has a significant improvement in mAP performance, increasing from 0.3703 to 0.3991, indicating the effectiveness of DSQ and BDQ. In addition, the complete version also reduces the prediction errors, namely minADE and minFDE, further demonstrating the effectiveness of the method in this chapter. It should be noted that BDQ will increase the error rate to a certain extent. The possible reason is that the dense connection introduces the sub-optimal intermediate layer dynamic queries into the final dynamic query, resulting in more prediction endpoints deviating from the threshold, thus leading to an increase in the error rate.
[0122] The above content is only to illustrate the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.
Claims
1. A method for predicting vulnerable road user trajectories based on motion query enhancement, characterized in that: The following steps are involved: The current position and direction of each motion sample are aligned by normalization, the endpoints of the normalized motion samples are clustered, the cluster centers obtained by clustering are used as the generated static query, and the observation sequence of the motion sample is added to the generation process of the static query to obtain the extended static query; Based on the cascaded Transformer decoding layer, a cross-connection dynamic query is used to establish a connection between the dynamic queries output by the shallow decoder and the deep decoder. The cross-attention module is combined to aggregate the enhanced trajectory features O and map features C to obtain the updated query features of each layer, and the optimized dynamic query is obtained. The optimized dynamic query is used as the input of the motion decoder module. In the motion decoder module, an MLP layer is used for each decoding layer to predict the probability and parameters of each Gaussian component using the Gaussian mixture model at each time step, and the predicted center of the Gaussian component is extracted to obtain the predicted trajectory; The predicted trajectories of multiple homogeneous motion decoder modules are integrated, and an integration strategy of similar motion patterns is adopted. Based on the predicted trajectories of different motion decoder modules and the scores corresponding to each predicted trajectory, the distance between all predicted trajectories is calculated using non-maximum suppression, and the one with a distance less than a set threshold is selected as the prediction result.
2. The method for predicting vulnerable road user trajectories based on motion query enhancement according to claim 1, characterized in that: The current position and direction of each motion sample are aligned by normalization, and the endpoints of the normalized motion samples are clustered including: For a given motion sample The current position and direction of each motion sample are aligned through the following process: in, and are the position and velocity of VRU at time t; T is the length of the observation sequence; R(θ T ) is the rotation matrix, whose rotation angle is θ T , that is, the angle between the position at time T and the coordinate axis. Through the above transformation, the current position of all motion samples is aligned to the origin, and its motion direction is aligned to a relatively restricted space. The distance from the target of the motion sample to the origin is called the effective range d. In the normalization process, d = T p , cluster the endpoints of the normalized motion samples using a clustering algorithm, and use the resulting cluster centers as the generated static query K is the number of static queries.
3. The method for predicting vulnerable road user trajectories based on motion query enhancement according to claim 1, characterized in that: The observation sequence of motion samples is added to the generation process of static queries, and the extended static queries include: The observation sequence of motion samples is added to the generation process of static queries, and the effective range d is changed from T p Expand to T+T p ; Translate the starting position of the observation sequence of all motion samples to the origin and rotate around the set motion direction. Given a motion sample The normalization process is as follows: in, is the starting point of the observation sequence.
4. The method for predicting vulnerable road user trajectories based on motion query enhancement according to claim 1, characterized in that: The acquisition of enhanced trajectory feature O includes the following steps: based on the trajectory data and map data after vectorization processing, VRU trajectory features and VRU map features are extracted and spliced and cross-attention calculated to obtain map features C, and the interactive features of future trajectories are extracted by combining VRU trajectory features, map features and trajectory history features; the enhanced trajectory feature O is obtained by combining the interactive features of VRU trajectory features and future trajectories through feature splicing and three MLP layers.
5. The method for predicting vulnerable road user trajectories based on motion query enhancement according to claim 4, characterized in that: Based on the trajectory data and map data after vectorization processing, the VRU trajectory features and VRU map features are extracted and concatenated and cross-attention calculated to obtain the map features C including: Transformer-based historical trajectory encoding is used, and a broken line encoder is used to encode each broken line as an input tag feature of the Transformer encoder, which is used to aggregate the features of each broken line into VRU trajectory features. Transformer-based map encoding similar to trajectory encoding is used to encode the road map, which is used to aggregate the features of each broken line into VRU map features. The obtained VRU map feature Cp is consistent with the VRU trajectory feature O p After splicing, the data is input into the Transformer encoder for cross attention calculation to obtain the map feature C.
6. The method for predicting vulnerable road user trajectories based on motion query enhancement according to claim 1, characterized in that: Based on the cascaded Transformer decoding layer, a cross-connection dynamic query is used to establish a connection between the dynamic queries output by the shallow decoder and the deep decoder. The cross-attention module is combined to aggregate the enhanced trajectory features O and map features C to obtain the updated query features of each layer. The optimized dynamic query includes: The dynamic query for each decoder output is updated according to the following formula: Among them, D l (q, k, v) is the l-th decoder layer with cross-attention mechanism in Transformer, 0≤l≤L-1; F c is the characteristic of the VRU to be predicted; is the interaction feature between the VRU to be predicted, the neighboring VRU and the map line; N is the total number of interaction objects, is the kth dynamic query in the lth decoder layer, dynamic query Initialized by static query; Dynamic query is achieved through convergence connection or cross connection through dense connection.
7. The method for predicting vulnerable road user trajectories based on motion query enhancement according to claim 6, characterized in that: The specific method of convergence connection is to aggregate all dynamic queries of the intermediate layers to the last decoder layer. For the final layer L, its expression is: The dense connection method connects each layer to all previous layers through dynamic queries, expressed as: D l+1 is the l+1th decoder layer with cross attention mechanism in Transformer, w p is the balance coefficient, is the kth dynamic query in the pth decoder layer.
8. The method for predicting vulnerable road user trajectories based on motion query enhancement according to claim 1, characterized in that: The predicted trajectories of multiple homogeneous motion decoder modules are integrated, and the integration strategy of similar motion patterns is adopted. The predicted trajectories based on different motion decoder modules and the scores corresponding to each predicted trajectory include: Given a model set {M1,M2,…,M J }, each model gives K prediction trajectories and corresponding scores; prediction trajectory and the corresponding ratings It is model M j The k-th predicted trajectory of p is the length of the future trajectory, k∈{1,2,…,K}, j∈{1,2,…,J}, let the kth predicted trajectory of each model have similar motion patterns, and their corresponding prediction scores are For confidence ensemble, calculate the average confidence of the kth trajectory given by all J models As the confidence of the kth predicted trajectory after integration; For trajectory ensemble: by calculating the end point of the k-th predicted trajectory given by J models Get the average trajectory end point Take and The nearest destination As the final integration target point, The corresponding predicted trajectory Then it is selected as the kth prediction trajectory of the ensemble. For K predictions, K trajectories are obtained. And the confidence corresponding to K trajectories 9. The method for predicting vulnerable road user trajectories based on motion query enhancement according to claim 1, characterized in that: When the motion decoder module is trained: Use an MLP layer for each decoding layer to predict the probability p of each Gaussian component using a Gaussian mixture model at each time step k and parameters (μ x ,σ x ,μ y ,σ y ,ρ), and then use the predicted center of the extracted Gaussian component to obtain the predicted trajectory. The dynamic query of each decoder layer is optimized with the true value through a hard assignment strategy, and the probability of the selected positive Gaussian component is maximized using the cross entropy loss. The final loss is the weighted sum of all Gaussian regression losses of each decoder layer, and the weight of each decoder layer is equal.
10. A vulnerable road user trajectory prediction system based on motion query enhancement, characterized in that: It includes a static query acquisition module, a dynamic query acquisition module and a prediction module; The static query acquisition module is used to align the current position and direction of each motion sample by normalization, cluster the endpoints of the normalized motion samples, use the clustering centers obtained by clustering as the generated static query, and add the observation sequence of the motion sample to the generation process of the static query to obtain an extended static query; The dynamic query acquisition module is based on the cascaded Transformer decoding layer. It uses cross-connected dynamic queries to establish connections between the dynamic queries output by the shallow decoder and the deep decoder. It combines the cross-attention module to aggregate the enhanced trajectory features O and map features C to obtain the updated query features of each layer, and obtain the optimized dynamic query. The prediction module is used to take the optimized dynamic query as the input of the motion decoder module. In the motion decoder module, an MLP layer is used for each decoding layer to predict the probability and parameters of each Gaussian component of the Gaussian mixture model at each time step, and the predicted trajectory is obtained by extracting the prediction center of the Gaussian component; the predicted trajectories of multiple homogeneous motion decoder modules are integrated, and an integration strategy of similar motion patterns is adopted. Based on the predicted trajectories of different motion decoder modules and the corresponding scores of each predicted trajectory, the distance between all predicted trajectories is calculated using non-maximum suppression, and the one with a distance less than a set threshold is selected as the prediction result.
Citation Information
Patent Citations
Multi-modal pedestrian trajectory prediction method and system based on predefined tree
CN114511594A
Multi-modal trajectory prediction method, medium and system
CN118132981A
Multi-target joint detection tracking method based on time sequence representation enhancement and trajectory correction
CN118570754A
Automatic driving track prediction method and system based on track decoupling and Mama
CN119348649A
Sequential pedestrian trajectory prediction using step attention for collision avoidance
US20230038673A1
Cited By
Multi-mode pedestrian target trajectory prediction system and method
CN120632828A