Map-free long-time-domain vehicle trajectory prediction method based on end point clustering prior
Through the map-free trajectory prediction method with endpoint clustering prior, multimodal candidate endpoints are generated using the hierarchical map neural network and attention mechanism, which solves the problem of dependence on high-precision maps and low prediction efficiency in the existing technology, and achieves high-precision and real-time vehicle trajectory prediction under map-free conditions.
Patent Information
- Application Number
- CN202510395618.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
The existing trajectory prediction technology relies heavily on high-precision maps, resulting in high cost of switching between different scenarios, weak cross-region generalization capabilities, and low prediction efficiency cannot meet the real-time requirements.
A map-free long-time vehicle trajectory prediction method based on endpoint clustering prior is adopted. The interaction characteristics of traffic participants and the environment are extracted through a hierarchical map neural network, and the probability estimate of endpoint clustering and attention mechanisms are carried out. A non-maximum suppression algorithm is used to screen multimodal candidate endpoints, and a vehicle predicted trajectory is generated in combination with the trajectory completion module.
In driving scenarios without high-precision map support, accurately predicting the future long-distance motion trajectory of traffic participants, reducing the inference cost, meeting the real-time requirements of on-board terminals, and improving the generalization ability and prediction accuracy of the model.
Smart Images

Figure CN120260273A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of autonomous driving and intelligent transportation, and in particular, to a map-free long-time vehicle trajectory prediction method based on end-point clustering prior knowledge. Background Art
[0002] In autonomous driving and intelligent transportation systems, trajectory prediction technology, as a core link of environmental perception and decision-making planning, its accuracy is directly related to the safety and reliability of the autonomous driving system. Traditional methods mostly construct road topologies based on high-precision maps and predict the movement trajectories of traffic participants through rule engines or deep learning models. However, in some practical application scenarios, due to difficulties in precise positioning, signal interference caused by high-rise buildings and elevated structures, and loss of satellite signals in tunnels, the use of high-definition maps (HD maps) is severely restricted. At the same time, the dependence on map data hinders the transferability of the model, greatly increasing the cost of switching the model for use in different scenarios. Many bottlenecks have gradually emerged in the existing technology system in terms of map data dependence, computational efficiency, and scene generalization ability.
[0003] Current mainstream trajectory prediction solutions jointly model by fusing the lane topology information of high-precision maps and the historical trajectories of traffic participants. Although such methods can achieve high prediction accuracy, their technical implementation is restricted and highly dependent on high-precision maps. Currently, the map maintenance cost in the transportation field is increasing exponentially. The characteristic of continuous updating of high-precision maps is fundamentally contradictory to dynamic scenarios such as road construction and temporary traffic control. The cross-regional generalization ability of mainstream trajectory prediction solutions is weak, and the performance of models trained based on specific maps will drop sharply in un-surveyed areas. As pointed out in Waymo's 2022 technical report, its map maintenance cost has accounted for 17% of the total cost of the autonomous driving system, and the proportion of prediction errors caused by map update delays in urban scenarios reaches 23%. At the same time, existing models such as VectorNet need to perform piecewise vector coding on lane lines, and the encoding process of map semantic information consumes more than 30% of the inference computing power, and the performance drops sharply when deployed across regions. For example, it is pointed out in "Map-Free Trajectory Prediction in Traffic With Multi-Level Spatial-Temporal Modeling" that models relying on high-precision maps (HD maps) (such as LaneGCN, VectorNet, MMTransformer, AgentFormer, HiVT, AutoBots, etc.) are significantly restricted in performance or cannot run in map-free scenarios.
[0004] Meanwhile, most mainstream trajectory prediction methods cannot meet the real-time requirements of practical applications. Some mainstream trajectory prediction models mentioned in "Map-Free Trajectory Prediction in Traffic With Multi-Level Spatial-Temporal Modeling", such as DSP (a conference paper in IROS 2022), have a single prediction task duration of 1244 ms, while some map-free prediction methods such as NN (a conference paper in CVPR 2019) and CRAT-Pred (a conference paper in ICRA 2022) also have a prediction duration of around 200 ms, which cannot meet the real-time requirements of practical applications. Summary of the Invention
[0005] The object of the present invention is to overcome the above problems of highly relying on high-precision maps and low prediction efficiency that cannot meet the application requirements in actual engineering, and to provide a lightweight map-free trajectory prediction method.
[0006] The object of the present invention can be achieved by the following technical solutions:
[0007] A map-free long-term vehicle trajectory prediction method based on end-point clustering prior. The method constructs a trajectory prediction model, and after quantifying the trajectory prediction model, it is deployed on a vehicle terminal for vehicle trajectory prediction. The trajectory prediction steps specifically include:
[0008] Using a hierarchical graph neural network to encode the context information of the driving scenario and extract the features of the interaction between traffic participants and the surrounding environment;
[0009] By performing coordinate transformation and end-point extraction on historical trajectory data, end-point clustering is carried out in different speed domains to construct a set of dense candidate target points;
[0010] Performing probability estimation on the dense candidate target points through an attention mechanism, and using the non-maximum suppression algorithm to select multi-modal candidate end points;
[0011] Based on the multi-modal candidate end points, using a trajectory completion module to generate multi-modal long-term vehicle prediction trajectories.
[0012] As a preferred technical solution, the extraction of the features of the interaction between traffic participants and the surrounding environment is specifically as follows:
[0013] Vectorize the trajectories of traffic participants, and each vector node feature includes start / end coordinates, gender semantic labels, timestamps, and corresponding IDs;
[0014] Abstract the trajectories of each traffic participant as a polyline representation, and perform non-linear transformation on a single vector node through an MLP to generate preliminary embedding features;
[0015] Construct a fully connected subgraph for all nodes within the same polyline, and aggregate adjacent node information through max pooling;
[0016] Extract the local features of each subgraph, and adopt an attention mechanism to capture the interaction information between traffic participants to generate two-dimensional features.
[0017] As a preferred technical solution, the construction of a dense set of candidate target points is as follows:
[0018] Cut the continuous trajectory data into multiple trajectory segments of a set duration. For a trajectory segment, set a division moment. In the trajectory segment, the time before the division moment is the observation history, and the time after the division moment is the prediction target;
[0019] Taking the vehicle position at the division moment in the trajectory segment as the origin and the heading angle θ at this moment as the positive direction of the coordinate system, perform translation and rotation transformations on all points in the trajectory segment to convert them to the ego-vehicle coordinate system; for each trajectory segment, intercept the position coordinates at the end moment to form the original set of target points;
[0020] Perform speed stratification, calculate the average driving speed of the vehicle trajectory segment before the division moment, divide the trajectory segment into different speed domains based on a set speed threshold, and correspondingly divide the trajectory end points into the corresponding speed subsets; perform clustering on each speed subset respectively to generate multiple candidate target points.
[0021] As a preferred technical solution, the clustering of the speed subsets to generate multiple candidate target points is as follows:
[0022] In the initialization stage, randomly select the first clustering center from the speed subset as the initial seed point; calculate the shortest distance D(p i ) from each point to the selected center, and select the next center point according to the probability distribution P(p i ) ∝ D(p i ), and repeat this process until the speed subset is traversed; 2 In the iterative optimization stage, in the assignment stage, assign each trajectory end point to the nearest clustering center, and in the update stage, recalculate the cluster centers until the difference between the recalculated cluster centers and the original cluster centers is less than the set threshold or a certain number of iterations is reached;
[0023] Output the positions of the dense candidate target points obtained based on the end point clustering prior to obtain the dense set of candidate target points.
[0024] As a preferred technical solution, the selection of multi-modal candidate end points is as follows:
[0025] As a preferred technical solution, the selection of multi-modal candidate end points is as follows:
[0026] For the set of dense candidate target points, the attention mechanism is used to perform probability encoding on each dense candidate target point. The initial feature matrix E is obtained by encoding the 2D coordinates of the target points through an MLP, and the local information between the target points is obtained through the attention mechanism:
[0027] U = EW u , R = HW r , O = HW o
[0028]
[0029] Among them, is the linear projection matrix, H is the feature of the interaction between the extracted traffic participants and the surrounding environment, and each element B of matrix B ij represents the attention degree of target point i to j, and d k is the dimension of the query / key / value vector;
[0030] After multiplying matrix B by matrix O, the enhanced feature E' is obtained;
[0031] The predicted target scores of each dense candidate target point are calculated through the attention mechanism:
[0032]
[0033] Among them, the trainable function h(·) is implemented using an MLP
[0034] A probability heat map is formed according to the predicted scores of all dense candidate target points;
[0035] Using non-maximum suppression, the predicted end points with multi-modal features are selected according to the probability information of each candidate target point. A distance threshold is set to retain the point clusters with probability peaks and no overlap. Each cluster represents a possible driving intention, and other higher-probability points in the same area are suppressed. The top k dense candidate target points with multi-modal features that can cover multiple driving intentions are selected as the multi-modal candidate end points.
[0036] As an optimal technical solution, the trajectory completion module is specifically implemented as follows:
[0037] For each multi-modal candidate end point p k , the enhanced feature E' and the spatial coordinates obtained through the attention mechanism are extracted and concatenated with the feature of the interaction between the traffic participant and the surrounding environment to form a composite feature vector z t ;
[0038] The composite feature vector is unfolded in time through a 2-layer MLP decoder:
[0039]
[0040] Among them, σ is the ReLU activation function, MLP1 maps the input to the hidden space, and MLP2 outputs 2D coordinates to finally obtain the complete trajectory.
[0041] As a preferred technical solution, the vehicle trajectory prediction method uses a transfer learning strategy to train the trajectory prediction model:
[0042] Use a public dataset to pre-train a benchmark model; use the parameters of the pre-trained model as the initial values, reduce the learning rate and perform fine-tuning training on the target dataset; use the teacher forcing technique to supervise and train the trajectory completion module to optimize the model parameters.
[0043] As a preferred technical solution, the target dataset uses private trajectory data, and the private trajectory data is cleaned and trajectory smoothed to construct the target dataset.
[0044] As a preferred technical solution, for the trained trajectory prediction floating-point model, the vehicle trajectory prediction method is converted into a fixed-point model through quantization processing and deployed on the in-vehicle board side, specifically as follows:
[0045] Perform quantization calibration on the trajectory prediction model, and adjust the quantization parameters to obtain the calibration stage model:
[0046] Quantize the weights and activation values of the model and initialize the quantization parameters:
[0047]
[0048] Among them, x float is the floating-point value, x quant is the quantized integer value, s is the scaling factor, and z is the zero point;
[0049] Use the calibration dataset to perform inference on the model, statistically analyze the distribution of activation values, and optimize the scaling factor s and zero point z by minimizing the quantization error:
[0050]
[0051] Among them, b is the quantization bit number;
[0052] Perform quantization-aware training on the calibration stage model, and improve the accuracy of the calibration stage model by simulating the quantization error:
[0053] Insert a pseudo-quantization node in the floating-point model to simulate the quantization process:
[0054]
[0055] Among them, the clip function is used to limit the quantized value within the valid range;
[0056] The gradient norm is introduced to weight the loss function to improve the modeling ability for tail samples, and the gradient coordination training is performed on the model with inserted pseudo-quantization nodes to obtain the QAT quantization model;
[0057] Convert the quantized QAT quantization model into a fixed-point model and generate a computation graph file.
[0058] As a preferred technical solution, after the trajectory prediction model is deployed on the vehicle terminal, the model performance is tested as follows:
[0059] Use a data collection vehicle to collect trajectory data on actual roads. The collected trajectory data includes: timestamp, trajectory ID, position in the UTM coordinate system, speed, heading angle, types of traffic participants, and target category labels;
[0060] Use data playback technology to perform inference on the trajectory prediction model with actual road data and calculate performance evaluation, specifically including:
[0061] Calculate the minimum value of the average displacement error of all time steps between the predicted trajectory and the true trajectory to evaluate the overall fitting ability of the trajectory prediction model for the trajectory;
[0062] Calculate the minimum value of the error between the predicted trajectory and the true trajectory at the final time step to evaluate the accuracy of the trajectory prediction model in end-point prediction.
[0063] Compared with the prior art, the present invention has the following beneficial effects:
[0064] 1) The present invention designs and implements a mapless trajectory prediction algorithm. The adopted end-point clustering prior strategy uses the trajectory slicing technology to divide the continuous trajectory into segments with a fixed duration, and then extracts their end-point positions to form a set of true end-points. And cluster the trajectory end-point data in different speed domains to construct a set of dense candidate target points, and perform probability encoding on the candidate target points through the attention mechanism. Finally, use the non-maximum suppression algorithm to screen out the multi-modal predicted end-points and complete the trajectory. It can accurately predict the long-time domain motion trajectory of traffic participants in the driving scenario without the support of high-precision maps, providing a basis for the autonomous driving decision-making algorithm.
[0065] 2) The present invention applies quantization calibration and quantization-aware training to reduce the inference cost. The trained floating-point model is converted into a fixed-point model through quantization processing. The model is initially quantized using quantization calibration technology; subsequently, quantization-aware training is adopted to simulate quantization errors during the training process, further reducing the accuracy loss after quantization. Finally, a fixed-point model that meets the deployment requirements is generated, thereby significantly improving the operation efficiency while maintaining a high prediction accuracy. The single prediction duration of the fixed-point model on the vehicle board reaches within 100 ms, ensuring that the quantized model meets the real-time requirements of the board end while guaranteeing the model accuracy requirements. Description of the Drawings
[0066] Figure 1 It is a flowchart of a mapless long-time vehicle trajectory prediction method based on end-point clustering prior of the present invention.
[0067] Figure 2 It is a schematic diagram of a dense candidate target point set obtained by end-point clustering prior in a specific embodiment of the present invention; 2a) Dense candidate target points of the low-speed subset, 2b) Dense candidate target points of the medium-speed subset, 2c) Dense candidate target points of the high-speed subset;
[0068] Figure 3 It is a probability heat map of generating dense candidate target points in a specific embodiment of the present invention;
[0069] Figure 4 It is an example of a predicted trajectory completed on the candidate target point set in a specific embodiment of the present invention; 4a) Trajectory of the vehicle to be predicted in the first two seconds in a turning scenario, 4b) Predicted result of trajectory completion in a turning scenario, 4c) Trajectory of the vehicle to be predicted in the first two seconds in a straight scenario, 4d) Predicted result of trajectory completion of the model in a straight scenario;
[0070] Figure 5 It is a visualization comparison diagram of the predicted trajectory and the real trajectory without map support in a specific embodiment of the present invention, 5a) Turning scenario, 5b) Straight line scenario. Detailed Embodiment
[0071] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.
[0072] Embodiment 1
[0073] Aiming at the problems that the existing trajectory prediction model highly depends on high-precision maps and the prediction efficiency is low and cannot meet the application requirements in actual engineering, the present invention designs and implements a mapless trajectory prediction algorithm. As Figure 1 shown, the mapless long-time vehicle trajectory prediction method based on end-point clustering prior proposed by the present invention includes the following steps:
[0074] Step 1: Generation of candidate target point set and model construction: In this stage, first, the historical trajectory data and relevant scenario information are preprocessed, and the sparse context encoding method is used to extract the features of the interaction between traffic participants and the surrounding environment. After converting the trajectory data into the ego-vehicle coordinate system, the trajectory slicing technology is used to divide the continuous trajectory into segments with a fixed duration, and then the end positions are extracted to form the real end point set. In order to get rid of the algorithm's dependence on the map, different from the traditional method of selecting end points based on the lane line position, the present invention uses a clustering algorithm to cluster the trajectory end point data in different speed domains to construct a dense set of candidate target points (Goals), and the candidate target points are probabilistically encoded through an attention mechanism. Finally, the non-maximum suppression algorithm is used to screen out the multi-modal predicted end points and complete the trajectory complementation to complete the mapless trajectory prediction task.
[0075] Step 2: Construction of the target data set and model transfer training: First, the private trajectory data is cleaned and the trajectory is smoothed to ensure the quality of the input data and construct the target data set; then, the transfer learning strategy is adopted, and a benchmark model is pre-trained using the public data set. The parameters of the pre-trained model are used as the initial values and fine-tuned on the target data set; at the same time, the teacher forcing technology is used to supervise the training of the trajectory complementation module, enabling the model to accurately predict the future trajectory from the historical observations, and continuously optimizing the model parameters through the smoothing loss function to improve the prediction accuracy and robustness. Finally, the floating-point model after transfer learning is tested in multiple scenarios to ensure that its prediction accuracy is better than that of the model before transfer in the mapless case.
[0076] Step 3: Model quantization-aware training: To meet the requirements of real-time deployment on in-vehicle terminals, the trained floating-point model is converted into a fixed-point model through quantization processing in this stage. First, statistical analysis is performed on the weights and activation values of the model to determine the appropriate scaling factor and zero point parameters, and the model is initially quantized using quantization calibration technology; then, quantization-aware training (QAT) is used to simulate quantization errors during the training process to further reduce the accuracy loss after quantization, and finally an INT8 fixed-point model that meets the deployment requirements is generated, thus significantly improving the operation efficiency while maintaining a high prediction accuracy. At the same time, the accuracy metrics (minimum average displacement error minADE, minimum final displacement error minFDE, etc.) of the fixed-point model and the floating-point model are compared to ensure that the quantized model meets the real-time requirements of the board while ensuring the model accuracy requirements.
[0077] Step 4: Deployment and Performance Testing on the Model Board: Use a data acquisition vehicle to collect actual road vehicle trajectory data to generate trajectory data packets; after completing model quantization, perform in-vehicle board deployment, replay the actual road vehicle trajectory data packets, and respectively test the inference speed and accuracy of the model on the development machine and the target board. Compare the differences in prediction performance (minADE, minFDE) between the target end point scattering algorithm based on clustering and uniform distribution scattering; at the same time, verify the stability and generalization ability of the model under different scenarios, and visually compare the prediction results with the real trajectories through the prediction results to ensure that the model can efficiently and accurately complete the trajectory prediction task in real vehicle deployment.
[0078] As shown in the attached drawings, the detailed processes of each step are as follows:
[0079] Step 1.1: According to the spatial structure and interaction relationships of each traffic participant in the scenario, use a hierarchical graph neural network to encode the traffic participant structured information collected in vectorized form. The specific implementation method first vectorizes the trajectories of traffic participants. Each vector node feature includes start / end coordinates, attribute information (semantic label, timestamp), and corresponding ID. Abstract the trajectories of each traffic participant as a polyline representation, perform a non-linear transformation on a single vector node through an MLP to generate preliminary embedding features. Construct a fully connected subgraph for all nodes within the same polyline, and aggregate adjacent node information through max-pooling. The formula is:
[0080]
[0081] where, is the l-th layer polyline-level feature. Through this process, by stacking multiple subgraph networks, the geometric and semantic associations within the polyline are gradually fused. Each output polyline local feature includes geometric structures such as the trajectory direction and category of the traffic participant, as well as semantic attributes. Then, use the polyline-level feature output by the subgraph module as global graph nodes. The total number of nodes covers all traffic participants to construct a fully connected interaction graph, and there are potential interaction edges between nodes. Use the attention mechanism to capture the interaction information between traffic participants, specifically manifested as competitive relationships such as the preemption intention of adjacent vehicles in the same lane and cooperative behaviors such as deceleration associations caused by pedestrians or adjacent vehicles. Concatenate the low-level features of the subgraph module, such as the original polyline geometric information, with the high-level interaction features of the global graph. The formula is:
[0082]
[0083] where, is the initial output of the subgraph module, This is the global final output. The concatenated features are projected onto a unified dimension d′ through a fully connected layer, and layer normalization (LayerNorm) is applied to stabilize the training process, finally obtaining a two-dimensional feature matrix Each row has H i which corresponds to the comprehensive representation of the i-th entity.
[0084] Step 1.2: Based on historical trajectory data, generate a dense set of candidate target points without the assistance of a high-precision map. The specific implementation steps are as follows:
[0085] 1) First, trajectory coordinate transformation and endpoint extraction are required. According to the prediction task objective of predicting the trajectory in the next five seconds based on the real trajectory in the previous two seconds, the continuous trajectory data is cut into 7-second segments, where the first 2 seconds are the observed history and the next 5 seconds are the prediction targets. Taking the vehicle position at the second second in the trajectory segment as the origin (x0, y0), and the heading angle θ (in radians) at this moment as the positive direction of the coordinate system. For all points (x t , y t ) in the trajectory, translation and rotation transformations are performed to convert them to the ego-vehicle coordinate system:
[0086]
[0087] where (x′, y′) are the transformed coordinates.
[0088] Finally, for each 7-second trajectory segment, the position coordinates (x′7, y′7) at the 7th second are intercepted, such as Figure 1 , to form the original set of target points:
[0089] P = {p1, p2,..., p N}, p i = (x′ 7,i , y′ 7,i )
[0090] 2) First, perform speed stratification and calculate the average driving speed of the vehicle in the previous 2 seconds Taking the data of a certain urban road as an example, speed thresholds of 5 m / s and 20 m / s are used to divide different speed domains of low, medium, and high speeds respectively, and the trajectory endpoints are accordingly assigned to three subsets P low , P mid , P high . Among them, the actual number of endpoints in P low is 10833, the actual number of endpoints in P mid is 8856, and the actual number of endpoints in P high is 9321. For each speed subset P s(s∈{low, mid, high}) respectively perform k-means clustering to generate 2048 candidate target points to meet the model scale requirements. In the initialization stage, first select the set P s The first cluster center μ1 is randomly selected as the initial seed point, and then each point p is calculated i The shortest distance D(p i ), and according to the probability distribution P(p i )∝D(p i ) 2 Select the next center point and repeat this process until 2048 initial centers {μ1,...,μ k Then in the iterative optimization phase, firstly, each trajectory endpoint p is assigned i Assign to the nearest cluster center, and then in the update phase according to the formula Recalculate the centers of each cluster, where C j Represents the point set in the current cluster until Or until a certain number of iterations is reached. The final output result is the candidate target point set G s (s∈{low, mid, high})={μ1,...,μ 2048}, three groups of dense candidate target point locations obtained based on endpoint clustering priors are as follows Figure 2 shown.
[0091] Step 1.3: To further refine the candidate target point information, for the dense candidate target point set generated in step 1.2, the attention mechanism is used to probabilistically encode each candidate target point. First, the 2D coordinates of the target are encoded by MLP to obtain the initial feature matrix where d h is the hidden layer dimension. The local information between target points is obtained through the attention mechanism:
[0092] U=EW u ,R=HW r ,O=HW o
[0093]
[0094] in, is a linear projection matrix, each element B of the matrix B ij represents the attention degree of target point i to j, d k is the dimension of the query / key / value vector, and the enhanced features are obtained by multiplying B with the matrix O.
[0095] The prediction score of the i-th target can be expressed as:
[0096]
[0097] Among them, the trainable function h(·) is implemented using an MLP. The binary cross-entropy loss between the predicted target score φ and the true target score ψ is used:
[0098] L goal = L CrossEntropy (φ, ψ)
[0099] The predicted target score reflects both the coordinate features of the target point itself (through E) and the neighborhood context relationship (through B), forming a probability heatmap for the predicted scores of all candidate target points in step 1.2. The final candidate target point probability heatmap is as Figure 3 .
[0100] Subsequently, non-maximum suppression (NMS) is used to select the predicted end points with multi-modal features based on the obtained probability information of each candidate target point. First, NMS sets a distance threshold (5 meters) to retain clusters of points with probability peaks that do not overlap. Each cluster represents a possible driving intention (such as going straight, turning left). Other higher-probability points in the same area are suppressed to ensure the independence of different mode end points. By this method, k top predicted end points with multi-modal features are finally selected.
[0101] Step 1.4: Finally, trajectory completion is performed based on the obtained predicted target points. From the set of multi-modal target points screened by non-maximum suppression (NMS) in step 1, for each target point p k , extract the enhanced features obtained through the attention mechanism in step 1.3 and the spatial coordinates (x k , y k ). At the same time, combine the historical trajectory, global context features encoded in step 1.1, and the temporal encoding generated using the sine function, and splice them to form a composite feature vector z t , realizing joint spatio-temporal feature modeling. Unfold it over time through a 2-layer MLP decoder:
[0102]
[0103] Among them, σ is the ReLU activation function. MLP1 maps the input to the hidden space, and MLP2 outputs the 2D coordinates, finally obtaining the complete trajectory During training, the teacher forcing technique is adopted, and the true target is input into the decoder. The completion loss is:
[0104]
[0105] Among them, L reg is the smooth Loss is the predicted coordinate of the trajectory at time step k, t k is the true coordinate of the trajectory at time step k.
[0106] Step 2.1: Data preprocessing is a fundamental part of model training. Using the trajectory data that has undergone ego-vehicle coordinate transformation and trajectory slicing in Step 1.2, the data quality is improved through cleaning operations. First, static objects are eliminated, and a displacement threshold δ is defined to determine the static state of an object. Calculate the maximum activity range of each object in the UTM coordinate system:
[0107] R = max(max(x) - min(x), max(y) - min(y))
[0108] where, if R < 0.5 m, mark the object as in a static state and remove it.
[0109] At the same time, errors in in-vehicle perception data need to be processed. For example, data with multiple class attributes for the same track_id is marked as Defect and prohibited from being used as modeling data; process error data with unreasonable displacement per unit time, detect displacement jumps between adjacent frames. Let the position at time t be (x t , y t ), and the position at time t + Δt be (x t+Δt , y t+Δt ). If, within a short time Δt, Δx or Δy changes significantly. For example, when Δt = 0.1 second, if Δx > 5 m or Δy > 5 m, it is determined as a jump; when Δt = 0.04 second, if Δx > 3 m or Δy > 3 m, at this time the speed far exceeds the reasonable range, then it is error data and not used as modeling data.
[0110] Step 2.2: After completing the data cleaning in Step 2.1, to further improve the continuity and stability of the data, the trajectory data is smoothed. This step uses the method of moving average filter, and reduces noise and smooths the trajectory by calculating the mean of adjacent data points. First, set a sliding window, and use a moving average filter with a window size of 3, that is, each trajectory point is determined by itself and its adjacent front and back points. For the point sequence in the trajectory:
[0111] (x1, y1), (x2, y2),...,(x n , y n )
[0112] Smoothing calculations are performed on each non-boundary point. For the i-th point (x i , y i ) in the trajectory, if there are data points before and after it (x i-1 , yi-1 ) and (x i+1 , y i+1 ), then update its coordinates to:
[0113]
[0114] where (x i ′, y i ′) are the smoothed coordinates. For boundary point processing, the first and last points remain unchanged, i.e.:
[0115] x 1′ = x1, y1′ = y1, x n ′ = x n , y n ′ = y n
[0116] The trajectory data (x i ′, y i ′) after mean filtering is used as the finally smoothed data for subsequent model training.
[0117] Step 2.3: Pretrain a baseline model using the public trajectory dataset and improve the generalization ability of the model in the new scenario through transfer learning. The specific approach is to first load the checkpoint file of the baseline model and inherit its parameters. Reduce the initial learning rate to 5e-4. The low learning rate strategy can avoid drastic fluctuations in weights and enable the model to converge stably during the fine-tuning stage.
[0118] The purpose of quantization calibration is to reduce the accuracy loss during the conversion of the model from floating-point (FP32) to fixed-point (INT8) by adjusting the quantization parameters. The specific steps are as follows:
[0119] 1) Quantization parameter initialization
[0120] Quantize the weights and activation values of the model and initialize the quantization parameters (scaling factor s and zero point z):
[0121]
[0122] where x float is the floating-point value and x quant is the quantized integer value.
[0123] 2) Calibration dataset
[0124] Use the calibration dataset to perform inference on the model and statistically analyze the distribution of activation values.
[0125] Optimize the scaling factor s and zero point z by minimizing the quantization error:
[0126]
[0127] Among them, b is the quantization bit width (b = 8 in INT8), and the model in the calibration stage is obtained through quantization calibration.
[0128] Step 3.2: Quantization-aware training improves the accuracy of the calibration-stage model by simulating quantization errors during the training process. The specific steps are as follows:
[0129] 1) Insert pseudo-quantization nodes
[0130] Insert pseudo-quantization nodes into the floating-point model to simulate the quantization process:
[0131]
[0132] Among them, the clip function is used to limit the quantized value within the valid range.
[0133] 2) Gradient coordination training
[0134] Introduce the Gradient Norm (GN) to weight the loss function to improve the modeling ability for tail samples:
[0135]
[0136] where w i is the weight of sample i, and L i is the loss of sample i. The QAT quantization model is obtained through quantization-aware training.
[0137] Step 3.3: Convert the quantized QAT quantization model into a fixed-point model (INT8), generate a deployable computation graph file (.hbm file), and perform testing and accuracy verification.
[0138] Compare the accuracy metrics (such as minADE and minFDE) of the fixed-point model and the floating-point model to ensure that the performance of the quantized model meets the requirements. The specific performance test is shown in Table 1.
[0139] Table 1 Performance comparison at different quantization stages
[0140] Model Quantization Phase minADE minFDE MR Floating-point Model (FP32) 1.865 4.501 0.603 Quantization Calibration 4.974 10.694 0.732 Quantization-Aware Training (QAT) 2.220 5.267 0.630 Fixed-point Model (Int8) 2.255 5.372 0.630
[0141] It can be found that during the quantization-aware training stage, the model performance is significantly restored, as shown in the performance comparison between the quantization calibration (Calibration) stage and the quantization-aware training (QAT) stage in the above table. At the same time, when converting to the fixed-point model, not too much model performance is lost. As shown in the comparison between the quantization-aware training stage and the fixed-point model in the above table, the performance of the finally quantized model compared to the floating-point model has no significant performance loss on the basis of lightweighting. At the same time, the single prediction duration of the fixed-point model on the in-vehicle board side reaches within 100 ms, meeting the real-time requirement.
[0142] Step 4.1: Use a data acquisition vehicle to collect trajectory data on actual roads. Install a multi-modal sensor array on the data acquisition vehicle, including a high-precision lidar, front / rear view cameras, and a GPS / IMU integrated navigation system, and complete the sensor time synchronization calibration. Drive in diverse actual road scenarios (urban roads, highways, intersections, etc.), and use the sensors to record the trajectory data of the vehicle itself and surrounding traffic participants (vehicles, pedestrians, etc.) in real time. The data fields include timestamp, trajectory ID, position in the UTM coordinate system, speed, city name, heading angle, type of traffic participant, and target category label.
[0143] Step 4.2: After the data acquisition process in Step 4.1, use data playback technology to infer the model on the actual road data and calculate the performance evaluation. The specific metrics include the minimum average displacement error (minADE), minimum final displacement error (minFDE), miss rate (MR), and the minADE in the lateral and longitudinal directions. Use the fixed-point model deployed on the vehicle board for testing, and visualize the predicted trajectory results and the true trajectories, such as Figure 5 , and it is found by comparison that the predicted results fit the true trajectories.
[0144] The minimum average displacement error (minADE) is defined as the minimum value of the average displacement error of all time steps between the predicted trajectory and the true trajectory, reflecting the overall fitting ability of the model to the trajectory. Its calculation formula is as follows:
[0145]
[0146] where T is the number of time steps, is the coordinate of the k-th predicted trajectory at time step t, and s t is the coordinate of the true trajectory at time step t.
[0147] The minimum final displacement error (minFDE) is defined as the minimum value of the error between the predicted trajectory and the true trajectory at the final time step, reflecting the accuracy of the model in predicting the end point:
[0148]
[0149] where, is the coordinate of the k-th predicted trajectory at the final time step T, and s T is the coordinate of the true trajectory at the final time step.
[0150] The comparison results of the model performance are shown in Table 2. Here, the advantages of the end-point clustering prior strategy are demonstrated. Compared with the most commonly used uniform distribution of scatter points mode, the experimental results show that the model adopting the end-point clustering prior strategy has improved in multiple performance indicators compared with the uniform distribution of scatter points strategy. Specifically, the minADE has decreased by 3.3% (from 1.928 m to 1.865 m), and the minFDE has decreased by 3.9% (from 4.684 m to 4.501 m). At the same time, the lateral minADE has a significant decrease of 15.6% (from 0.352 m to 0.297 m), and the longitudinal minADE has decreased by 2.1% (from 1.765 m to 1.728 m). In addition, the trajectory mismatch rate (MR) has also decreased from 0.623 to 0.603, with an optimization amplitude of 3.2%.
[0151] Table 2 Comparison of Model Performance Indicators
[0152]
[0153]
[0154] Deploy and test the fixed-point model on the Horizon Journey 5 (J5) board end. The results are as Figure 4 and Figure 5 shown Figure 4 a The red square trajectory points represent the trajectory of the vehicle to be predicted in the first two seconds. In this scenario, other traffic participants have an obvious tendency to turn right, and there is also a possibility of some left turns. Figure 4 b Visualization of the prediction results after trajectory completion with a non-maximum suppression parameter of 150 in this scenario, indicating that the trajectory prediction algorithm based on the end-point clustering prior can accurately judge the multi-modal driving intentions. Figure 4 c, 4d The prediction results of the model in the straight-line scenario. Under the condition of missing map, the algorithm of the present invention also accurately locates the end point of the trajectory and completes the trajectory completion. Figure 5 It is a comparison chart of the predicted trajectory and the real trajectory. The blue triangles mark the movement trajectory of the autonomous vehicle in the first 2 seconds, and the red squares mark the trajectory of the measured target traffic participant in the same time period. For comparative analysis of the prediction effect, the purple dotted line in the figure shows the real trajectory of the target traffic participant in the subsequent 5 seconds, and the best predicted trajectory generated by the algorithm is marked with a black star. The experimental results show that whether in the turning scenario or the straight-line driving scenario, the trajectory prediction results based on the mapless condition are highly consistent with the real trajectory, demonstrating excellent prediction accuracy.
[0155] In summary, aiming at the problems that the existing trajectory prediction models highly rely on high-precision maps and the prediction efficiency is low and cannot meet the application requirements in practical engineering, the present invention designs and implements a mapless trajectory prediction algorithm. In the key method for generating a set of dense candidate target points without relying on maps, the end-point clustering prior strategy adopted by the present invention is superior to the uniform distribution point-scattering strategy, and the minADE decreases by 7.9%. At the same time, quantization calibration and quantization-aware training (QAT) are applied to reduce the inference cost, and the model is deployed on a vehicle-mounted terminal, and the single prediction duration reaches within 100 ms, meeting the real-time requirement.
[0156] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art shall fall within the protection scope determined by the claims.
Claims
1. A mapless long-term vehicle trajectory prediction method based on end-point clustering prior, characterized in that The method constructs a trajectory prediction model, quantizes the trajectory prediction model and then deploys it on a vehicle terminal for vehicle trajectory prediction. The trajectory prediction steps specifically include: Encoding the driving scenario context information using a hierarchical graph neural network to extract the features of the interaction between traffic participants and the surrounding environment; Performing coordinate transformation and endpoint extraction on the historical trajectory data, clustering the endpoints in different speed domains, and constructing a set of dense candidate target points; Performing probability estimation on the dense candidate target points through an attention mechanism, and using the non-maximum suppression algorithm to select multi-modal candidate endpoints; Based on the multi-modal candidate endpoints, using a trajectory completion module to generate multi-modal long-time-domain vehicle prediction trajectories.
2. The method for predicting a long-time vehicle trajectory without a map based on end-point clustering prior according to claim 1, wherein The extraction of the features of the interaction between traffic participants and the surrounding environment is specifically as follows: Vectorize the trajectories of traffic participants. Each vector node feature includes start / end coordinates, gender semantic labels, timestamps, and corresponding IDs; Abstract the trajectories of each traffic participant as a polyline representation, and perform non-linear transformation on a single vector node through an MLP to generate preliminary embedding features; Construct a fully connected subgraph for all nodes within the same polyline, and aggregate adjacent node information through max pooling; Extract the local features of each subgraph, and use an attention mechanism to capture the interaction information between traffic participants to generate two-dimensional features.
3. A mapless long-time vehicle trajectory prediction method based on end-point clustering prior according to claim 1, characterized in that, The construction of the set of dense candidate target points is specifically as follows: Cut the continuous trajectory data into multiple trajectory segments of a set duration. For a trajectory segment, set a division moment. In the trajectory segment, the time before the division moment is the observation history, and the time after the division moment is the prediction target; Taking the vehicle position at the division moment in the trajectory segment as the origin and the heading angle θ at this moment as the positive direction of the coordinate system, perform translation and rotation transformation on all points in the trajectory segment to convert them to the ego-vehicle coordinate system; for each trajectory segment, intercept the position coordinates at the end moment to form the original target point set; Perform speed stratification, calculate the average driving speed of the vehicle trajectory segment before the division moment, divide the trajectory segment into different speed domains based on a set speed threshold, and correspondingly divide the trajectory endpoints into the corresponding speed subsets; perform clustering on each speed subset respectively to generate multiple candidate target points.
4. A mapless long-time vehicle trajectory prediction method based on end-point clustering prior according to claim 3, characterized in that Performing clustering on the speed subset to generate multiple candidate target points is specifically as follows: In the initialization phase, randomly select the first cluster center from the velocity subset as the initial seed point; calculate the shortest distance D(p i ) from each point to the selected center, and select the next center point according to the probability distribution P(p i ) ∝ D(p i ). Repeat this process until the velocity subset is traversed; 2 In the iterative optimization stage, in the assignment stage, assign each trajectory endpoint to the nearest cluster center, and in the update stage, recalculate the cluster centers until the difference between the recalculated cluster center and the original cluster center is less than the set threshold or a certain number of iterations is reached; Output the positions of the dense candidate target points obtained based on the endpoint clustering prior to obtain the set of dense candidate target points.
5. A mapless long-time vehicle trajectory prediction method based on end-point clustering prior according to claim 1, characterized in that The selection of multi-modal candidate endpoints is specifically as follows: For the set of dense candidate target points, use an attention mechanism to perform probability encoding on each dense candidate target point, obtain the initial feature matrix E by encoding the 2D coordinates of the target points through an MLP, and obtain the local information between target points through an attention mechanism: U = EW u , R = HW r , O = HW o Among them, is a linear projection matrix, H is the feature of the extracted traffic participant interacting with the surrounding environment, and each element B of matrix B ij represents the attention degree of target point i to j, and d k is the dimension of the query / key / value vector; Multiply matrix B by matrix O to obtain the enhanced feature E'; Calculate the prediction target scores of each dense candidate target point through an attention mechanism: Among them, the trainable function h(·) is implemented using an MLP Form a probability heat map based on the predicted scores of all dense candidate target points; Use non-maximum suppression to select the predicted end points with multi-modal characteristics according to the probability information of each candidate target point. Set a distance threshold to retain the point clusters with probability peaks and no overlapping. Each cluster represents a possible driving intention, where other higher-probability points in the same area are suppressed. Select the top k dense candidate target points that can cover multiple driving intentions and have multi-modal characteristics as the multi-modal candidate end points.
6. A mapless long-time vehicle trajectory prediction method based on end-point clustering prior according to claim 5, characterized in that The trajectory completion module is specifically implemented as follows: For each multimodal candidate endpoint p k , the enhanced feature E' obtained through the attention mechanism and the spatial coordinates are extracted, and are concatenated with the features of the traffic participant's interaction with the surrounding environment to form a composite feature vector z t ; Unfold the composite feature vector over time through a 2-layer MLP decoder: Among them, σ is the ReLU activation function. MLP1 maps the input to the hidden space, and MLP2 outputs 2D coordinates, finally obtaining the complete trajectory.
7. A mapless long-time vehicle trajectory prediction method based on end-point clustering prior according to claim 1, characterized in that The vehicle trajectory prediction method uses a transfer learning strategy to train the trajectory prediction model: Pre-train a baseline model using a public dataset; use the parameters of the pre-trained model as the initial values, reduce the learning rate, and perform fine-tuning training on the target dataset; Use teacher forcing technology to supervise the training of the trajectory completion module and optimize the model parameters.
8. A mapless long-time vehicle trajectory prediction method based on end-point clustering prior according to claim 7, characterized in that The target dataset uses private trajectory data, and the private trajectory data is cleaned and trajectory smoothed to construct the target dataset.
9. A mapless long-time vehicle trajectory prediction method based on end-point clustering prior according to claim 1, wherein For the trained trajectory prediction floating-point model, the vehicle trajectory prediction method converts it into a fixed-point model through quantization processing and deploys it on the in-vehicle board. Specifically as follows: Perform quantization calibration on the trajectory prediction model, and adjust the quantization parameters to obtain the calibration-phase model: Quantize the weights and activation values of the model and initialize the quantization parameters: where x float is a floating-point value, x quant is a quantized integer value, s is a scaling factor, and z is a zero point; Use the calibration dataset to perform inference on the model, statistically analyze the distribution of activation values, and optimize the scaling factor s and zero point z by minimizing the quantization error: where b is the quantization bit number; Perform quantization-aware training on the calibration-phase model, and improve the accuracy of the calibration-phase model by simulating the quantization error: Insert pseudo-quantization nodes in the floating-point model to simulate the quantization process: where the clip function is used to limit the quantized values within the valid range; Introduce the gradient norm to weight the loss function to improve the modeling ability for tail samples, and perform gradient coordination training on the model with inserted pseudo-quantization nodes to obtain the QAT quantization model; Convert the quantized QAT quantization model into a fixed-point model and generate a computation graph file.
10. A mapless long-time vehicle trajectory prediction method based on end-point clustering prior according to claim 1, characterized in that, After the trajectory prediction model is deployed on the in-vehicle terminal, test the model performance. Specifically as follows: Use a data collection vehicle to collect trajectory data on actual roads. The collected trajectory data includes: timestamp, trajectory ID, position in the UTM coordinate system, speed, heading angle, traffic participant type, and target category label; Use the data playback technology to perform inference on the trajectory prediction model on the actual road data and calculate the performance evaluation. Specifically include: Calculate the minimum value of the average displacement error at all time steps between the predicted trajectory and the true trajectory to evaluate the overall fitting ability of the trajectory prediction model for the trajectory; Calculate the minimum value of the error between the predicted trajectory and the true trajectory at the final time step to evaluate the accuracy of the trajectory prediction model in end point prediction.
Citation Information
Cited By
Method, application and equipment for generating anthropomorphic foot end track of biped robot based on full-connection network
CN120993947A