Multi-mode pedestrian street-crossing long-time trajectory prediction method based on space-time coupling and anchor point driving
By adopting a multimodal method based on space-time coupling and anchor driving in pedestrian trajectory prediction, the problem of pedestrian crossing path prediction in signal-free control environment is solved, and trajectory prediction with higher accuracy and robustness is achieved, and the safety of autonomous vehicles is enhanced.
Patent Information
- Application Number
- CN202510308226.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-27
AI Technical Summary
The existing pedestrian trajectory prediction methods are difficult to accurately predict pedestrian crossing paths in a signal-free control environment, especially when pedestrians are free to behave, fuzzy right of road and complex human-vehicle interactions.
A multimodal pedestrian crossing long-term trajectory prediction method based on space-time coupling and anchor driving is adopted. By combining social and time encoders to integrate environmental information, a social interaction model of people-vehicle-road is constructed, and a self-attention mechanism and cross-attention mechanism are used to achieve interaction within the trajectory sequence and between the trajectory and the environment is proposed. A trajectory correction method based on anchor points and a multimodal trajectory output and loss function design are proposed.
It improves the accuracy and robustness of pedestrian trajectory prediction, can more accurately capture pedestrian behavior characteristics and individual differences, and enhances the ability of autonomous vehicles to predict pedestrian intentions.
Smart Images

Figure CN120220026A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving safety, and particularly relates to a multi-modal pedestrian cross-street long-term trajectory prediction method based on spatio-temporal coupling and anchor point driving. Background Art
[0002] Pedestrian trajectory prediction refers to the process of predicting the movement path and position change of a pedestrian within a certain period in the future by analyzing the current and historical states of the pedestrian and combining environmental information, using methods such as physical models, machine learning, or deep learning. The issue of pedestrian safety is an important part of the safety issues of autonomous vehicles. Whether an autonomous vehicle can reasonably anticipate the intentions of pedestrians and perform trajectory prediction on them is crucial for the safe and stable driving of intelligent vehicles.
[0003] Using the physics-based method based on physical models in pedestrian trajectory prediction requires relying on a large amount of prior knowledge and parameter adjustment, making it difficult to capture pedestrian behavior characteristics and individual differences, and being limited by hand-designed functions with limited prediction ability.
[0004] Due to the limitations of the physics-based trajectory prediction method, deep learning methods have been widely used in recent years. These models comprehensively consider pedestrian trajectory characteristics, behavioral characteristics, scene semantics, and actor interaction characteristics using deep learning techniques. They can consider interaction factors and adapt to more complex scenarios. Summary of the Invention
[0005] When there is no signal control for pedestrians to cross the street, pedestrians have free behavior, ambiguous road rights, and complex human-vehicle interactions, resulting in a significant increase in the difficulty of their trajectory prediction. To fully analyze the factors affecting the path selection of pedestrians crossing the street and improve the accuracy of trajectory prediction, the specific steps of the present invention are as follows:
[0006] S1: A clear video stream will be captured; the video stream will be frame-dropped using the high-performance video codec of the NVIDIA GPU, and frames will be extracted from the frame-dropped video to reduce the data volume;
[0007] S2: Build a joint social and temporal encoder, fuse environmental information and flatten the trajectory data of all pedestrians and vehicles, break the traditional phased modeling method, and achieve synchronous in-depth mining of temporal features and social interaction features;
[0008] S3: Build a social interaction model of the person-vehicle-road triad, realize the interaction between traffic participants within the trajectory sequence through the self-attention mechanism, and at the same time use the cross-attention mechanism to realize the interaction between the trajectory and the environment;
[0009] S4: Propose an anchor-based trajectory correction method, divide the trajectory prediction into two stages for processing, and realize the correction of the pedestrian trajectory through an iterative method;
[0010] S5: Propose a method for multi-modal trajectory output and loss function design, use the Laplace mixture distribution to simulate the various possibilities of pedestrians crossing the street, and use the trajectory proposal loss L of end-to-end training propose , trajectory improvement loss L refine and classification loss L cls to accelerate model convergence.
[0011] Furthermore, in the step S1, it includes:
[0012] S101: Use an in-vehicle high-definition IPC camera to capture a high-definition video stream of the area in front of the vehicle during driving.
[0013] S102: To reduce the data volume and improve the processing efficiency, use the high-performance video codec of NVIDIA GPU to perform frame extraction on the video stream. First, ensure that the system has installed an NVIDIA graphics card and driver supporting CUDA; second, obtain and compile the FFmpeg source code, configure parameters to enable CUDA support; finally, use the NVIDIA hardware codec to complete video processing.
[0014] Furthermore, in the step S2, it includes:
[0015] S201: Discretize the trajectory features of traffic participants. Represent the historical trajectory time series Y of traffic participants. Based on the respective trajectory time series of each traffic participant, the change in the speed of traffic participants over time can be obtained by calculating the Euclidean distance difference between adjacent frames through the following formula.
[0016]
[0017] where represents the abscissa of traffic participant i at time T, represents the ordinate of traffic participant i at time T. -Obs to 0 represents the time period for observing the historical trajectories of traffic participants.
[0018] By calculating the angle θ between the trajectory points of adjacent frames and the positive direction of the X-axis, the time series of the crossing heading angle of traffic participants can be obtained:
[0019]
[0020] By calculating the speed difference between adjacent frames, the time series of the acceleration characteristics of traffic participants can be obtained:
[0021]
[0022] Finally, all the trajectory features of traffic participant i at time T can be obtained:
[0023]
[0024] S202. Road environment information extraction. Use the ResNet-34 network to perform pre-training on the ImageNet large-scale dataset. By means of transfer learning, input the image into the network model, and obtain the road environment features using the following formula:
[0025]
[0026]
[0027] where I t represents the map I at time t, CNN represents the ResNet-34 network model, flatten represents the flattening operation, FC represents the fully connected layer, and finally, the high-dimensional feature input vector of the road environment can be obtained through the above steps.
[0028] S203. Input into the feature vector embedding layer to obtain the input information of the trajectory prediction model:
[0029]
[0030] where W1 represents the weight parameter, b1 represents the bias, is the model input that fuses all interaction information of traffic participant i at time T.
[0031] Further, in step S3, it includes:
[0032] S301. Interaction between traffic participants. Introduce the multi-head attention mechanism to process the input sequence, learn the attention information of different aspects of the sequence, and integrate the information so as to more comprehensively capture the semantic relationships or feature associations in the input sequence, simulating the process in which traffic participants affect each other's future trajectories:
[0033]
[0034] attention = Softmax(Score)
[0035] where Q represents the query vector, K T represents the transpose of the key vector, d k represents the dimension of the key vector, and Softmax represents normalization.
[0036] S302. Interaction between traffic participants and the road environment. By introducing the cross-attention mechanism, interact the feature vector of traffic participants and the feature vector of the road environment, and the calculation formula is as follows:
[0037] Q = X1W Q
[0038] K = V = X2W K
[0039]
[0040] where X1 and X2 respectively represent the feature vectors of traffic participants and the feature vector of the road environment W Q and W K respectively represent learnable parameter matrices
[0041] Further, in step S4, it includes
[0042] S401. Randomly initialize the trajectory path. First, randomly initialize an anchor-free trajectory queries vector as the future trajectory end point of the target traffic participant to be predicted
[0043] S402. Interaction between traffic participants and the road environment. Re-weight and focus on the internal features of the trajectory by using the multi-head attention mechanism for CrossAttention(X1, X2) obtained in step S302 and the anchor-free trajectory queries in step S401, and assign weights by calculating the correlation degree between features to extract more representative and discriminative trajectory features
[0044] S403. Long-term trajectory prediction. Predicting the pedestrian crossing trajectory for a relatively long time only based on the observed time steps will have serious prediction errors. Therefore, the present invention converts the long-term prediction into several short-term predictions, and the short-term prediction calculation formula is as follows
[0045]
[0046] where f is the prediction model, θ is the model parameter is the information including all traffic participants and the environment in the current scene is the historical trajectory of the target pedestrian to be predicted
[0047] After each short-term prediction, update the input data to reflect the latest input and scene information
[0048]
[0049] where represents the result of the current short-term prediction represents the result of the next short-term prediction
[0050] Final long-term prediction trajectory It is composed of splicing all short-term prediction results:
[0051]
[0052] Finally, the trajectory direction anchor point (end point) of the pedestrian in the future can be obtained.
[0053] S404. Correction of the pedestrian trajectory based on the anchor point. The correction process can be abstractly summarized by the formula as follows:
[0054] Y = F(X)
[0055]
[0056] Among them, F can be regarded as a neural network, X is the input, and Y obtained is the output.
[0057] However, previous research ended here. On this basis, the present invention further calculates the deviation amount Δy of the trajectory and uses the corrected trajectory as the final output. Using the predicted pedestrian future trajectory anchor point as the trajectory end point, the start point and end point of the known trajectory, and the parameters for narrowing the interaction range to exclude irrelevant interferences.
[0058] Furthermore, in the step S5, it includes:
[0059] S501. The multi-modal trajectory of the i-th predicted pedestrian is obtained by using a mixture of Laplace distributions:
[0060]
[0061] Among them, is the mixing coefficient, and the Laplace density of the k-th mixing component at the time step T is determined by the position and the scale
[0062] S502. Then use the classification loss L cls to optimize the mixing coefficient predicted by the improvement module. This loss minimizes the negative log-likelihood of, and the present invention stops the gradients of the position and scale to optimize the mixing coefficient. Only the best proposed proposal and its refinement are backpropagated. For stability, the refinement module stops the gradients of the proposed trajectory anchor points. The final loss function combines the trajectory proposal loss L propose for end-to-end training, the trajectory improvement loss L refine and the classification loss L cls :
[0063] L = L propose + L refine+L cls
[0064] Among them, the present invention uses λ to balance regression and classification.
[0065] Beneficial benefits of the present invention:
[0066] 1. Use NVIDIA GPU high-performance video codec to process the collected high-definition video stream, which can improve the speed and efficiency of video processing.
[0067] 2. In pedestrian trajectory prediction, space-time interaction is crucial. Spatially, building layout, street characteristics and traffic dynamics affect the starting point, route and speed of pedestrians; temporally, historical motion characteristics (such as speed changes, pause patterns) have an impact on current decisions. Existing models usually process space-time characteristics separately, resulting in information isolation, failure to capture dynamic correlations, and reduced prediction accuracy and robustness. To this end, the present invention breaks the traditional staged modeling approach by fusing environmental information and integrating the trajectory data of pedestrians and vehicles, and realizes the simultaneous in-depth mining of temporal characteristics and social interaction characteristics, thereby more comprehensively capturing the key elements in trajectory prediction.
[0068] 3. Long-term trajectory prediction is one of the difficulties in pedestrian trajectory prediction. Long-term prediction refers to the future trajectory of pedestrians greater than or equal to 5 seconds. Although long-term prediction can provide more decision-making time for autonomous driving, due to the high randomness and uncertainty of pedestrian behavior, the difficulty of prediction increases exponentially with the increase of time. The present invention decomposes the long-term prediction into multiple short-term predictions and dynamically updates the model input to capture scene changes.
[0069] 4. Common multimodal methods usually rely on Gaussian mixture models or implicit spatial models to generate a series of trajectories, but this method may lead to modal collapse problems and ignore less likely but equally important pedestrian motion patterns. Based on the consideration of spatiotemporal characteristics, this paper proposes a two-stage trajectory prediction method based on anchor points to eliminate irrelevant interference factors and improve the accuracy of trajectory prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0071] Figure 1 It is a flow chart of a multi-modal pedestrian crossing long-term trajectory prediction method based on spatiotemporal coupling and anchor point driving of the present invention;
[0072] Figure 2 It is the discrete sampling of trajectory features proposed by the present invention;
[0073] Figure 3It is a modeling method for integrating spatio-temporal information proposed by the present invention;
[0074] Figure 4 It is the multi-modal trajectory prediction and long-term prediction framework of the present invention. Specific implementation manner
[0075] The following combines the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0076] As Figure 1 shown: An example of the present invention proposes a multi-modal pedestrian cross-street long-term trajectory prediction method based on spatio-temporal coupling and anchor point driving, and the specific implementation manner is as follows:
[0077] Step S1: Obtain the trajectory data of pedestrians and vehicles
[0078] (1) During the process of image data acquisition, first use an in-vehicle high-definition IPC camera to capture a high-resolution video stream to ensure that the details in the image are clear enough for subsequent processing and analysis. The video clarity that the camera can capture is 1080P, that is, a high-definition video of 1920×1080 pixels. This resolution can provide a clear picture, enabling accurate capture of details such as vehicles, pedestrians, and road surfaces.
[0079] (2) The collected video is 30 frames per second. Since the speed of pedestrians crossing the street is generally slow and the movement range between each frame is very small, the detection results are taken every 10 frames. Sort the collected results according to the ID, and filter out the data with a pedestrian trajectory sequence of 12 frames or more.
[0080] (3) Divide these data into two groups. One group uses the observed 6-frame (2-second) trajectory to predict the future 6-frame (2-second) trajectory as short-term prediction data, and the other group uses the observed 6-frame (2-second) trajectory to predict the future 15-frame (5-second) trajectory as long-term prediction data. Finally, the number of different pedestrian continuous trajectories greater than 12 frames is 2323, and the number of different pedestrian continuous trajectories greater than 21 frames is 1864.
[0081] (4) Divide the number of trajectories into a training set, a test set, and a validation set in the ratio of 7:2:1.
[0082] Step S2: Set the model parameters.
[0083] (1) In the trajectory prediction model, the hyperparameter settings of the model are as follows: the dimensions of all features are set to 256. Both the encoder and the decoder adopt a Transformer structure with N = 6 layers. At the same time, 8 multi-head attention mechanisms are used in the attention mechanism to enhance the model's feature extraction and expression capabilities.
[0084] (2) For the evaluation metrics Top-k ADE and Top-k FDE, the value of K is usually selected as 5 to measure the average displacement error and the final displacement error of the model under different top results, and to more comprehensively evaluate the prediction performance of the model. The batch size for model training is set to 32, the initial learning rate is 0.0001, and truncation is performed when the learning rate decays to 0.000001 to prevent the training from stagnating due to too small a learning rate.
[0085] (3) In addition, if the performance metrics on the validation set do not improve after 5 consecutive training epochs, the training process is stopped to avoid overtraining and resource waste. Under this condition, the total number of training times of the model is set to 100 times. To prevent the model from overfitting, a random inactivation mechanism is introduced at a specific stage of the model, and the random inactivation rate is set to 0.1.
[0086] Step S3: Trajectory feature extraction
[0087] Using the object detection and tracking model, extract all pedestrian, vehicle, and environmental information in the image and video. The continuous trajectories of the obtained pedestrians and vehicles are passed through the following formula:
[0088]
[0089] Discretize the continuous trajectory information of pedestrians and vehicles into two-dimensional coordinates to obtain the motion states of pedestrians and vehicles.
[0090] Step S4: Environmental information extraction
[0091] Send the image at time t into the convolutional neural network ResNet-34 network, which is the initial module for preprocessing the input image, including a 7×7 convolutional layer, followed by batch normalization and ReLU activation, and then downsampling through a 3×3 max pooling layer; next are four residual layers, each of which consists of multiple BasicBlock residual blocks, the number of output channels increases sequentially, and the first block of each residual layer realizes downsampling through a convolutional layer with a stride of 2. Each residual block contains two 3×3 convolutional layers, batch normalization, and ReLU activation, and at the same time, skip connections are used to achieve residual learning; finally, there is a global average pooling layer, which compresses the feature map into a 1×1 vector, and then outputs the final classification result through a fully connected layer.
[0092] Step S5: Encoder for Joint Spatiotemporal Interaction
[0093] (1) Input Feature Mapping
[0094] First, map the motion features of pedestrians and vehicles and environmental features into high-dimensional vectors respectively. Here, T represents the number of time steps, and d m and d e represent the dimensions of motion features and environmental features respectively. Through linear transformation, map the motion features and environmental features to the same dimension d:
[0095] H m = X m W m + b m
[0096] H e = X e W e + b e
[0097] where and are learnable weight matrices, and b m and b e are bias terms.
[0098] (2) Feature Concatenation
[0099] Concatenate the mapped environmental feature H e behind the motion feature H m to obtain the joint feature
[0100] H = concat(H m , H e )
[0101] (3) Temporal Encoding
[0102] To retain the temporal information in the sequence, this paper uses a temporal encoder similar to the positional encoding used in the original Transformer. Instead of encoding the position based on the index of each element in the sequence, the timestamp is calculated based on the time step t of the element. The timestamp uses the same sine wave design as the positional encoding. Taking the pedestrian trajectory sequence f as an example, for each element define its timestamp as as shown in Equation (4.19):
[0103]
[0104] where \(t\in(-Obs,-Obs + 1,\cdots,0)\) represents the time step, \(M\) represents a constant, \(k\) represents the index of the feature dimension, and \(d\) τ is the dimension of the timestamp. By concatenating the timestamp corresponding to each element with itself, the trajectory feature with time information can be obtained. After passing through the time encoder, the output of the model can be expressed as follows:
[0105]
[0106] where represents the learnable weight parameter, similar to \(W1\), and represents the concatenation operation.
[0107] Step S6 Social Interaction of the Target Pedestrian
[0108] (1) For the input \(H\), first calculate the query (Query), key (Key), and value (Value) matrices:
[0109] \(Q = HW\) Q , \(K = HW\) K , \(V = HW\) V
[0110] where is the learnable weight matrix, and \(d\) k is the dimension of each attention head. Then, calculate the attention scores:
[0111]
[0112] The multi - head self - attention mechanism concatenates the outputs of multiple attention heads and obtains the final output through a linear transformation:
[0113] \(MHSA(H)=concat(head1,head2,\cdots,head\) h )W O
[0114] where \(head\) h is the number of attention heads, and is the output weight matrix.
[0115] (2) The output of the multi - head self - attention mechanism passes through a feed - forward neural network:
[0116] \(FFN(x)=max(0,xW1 + b1)W2 + b2\)
[0117] where and are the learnable weight matrices, and \(d\) ff is the hidden layer dimension of the feed - forward neural network.
[0118] Step S7 Anchor-based Multimodal Trajectory Prediction
[0119] (1) Long-term Trajectory Prediction
[0120] Assume predicting the future trajectory of a pedestrian within T pre time steps, and using T rec recurrent steps during the decoding process, that is, the time for each prediction is If only predicting the pedestrian's crossing trajectory for a relatively long time based on the observed time steps, there will be serious prediction errors. Therefore, this paper converts the long-term prediction into several short-term predictions, dynamically updates the input of the model, and updates the changes in the scene information at all times, so as to achieve accurate prediction.
[0121] (2) Anchor-free Pedestrian Trajectory Prediction
[0122] After all short-term predictions are completed, the model generates specific path anchors through the interaction between traffic participants and environmental information (Tra2Scene CrossAttn) and the internal interaction between traffic participants (Tra2Tra SelfAttn), and decodes through a multi-layer perceptron (MLP). MLP is a feedforward artificial neural network composed of multiple layers (usually two or more fully connected layers). It mainly includes a BN layer, two 1*1 convolutions, and activation functions.
[0123] (3) Anchor-based Pedestrian Trajectory Correction
[0124] In the first stage of trajectory prediction, the model has predicted different future crossing paths of the pedestrian. The predicted crossing paths are used to embed the trajectory features of Y in the first step through a gated recurrent unit (GRU), and the embedded features are used as the query vectors of the Transformer. The operations of Tra2Scene CrossAttn and Tra2Tra SelfAttn are repeated, but the interaction range is reduced during the interaction process, so as to achieve anchor-based pedestrian trajectory correction.
[0125] Step S8: Multimodal Trajectory Output and Loss Function Design
[0126] Parameterize the future trajectory of the i-th predicted pedestrian as a mixture of Laplace distributions:
[0127]
[0128] where, is the mixing coefficient, and the Laplace density of the k-th mixture component at time step T is given by the location and scale
[0129] Then use the classification loss Lcls to optimize the mixing coefficients predicted by the improvement module. This loss minimizes the negative log-likelihood, and the present invention stops the gradients of the location and scale to optimize the mixing coefficients. Backpropagation is only performed on the best predicted proposals and their refinements. For stability, the refinement module stops the gradients of the proposed trajectory anchors. The final loss function combines the trajectory proposal loss L for end-to-end training, the trajectory improvement loss L propose , and the classification loss L refine : cls :
[0130] L = L propose + L refine + L cls
[0131] wherein, the present invention uses λ to balance regression and classification.
Claims
1. A multimodal pedestrian long-term trajectory prediction method based on spatiotemporal coupling and anchor point driving, characterized in that: The following steps are involved: S1: Captures a clear video stream; uses the NVIDIA GPU's high-performance video codec to downscale the video stream and extract frames from the downscaled video to reduce the amount of data; S2: Build a joint social and temporal encoder to integrate environmental information and flatten the trajectory data of all pedestrians and vehicles, breaking the traditional stage-by-stage modeling approach and achieving simultaneous deep mining of temporal features and social interaction features; S3: Construct a social interaction model among people, vehicles and roads, realize the interaction between traffic participants in the trajectory sequence through the self-attention mechanism, and realize the interaction between the trajectory and the environment through the cross-attention mechanism; S4: A trajectory correction method based on spatiotemporal coupling and anchor points is proposed. The trajectory prediction is divided into two stages and the pedestrian trajectory is corrected in an iterative manner. S5: Propose a method for multimodal trajectory output and loss function design, using Laplace mixture distribution to simulate the possibility of pedestrians crossing the street in various ways and using the trajectory suggestion loss L trained end-to-end propose , trajectory improvement loss L refine and classification loss L cls Accelerate model convergence.
2. The multimodal pedestrian crossing long-term trajectory prediction method based on spatiotemporal coupling and anchor point driving according to claim 1 is characterized in that: In the step S1: using the high-performance video codec of the NVIDIA GPU to perform frame reduction and frame extraction processing on the video stream can reduce the overall data volume and improve the video processing speed and efficiency.
3. The multimodal pedestrian crossing long-term trajectory prediction method based on spatiotemporal coupling and anchor point driving according to claim 1 is characterized in that: In step S2: the historical trajectory time series of the traffic participants is processed by discretization to calculate the speed, acceleration, heading angle, etc. Features are combined with road environment information to build a joint social and time encoder to achieve synchronous deep mining of time features and social interaction features. The calculation formula is as follows: in, represents the horizontal coordinate of traffic participant i at time T, Represents the ordinate of traffic participant i at time T. -Obs to 0 represents the time period for observing the historical trajectories of traffic participants. represents the heading angle of traffic participant i at time T, represents the acceleration of traffic participant i at time T.
4. The multimodal pedestrian crossing long-term trajectory prediction method based on spatiotemporal coupling and anchor point driving according to claim 1, characterized in that: In step S3: the interaction between traffic participants is realized through the self-attention mechanism, and the interaction between traffic participants and the road environment is realized through the cross-attention mechanism, so as to simulate the influence of traffic participants on future trajectories.
5. The multimodal pedestrian crossing long-term trajectory prediction method based on spatiotemporal coupling and anchor point driving according to claim 1, characterized in that: In step S4, a trajectory correction method based on anchor points is proposed, and trajectory prediction is divided into two stages. The pedestrian trajectory is corrected in an iterative manner, the interaction range is reduced, irrelevant interference is eliminated, and the accuracy of trajectory prediction is improved. The process can be abstracted as follows: Y=F(X) Among them, F can be regarded as a neural network, X is the input, and Y is the output.
6. The multimodal pedestrian crossing long-term trajectory prediction method based on spatiotemporal coupling and anchor point driving according to claim 1 is characterized in that: In step S4: the long-term trajectory prediction is decomposed into multiple short-term predictions through the trajectory correction method based on anchor points, and the model input is dynamically updated to capture scene changes. The formula is as follows: Where f is the prediction model, θ is the model parameter, It includes information about all traffic participants and the environment in the current scene. The goal is to predict the historical trajectory of the target pedestrian. Represents the result of the current short-term forecast, Represents the result of the next short-term forecast.
7. The multimodal pedestrian crossing long-term trajectory prediction method based on spatiotemporal coupling and anchor point driving according to claim 1 is characterized in that: In step S5: using Laplace mixture distribution to simulate the possibility of pedestrians crossing the street and using the trajectory suggestion loss L of end-to-end training propose , trajectory improvement loss L refine and classification loss L cls Accelerate model convergence. L=L propose +L refine +L cls in, is the mixing coefficient, and the Laplace density of the k-th mixing component at time step T is given by the position and scale
Citation Information
Cited By
Multi-modal dynamic trajectory prediction method and device based on space-time invariance and medium
CN120808314A