Pedestrian trajectory prediction method based on progressive transfer learning

By dividing the pedestrian trajectory prediction task into three stages, using the progressive transfer learning architecture and the TALSTM module, the problem of distinguishing between short-term and long-term prediction is solved, and high-precision prediction of pedestrian trajectory is achieved.

CN120429564APending Publication Date: 2025-08-05HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510356255.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing pedestrian trajectory prediction methods fail to effectively distinguish between short-term prediction and long-term prediction, resulting in insufficient prediction accuracy, especially when capturing the interaction between pedestrians.

Method used

The progressive transfer learning architecture is adopted to divide the trajectory prediction task into three stages. Each stage uses transfer learning technology to transmit information, capture the characteristics of different stages through the TALSTM module and the fit encoder module, process short-term and long-term prediction tasks respectively, and fit the complete trajectory in the final stage.

Benefits of technology

The accuracy and stability of pedestrian trajectory prediction are improved, and accurate future trajectory prediction is achieved through phased processing, local spatial interaction relationships are captured and the characteristics of transfer learning are used.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429564A_ABST
    Figure CN120429564A_ABST
Patent Text Reader

Abstract

The invention discloses a pedestrian trajectory prediction method based on progressive transfer learning, and the method comprises the steps: segmenting an obtained pedestrian trajectory prediction data set, and carrying out the preprocessing; a progressive transfer learning network is constructed, three stages are included, and each stage comprises a TALSTM module and a fitting encoder module; the TALSTM module performs coding prediction on a time sequence and outputs a prediction result; cross-time-sequence interaction of the encoder module is fitted, and the capability of capturing the interaction relation of the module is enhanced; in each stage, information is transmitted to the next stage by utilizing a transfer learning technology; in the first stage, track information at the next moment is predicted by using track information at the previous moment; in the second stage, the long-term dependency relationship is learned on the basis of the short-term prediction capability, and a track end point is predicted; and a third stage, fitting a complete trajectory by using the predicted trajectory end point. The invention aims to utilize different distribution laws of trails captured in different stages to improve the accuracy of trajectory prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of trajectory prediction, and in particular relates to a pedestrian trajectory prediction method based on progressive transfer learning. Background Art

[0002] With the rapid development of society and the booming development of digitalization and intelligence, artificial intelligence has swept across various fields, reshaping every aspect of people's lives and work. The bustling crowds on city streets, as one of the most dynamic elements in the urban dynamic system, have become a frontier hotspot that crosses multiple disciplines. Pedestrian trajectory prediction is a key link in this exploration.

[0003] Currently, many researchers are using deep learning methods to study pedestrian trajectory prediction. For example, Alahi et al. proposed the Social-LSTM, which introduced a "social" pooling layer to address the interaction problem between pedestrians. Subsequent work has extended the Social-LSTM's visualization capabilities and introduced new pooling mechanisms to improve prediction accuracy. With the emergence of attention mechanisms, some researchers have introduced attention to trajectory prediction as a primary means of capturing interactive relationships. Vemula et al. proposed a social attention mechanism that captures the relative importance of other pedestrians in the scene to the current pedestrian's navigation. Velickovi et al. proposed a graph attention network (GAT), in which stacked nodes can focus on layers of features in their neighborhood and assign different weights to different nodes in the neighborhood. However, these methods fail to address the difference between short-term and long-term prediction. Short-term prediction requires predicting local changes from fine-grained trajectory changes, while long-term prediction requires capturing long-term dependencies in trajectories to infer future motion trends. To address these overlooked issues, this paper proposes a progressive transfer learning architecture that aims to capture long-term dependencies by learning from short-term dynamic changes. Summary of the Invention

[0004] Purpose of the invention: The present invention provides a pedestrian trajectory prediction method based on progressive transfer learning. The method divides the prediction task into three stages and uses transfer learning to associate the three stages, thereby achieving a complete prediction task and ultimately generating an accurate and stable future trajectory.

[0005] The technical solution of the present invention is: a pedestrian trajectory prediction method based on progressive transfer learning described in the present invention comprises the following steps:

[0006] (1) Segment and preprocess the obtained pedestrian trajectory prediction dataset;

[0007] (2) Constructing a progressive transfer learning network, which is divided into three stages. Each stage uses transfer learning technology to pass information to the next stage. In the first stage, the trajectory information of the previous moment is used to predict the trajectory information of the next moment. In the second stage, long-term dependencies are learned based on the short-term prediction ability, and the trajectory endpoint is predicted. In the third stage, the predicted trajectory endpoint is used to fit the complete trajectory.

[0008] (3) The preprocessed training data is fed into the progressive transfer learning network for training, and the progressive transfer learning network is optimized by backpropagation updating the weights using the AdamW optimizer;

[0009] (4) Load the test set and send the test data to the trained progressive transfer learning network for testing. During the test phase, load the prediction results and true values, and use matplotlib to visualize the trajectory prediction to intuitively show the model prediction effect.

[0010] Furthermore, the preprocessing in step (1) is as follows: setting the number of trajectories to be processed each time, processing data in batches, and constructing an interaction matrix for each batch of data according to the pedestrian interaction radius, and finally saving the training data.

[0011] Furthermore, the three stages of the progressive transfer learning network in step (2) all include a TALSTM module and a fitting encoder module; the TALSTM module encodes and predicts the time series and outputs the prediction results; the fitting encoder module interacts across time series, enhancing the module's ability to capture interactive relationships.

[0012] Furthermore, the TALSTM module controls the fusion of historical information and input information through the ITB module and controls the output at the current moment through the gating mechanism. The specific calculation formula is as follows:

[0013]

[0014] F t =σ(ITB(concat(x t ,H t-1 ),[Mask])) (2)

[0015]

[0016] H t =F t tanh(C t ) (4)

[0017] Among them, θ, φ, All are linear layers, X q 、X k 、Xv All are input information x at the current moment t Compared with the output H at the previous moment t-1 The features after splicing, normalize is the normalization operation, concat represents the splicing operation, σ is the sigmoid operation, tanh is the activation function, mask represents the interaction matrix, F t The gate control unit is used to control the fusion of input information and output information at the previous moment, H t and C t They are the output and storage information at the current moment respectively.

[0018] Furthermore, the fitting encoder utilizes the attention mechanism to interact with TALSTM in the first and second stages to predict the output trajectory features; in the third stage, the training data and the endpoints generated in the second stage are input, and the complete trajectory is obtained by transfer learning and the fitting encoder.

[0019] Furthermore, the first stage implementation process of step (2) is as follows:

[0020] By encoding the trajectory information of the previous moment, the trajectory information of the next moment is predicted through the TALSTM module and the fitting encoder module, that is:

[0021] (x t+1 ,y t+1 )=F(TALSTM(x t ,y t )) (5)

[0022] Among them, TALSTM(·) is the TALSTM module, F(·) is the fitting encoder, (x t ,y t ) represents the trajectory coordinate characteristics at time t, (x t+1 ,y t+1 (represents the trajectory coordinate characteristics at time t+1;

[0023] The goal of the first stage is to help the network understand the changing patterns of short-term space and learn fine-grained motion features; its loss function is:

[0024]

[0025] in, is the predicted value of the i-th trajectory, y i is the i-th true trajectory.

[0026] Furthermore, the second stage implementation process of step (2) is as follows:

[0027] Using transfer learning technology, the TALSTM module and the fitting encoder module are used to achieve long-term prediction tasks based on the short-term prediction model training, that is, to predict the endpoint set

[0028] T D =split(TALSTM(X i ,T D )) (7)

[0029]

[0030] Among them, X i is the input time series, T D represents the long-term target Token of modeling, split(·) is the split operation, T′ D represents the long-term target Token after interaction, MLP(·) represents the multi-layer perceptron, represents the end point of the model prediction, represents the set of endpoint positions;

[0031] In order to predict the end target, the task is divided into two parts. The first part uses TALSM to predict the position T′ of 14 frames. D Since the end point is at 20 frames, the offset T is added in the second part. B Fit the end point position in the way of The endpoint closest to the true value is Right now:

[0032]

[0033] index=argmind(10)

[0034]

[0035] Among them, || ||2 represents the Euclidean distance, D represents the true target point, k represents the predicted k endpoints, d represents the set of distances between the true target point and the predicted point, and index represents the index of the predicted target point closest to the true target point. and The predicted point represents the nearest true target point;

[0036] The second stage is to predict k terminal locations with multimodal distribution. In order to ensure that the prediction has multimodal properties, we set L diversity The loss function is as follows:

[0037]

[0038] in, and are different prediction endpoints, L2 is the Euclidean distance, and d represents the scaling factor;

[0039] The model outputs k different endpoints. In order to ensure the prediction accuracy of the model output, the loss function L is designed. destinctions :

[0040]

[0041] In order to balance prediction accuracy and multimodality, the following loss function is designed:

[0042] L task2 =L destinctiions +λL diversity (14)

[0043] Among them, λ is a hyperparameter.

[0044] Furthermore, the third stage implementation process of step (2) is as follows:

[0045] The third stage uses the modeling capabilities of the first stage and the end position of the second stage, and its input sequence adds the The remaining time features are filled with learnable tokens, namely:

[0046]

[0047] Among them, X i is the input time series, T other represents the filled learnable token, split(·) is the segmentation operation, T represents the complete trajectory, and the future 12-frame trajectory to be predicted only needs to be segmented from T;

[0048] The loss in the third stage includes the basic loss function and distillation loss, so that the final stage can be guided by training:

[0049]

[0050] L trajectory =L predict +L distill (17)

[0051] Among them, task1 is the output feature of the first stage, Trajectory features generated for the third stage, is the endpoint of the third stage input, D is the actual endpoint of the trajectory, g(·) is a linear change, α, β are hyperparameters, L predict Generate the Euclidean distance between the predicted trajectory and the true value for the third stage.

[0052] Beneficial effects: Compared with the prior art, the present invention has the following beneficial effects: the present invention designs a progressive transfer learning architecture, utilizes the characteristics of transfer learning, divides the task into stages, designs different processing modes for different stages, and improves the prediction accuracy of the final stage; in order to capture the interactive relationship in the local space, the present invention designs the TALSTM module, introduces the interaction matrix, reduces unnecessary influences, and accurately captures the mutual influence between pedestrians; the present invention adopts the attention mechanism and designs a fitting module applicable to all stages. Through early learning to accumulate knowledge, transfer learning is used in the later stage to apply the early results to prediction, and achieve accurate fitting of the final trajectory. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 This is a three-stage architecture diagram proposed by the present invention;

[0054] Figure 2 1 is a schematic diagram of TALS™ proposed in the present invention;

[0055] Figure 3 Schematic diagram of the TALS™ cell structure proposed in the present invention;

[0056] Figure 4 Schematic diagram of the structure of the fitting encoder proposed by the present invention;

[0057] Figure 5 It is a schematic diagram of the long-term prediction architecture proposed by the present invention;

[0058] Figure 6 Schematic diagram of the full sequence prediction architecture proposed by the present invention.

[0059] Figure 7 It is the prediction result diagram of the model under single mode;

[0060] Figure 8 It is the prediction result diagram of the multimodal model. DETAILED DESCRIPTION

[0061] The present invention will be further described in detail below with reference to the accompanying drawings.

[0062] This paper proposes a pedestrian trajectory prediction method based on progressive transfer learning. The method first segments the acquired pedestrian trajectory prediction dataset into a training set and a validation set. The method then sets the number of trajectories to be processed at a time, processes the data in batches, and constructs an interaction matrix for each batch. Finally, the training data is saved. After data processing, a progressive transfer learning network is constructed. The network consists of three stages, each of which uses transfer learning techniques to pass information to the next stage. In the first stage, the trajectory information from the previous moment is primarily used to predict the trajectory information for the next moment. In the second stage, long-term dependencies are learned based on short-term prediction capabilities, and the trajectory endpoints are predicted. In the final stage, the predicted trajectory endpoints are used to fit the complete trajectory. The preprocessed training data is then fed into the progressive transfer learning network for training. The network is then optimized using the AdamW optimizer by backpropagation to update weights. Finally, a test set is loaded and fed into the trained progressive transfer learning network for testing. During the test phase, the predicted results and ground truth values are loaded. The trajectory predictions are visualized using Matplotlib, intuitively displaying the model's prediction performance.

[0063] The overall structure of the progressive transfer learning network proposed in this invention is as follows: Figure 1 As shown in Figure 2, the model consists of three stages, linked by transfer learning. Each stage consists of two modules: TALSTM and a fitted encoder. These two modules are used to set up different architectures and perform different tasks in each stage. Specifically, the first stage predicts short-term trajectories, the second stage builds a long-term prediction architecture based on the first stage, and the final stage builds a full-sequence prediction architecture.

[0064] In order to capture the local spatial interaction relationship and model the temporal relationship at the same time, the TALSTM module is designed, such as Figure 2 TALSTM is composed of TALSTM cells combined with attention mechanism and masking technology based on LSTM structure. The specific cell structure is as follows: Figure 3 shown.

[0065] The TALSTM module captures local spatial interactions while modeling temporal relationships. It mainly controls the fusion of historical information and input information through the ITB module and controls the output at the current moment through the gating mechanism. The specific calculation formula is as follows:

[0066]

[0067] F t =σ(ITB(concat*x t ,H t-1 ),[Mask])) (2)

[0068]

[0069] H t =F t tanh(C t ) (4)

[0070] Among them, θ, φ, All are linear layers, X q 、X k 、X v All are input information x at the current moment t Compared with the output H at the previous moment t-1 The features after splicing, normalize is the normalization operation, concat represents the splicing operation, σ is the sigmoid operation, tanh is the activation function, mask represents the interaction matrix, F t The gate control unit is used to control the fusion of input information and output information at the previous moment, H t and C t They are the output and storage information at the current moment respectively.

[0071] The specific structure of the fitting encoder is as follows Figure 4 As shown in the figure, this module has different functions in different stages. In the first and second stages, the attention mechanism is used to interact with TALSTM to predict the output trajectory features. In the third stage, the training data and the endpoints generated in the second stage are input, and the complete trajectory is obtained by transfer learning and fitting the encoder.

[0072] In the first stage, the trajectory information of the previous moment is encoded and predicted through the TALSTM module and the fitting encoder module to help the network capture short-term spatial changes and learn fine-grained features, laying the foundation for subsequent trajectory prediction. The tasks of the first stage are:

[0073] (x t+1 ,y t+1 )=F(TALSTM(x t ,y t )) (5)

[0074] Among them, TALSTM(·) is the TALSTM module, F(·) is the fitting encoder, (x t ,y t ) represents the trajectory coordinate characteristics at time t, (x t+1 ,y t+1 ) represents the trajectory coordinate characteristics at time t+1.

[0075] The goal of this stage is to help the network understand the changing patterns of short-term space and learn fine-grained motion features. Its loss function is:

[0076]

[0077] in, is the predicted value of the i-th trajectory, y i is the i-th true trajectory.

[0078] In the second stage, due to the potential correlation between short-term trajectories and predicted endpoints, the present invention applies the short-term prediction model training results to long-term prediction tasks through transfer learning. The input of this stage is the first 8 frames of data, and the purpose is to predict the location of the endpoint of the 20th frame trajectory. In order to improve the accuracy of the prediction and make full use of the ability of short-term prediction, the following settings are set: Figure 5 The long-term prediction architecture shown in the figure is divided into prediction and optimization parts. TALSTM is used to predict the trajectory position of the 14th frame, and a learnable token is added as a bias at the position of the 14th frame. The fitting module is combined with historical input to optimize the endpoint position to achieve the long-term prediction task, that is, to predict the endpoint set.

[0079] T′ D =split(TALSTM(X i ,T D )) (7)

[0080]

[0081] Among them, X i is the input time series, T D represents the long-term target Token of modeling, split(·) is the split operation, T′ D represents the long-term target Token after interaction, MLP(·) represents the multi-layer perceptron, represents the end point of the model prediction, Represents the set of endpoint positions. It is worth noting that in order to predict the endpoint target, the task is divided into two parts. The first part uses TALSM to predict the position T′ of 14 frames. D Since the end point is at 20 frames, the offset T is added in the second part. B The end point position is fitted in this way.

[0082] In the collection The endpoint closest to the true value is Right now:

[0083]

[0084] index=arg min d (10)

[0085]

[0086] Among them, || ||2 represents the Euclidean distance, D represents the true target point, k represents the predicted k endpoints, d represents the set of distances between the true target point and the predicted point, and index represents the index of the predicted target point closest to the true target point. and The predicted point represents the closest true target point.

[0087] The core function of the second stage is to predict k terminal locations with multimodal distribution. In order to ensure that the prediction has multimodal properties, this paper sets L diversity The loss function is as follows:

[0088]

[0089] in, and are different prediction endpoints, L2 is the Euclidean distance, and d represents the scaling factor.

[0090] The model outputs k different endpoints. In order to ensure the prediction accuracy of the model output, a loss function L is designed. destinctions ,Right now:

[0091]

[0092] Finally, in order to balance prediction accuracy and multimodality, the following loss function is designed, namely:

[0093] L task2 =L destinctiions +λL diversity (14)

[0094] Among them, λ is a hyperparameter.

[0095] In the third stage, once the starting point and end point of a trajectory are determined, the general movement trend of the trajectory can be determined, and the full sequence prediction architecture can be designed. The input of the third stage consists of 1 to 8 frames of historical trajectory, the predicted end point of the second stage, and the learnable token. The specific structure of the full sequence prediction architecture is as follows: Figure 6 As shown. By inputting the historical 8 frames of data to predict the future 12 frames of trajectory, that is:

[0096]

[0097] Among them, X i is the input time series, is the endpoint generated in the second stage, T other represents the filled learnable Token, split(·) is the segmentation operation, T represents the complete trajectory, and the future 12-frame trajectory to be predicted only needs to be segmented from T.

[0098] Since the cornerstone of the third stage is the short-term prediction ability of the first stage and the third stage relies on the endpoint of the second stage, when designing the loss function, in addition to setting the basic loss function, the distillation loss is additionally introduced so that the final stage can be guided by the training.

[0099]

[0100] L trajectory =L predict +L distill (17)

[0101] Among them, task1 is the output feature of the first stage, Trajectory features generated for the third stage, is the endpoint of the third stage input, D is the actual endpoint of the trajectory, g(·) is a linear change, α, β are hyperparameters, L predict Generate the Euclidean distance between the predicted trajectory and the true value for the third stage.

[0102] In this embodiment, the Stanford University UAV Dataset (SDD) is first obtained from the official website, and the data is processed and divided. Specifically, the first 8 frames of pedestrian motion data are used as input features and the last 12 frames are used as true label values to construct data samples and labels, and the data is divided into 8:2; the preprocessed training set data is input into the network for training, wherein the number of training rounds in the first stage is 3000 rounds with an initial learning rate of 0.001, the second stage is trained for 600 rounds with an initial learning rate of 0.003, and the third stage is trained for 800 rounds with an initial learning rate of 0.001; finally, the trained model is tested on the test set and the prediction results and true values are loaded in the test stage. The trajectory prediction is visualized through matplotlib to intuitively display the model prediction effect.

[0103] Specific visualization effects such as Figure 7 、 Figure 8 As shown, the prediction includes unimodal and multimodal effects, where the multimodal result is the 20 possible future trajectory positions predicted by the model. Figure 7 、 Figure 8 In the figure, green represents the real data and red represents the predicted data. The green five-pointed star represents the real end point and the red five-pointed star represents the predicted end point. In the figure, the x-axis represents the relative position of the trajectory in the x-coordinate and the y-axis represents the relative position of the trajectory in the y-coordinate. The relative position is based on the split point as the coordinate origin. Figure 7 、 Figure 8 It can be seen from the fitting effect that under the condition of single-modal prediction, the model designed by the invention can perceive the movement trend of the future trajectory and make predictions. Under the condition of multi-modal prediction, the model designed by the invention can perfectly cover the different positions where the future trajectory may appear.

[0104] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made in accordance with the spirit of the present invention are intended to be covered by the scope of protection of the present invention.

Claims

1. A pedestrian trajectory prediction method based on progressive transfer learning, characterized in that: The following steps are involved: (1) Segment and preprocess the obtained pedestrian trajectory prediction dataset; (2) Constructing a progressive transfer learning network, which is divided into three stages. Each stage uses transfer learning technology to transfer information to the next stage. In the first stage, the trajectory information of the previous moment is used to predict the trajectory information of the next moment. In the second stage, long-term dependencies are learned based on the short-term prediction capability, and the trajectory endpoint is predicted; In the third stage, the complete trajectory is fitted using the predicted trajectory endpoints; (3) The preprocessed training data is fed into the progressive transfer learning network for training, and the progressive transfer learning network is optimized by backpropagation updating the weights using the AdamW optimizer; (4) Load the test set and send the test data to the trained progressive transfer learning network for testing. During the test phase, load the prediction results and true values, and use matplotlib to visualize the trajectory prediction to intuitively show the model prediction effect.

2. The pedestrian trajectory prediction method based on progressive transfer learning according to claim 1, characterized in that: The preprocessing in step (1) is as follows: setting the number of trajectories to be processed each time, processing data in batches, and constructing an interaction matrix for each batch of data according to the pedestrian interaction radius, and finally saving the training data.

3. The pedestrian trajectory prediction method based on progressive transfer learning according to claim 1, characterized in that: The three stages of the progressive transfer learning network in step (2) all include a TALSTM module and a fitting encoder module; the TALSTM module encodes and predicts the time series and outputs the prediction results; the fitting encoder module interacts across time series and enhances the module's ability to capture interactive relationships.

4. The pedestrian trajectory prediction method based on progressive transfer learning according to claim 3, characterized in that: The TALSTM module controls the fusion of historical information and input information through the ITB module and controls the output at the current moment through the gating mechanism. The specific calculation formula is as follows: F t =σ(ITB(concat(x t ,H t-1 ),[Mask])) (2) H t =F t ·tanh(C t ) (4) Among them, θ, φ, All are linear layers, X q 、X k 、X v All are input information x at the current moment t Compared with the output H at the previous moment t-1 The features after splicing, normalize is the normalization operation, concat represents the splicing operation, σ is the sigmoid operation, tanh is the activation function, mask represents the interaction matrix, F t The gate control unit is used to control the fusion of input information and output information at the previous moment, H t and C t They are the output and storage information at the current moment respectively.

5. The pedestrian trajectory prediction method based on progressive transfer learning according to claim 3, characterized in that: The fitting encoder uses the attention mechanism to interact with TALSTM in the first and second stages to predict the output trajectory features; In the third stage, the training data and the endpoints generated in the second stage are input, and the complete trajectory is obtained by using transfer learning and fitting the encoder.

6. The pedestrian trajectory prediction method based on progressive transfer learning according to claim 1, characterized in that: Step (2) The first stage implementation process is as follows: By encoding the trajectory information of the previous moment, the trajectory information of the next moment is predicted through the TALSTM module and the fitting encoder module, that is: (x t+1 ,y t+1 )=F(TALSTM(x t ,y t )) (5) Among them, TALSTM(·) is the TALSTM module, F(·) is the fitting encoder, (x t ,y t ) represents the trajectory coordinate characteristics at time t, (x t+1 ,y t+1 ) represents the trajectory coordinate characteristics at time t+1; The goal of the first stage is to help the network understand the changing patterns of short-term space and learn fine-grained motion features; its loss function is: in, is the predicted value of the i-th trajectory, y i is the i-th true trajectory.

7. The pedestrian trajectory prediction method based on progressive transfer learning according to claim 1, characterized in that: Step (2) The second stage is implemented as follows: Using transfer learning technology, the TALSTM module and the fitting encoder module are used to achieve long-term prediction tasks based on the short-term prediction model training, that is, to predict the endpoint set T′ D =split(TALSTM(X i ,T D )) (7) Among them, X i is the input time series, T D represents the long-term target Token of modeling, split(·) is the split operation, T′ D represents the long-term target Token after interaction, MLP(·) represents the multi-layer perceptron, represents the end point of the model prediction, represents the set of endpoint positions; In order to predict the end target, the task is divided into two parts. The first part uses TALSM to predict the position T of 14 frames. D ′, since the end point is at the 20th frame, the offset T is added in the second part. B Fit the end point position in the way of The endpoint closest to the true value is Right now: index=argmind(10) Among them, || ||2 represents the Euclidean distance, D represents the true target point, k represents the predicted k endpoints, d represents the set of distances between the true target point and the predicted point, and index represents the index of the predicted target point closest to the true target point. and The predicted point represents the nearest true target point; The second stage is to predict k terminal locations with multimodal distribution. In order to ensure that the prediction has multimodal properties, we set L diversity The loss function is as follows: in, and are different prediction endpoints, L2 is the Euclidean distance, and d represents the scaling factor; The model outputs k different endpoints. In order to ensure the prediction accuracy of the model output, the loss function L is designed. destinctions : In order to balance prediction accuracy and multimodality, the following loss function is designed: THE task2 =L destinctiions +λL diversity (14) Among them, λ is a hyperparameter.

8. The pedestrian trajectory prediction method based on progressive transfer learning according to claim 1, characterized in that: The third stage implementation process of step (2) is as follows: The third stage uses the modeling capabilities of the first stage and the end position of the second stage, and its input sequence adds the The remaining time features are filled with learnable tokens, namely: Among them, X i is the input time series, T other represents the filled learnable token, split(·) is the segmentation operation, T represents the complete trajectory, and the future 12-frame trajectory to be predicted only needs to be segmented from T; The loss in the third stage includes the basic loss function and distillation loss, so that the final stage can be guided by training: L trajectory =L predict +L distill (17) Among them, task1 is the output feature of the first stage, Trajectory features generated for the third stage, is the endpoint of the third stage input, D is the actual endpoint of the trajectory, g(·) is a linear change, α, β are hyperparameters, L predict Generate the Euclidean distance between the predicted trajectory and the true value for the third stage.

Citation Information

Cited By

  • Track prediction method based on endpoint guided loop learning and one-time prediction

    CN121412654A