Pedestrian trajectory prediction method based on destination guidance

By adopting the destination dual-channel expression mechanism and conditional diffusion method in pedestrian trajectory prediction, the problem of inaccurate pedestrian future trajectory prediction in the prior art is solved, and a more accurate and robust trajectory prediction effect is achieved.

CN120125609APending Publication Date: 2025-06-10CHANGZHOU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510187402.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the future trajectory of pedestrians, especially the prediction results of diversified trajectory due to uncertainty in pedestrian subjective intentions and scene complexity.

Method used

A pedestrian trajectory prediction method based on the destination dual-channel expression mechanism and conditional diffusion method is adopted. A diverse future trajectory is generated through the destination prediction module and the spatiotemporal encoding module, and a destination information is used as a condition to accurately predict the trajectory.

Benefits of technology

It significantly improves the accuracy and robustness of pedestrian trajectory prediction, can better reflect the pedestrian's subjective intentions, and generate prediction results that are more consistent with the actual trajectory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125609A_ABST
    Figure CN120125609A_ABST
Patent Text Reader

Abstract

The invention relates to the field of trajectory prediction, in particular to a pedestrian trajectory prediction method based on destination guidance. The method comprises the following steps: constructing a pedestrian trajectory prediction model, training the pedestrian trajectory prediction model, and predicting a pedestrian trajectory by adopting the trained pedestrian trajectory prediction model; the working process of the pedestrian trajectory prediction model is as follows: through a destination prediction module, destination features including destination sampling and destination local context semantics are generated based on an input scene and an observation trajectory; through a space-time coding module, motion features containing pedestrian interaction information are generated based on the observation trajectory; inputting the destination features and the motion features into a conditional diffusion model to generate diversified future trajectories; and selecting an optimal trajectory from the diversified future trajectories as a pedestrian prediction trajectory. According to the invention, a destination dual-channel expression mechanism and a conditional diffusion method are adopted to more comprehensively describe the potential destination of the pedestrian, and accurate prediction of the trajectory is realized under the condition of the potential destination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of trajectory prediction, and particularly to a pedestrian trajectory prediction method based on destination orientation. Background Art

[0002] The rapid development of autonomous intelligent systems such as autonomous driving vehicles, automated guided vehicles, and guiding robots has brought convenience to people's lives. However, the smooth operation of autonomous intelligent systems requires accurate detection of the movement trajectories of pedestrians in the surrounding environment to prevent collisions with people. Affected by the uncertainty of pedestrians' subjective intentions and the complexity of objective scenarios, it is difficult to accurately predict future pedestrian trajectories. Therefore, conducting research on pedestrian trajectory prediction is of great significance for the safe operation of autonomous intelligent systems.

[0003] The uncertainty of pedestrians' subjective intentions leads to diverse trajectory prediction results. Therefore, trajectory prediction models need to accurately capture and model potential uncertainty factors to generate diverse future trajectories. Traditional research often uses a generative framework to achieve trajectory prediction. Some scholars use variational autoencoders to achieve diverse predictions of trajectories. SocialGAN introduces a generative adversarial strategy and random noise to generate multi-modal future trajectories. However, the generation results of variational autoencoders are not accurate enough, and the training process of generative adversarial networks is unstable and prone to mode collapse. At the same time, the above methods use random noise to generate diverse trajectories, which do not truly reflect pedestrians' subjective intentions, resulting in a large number of trajectories that do not conform to reality. Recent research usually predicts pedestrians' potential destinations as guiding clues for trajectory prediction, and many methods use a single expression of potential destinations, without fully utilizing destination-related information in the scenario.

[0004] Therefore, how to generate future trajectories that conform to pedestrians' subjective intentions has become a problem to be solved. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide a pedestrian trajectory prediction method based on destination orientation, which adopts a destination dual-channel expression mechanism and a conditional diffusion method to more comprehensively describe pedestrians' potential destinations, and realizes accurate prediction of trajectories based on the potential destinations.

[0006] To solve the above technical problem, the technical solution of the present invention is: A pedestrian trajectory prediction method based on destination orientation, comprising:

[0007] Construct a pedestrian trajectory prediction model, train the pedestrian trajectory prediction model, and use the trained pedestrian trajectory prediction model to predict pedestrian trajectories; wherein, the working process of the pedestrian trajectory prediction model is:

[0008] Generate the destination feature U including destination sampling and destination local context semantics based on the input scene and the observed trajectory through the destination prediction module g ;

[0009] Generate the motion feature U containing pedestrian interaction information based on the observed trajectory through the spatio-temporal encoding module h ;

[0010] Input the destination feature U g and the motion feature U h into the conditional diffusion model to generate diverse future trajectories;

[0011] Select the optimal trajectory from the diverse future trajectories as the pedestrian prediction trajectory.

[0012] Furthermore, generate the destination feature U including destination sampling and destination local context semantics based on the input scene and the observed trajectory g ; Specifically including:

[0013] Use the lightweight semantic segmentation network MobileNet to extract the semantic map I of the input scene s , and generate the historical trajectory heatmap G based on the observed trajectory h ;

[0014] Align and splice the semantic map I s and the historical trajectory heatmap G h , and input them into the U-Net network to obtain the target distribution heatmap of the future frame T f

[0015] Sample the last frame probability map of the target distribution heatmap to generate N estimated destinations g N ; Crop the local trajectory heatmap in the target distribution heatmap based on the coordinates of the last frame of the observed trajectory and sequentially use the convolutional neural network and the fully connected layer to extract the context F related to the destination from the local trajectory heatmap s ;

[0016] Calculate the velocity V and acceleration A according to the observed trajectory, and subtract the position vector x of each frame in the observed trajectory t from the estimated destination g ∈ g N , then splice the velocity V, acceleration A and the observed trajectory to obtain the augmented state of each estimated destination

[0017] The augmented state ​First, perform feature extraction through a one-dimensional convolutional network, and then send it to the LSTM encoder to extract the temporal information therein to obtain the destination location feature F g ;

[0018] Input the destination location feature F g and the destination semantic feature F s into the attention mechanism for fusion to obtain a more comprehensive destination feature U g 。

[0019] Furthermore, select the optimal trajectory from the diverse future trajectories as the pedestrian prediction trajectory; specifically:

[0020] Use the average displacement error and the final displacement error to select the trajectory with the smallest error from the diverse future trajectories as the pedestrian prediction trajectory.

[0021] Furthermore, the destination prediction module is trained using binary cross-entropy loss to measure the dissimilarity between the future true trajectory heatmap G f and the predicted target distribution heatmap as:

[0022]

[0023] The conditional diffusion model is trained using mean squared error loss, expressed as:

[0024]

[0025] In the formula, the error is calculated through the expected value ; the contaminated trajectory state after K steps of forward diffusion is x is the observed trajectory; θ and φ are the parameters of the destination prediction module and the spatio-temporal encoding module respectively;

[0026] The final loss is expressed as:

[0027]

[0028] In the formula, λ represents the hyperparameter for balancing different losses.

[0029] Furthermore, to improve the generalization performance of the global model, the pedestrian trajectory prediction model is trained; specifically:

[0030] Perform federated multi-factor aggregation training on the pedestrian trajectory prediction model, and use three factors, namely the data volume, data distribution quality, and training loss of the client, as indicators to evaluate the client contribution degree, and determine the weight of its local model in the aggregation process.

[0031] Furthermore, the client contribution degree is evaluated with the client data volume as an indicator, expressed as:

[0032]

[0033] Wherein, M represents the total number of clients; represents the local data volume of the m-th client; represents the total data volume of all clients, The local model of the m-th client in the r-th round is expressed as

[0034] Taking the data distribution quality as an index to evaluate the client contribution degree, it is expressed as:

[0035]

[0036] Wherein, represents the gradient of the global model parameters in the previous round; represents the local model parameter gradient of the m-th client in this round of training Sig() represents the Sigmoid activation function;

[0037] Taking the training loss as an index to evaluate the client contribution degree, it is expressed as:

[0038]

[0039] Wherein, represents the model ω r The loss on the m-th client, Sig() is the sigmoid function, indicating that is normalized;

[0040] Using the three factors of the client's data volume, data distribution quality, and training loss as indicators to evaluate the client contribution degree, and determining the weight of its local model in the aggregation process, it is expressed as:

[0041]

[0042] Wherein,

[0043] After adopting the above technical solution, the present invention uses the destination dual-channel expression mechanism to more comprehensively describe the potential destinations of pedestrians. The destination features include potential destination sampling and destination local context semantics. The former provides possible destination locations, and the latter provides semantic information related to the destination. Compared with only using potential destination sampling as a representation, the method of the present invention provides a more comprehensive destination description, significantly improving the robustness of the destination for guiding pedestrian trajectory generation. Moreover, it also uses the destination information as a condition to guide the generation of the diffusion model, suppressing the uncertainty of future trajectories and further improving the accuracy of pedestrian trajectory prediction. Description of the Drawings

[0044] Figure 1 It is the framework diagram of the pedestrian trajectory prediction model of the present invention;

[0045] Figure 2 It is the framework diagram of the destination prediction module of the present invention;

[0046] Figure 3 It is the visualization comparison diagram of the method proposed by the present invention and the MID method on the dataset SDD. Detailed implementation manners

[0047] In order to make the content of the present invention easier to be clearly understood, the present invention will be further described in detail below according to specific embodiments in conjunction with the accompanying drawings.

[0048] As Figure 1 and Figure 2 shown, a destination-oriented pedestrian trajectory prediction method includes:

[0049] Construct a pedestrian trajectory prediction model, train the pedestrian trajectory prediction model, and use the trained pedestrian trajectory prediction model to predict the pedestrian trajectory; wherein, the working process of the pedestrian trajectory prediction model is:

[0050] Step S1, through the destination prediction module, generate the destination feature U including destination sampling and destination local context semantics based on the input scene and the observation trajectory X g ; specifically including:

[0051] Step S11, use the lightweight semantic segmentation network MobileNet to extract the semantic map of the input scene H and W are the height and width of the image, respectively, and I s contains various scene elements, such as sidewalks, buildings, grasslands, etc. For the observation trajectory X, calculate the historical trajectory heat map G using the same method as Y-Net h ;

[0052] Step S12, after aligning and splicing the semantic map I s and the historical trajectory heat map G h , input them into the U-Net network to obtain the target distribution heat map of the future frame T f ;

[0053] Step S13, in order to obtain a more comprehensive destination representation, use the Test-Time-Sampling-Trick (TTST, test time sampling trick) to sample the probability map of the last frame of the target distribution heat map to generate N estimated destinations g N ;

[0054] Step S14, considering that the scene context contained in can assist g N , so based on the coordinates of the last frame of the observation trajectory X, a local trajectory heatmap is cropped from the target distribution heatmap and the context F related to the destination is extracted from the local trajectory heatmap in turn using a convolutional neural network (CNN) and a fully connected layer (FC); It is expressed as: s ; expressed as:

[0055]

[0056] In the formula, θ c and θ f are the parameters of the CNN and FC respectively.

[0057] Step S15, calculate the velocity V and acceleration A according to the observation trajectory X, and subtract the position vector x of each frame in the observation trajectory X t from the estimated destination g ∈ g N , then splice the velocity V, acceleration A and the observation trajectory X to obtain the augmented state of each estimated destination It is expressed as:

[0058] J = {x t - g|t = 1, 2,..., T obs}

[0059]

[0060] Step S16, first extract features from the augmented state through a one-dimensional convolutional network (ID-CNN), and then send it to an LSTM encoder to extract the temporal information therein to obtain the destination location feature F g ; It is expressed as:

[0061]

[0062] Step S17, input the destination location feature F g and the destination semantic feature F s into the attention mechanism for fusion to obtain a more comprehensive destination feature U g ; It is expressed as:

[0063] U g = softmax(F g * ε) * F g + softmax(F s * (1 - ε)) * F s

[0064] where ε is the attention weight, used to measure F g and F s 's importance, set to 0.6 through experiments, and softmax represents the softmax() operation.

[0065] Step S2, through the spatio-temporal encoding module, generate the motion feature U containing pedestrian interaction information based on the observed trajectory X h ;

[0066] Step S3, input the destination feature U g and the motion feature U h into the conditional diffusion model to generate diverse future trajectories;

[0067] Specifically, diffusion-based trajectory generation can obtain more accurate results and the training process is more stable. Figure 1 is the framework diagram of the pedestrian trajectory prediction model after introducing the conditional diffusion model, FC represents the fully connected layer, represents concatenation.

[0068] The diffusion generation of pedestrian trajectories consists of two processes: forward diffusion and reverse diffusion. Forward diffusion contaminates the future real trajectory by introducing random noise to obtain the contaminated trajectory state. Given the real future trajectory Y 0 , the contaminated trajectory state after diffusing K steps is Y K , and the diffusion process Y (0:K) is defined as follows:

[0069]

[0070] where β 1 , β 2 ,..., β k is a fixed variance scheduler that controls the level of injected noise. Since Gaussian distributed random noise is added at each step, represents the noise mean, and β k I represents the variance of the noise. Let α k = 1 - β k , The diffusion process at any step k is calculated as follows:

[0071]

[0072] where when the diffusion step K is large enough, it will approximately obtain

[0073] Reverse diffusion is used to recover the real trajectory from the contaminated trajectory state, so as to learn how to generate pedestrian trajectories from the noise distribution. The reverse diffusion process is defined as follows:

[0074]

[0075] Wherein, θ are the parameters of the Transformer-based diffusion network, and U is the guiding condition for reverse diffusion. The variance term of the Gaussian distribution is ∑ θ (Y K , k) = β k I. In other words, during the trajectory generation process, the diffusion network predicts the noise conditioned on k and U, and uses the predicted noise to iteratively denoise to obtain the future predicted trajectory. By training the network, the predicted trajectory tends to the true future trajectory Y 0 .

[0076] Step S4: Select the optimal trajectory from the diverse future trajectories as the pedestrian prediction trajectory; specifically:

[0077] Use the average displacement error and the final displacement error to select the trajectory with the minimum error from the diverse future trajectories as the pedestrian prediction trajectory.

[0078] In this embodiment, during the training process, the destination prediction module is trained using binary cross-entropy loss to measure the dissimilarity between the future true trajectory heatmap G f and the predicted target distribution heatmap which is expressed as:

[0079]

[0080] The conditional diffusion model is trained using mean squared error loss, which is expressed as:

[0081]

[0082] Wherein, the error is calculated through the expected value The contaminated trajectory state after K steps of forward diffusion is x is the historical trajectory; θ and φ are the parameters of the destination prediction module and the spatio-temporal encoding module respectively;

[0083] The final loss is composed of the above two losses weighted and optimized by the Adam optimizer. At the same time, the client local training loss, data distribution quality, and data volume are calculated. The final loss is expressed as:

[0084]

[0085] Wherein, λ represents the hyperparameter for balancing different losses.

[0086] In this embodiment, in step S3, in the conditional diffusion model, the destination feature Ug and motion feature U h First, through the fully connected layer, the data is mapped to a higher-dimensional feature space so that the model can capture more complex features. By introducing noise, a noisy trajectory state Y is generated k , after passing through the fully connected layer, it is concatenated with the destination feature and the trajectory feature, and then through the Transformer module, the model can process the pedestrian trajectory data and generate future predicted trajectories. The generated trajectory will pass through the fully connected layer and be compared with the real trajectory through the mean square error (MSE) loss function. The purpose of this process is to minimize the error between the predicted trajectory and the actual trajectory, so that the output of the model is as close as possible to the real trajectory.

[0087] It should be noted that in the test phase, there is no forward process, only the denoising process. During testing, samples are taken from Gaussian noise, conditioned on the historical trajectory, and then denoised to generate the predicted trajectory.

[0088] In this embodiment, during the training process, the pedestrian trajectory prediction model is trained by federated multi-factor aggregation. The data volume, data distribution quality, and training loss of the client are used as three factors to evaluate the client contribution degree, and the weight of its local model in the aggregation process is determined.

[0089] Specifically, the pedestrian trajectory prediction model needs to use a large amount of rich pedestrian trajectory data in different scenarios for collaborative training, but it also brings the risk of data leakage. To better protect the privacy of pedestrians in different scenarios, this embodiment introduces a federated learning framework for collaborative training of the trajectory prediction model. The aggregation of the global model is the core of federated learning. Since the data is distributed on multiple local clients, the training and update of the global model are completed by aggregating the client model parameters. Traditional FedAvg performs weighted average aggregation on the model parameters according to the client data volume, which has certain limitations. The aggregation of the global model needs to analyze the importance of client data from multiple perspectives. Therefore, this embodiment proposes a multi-factor aggregation method FedMA, which uses the data volume, data distribution quality, and training loss of the client as three factors to evaluate the client contribution degree, and then determines the weight of its local model in the aggregation process. The details of different factors are as follows:

[0090] ① Data volume

[0091] Suppose there are a total of M scenario clients (S 1 , S 2 ,..., S M ), and the local data set of the client in the r-th round is defined as where the local data volume of the m-th client is The total data volume of all clients is expressed as FedAvg performs weighted aggregation based on the proportion of the local data volume of the client in the total data volume. Suppose the local model of client m in the r-th round is represented as Then the global model ω in the r-th round r is defined as follows:

[0092]

[0093] The proportion of the data volume occupied by client m is calculated as follows:

[0094]

[0095] ② Data distribution quality

[0096] Considering the data heterogeneity in different scenarios, this embodiment introduces a data distribution quality index to more comprehensively evaluate the contribution of clients to the global model. By measuring the cosine similarity between the global model and the local model of the client in the parameter update direction, the data distribution difference is reflected, so as to optimize the weight allocation.

[0097] Specifically, in the r-th round of global training, the data distribution quality index uses the gradient of the global model parameters in the previous round and the gradient of the local model parameters of client m in this round of training to calculate the cosine similarity, as shown below:

[0098]

[0099] By measuring the similarity between the global model parameter gradient and the client local model parameter gradient, the contribution degree of client m to the global model can be evaluated. If the similarity is high, it means that the local model of client m is similar to the optimization direction of the global model, indicating that the data distribution of this client is relatively consistent with the global data. Therefore, it has a greater contribution to the update of the global model. On the contrary, a low similarity means that the contribution of this client to the global model update is relatively small. Sig() represents the Sigmoid activation function, which is used to prevent the client contribution from being negative.

[0100] ③ Training loss

[0101] The loss of client local training reflects the prediction ability of the global model for local data. The higher the loss, the weaker the prediction ability of the model, and the greater the error between the prediction trajectory and the real trajectory. Therefore, in this embodiment, by assigning higher weights to clients with higher loss values, while the global model accelerates convergence, it pays more attention to the data distribution of these clients, thereby enhancing its generalization ability in non-independent and identically distributed scenarios. The local training loss of client m in the r-th round is defined as and is defined as follows:

[0102]

[0103] wherein is the loss of model ω r on client m, and the Sigmoid function is used to perform normalization to ensure that the three metrics are in the same range. Then, the optimal weights are calculated by comprehensively considering the three metrics as follows:

[0104]

[0105] wherein, is the sum of the three metric parameters of client m, represents the model weight of client m at the r-th round of update. Thus, the new weights can be used to perform weighted aggregation on the local model to obtain a new round of global model.

[0106] This embodiment proposes a multi-factor aggregation federated learning, which can effectively utilize pedestrian trajectory data in different scenarios and protect pedestrian privacy. Under the federated learning framework, aiming at the heterogeneity of trajectory data in different scenarios, three factors, namely data volume, data distribution quality, and training loss, are introduced as weighted aggregation metrics. By dynamically evaluating the importance of each scenario client, the generalization performance of the global model is improved.

[0107] A quantitative analysis is carried out on the method involved in the above embodiment and the traditional method as follows.

[0108] Quantitative analysis of the SDD dataset (T represents trajectory information, T+S represents trajectory and scenario information)

[0109]

[0110] In terms of method comparison, according to the input information, the methods can be divided into two types: the first type is to only input trajectory information (represented by the symbol T); the second type is to input trajectory and scenario environment information (represented by the symbol T+S).

[0111] Social-GAN only relies on the observed trajectory information to predict future trajectories, ignoring the influence of the scene environment. Therefore, its prediction performance is relatively poor. SoPhie improves the prediction performance by using CNN to encode the scene environment information, and the prediction errors ADE20 and FDE20 are 40% and 29% lower than those of Social-GAN respectively. PECNet introduces destination information and predicts trajectories based on the end-point guidance, thus significantly improving the prediction accuracy. IRLSOT further explores complex scenes through inverse reinforcement learning and effectively learns the correlation between pedestrian trajectories and scene information by combining the attention mechanism, reducing the prediction errors ADE20 and FDE20 to 9.66 and 13.05 respectively. The Trajectron++ model focuses on multi-agent behavior prediction, successfully fuses heterogeneous data and considers dynamic constraints, further reducing the ADE20 error. Y-net uses the U-net network to perform semantic segmentation on the RGB image of the scene, aligns the trajectory information and scene information at the pixel level, and then predicts the destination probability distribution of future trajectories in the scene-based trajectory heat map, further improving the prediction accuracy. Different from these methods, MID does not process image data. Instead, by introducing a diffusion model, it learns a parameterized Markov chain conditional on the observed trajectories and generates the desired future trajectories through a denoising process. The prediction errors ADE20 and FDE20 are 7.61 and 14.30 respectively. Compared with MID, this embodiment uses a scene semantic segmentation network to deeply understand the environmental information in the scene, combines it with the trajectory information, and further obtains future destination information. Combining the pedestrian spatio-temporal interaction features as a condition, this embodiment guides the diffusion network to generate more accurate future trajectories, significantly improving the prediction performance, and reducing ADE20 and FDE20 to 7.26 and 12.87 respectively. It is worth mentioning that if the destination information is removed as a guide, the prediction errors ADE20 and FDE20 will increase to 7.55 and 14.12 respectively, indicating that the destination information plays an important role in improving the trajectory prediction accuracy.

[0112] Inspired by the above ideal embodiments of the present invention, through the above description, relevant staff can make various changes and modifications within the scope not deviating from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. A pedestrian trajectory prediction method based on destination guidance, characterized in that: include: A pedestrian trajectory prediction model is constructed and trained, and the pedestrian trajectory prediction model is used to predict pedestrian trajectories. The working process of the pedestrian trajectory prediction model is as follows: The destination prediction module generates the destination feature U including the destination sample and the local context semantics of the destination based on the input scene and the observed trajectory. g ; Through the spatiotemporal encoding module, the motion feature U containing pedestrian interaction information is generated based on the observed trajectory. h ; Set the destination feature U g and motion characteristics U h Input the conditional diffusion model to generate diverse future trajectories; The optimal trajectory is selected from the diverse future trajectories as the pedestrian prediction trajectory.

2. The method for predicting pedestrian trajectories based on destination guidance according to claim 1, characterized in that: Generate destination features U including destination sampling and destination local context semantics based on input scene and observation trajectory g ; Specifically include: Extracting semantic graph of input scene using lightweight semantic segmentation network MobileNet s , based on the observed trajectory, generate the historical trajectory heat map G h ; The semantic graph I s and historical trajectory heat map G h After alignment and splicing, it is input into the U-Net network to obtain the future frame T f Target distribution heat map Target distribution heat map The last frame probability map is sampled to generate N estimated destinations g N ; Based on the coordinates of the last frame of the observed trajectory, the target distribution heat map Mid-cropped local trajectory heatmap And use convolutional neural network and fully connected layer to extract local trajectory heat map Extract the context F related to the destination s ; Calculate the velocity V and acceleration A according to the observed trajectory, and convert the position vector x of each frame in the observed trajectory t With the estimated destination g∈g N Then, we concatenate the velocity V, acceleration A, and observed trajectory to obtain the augmented state of each estimated destination. The state will be augmented First, the feature is extracted through a one-dimensional convolutional network, and then sent to the LSTM encoder to extract the timing information to obtain the destination location feature F. g ; The destination location feature F g and the destination semantic feature F s Input the attention mechanism for fusion to obtain a more comprehensive destination feature U g .

3. The method for predicting pedestrian trajectories based on destination guidance according to claim 1, characterized in that: Select the best trajectory from various future trajectories as the pedestrian prediction trajectory; Specifically: The average displacement error and the final displacement error are used to select the trajectory with the smallest error from the diverse future trajectories as the pedestrian prediction trajectory.

4. The method for predicting pedestrian trajectories based on destination guidance according to claim 1, characterized in that: The destination prediction module is trained using binary cross entropy loss to measure the future true trajectory heat map G f And predicted target distribution heat map The dissimilarity between them is expressed as: The conditional diffusion model is trained using mean square error loss, expressed as: In the formula, the expected value To calculate the error; the state of the contaminated trajectory after k steps of forward diffusion is x is the observation trajectory; θ and φ are the parameters of the destination prediction module and the spatiotemporal encoding module, respectively; The final loss is expressed as: Where λ represents a hyperparameter that balances different losses.

5. The method for predicting pedestrian trajectories based on destination guidance according to claim 1, characterized in that: Train the pedestrian trajectory prediction model; specifically: The pedestrian trajectory prediction model is trained through federated multi-factor aggregation. The client's data volume, data distribution quality and training loss are used as indicators to evaluate the client's contribution and determine the weight of its local model in the aggregation process.

6. The method for predicting pedestrian trajectories based on destination guidance according to claim 5, characterized in that: The client data volume is used as an indicator to evaluate the client contribution, which is expressed as: Where M represents the total number of clients; Indicates the amount of local data of the mth client; Indicates the total data volume of all clients. The local model of client m in round r is expressed as The data distribution quality is used as an indicator to evaluate the client contribution, which is expressed as: In the formula, Represents the gradient of the global model parameters in the previous round; Represents the local model parameter gradient of client m in this round of training Sig() represents the Sigmoid activation function; The training loss is used as an indicator to evaluate the client contribution, which is expressed as: In the formula, Representation model ω r The loss on client m, Sig() is the sigmoid function, which represents the To standardize; The client's data volume, data distribution quality, and training loss are used as indicators to evaluate the client's contribution and determine the weight of its local model in the aggregation process, which is expressed as: In the formula,

Citation Information

Cited By

  • Pedestrian trajectory prediction method based on interactive perception diffusion model

    CN121438409A

  • A pedestrian trajectory prediction method based on interactive perception diffusion model

    CN121438409B