Pedestrian trajectory prediction method based on social relation grouping diffusion
By adopting the social relationship grouping diffusion method in pedestrian trajectory prediction, the multimodal distribution of future trajectories is directly learned, and the problems of ignoring integrity and inefficiency in the existing technology are solved, and high-precision and high-efficiency pedestrian trajectory prediction are achieved.
Patent Information
- Application Number
- CN202510002612.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art ignores the integrity of pedestrian motion in pedestrian trajectory prediction and cannot capture the diversity in the underlying distribution of the model, resulting in an increase in the final offset error of the prediction results, while the prediction efficiency is limited by a large number of denoising steps.
The pedestrian trajectory prediction method based on social relationship grouping diffusion is adopted. By constructing a pedestrian grouping network and a trajectory distribution initialization network, the multimodal distribution of future trajectories is directly learned, and a large number of denoising steps are skipped to improve prediction efficiency.
High-precision and high-efficiency pedestrian trajectory prediction are achieved, which reduces the average offset error and the final offset error, and improves the accuracy and speed of prediction.
Smart Images

Figure CN119942589A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image processing, and in particular relates to a pedestrian trajectory prediction method. Background Art
[0002] A multimodal pedestrian trajectory prediction network based on conditional autoencoders discloses a multimodal pedestrian trajectory prediction method based on a latent diffusion model. The pedestrian trajectory distribution is denoised through a diffusion model to obtain a noise representation and train the diffusion model according to the sample conditional state and a preset loss function. The trained diffusion model and variational inference are used to generate prediction results. When modeling social relationships, this method only considers the mutual influence relationship between individual pedestrians, ignoring the integrity of pedestrian movement, resulting in the inability to capture the diversity in the underlying distribution of the model and omitting important possibilities of pedestrians' future movement directions, thereby improving the final offset error FDE of the prediction result. At the same time, this method is limited by the fact that the diffusion model requires a large number of denoising steps to ensure the accuracy of pedestrian trajectory prediction, which severely limits its prediction efficiency. Summary of the invention
[0003] The technical problem to be solved by the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a pedestrian trajectory prediction method based on social relationship group diffusion with high accuracy and efficiency.
[0004] The technical solution adopted to solve the above technical problems consists of the following steps:
[0005] (1) Processing the training data set
[0006] The public dataset ETH-UCY contains five different scenes: ZARA1, ZARA2, ETH, HOTEL, and UNIV. The four scenes of ZARA1, ZARA2, ETH, and HOTEL are selected as training sets, and the UNIV scene is selected as the test set. The trajectories are sampled at intervals of 0.4 seconds. The first 3.2 seconds of the trajectory are observation data, and the trajectory after 4.8 seconds is the future trajectory to be predicted. The trajectory coordinate sequence X of pedestrian i i According to formula (1), it can be expressed as:
[0007]
[0008] in, is the world coordinate of the target pedestrian i at time t, i∈{1,2,…,N}, N is the total number of pedestrians in the scene, which is a finite positive integer. obs}, t obs is the observation sequence length, t obs ∈[1,12].
[0009] Determine the pedestrian’s social graph P as follows: G :
[0010] P G =(V p ,E p ),
[0011] V p ={X i},
[0012] E p ={e i,j},
[0013] Among them, V p is the vertex set of the graph, representing the trajectories of all pedestrians in the scene, E p is the edge set of the graph, e i,j Represents the social relationship between two people, i,j∈{1,2,…,N}.
[0014] (2) Building a pedestrian trajectory prediction network
[0015] The pedestrian trajectory prediction network is composed of a pedestrian grouping network, a pedestrian trajectory distribution initialization network, and a trajectory distribution denoising module connected in series.
[0016] (3) Adding noise to pedestrian trajectory distribution
[0017] According to formula (2), add noise y to the pedestrian trajectory distribution k :
[0018]
[0019]
[0020] Among them, α i is a parameter that controls the intensity of added noise, y0 is the actual trajectory of the pedestrian, n is the standard Gaussian noise, α start is the initial noise intensity, α start ∈[0.5,0.9], K is the total number of diffusion steps, K∈[100,200].
[0021] (4) Training the noise estimation network
[0022] 1) Constructing noise estimation loss function
[0023] The noise estimation loss function L1 is determined as follows:
[0024]
[0025] in, is the prediction noise of the noise estimation network.
[0026] 2) Training the noise estimation network
[0027] The training set is input into the pedestrian trajectory prediction network for training. During the training process, the training parameters are: the initial learning rate is 0.01, and it decays by half every 10 epochs, the training rounds are 100, the data batch is 64, the number of denoising steps γ∈[20,50], the multi-head attention network is set to 8 heads, the number of hidden units in the fully connected layer 1 is 512, the number of hidden units in the fully connected layer 2 is 1024, and the number of hidden units in the fully connected layer 3 is 2048.
[0028] (5) Training pedestrian trajectory prediction network
[0029] 1) Construct pedestrian trajectory loss function
[0030] The pedestrian trajectory loss function L2 is determined as follows:
[0031]
[0032] Among them, Y is the actual trajectory of the pedestrian, is the predicted result of the pedestrian trajectory.
[0033] 2) Training human trajectory prediction network
[0034] The training set is input into the pedestrian trajectory prediction network for training. During the training process, the training parameters are: learning rate is 0.0015, training rounds are 400, data batch is 64, multi-head attention network is set to 8 heads, the number of hidden units in the fully connected layer 1 is 512, the number of hidden units in the fully connected layer 2 is 1024, and the number of hidden units in the fully connected layer 3 is 2048.
[0035] (6) Testing the Pedestrian Trajectory Prediction Network
[0036] The test set is input into the trained pedestrian trajectory prediction network for testing, and the trajectory coordinate sequence X of pedestrian i is i Input data and learn the group feature G of pedestrian i as follows:
[0037] G=F g (X i )
[0038] Among them, F g (·) Pedestrian grouping network, learn X as follows i The time series characteristic information
[0039] Push-to-learn X i The time series characteristic information
[0040]
[0041] Among them, Ft (·) is the temporal feature learning module, which transforms the pedestrian’s social graph P G The group features G of pedestrians are used as the input of the social relationship feature learning module, and the social relationship feature information of pedestrian i is learned as follows
[0042]
[0043] Among them, F s (·) is the social relationship feature learning module. and As the input of the feature fusion module, the mean μ of the trajectory distribution is calculated as follows: θ , standard deviation σ θ , predict sample S θ :
[0044]
[0045] in, is the mean gated recurrent unit, is the standard deviation of the gated recurrent unit, is the gated recurrent unit of the predicted sample, F fusion (·) is the feature fusion module. The initial distribution of pedestrian trajectories is determined as follows: and prediction noise
[0046]
[0047] Among them, F n (·) is the noise estimation network, γ is the current denoising step, and the denoising process of the γ-times standard diffusion model is performed according to formula (3) d (·):
[0048]
[0049] Where z is a random variable with standard Gaussian distribution, α k is a parameter that controls the intensity of noise subtraction, α k ∈[0.7,0.9], n γ yes The noise that needs to be removed, is the result of the previous denoising process. Formula (3) is executed γ times repeatedly until the pedestrian trajectory prediction result is predicted. The predicted trajectory of the pedestrian is determined according to formula (4):
[0050]
[0051] in, For the target pedestrian i in the future t pThe predicted value of the world coordinates at time t p ∈{t obs ,t obs +1,…,t obs +t pred}, t pred is the predicted trajectory step length, t pred ∈[8,12], each step is 0.48 seconds; the actual trajectory Y of the pedestrian is determined by the following formula:
[0052]
[0053] in, For the target pedestrian i in the future t p The true value of the world coordinates at the moment is determined by the average offset error ADE as follows:
[0054]
[0055] The final offset error FDE is determined as follows:
[0056]
[0057] in, The final coordinates of the pedestrian trajectory prediction sequence, The final coordinates of the pedestrian trajectory observation sequence are output, and the pedestrian trajectory test results are output.
[0058] In step (2) of constructing a pedestrian trajectory prediction network of the present invention, the pedestrian grouping network is composed of a group allocation module and a pooling module connected in series.
[0059] In step (2) of the present invention, the pedestrian trajectory prediction network is constructed, and the pedestrian trajectory distribution initialization network is composed of a time series feature learning module, a social relationship feature learning module, a feature fusion module, and a noise estimation network. The output ends of the trajectory time series feature learning module and the social relationship feature learning module are connected in series with the feature fusion module and the noise estimation network in turn.
[0060] The group allocation module of the present invention is composed of a convolution layer 1, a Relu layer 1, a normalization layer 1, and a convolution layer 2 connected in series in sequence.
[0061] The convolution kernel size of the convolution layer 1 of the present invention is 5×5 and the step size is 2, and the convolution kernel size of the convolution layer 2 is 3×3 and the step size is 1.
[0062] The temporal feature learning module of the present invention is composed of a multi-head self-attention layer and a fully connected layer 1, a fully connected layer 2, and a normalization layer 2 connected in series in sequence.
[0063] The social feature learning module of the present invention is composed of a graph convolution layer 1, a graph convolution layer 2, and a fully connected layer 3 connected in series in sequence.
[0064] The present invention groups social interactions into models, fully learns the social relationship characteristics of pedestrians, and directly learns the multimodal distribution of future trajectories by constructing a pedestrian trajectory distribution initialization network. This method avoids the use of a standard Gaussian distribution with poor representation ability, thereby enhancing the model's ability to represent the diversity of pedestrian movements. At the same time, the present invention skips a large number of denoising steps, improves the efficiency of the diffusion model in predicting pedestrian trajectories, and achieves accurate and rapid prediction of pedestrian trajectories. The present invention and the prior art were compared in experiments. The experiments show that compared with the comparative experiments, the average offset error and final offset error of the present invention are better than those of the comparative experiments. The present invention has the advantages of high prediction accuracy and fast speed, and can be promoted and used in the field of image processing technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 This is a flow chart of Example 1 of the present invention.
[0066] Figure 2 It is a structural diagram of the pedestrian trajectory prediction network.
[0067] Figure 3 yes Figure 2 Schematic diagram of the structure of the group allocation module.
[0068] Figure 4 yes Figure 2 Schematic diagram of the structure of the temporal feature learning module.
[0069] Figure 5 yes Figure 2 Schematic diagram of the structure of the social feature learning module. DETAILED DESCRIPTION
[0070] The present invention will be further described in detail below in conjunction with the accompanying drawings and examples, but the present invention is not limited to the following embodiments.
[0071] Example 1
[0072] The pedestrian trajectory prediction method based on social relationship group diffusion of this embodiment consists of the following steps (see Figure 1 ):
[0073] (1) Processing the training data set
[0074] The public dataset ETH-UCY contains five different scenes: ZARA1, ZARA2, ETH, HOTEL, and UNIV. The four scenes of ZARA1, ZARA2, ETH, and HOTEL are selected as training sets, and the UNIV scene is selected as the test set. The trajectories are sampled at intervals of 0.4 seconds. The first 3.2 seconds of the trajectory are observation data, and the trajectory after 4.8 seconds is the future trajectory to be predicted. The trajectory coordinate sequence X of pedestrian ii According to formula (1), it can be expressed as:
[0075]
[0076] in, is the world coordinate of the target pedestrian i at time t, i∈{1,2,…,N}, N is the total number of pedestrians in the scene, which is a finite positive integer, and t∈{1,2,…,t obs}, t obs is the observation sequence length, t obs ∈[1,12], t in this embodiment obs The value is 6.
[0077] Determine the pedestrian’s social graph P as follows: G :
[0078] P G =(V p ,E p ),
[0079] V p ={X i},
[0080] E p ={e i,J}, where V p is the vertex set of the graph, representing the trajectories of all pedestrians in the scene, E p is the edge set of the graph, e i,j Represents the social relationship between two people, i,j∈{1,2,…,N}.
[0081] (2) Building a pedestrian trajectory prediction network
[0082] Figure 2 The structural diagram of the pedestrian trajectory prediction network is given. Figure 2 In the embodiment, the pedestrian trajectory prediction network is composed of a pedestrian grouping network, a pedestrian trajectory distribution initialization network, and a trajectory distribution denoising module connected in series.
[0083] The pedestrian grouping network of this embodiment is composed of a group allocation module and a pooling module connected in series.
[0084] The pedestrian trajectory distribution initialization network of this embodiment is composed of a temporal feature learning module, a social relationship feature learning module, a feature fusion module and a noise estimation network; the output ends of the temporal feature learning module and the social feature learning module are connected in series with the feature fusion module and the noise estimation network in sequence.
[0085] Figure 3 Given Figure 2 Schematic diagram of the structure of the group allocation module. Figure 3In the embodiment, the group allocation module is composed of a convolution layer 1, a Relu layer 1, a normalization layer 1, and a convolution layer 2 connected in series. The convolution kernel size of the convolution layer 1 is 5×5 and the step size is 2, and the convolution kernel size of the convolution layer 2 is 3×3 and the step size is 1.
[0086] Figure 4 Given Figure 2 Schematic diagram of the structure of the temporal feature learning module. Figure 4 In the embodiment, the temporal feature learning module is composed of a multi-head self-attention layer and a fully connected layer 1, a fully connected layer 2, and a normalization layer 2 connected in series in sequence.
[0087] Figure 5 Given Figure 2 Schematic diagram of the structure of the social feature learning module. Figure 5 In the embodiment, the social feature learning module is composed of a graph convolution layer 1, a graph convolution layer 2, and a fully connected layer 3 connected in series.
[0088] (3) Adding noise to pedestrian trajectory distribution
[0089] According to formula (2), add noise y to the pedestrian trajectory distribution k :
[0090]
[0091] Among them, α i The parameter that controls the intensity of the added noise, y0 is the actual trajectory of the pedestrian, α start is the initial noise intensity, n is the standard Gaussian noise, α start ∈[0.5,0.9], α in this example start The value is 0.7, K is the total number of diffusion steps, K∈[100,200], and the value of K in this embodiment is 150.
[0092] (4) Training the noise estimation network
[0093] 1) Constructing noise estimation loss function
[0094] The noise estimation loss function L1 is determined as follows:
[0095]
[0096] in, is the prediction noise of the noise estimation network.
[0097] 2) Training the noise estimation network
[0098] The training set is input into the pedestrian trajectory prediction network for training. During the training process, the training parameters are as follows: the initial learning rate is 0.01, and it decays by half every 10 epochs, the training round is 100, the data batch is 64, γ is the number of denoising steps, γ∈[20,50], and the value of γ in this embodiment is 35. The multi-head attention network is set to 8 heads, the number of hidden units in the fully connected layer 1 is 512, the number of hidden units in the fully connected layer 2 is 1024, and the number of hidden units in the fully connected layer 3 is 2048;
[0099] (5) Training pedestrian trajectory prediction network
[0100] 1) Construct pedestrian trajectory loss function
[0101] The pedestrian trajectory loss function L2 is determined as follows:
[0102]
[0103] Among them, Y is the actual trajectory of the pedestrian, is the predicted result of pedestrian trajectory;
[0104] 2) Training human trajectory prediction network
[0105] The training set is input into the pedestrian trajectory prediction network for training. During the training process, the training parameters are: learning rate is 0.0015, training rounds are 400, data batch is 64, multi-head attention network is set to 8 heads, the number of hidden units in the fully connected layer 1 is 512, the number of hidden units in the fully connected layer 2 is 1024, and the number of hidden units in the fully connected layer 3 is 2048;
[0106] (6) Testing the Pedestrian Trajectory Prediction Network
[0107] The test set is input into the trained pedestrian trajectory prediction network for testing, and the trajectory coordinate sequence X of pedestrian i is i Input data and learn the group feature G of pedestrian i as follows:
[0108] G=F g (X i ), where F g (·) Pedestrian grouping network, learn X as follows i The time series characteristic information
[0109]
[0110] Among them, F t (·) is the temporal feature learning module, which transforms the pedestrian’s social graph P GThe group features G of pedestrians are used as the input of the social relationship feature learning module, and the social relationship feature information of pedestrian i is learned as follows
[0111]
[0112] Among them, F s (·) is the social relationship feature learning module. and As the input of the feature fusion module, the mean μ of the trajectory distribution is calculated as follows: θ , standard deviation σ θ , predict sample S θ :
[0113]
[0114] in, is the mean gated recurrent unit, is the standard deviation of the gated recurrent unit, is the gated recurrent unit of the predicted sample, F fusion (·) is the feature fusion module, and the initial distribution of pedestrian trajectories is determined by the following formula and prediction noise
[0115]
[0116] Among them, F n (·) is the noise estimation network, γ is the current denoising step, and the denoising process of the γ-times standard diffusion model is performed according to formula (3) d (·):
[0117]
[0118]
[0119] Where z is a random variable with standard Gaussian distribution, α k is a parameter that controls the intensity of noise subtraction, α k ∈[0.7,0.9], α in this example k The value is 0.8, n γ yes The noise that needs to be removed, is the result of the previous denoising process. Formula (3) is executed γ times repeatedly until the pedestrian trajectory prediction result is predicted. The predicted trajectory of the pedestrian is determined according to formula (4).
[0120]
[0121] in, For the target pedestrian i in the future t p The predicted value of the world coordinates at time t p ∈{t obs ,t obs +1,…,t obs +t pred}, t pred is the predicted trajectory step length, t pred ∈[8,12], t in this embodiment pred The value is 10, and each step is 0.48 seconds. The actual trajectory Y of the pedestrian is determined as follows:
[0122]
[0123] in, For the target pedestrian i in the future t p The true value of the world coordinates at the moment is determined by the average offset error ADE as follows:
[0124]
[0125] The final offset error FDE is determined as follows:
[0126]
[0127] in, The final coordinates of the pedestrian trajectory prediction sequence, The final coordinates of the pedestrian trajectory observation sequence are output, and the pedestrian trajectory test results are output.
[0128] Complete the pedestrian trajectory prediction method based on social relationship group diffusion.
[0129] Example 2
[0130] The pedestrian trajectory prediction method based on social relationship group diffusion of this embodiment consists of the following steps.
[0131] (1) Processing the training data set
[0132] The public dataset ETH-UCY contains five different scenes: ZARA1, ZARA2, ETH, HOTEL, and UNIV. The four scenes of ZARA1, ZARA2, ETH, and HOTEL are selected as training sets, and the UNIV scene is selected as the test set. The trajectories are sampled at intervals of 0.4 seconds. The first 3.2 seconds of the trajectory are observation data, and the trajectory after 4.8 seconds is the future trajectory to be predicted. The trajectory coordinate sequence X of pedestrian i i According to formula (1):
[0133] The expression of formula (1) is the same as that of Example 1.
[0134] In formula (1), is the world coordinate of the target pedestrian i at time t, i∈{1,2,…,N}, N is the total number of pedestrians in the scene, which is a finite positive integer, and t∈{1,2,…,t obs}, t obs is the observation sequence length, t obs ∈[1,12], t in this embodiment obs The value is 1. The meanings and value ranges of other parameters and variables are the same as those in Example 1.
[0135] The other steps of this step are the same as those in Example 1.
[0136] (2) Building a pedestrian trajectory prediction network
[0137] This step is the same as Example 1
[0138] (3) Noise the pedestrian trajectory distribution
[0139] According to formula (2), add noise y to the pedestrian trajectory distribution k :
[0140] The expression of formula (2) is the same as that of Example 1.
[0141] In formula (2), α start is the initial noise intensity, α start ∈[0.5,0.9], α in this example start The value is 0.5, K is the total number of diffusion steps, K∈[100,200], and the value of K in this embodiment is 100. The meanings and value ranges of other parameters and variables are the same as those in Embodiment 1.
[0142] This step is the same as in Example 1.
[0143] (4) Training the noise estimation network
[0144] 1) Constructing noise estimation loss function
[0145] This step is the same as in Example 1.
[0146] 2) Training the noise estimation network
[0147] In this step, the training parameter γ is the number of denoising steps, γ∈[20,50], the value of γ in this embodiment is 20, and the meanings and value ranges of other parameters and variables are the same as those in Example 1.
[0148] (5) Training pedestrian trajectory prediction network
[0149] This step is the same as in Example 1.
[0150] (6) Testing the Pedestrian Trajectory Prediction Network
[0151] In this step, the expression of formula (3) is the same as that of Example 1.
[0152] In formula (3), z is a random variable with standard Gaussian distribution, α k is a parameter that controls the intensity of noise subtraction, α k ∈[0.7,0.9], α in this example k The value is 0.7. The meanings and value ranges of other parameters and variables are the same as those in Example 1.
[0153] The expression of formula (4) is the same as that of Example 1.
[0154] In formula (4), For the target pedestrian i in the future t p The predicted value of the world coordinates at time t p ∈{t obs ,t obs +1,…,t obs +t pred}, t pred is the predicted trajectory step length, t pred ∈[8,12], t in this embodiment pred The value is 8, and each step is 0.48 seconds.
[0155] The other steps of this step are the same as those in Example 1.
[0156] Complete the pedestrian trajectory prediction method based on social relationship group diffusion.
[0157] Example 3
[0158] The pedestrian trajectory prediction method based on social relationship group diffusion of this embodiment consists of the following steps.
[0159] (1) Processing the training data set
[0160] The public dataset ETH-UCY contains five different scenes: ZARA1, ZARA2, ETH, HOTEL, and UNIV. The four scenes of ZARA1, ZARA2, ETH, and HOTEL are selected as training sets, and the UNIV scene is selected as the test set. The trajectories are sampled at intervals of 0.4 seconds. The first 3.2 seconds of the trajectory are observation data, and the trajectory after 4.8 seconds is the future trajectory to be predicted. The trajectory coordinate sequence X of pedestrian i i According to formula (1):
[0161] The expression of formula (1) is the same as that of Example 1.
[0162] In formula (1), is the world coordinate of the target pedestrian i at time t, i∈{1,2,…,N}, N is the total number of pedestrians in the scene, which is a finite positive integer, and t∈{1,2,…,t obs}, t obs is the observation sequence length, t obs ∈[1,12], t in this embodiment obs The value is 12. The meanings and value ranges of other parameters and variables are the same as those in Example 1.
[0163] The other steps of this step are the same as those in Example 1.
[0164] (2) Building a pedestrian trajectory prediction network
[0165] This step is the same as Example 1
[0166] (3) Noise the pedestrian trajectory distribution
[0167] According to formula (2), add noise y to the pedestrian trajectory distribution k :
[0168] The expression of formula (2) is the same as that of Example 1.
[0169] In formula (2), α start is the initial noise intensity, α start ∈[0.5,0.9], α in this example start The value is 0.9, K is the total number of diffusion steps, K∈[100,200], and the value of K in this embodiment is 200. The meanings and value ranges of other parameters and variables are the same as those in Embodiment 1.
[0170] This step is the same as in Example 1.
[0171] (4) Training the noise estimation network
[0172] 1) Constructing noise estimation loss function
[0173] This step is the same as in Example 1.
[0174] 2) Training the noise estimation network
[0175] In this step, the training parameter γ is the number of denoising steps, γ∈[20,50], the value of γ in this embodiment is 50, and the meanings and value ranges of other parameters and variables are the same as those in Example 1.
[0176] (5) Training pedestrian trajectory prediction network
[0177] This step is the same as in Example 1.
[0178] (6) Testing the Pedestrian Trajectory Prediction Network
[0179] In this step, the expression of formula (3) is the same as that of Example 1.
[0180] In formula (3), z is a random variable with standard Gaussian distribution, α k is a parameter that controls the intensity of noise subtraction, α k ∈[0.7,0.9], α in this example k The value is 0.9. The meanings and value ranges of other parameters and variables are the same as those in Example 1.
[0181] The expression of formula (4) is the same as that of Example 1.
[0182] In formula (4), For the target pedestrian i in the future t p The predicted value of the world coordinates at time t p ∈{t obs ,t obs +1,…,t obs +t pred}, t pred is the predicted trajectory step length, t pred ∈[8,12], t in this embodiment pred The value is 12, and each step is 0.48 seconds.
[0183] The other steps of this step are the same as those in Example 1.
[0184] Complete the pedestrian trajectory prediction method based on social relationship group diffusion.
[0185] In order to verify the beneficial effects of the present invention, the pedestrian trajectory prediction method based on social group diffusion in Example 1 of the present invention (hereinafter referred to as Example 1) is compared with “Qi Zhu, Siheng Chen, Yanfeng Wang, Weibo Mao, Chenxin Xu. Leapfrog diffusion model for stochastic trajectory prediction. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 5517–5526, 2023.” (hereinafter referred to as Comparative Experiment 1), “Zhenyang Ni, Ya Zhang, Chenxin Xu, Maosen Li and Siheng Chen. Groupnet: Multiscale hypergraph neural networks for trajectory prediction with relational reasoning. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 6498–6507, 2022.” (hereinafter referred to as Comparative Experiment 2), “Inhwan Bae, Jin-Hwi Park, Hae-Gon Jeon. Non-Probability Sampling Network for Stochastic Human Trajectory Prediction. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 6477-6487, 2022. "(referred to as comparative experiment 3) was conducted a comparative experiment, and the evaluation indicators used the average offset error ADE and the final offset error FDE.
[0186] The experimental results are shown in Table 1.
[0187] Table 1 Experimental results of Example 1 method and comparative experiment
[0188] Experimental Group Final offset error FDE Average offset error ADE Comparative experiment 1 0.33 0.21 Comparative experiment 2 0.44 0.25 Comparative experiment 3 0.38 0.21 Example 1 0.28 0.18
[0189] As can be seen from Table 1, compared with comparative experiments 1-3, the performance of Example 1 is greatly improved. The final offset error and average offset error of Example 1 are reduced by 15% and 14% respectively compared with comparative experiment 1, 36% and 28% respectively compared with comparative experiment 2, and 26% and 14% respectively compared with comparative experiment 3. The above experiments show that the present invention can better predict pedestrian trajectories.
Claims
1. A pedestrian trajectory prediction method based on social relationship group diffusion, characterized by It consists of the following steps: (1) Processing the training data set The public dataset ETH-UCY contains five different scenes: ZARA1, ZARA2, ETH, HOTEL, and UNIV. The four scenes of ZARA1, ZARA2, ETH, and HOTEL are selected as training sets, and the UNIV scene is selected as the test set. The trajectories are sampled at intervals of 0.4 seconds. The first 3.2 seconds of the trajectory are observation data, and the trajectory after 4.8 seconds is the future trajectory to be predicted. The trajectory coordinate sequence X of pedestrian i i According to formula (1), it can be expressed as: in, is the world coordinate of the target pedestrian i at time t, i∈{1,2,…,N}, N is the total number of pedestrians in the scene, which is a finite positive integer. obs }, t obs is the observation sequence length, t obs ∈[1,12]; Determine the pedestrian’s social graph P as follows: G : P G =(V p ,E p ), V p ={X i }, AND p ={and i,j }, Among them, V p is the vertex set of the graph, representing the trajectories of all pedestrians in the scene, E p is the edge set of the graph, e i,j Represents the social relationship between two people, i,j∈{1,2,…,N}; (2) Building a pedestrian trajectory prediction network The pedestrian trajectory prediction network is composed of a pedestrian grouping network, a pedestrian trajectory distribution initialization network, and a trajectory distribution denoising module connected in series. (3) Adding noise to pedestrian trajectory distribution According to formula (2), add noise y to the pedestrian trajectory distribution k : Among them, α i is a parameter that controls the intensity of added noise, y0 is the actual trajectory of the pedestrian, n is the standard Gaussian noise, α start is the initial noise intensity, α start ∈[0.5,0.9], K is the total number of diffusion steps, K∈[100,200]; (4) Training the noise estimation network 1) Constructing noise estimation loss function The noise estimation loss function L1 is determined as follows: in, is the prediction noise of the noise estimation network; 2) Training the noise estimation network The training set is input into the pedestrian trajectory prediction network for training. During the training process, the training parameters are as follows: the initial learning rate is 0.01, and it decays by half every 10 epochs, the training rounds are 100, the data batch is 64, the number of denoising steps γ∈[20,50], the multi-head attention network is set to 8 heads, the number of hidden units in the fully connected layer 1 is 512, the number of hidden units in the fully connected layer 2 is 1024, and the number of hidden units in the fully connected layer 3 is 2048; (5) Training pedestrian trajectory prediction network 1) Construct pedestrian trajectory loss function The pedestrian trajectory loss function L2 is determined as follows: Among them, Y is the actual trajectory of the pedestrian, is the predicted result of pedestrian trajectory; 2) Training human trajectory prediction network The training set is input into the pedestrian trajectory prediction network for training. During the training process, the training parameters are: learning rate is 0.0015, training rounds are 400, data batch is 64, multi-head attention network is set to 8 heads, the number of hidden units in the fully connected layer 1 is 512, the number of hidden units in the fully connected layer 2 is 1024, and the number of hidden units in the fully connected layer 3 is 2048; (6) Testing the Pedestrian Trajectory Prediction Network The test set is input into the trained pedestrian trajectory prediction network for testing, and the trajectory coordinate sequence X of pedestrian i is i Input data and learn the group feature G of pedestrian i as follows: G=F g (X i ), Among them, F g (·) Pedestrian grouping network, learn X as follows i The time series characteristic information Among them, F t (·) is the temporal feature learning module, which transforms the pedestrian’s social graph P G The group features G of pedestrians are used as the input of the social relationship feature learning module, and the social relationship feature information of pedestrian i is learned as follows Among them, F s (·) is the social relationship feature learning module. and As the input of the feature fusion module, the mean μ of the trajectory distribution is calculated as follows: θ , standard deviation σ θ , predict sample S θ : in, is the mean gated recurrent unit, is the standard deviation of the gated recurrent unit, is the gated recurrent unit of the predicted sample, F fusion (·) is the feature fusion module, and the initial distribution of pedestrian trajectories is determined by the following formula and prediction noise Among them, F n (·) is the noise estimation network, γ is the current denoising step, and the denoising process of the γ-times standard diffusion model is performed according to formula (3) d (·): Among them, z is a random variable with standard Gaussian distribution, α k is a parameter that controls the intensity of noise subtraction, α k ∈[0.7,0.9], n γ yes The noise that needs to be removed, is the result of the previous denoising process. Formula (3) is executed γ times repeatedly until the pedestrian trajectory prediction result is predicted. The predicted trajectory of the pedestrian is determined according to formula (4). in, For the target pedestrian i in the future t p The predicted value of the world coordinate at time t p ∈{t obs ,t obs +1,…,t obs +t pred }, t pred is the predicted trajectory step length, t pred ∈[8,12], each step is 0.48 seconds; the actual trajectory Y of the pedestrian is determined by the following formula: in, For the target pedestrian i in the future t p The true value of the world coordinates at the moment is determined by the average offset error ADE as follows: The final offset error FDE is determined as follows: in, The final coordinates of the pedestrian trajectory prediction sequence, The final coordinates of the pedestrian trajectory observation sequence are output, and the pedestrian trajectory test results are output.
2. The pedestrian trajectory prediction method based on social relationship group diffusion according to claim 1 is characterized by: In step (2) of constructing a pedestrian trajectory prediction network, the pedestrian grouping network is composed of a group allocation module and a pooling module connected in series.
3. The pedestrian trajectory prediction method based on social relationship group diffusion according to claim 1 is characterized by: In step (2) of constructing a pedestrian trajectory prediction network, the pedestrian trajectory distribution initialization network is composed of a time series feature learning module, a social relationship feature learning module, a feature fusion module, and a noise estimation network. The output ends of the trajectory time series feature learning module and the social relationship feature learning module are connected in series with the feature fusion module and the noise estimation network in turn.
4. The pedestrian trajectory prediction method based on social relationship group diffusion according to claim 2 is characterized by: The group allocation module is composed of a convolutional layer 1, a Relu layer 1, a normalization layer 1, and a convolutional layer 2 connected in series in sequence.
5. The pedestrian trajectory prediction method based on social relationship group diffusion according to claim 4 is characterized by: The convolution kernel size of the convolution layer 1 is 5×5 and the step size is 2, and the convolution kernel size of the convolution layer 2 is 3×3 and the step size is 1.
6. The pedestrian trajectory prediction method based on social relationship group diffusion according to claim 3 is characterized by: The temporal feature learning module is composed of a multi-head self-attention layer and a fully connected layer 1, a fully connected layer 2, and a normalization layer 2 connected in series in sequence.
7. The pedestrian trajectory prediction method based on social relationship group diffusion according to claim 3 is characterized by: The social feature learning module is composed of a graph convolution layer 1, a graph convolution layer 2, and a fully connected layer 3 connected in series.
Citation Information
Cited By
Trajectory prediction planning method, trajectory prediction planning model training method, trajectory prediction planning model training device, medium and equipment
CN121438254A
Water and soil pollution remediation result prediction model training method and device, equipment and medium
CN122412966A