Pedestrian trajectory prediction method based on space-time interactive perception

By constructing time and space graphs, extracting pedestrian features and learning spatiotemporal correlations, and combining them with a full-dimensional dynamic temporal convolutional network, the problem of the failure of existing technologies to effectively combine spatiotemporal interaction perception is solved, and more accurate pedestrian trajectory prediction is achieved.

CN120635495APending Publication Date: 2025-09-12GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510847644.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing pedestrian trajectory prediction methods fail to effectively combine spatiotemporal interactive perception, resulting in the inability to accurately predict pedestrian motion trajectories.

Method used

A method based on spatiotemporal interaction perception is adopted to extract the temporal and spatial features between pedestrians by constructing time graphs and space graphs. The spatiotemporal interaction perception module is used to learn the spatiotemporal correlation of pedestrian movements, and the trajectory prediction is performed in combination with a full-dimensional dynamic temporal convolutional network.

Benefits of technology

It achieves more realistic and accurate pedestrian trajectory prediction, and improves the accuracy and robustness of trajectory prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635495A_ABST
    Figure CN120635495A_ABST
Patent Text Reader

Abstract

The invention discloses a pedestrian trajectory prediction method based on space-time interaction perception, and aims to solve the problem of complex space-time interaction behaviors of pedestrians and improve trajectory prediction precision. The method comprises the following steps: firstly, processing a pedestrian trajectory data set, and constructing a time diagram and a space diagram as input; extracting space and time features among pedestrians through a self-attention mechanism, and inputting the space and time features into a space-time interaction perception module to learn a dynamic coupling relationship between motion and interaction to obtain interaction perception features; and then inputting the interactive features into a graph convolutional network to obtain trajectory representation features, finally predicting two-dimensional Gaussian distribution of future trajectories through a full-dimensional dynamic time sequence convolutional network, and outputting a high-precision trajectory prediction result. According to the invention, through a spatio-temporal joint modeling framework, spatio-temporal feature dynamic fusion of pedestrian motion is realized, and the accuracy and robustness of pedestrian trajectory prediction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pedestrian trajectory prediction in autonomous driving, and in particular to a pedestrian trajectory prediction method based on spatiotemporal interactive perception. Background Art

[0002] Pedestrian trajectory prediction is a key technology that aims to accurately predict a pedestrian's future coordinate sequence based on their historical trajectory and current state. This technology has broad application prospects in intelligent monitoring systems and autonomous driving. In intelligent monitoring systems, pedestrian trajectory prediction can be used to detect anomalies and help identify potential safety issues. In autonomous driving, pedestrian trajectory prediction provides critical pedestrian movement information, providing essential reference for vehicle path planning and driving safety.

[0003] Current pedestrian trajectory prediction methods can be broadly categorized into two main categories: traditional model-based and data-driven. Traditional pedestrian trajectory prediction methods use rules and handcrafted features to model pedestrian behavior. These models estimate changes in walking motion states and ultimately derive walking path predictions. However, pedestrians are highly mobile and their movements are highly variable, making them difficult to represent using handcrafted rules. In recent years, the development of deep learning technologies, particularly the introduction of spatiotemporal graph neural networks, spatiotemporal transformers, and generative adversarial networks, has led to significant progress in pedestrian trajectory prediction. Data-driven methods automatically extract spatiotemporal features from large-scale trajectory data through end-to-end learning, effectively capturing the key characteristics of pedestrian motion. However, most current pedestrian trajectory prediction methods independently address the interactions between pedestrians in the spatiotemporal dimensions, rather than capturing the correlation between the two through joint spatiotemporal modeling. This makes it difficult to accurately predict pedestrian trajectories. The key challenge in improving trajectory prediction accuracy is how to combine spatiotemporal interaction perception modules for trajectory prediction and achieve dynamic fusion of motion features. Summary of the Invention

[0004] In response to the shortcomings of the existing technology, the present invention proposes a pedestrian trajectory prediction method based on spatiotemporal interactive perception, which considers the impact of spatiotemporal correlation in pedestrian motion on trajectory prediction results, and achieves more realistic and accurate trajectory prediction.

[0005] The present invention provides the following technical solution: a pedestrian trajectory prediction method based on spatiotemporal interactive perception, which is implemented by the following steps:

[0006] Step 1: Process the pedestrian trajectory prediction dataset;

[0007] Step 2: Use pedestrian trajectories to construct time graphs and spatial graphs;

[0008] Step 3: The time graph and space graph are used in a feature extraction module to obtain the temporal and spatial features between pedestrians;

[0009] Step 4: Input the obtained temporal and spatial features into the spatiotemporal interaction perception module to learn the spatiotemporal correlation of pedestrian motion and obtain interactive perception features;

[0010] Step 5: Input the interaction perception features into graph convolution to obtain trajectory representation features;

[0011] Step 6: Use the full-dimensional dynamic temporal convolutional network to predict the two-dimensional Gaussian distribution of future trajectories.

[0012] Preferably, in step 1, the public data sets ETH and UCY are processed. ETH contains two scenes, ETH and HOTEL, and UCY contains three scenes, ZARA1, ZARA2, and UNIV. The original data is stored in 8 independent text files, and the data is distributed using a 5-fold cross-validation strategy. One of the files is selected as the test set for each experiment, and the remaining data is divided into a training set and a validation set. A fixed-frequency sampling mechanism is used, and a frame of trajectory coordinate points is intercepted every 0.4 seconds. The model input is a continuous 3.2-second observation sequence, corresponding to 8 frames, and the output is a predicted trajectory for the subsequent 4.8 seconds, corresponding to 12 sampling frames.

[0013] Preferably, in step 2, a graph structure modeling is performed on the trajectory in the time dimension and the space dimension, where pedestrians are represented as vertices, the position coordinates of pedestrians are used as node attributes, and social interactions are used as edges in the graph. The specific operations are:

[0014] Given an input trajectory Where D represents the dimension of the spatial coordinates, construct a spatial graph and a time graph. The spatial graph G at time step t s Represents the location of the pedestrian, the time graph G at time step n t Represents the trajectory of the corresponding pedestrian. The definitions of the spatial graph and the temporal graph are respectively expressed as:

[0015] G s =(V t ,U t )

[0016] G t =(V n ,U n )

[0017] in, and Represents G s and G t Node, It is represented by the coordinates (x′) of the nth pedestrian at time step t n ,y′n ). and Represents G s and G t edge.

[0018] Preferably, in step 3, the constructed time graph and space graph are used to extract the temporal interaction features R between pedestrians through the self-attention mechanism. t and spatial interaction features R s Since the same structure is used when processing spatial graphs and time graphs, the processing of spatial graphs is explained one by one. The specific implementation is as follows:

[0019]

[0020]

[0021] Where φ represents the nonlinear transformation, E s represents graph embedding, Q s and K s Represents the query and key of the self-attention mechanism in the spatial dimension, represents the weight of the linear transformation, Expressed as a scaling factor to ensure numerical stability.

[0022] Preferably, in step 4, the spatial interaction feature R s and time interaction feature R t The specific steps of outputting the spatiotemporal interaction perception module to obtain the interaction perception features are as follows:

[0023] (1) First, use maximum pooling and average pooling to process the spatial interaction feature R s and time interaction feature R t , obtain the time feature vector VR t and spatial eigenvector VR s , as shown in the equation

[0024] VR t =AvgPooling(R t ) * λ+MaxPooling(R t ) * (1-λ)

[0025] VR s =MaxPooling(R s ) * λ+MaxPooling(R s ) * (1-λ)

[0026] Max Pooling represents maximum pooling, Avg Pooling represents average pooling, and λ represents a weight between 0 and 1, which is a learnable parameter.

[0027] (2) The attention score matrix is ​​calculated through the cross attention mechanism, which is expressed as:

[0028]

[0029] in is the weight matrix.

[0030] Secondly, the correlation matrix is ​​built through the dot product operation, and then the correlation score is normalized by the softmax function. Then, the correlation matrix is ​​multiplied by the vector to obtain the vector, which is expressed as:

[0031]

[0032] By nonlinear transformation, the vector Z t It is reprojected back to the original space and added to the input sequence through a residual connection. A learnable coefficient is applied to each branch of the residual connection of the equation to adaptively learn data from different branches to achieve performance gains. Furthermore, a feedforward network with two fully connected layers is used to further refine the global information to improve the robustness and accuracy of the model, and the enhanced feature representation is output as:

[0033] VR′ t =α·Z t W o +β·(VR t )

[0034]

[0035] in represents the output weight matrix before the FFN layer, and α, β, γ, and δ are learnable parameters initialized to 1 during training.

[0036] The depth of the network is deepened through multiple iterations, and the complementary information of spatiotemporal information is gradually refined.

[0037]

[0038] in, Represents the spatiotemporal interactive perception module. {(VR t ),(VR s )} represents the input of the spatiotemporal interaction perception module, It represents the output after n iterative operations, and the output of each iterative operation serves as the input for the next round.

[0039] (3) Then the above Fusion is performed to obtain interactive perception feature A c .

[0040]

[0041] Preferably, in step five, the spatial graph G s , time graph G t and interactive perception feature A c The trajectory feature representation is obtained through graph convolution operation, which is specifically expressed as follows:

[0042] H c =δ(A c C s W Gs )+δ(A c G t W Gt )

[0043] Among them, δ represents the PReLU activation function, W Gs , W Gt is the weight of the graph convolution.

[0044] Preferably, in step 6, the designed full-dimensional dynamic temporal convolution module includes a layer of full-dimensional dynamic convolution and four convolution blocks. The full-dimensional dynamic convolution is the input end, which is used to receive the feature representation obtained in the previous step as input, and then sequentially processed by the four convolution blocks, and finally output the predicted features.

[0045] The bivariate Gaussian distribution parameters of trajectory prediction are obtained through the full-dimensional dynamic temporal convolution module, and the network model is trained by minimizing the negative log-likelihood loss function, which is expressed as:

[0046]

[0047] in represents the average value, represents the standard deviation, represents the correlation coefficient of the prediction.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] 1. Pedestrians not only move independently in time but also interact with each other in space. In multi-agent scenarios, spatial and temporal information are deeply coupled. This paper introduces a spatiotemporal interaction perception module, combined with temporal and spatial feature extraction modules, to learn the complex spatiotemporal characteristics of pedestrian trajectories, achieving more realistic trajectory prediction.

[0050] 2. Based on the structure of the time-extrapolated convolutional network, this paper designs a full-dimensional dynamic temporal convolutional network. The full-dimensional dynamic temporal convolutional network includes a full-dimensional attention mechanism and a complementary attention mechanism. It can flexibly adjust the weights of the convolution kernel according to the dynamic changes of input features, adaptively extract more representative contextual information, and achieve accurate modeling of complex spatiotemporal information. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for describing the embodiments or the prior art.

[0052] Figure 1 3 is a flow chart of a pedestrian trajectory prediction method based on spatiotemporal interactive perception in an embodiment of the present invention.

[0053] Figure 2 It is a structural block diagram of a pedestrian trajectory prediction method based on spatiotemporal interactive perception in an embodiment of the present invention.

[0054] Figure 3 It is a full-dimensional dynamic temporal convolution module of the pedestrian trajectory prediction method based on spatiotemporal interactive perception in an embodiment of the present invention. DETAILED DESCRIPTION

[0055] The present invention will be further described below with reference to the accompanying drawings and embodiments, but the present invention is not limited thereto.

[0056] See also Figure 1 and Figure 2 The present invention provides a technical solution: a pedestrian trajectory prediction method based on spatiotemporal interactive perception, the method comprising the following steps:

[0057] Step 1: Process the pedestrian trajectory prediction dataset;

[0058] Step 2: Use pedestrian trajectories to construct time graphs and spatial graphs;

[0059] Step 3: The time graph and space graph are used in a feature extraction module to obtain the temporal and spatial features between pedestrians;

[0060] Step 4: Input the obtained temporal and spatial features into the spatiotemporal interaction perception module to learn the spatiotemporal correlation of pedestrian motion and obtain interactive perception features;

[0061] Step 5: Input the interaction perception features into graph convolution to obtain trajectory representation features;

[0062] Step 6: Use the full-dimensional dynamic temporal convolutional network to predict the two-dimensional Gaussian distribution of future trajectories.

[0063] In step 1, the public datasets ETH and UCY are selected and processed. ETH contains two scenes, ETH and HOTEL, and UCY contains three scenes, ZARA1, ZARA2, and UNIV. The original data is stored in eight independent text files and distributed using a five-fold cross-validation strategy. One of the files is selected as the test set for each experiment, and the remaining data is divided into a training set and a validation set. A fixed-frequency sampling mechanism is used, capturing a frame of trajectory coordinate points every 0.4 seconds. The model input is a continuous 3.2-second observation sequence, corresponding to 8 frames, and the output is a predicted trajectory for the subsequent 4.8 seconds, corresponding to 12 sampling frames.

[0064] In the step 2, the specific steps are as follows:

[0065] Given an input trajectory Where D represents the dimension of the spatial coordinates, construct a spatial graph and a time graph. The spatial graph G at time step t s Represents the location of the pedestrian, the time graph G at time step n t Represents the trajectory of the corresponding pedestrian. The definitions of the spatial graph and the temporal graph are respectively expressed as:

[0066] G s =(V t ,U t )

[0067] G t =(V n ,U n )

[0068] in, and Represents G s and G t Node, It is represented by the coordinates (x′) of the nth pedestrian at time step t n ,y′ n ). and Represents G s and G t edge.

[0069] In the step three, specifically:

[0070] The constructed time graph and space graph are used to extract the temporal interaction features R between pedestrians through the self-attention mechanism. t and spatial interaction features R s Since the same structure is used when processing spatial graphs and time graphs, the processing of spatial graphs is explained one by one. The specific implementation is as follows:

[0071]

[0072] Where φ represents the nonlinear transformation, E s represents graph embedding, Q s and K s Represented as Query and Key of the self-attention mechanism in the spatial dimension, represents the weight of the linear transformation, Expressed as a scaling factor to ensure numerical stability.

[0073] The step 4 is specifically as follows:

[0074] (1) First, use maximum pooling and average pooling to process the spatial interaction feature R s and time interaction feature R t , obtain the time feature vector VR t and spatial eigenvector VR s , as shown in the equation

[0075] VR t =AvgPooling(R t ) * λ+MaxPooling(R t ) * (1-λ)

[0076] VR s =MaxPooling(Rx) * λ+MaxPooling(R s ) * (1-λ)

[0077] Max Pooling represents maximum pooling, Avg Pooling represents average pooling, and λ represents a weight between 0 and 1, which is a learnable parameter.

[0078] (2) The attention score matrix is ​​calculated through the cross attention mechanism, which is expressed as:

[0079]

[0080] in is the weight matrix.

[0081] Secondly, the correlation matrix is ​​built through the dot product operation, and then the correlation score is normalized by the softmax function. Then, the correlation matrix is ​​multiplied by the vector to obtain the vector, which is expressed as:

[0082]

[0083] By nonlinear transformation, the vector Z tIt is reprojected back to the original space and added to the input sequence through a residual connection. A learnable coefficient is applied to each branch of the residual connection of the equation to adaptively learn data from different branches to achieve performance gains. Furthermore, a feedforward network with two fully connected layers is used to further refine the global information to improve the robustness and accuracy of the model, and the enhanced feature representation is output as:

[0084] VR′ t =α·Z t W o +β·(VR t )

[0085]

[0086] in represents the output weight matrix before the FFN layer, and α, β, γ, and δ are learnable parameters initialized to 1 during training.

[0087] The depth of the network is deepened through multiple iterations, and the complementary information of spatiotemporal information is gradually refined.

[0088]

[0089] in, Represents the spatiotemporal interactive perception module. {(VR t ),(VR s )} represents the input of the spatiotemporal interaction perception module, It represents the output after n iterative operations, and the output of each iterative operation serves as the input for the next round.

[0090] (3) Then the above Fusion is performed to output interactive perception feature A c .

[0091]

[0092] The step five is specifically as follows:

[0093] The spatial graph G s , time graph G t and interactive perception feature A c The trajectory feature representation is obtained through graph convolution operation, and the calculation method is shown in the equation:

[0094] H c =δ(A c G s W Gs )+δ(A c G t W Gt )

[0095] Among them, δ represents the PReLu activation function, W Gs , W Gt is the weight of the graph convolution.

[0096] In step 6, a full-dimensional dynamic temporal convolution module is constructed, such as Figure 3 As shown, specifically:

[0097] The full-dimensional dynamic temporal convolution module consists of a full-dimensional dynamic convolution layer and four convolution blocks. The full-dimensional dynamic convolution layer is the input, which receives the feature representation obtained in the previous step as input. It is then processed by the four convolution blocks in sequence and finally outputs the predicted features.

[0098] The bivariate Gaussian distribution parameters of trajectory prediction are obtained through the full-dimensional dynamic temporal convolution module, and the network model is trained by minimizing the negative log-likelihood loss function, which is expressed as:

[0099]

[0100] in represents the average value, represents the standard deviation, represents the correlation coefficient of the prediction.

Claims

1. A pedestrian trajectory prediction method based on spatiotemporal interactive perception, characterized in that: The steps include: Step 1: Process the pedestrian trajectory prediction dataset; Step 2: Use pedestrian trajectories to construct time graphs and spatial graphs; Step 3: The time graph and space graph are used in a feature extraction module to obtain the temporal and spatial features between pedestrians; Step 4: Input the obtained temporal and spatial features into the spatiotemporal interaction perception module to learn the spatiotemporal correlation of pedestrian motion and obtain interactive perception features; Step 5: Input the interaction perception features into graph convolution to obtain trajectory representation features; Step 6: Use the full-dimensional dynamic temporal convolutional network to predict the two-dimensional Gaussian distribution of future trajectories.

2. The pedestrian trajectory prediction method based on spatiotemporal interactive perception according to claim 1, characterized in that: In step 1, the public datasets ETH and UCY were processed. ETH contains two scenes, ETH and HOTEL, and UCY contains three scenes, ZARA1, ZARA2, and UNIV. The raw data was stored in eight independent text files. Data was distributed using a 50-fold cross-validation strategy, with one copy selected as the test set for each experiment, and the remaining data divided into training and validation sets. A fixed-frequency sampling mechanism was used, capturing one frame of trajectory coordinates every 0.4 seconds. The model input was a continuous 3.2-second observation sequence, corresponding to 8 frames, and the output was a predicted trajectory for the subsequent 4.8 seconds, corresponding to 12 sampled frames.

3. The pedestrian trajectory prediction method based on spatiotemporal interactive perception according to claim 1, characterized in that: In step 2, the trajectory is modeled as a graph structure in the time and space dimensions, with pedestrians representing vertices, their position coordinates as node attributes, and social interactions as edges in the graph. The specific operations are: Given an input trajectory Where D represents the dimension of the spatial coordinates, construct a spatial graph and a time graph. The spatial graph G at time step t s Represents the location of the pedestrian, the time graph G at time step n t Represents the trajectory of the corresponding pedestrian. The definitions of the spatial graph and the temporal graph are respectively expressed as: G s =(V t ,U t ) G t =(V n ,U n ) in, and Represents G s and G t Node, It is represented by the coordinates (x′) of the nth pedestrian at time step t n ,y′ n ). and Represents G s and G t edge.

4. The pedestrian trajectory prediction method based on spatiotemporal interactive perception according to claim 1, characterized in that: In step 3, the constructed time graph and space graph are used to extract the temporal interaction features R between pedestrians through the self-attention mechanism. t and spatial interaction features R s Since the same structure is used when processing spatial graphs and time graphs, the processing of spatial graphs is explained one by one. The specific implementation is as follows: Where φ represents the nonlinear transformation, E s represents graph embedding, Q s and K s Represents the query and key of the self-attention mechanism in the spatial dimension, represents the weight of the linear transformation, Expressed as a scaling factor to ensure numerical stability.

5. The pedestrian trajectory prediction method based on spatiotemporal interactive perception according to claim 1, characterized in that: In the step 4, the spatial interaction feature R s and time interaction feature R t The specific steps of outputting the spatiotemporal interaction perception module to obtain the interaction perception features are as follows: (1) First, use maximum pooling and average pooling to process the spatial interaction feature R s and time interaction feature R t , obtain the time feature vector VR t and spatial eigenvector VR s , as shown in the equation VR t = AvgPooling(R t ) * λ+MaxPooling(R t ) * (1-λ) VR s = Avgpooling(R s ) * λ+MaxPooling(R s ) * (1-λ) Max Pooling represents maximum pooling, Avg Pooling represents average pooling, and λ represents a weight between 0 and 1, which is a learnable parameter. (2) The attention score matrix is ​​calculated through the cross attention mechanism, which is expressed as: in is the weight matrix. Secondly, the correlation matrix is ​​built through the dot product operation, and then the correlation score is normalized by the softmax function. Then, the correlation matrix is ​​multiplied by the vector to obtain the vector, which is expressed as: By nonlinear transformation, the vector Z t It is reprojected back to the original space and added to the input sequence through a residual connection. A learnable coefficient is applied to each branch of the residual connection of the equation to adaptively learn data from different branches to achieve performance gains. Furthermore, a feedforward network with two fully connected layers is used to further refine the global information to improve the robustness and accuracy of the model, and the enhanced feature representation is output as: VR′ t =α·Z t W o +β·(VR t ) in represents the output weight matrix before the FFN layer, and α, β, γ, and δ are learnable parameters initialized to 1 during training. The depth of the network is deepened through multiple iterations, and the complementary information of spatiotemporal information is gradually refined. in, Represents the spatiotemporal interactive perception module. {(VR t ),(VR s )} represents the input of the spatiotemporal interaction perception module, It represents the output after n iterative operations, and the output of each iterative operation serves as the input for the next round. (3) Then the above Fusion is performed to obtain interactive perception feature A c .

6. The pedestrian trajectory prediction method based on spatiotemporal interactive perception according to claim 1, characterized in that: In the step 5, the spatial graph G s , time graph G t and interactive perception feature A c The trajectory feature representation is obtained through graph convolution operation, which is specifically expressed as follows: H c =δ(A c C s W Gs )+δ(A c G t W Gt ) Among them, δ represents the PReLU activation function, W Gs , W Gt is the weight of the graph convolution.

7. The pedestrian trajectory prediction method based on spatiotemporal interactive perception according to claim 1, characterized in that: In step 6, the designed full-dimensional dynamic temporal convolution module includes a layer of full-dimensional dynamic convolution and four convolution blocks. The full-dimensional dynamic convolution is the input end, which receives the feature representation obtained in the previous step as input, then processes it through the four convolution blocks in sequence, and finally outputs the predicted features. The bivariate Gaussian distribution parameters of trajectory prediction are obtained through the full-dimensional dynamic temporal convolution module, and the network model is trained by minimizing the negative log-likelihood loss function, which is expressed as: in represents the average value, represents the standard deviation, represents the correlation coefficient of the prediction.

Citation Information

Cited By

  • Multi-data-source driving safety evaluation method and system

    CN121350947A