A Pedestrian Trajectory Prediction Method Based on Sparse Spatiotemporal Graph Transformer Network

Through the method based on the sparse spat space-time graph Transformer network, the spatial interaction and time-dependent modeling of pedestrians is solved, which solves the problem of difficulty in taking into account both spatial and temporal interaction in the prior art, and achieves high accuracy in prediction of long-distance trajectory in a crowded environment and long-distance trajectory.

CN118629006BActive Publication Date: 2025-06-06GUANGZHOU TEAM-E DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410744217.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-11
Publication Date
2025-06-06
Estimated Expiration
2044-06-11

AI Technical Summary

Technical Problem

The existing pedestrian trajectory prediction technology is difficult to take into account the interactive modeling of pedestrian space and time, resulting in insufficient accuracy of long-distance trajectory prediction in crowded environments and long-distance trajectory prediction.

Method used

Using a method based on the sparse spatiotemporal map Transformer network, pedestrians are modeled through spatial Transformer and temporal Transformer, and sparsely map trajectory features using the sparse mapping function entmax15. Finally, a model framework for spatiotemporal blockchain is designed, combining dynamic spatial dependence and long-term time dependence.

Benefits of technology

It effectively improves the accuracy of prediction of long-distance trajectory in crowded environments and long-distance trajectory, captures pedestrian status information in real time, reduces interaction redundancy, and enhances the model's capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118629006B_ABST
    Figure CN118629006B_ABST
Patent Text Reader

Abstract

The present invention provides a pedestrian trajectory prediction method based on a sparse spatiotemporal graph Transformer network, which belongs to the field of computer vision technology. The technical problem of low efficiency in predicting pedestrian trajectories in crowded environments and over long distances is solved. The technical solution is as follows: it includes the following steps: S1: obtaining the position information of pedestrians from image frames; S2: using a dynamic spatial Transformer, and using a multi-head attention mechanism to jointly model multiple spatially dependent models; S3: using a self-attention mechanism to achieve bidirectional temporal dependency modeling across multiple time steps; S4: obtaining a spatiotemporal Transformer network with sparse transformations; S5: designing a model framework of a spatiotemporal block chain based on a spatiotemporal Transformer. The beneficial effects of the present invention are: the present invention improves the accuracy of pedestrian trajectory prediction in crowded environments and over long distances.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a pedestrian trajectory prediction method based on a sparse spatiotemporal graph Transformer network. Background Art

[0002] With the advancement of technology, driverless cars have attracted much attention, but to protect the safety of pedestrians, especially pedestrians on the road, driverless cars need to accurately predict the future trajectory of pedestrians and adjust the vehicle's movement strategy to avoid collisions with pedestrians.

[0003] Affected by the surrounding environment, pedestrians' behavior trajectories are highly uncertain. Modeling complex social interactions, establishing the spatiotemporal relationship between pedestrians and surrounding pedestrians and semantic environment, and allowing unmanned vehicles to learn potential social interactions are the key to accurately predicting pedestrian trajectories.

[0004] The movement of pedestrians is not only affected by the current moment, but also by the trajectories and behaviors in the past. Therefore, by modeling the time dependence of pedestrians, the future movement trajectory of pedestrians can be predicted more accurately, thereby improving the accuracy and reliability of pedestrian trajectory prediction.

[0005] Existing pedestrian trajectory prediction technologies are mostly based on traditional models and cannot simultaneously take into account the spatial and temporal interaction modeling of pedestrians. Summary of the invention

[0006] The purpose of the present invention is to provide a pedestrian trajectory prediction method based on a sparse spatiotemporal graph Transformer network, which can effectively improve the accuracy of pedestrian trajectory prediction in crowded environments and over long distances.

[0007] The inventive idea of ​​the present invention is: the present invention provides a pedestrian trajectory prediction method based on a sparse spatiotemporal graph Transformer network. Firstly, based on the Transformer network framework, a spatial Transformer and a temporal Transformer are proposed, and the spatial interaction and temporal dependency of pedestrians are modeled at the same time; secondly, the use of a sparse mapping function entmax15 applied to the Transformer network is proposed to perform sparse mapping on the trajectory characteristics of pedestrians; finally, an effective spatiotemporal block chain model framework is designed to effectively process the trajectory characteristics of pedestrians. Through the above steps and methods, the accuracy of pedestrian trajectory prediction in crowded environments and over long distances can be effectively improved.

[0008] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is as follows: a pedestrian trajectory prediction method based on a sparse spatiotemporal graph Transformer network, the overall process of the method is as shown in the attached Figure 1The technical solution of the present invention is further described below in conjunction with specific embodiments:

[0009] S1: Obtain the location information of pedestrians from the image frame, perform data preprocessing on the pedestrian trajectory information, obtain the trajectory coordinates of each pedestrian in the current frame, and obtain the trajectory corresponding to all frames of the pedestrian by traversing the image frame.

[0010] S2: Use dynamic spatial Transformer, use self-attention mechanism to dynamically model spatial dependencies, and use multi-head attention mechanism to jointly model multiple models of spatial dependencies. The flowchart and structure diagram of the method are as follows Figure 2 , attached Figure 3 The specific method of this step is:

[0011] S2.1: Dynamic graph convolutional layer.

[0012] The spatial embedding feature projects the input features of each node into a high-dimensional latent subspace by learning a linear mapping, and captures the dynamic spatial dependencies through training and modeling in the high-dimensional latent subspace. First, the input feature X of the sparse spatial Transformer is s After spatial embedding, spatial embedding features are obtained Then embed the features of each time step Projected into a high-dimensional latent subspace. The projection mapping is implemented by a feed-forward neural network consisting of multiple fully connected layers. For each spatial dependency pattern, we embed the feature of each node at each time step Train three latent subspaces, including the query subspace Q s , key subspace K s Sum value subspace V s .

[0013]

[0014] in, Q s , K s , V s The weight matrix of .

[0015] Next, use Q s and K s The dot product of the dynamic spatial dependency S between nodes is calculated s .

[0016]

[0017] The sparse attention function 1.5-entmax is used to normalize the spatial dependencies. Therefore, S s Update the node feature and get the new feature of the node as As , the new feature node passes through the feedforward neural network and learns A s further improve the prediction.

[0018] A s (Q s , K s , V s )=S s V s

[0019] S2.2: Extracting interactions between pedestrians.

[0020] The self-attention mechanism is used for message passing on the graph, and the multi-head attention function MultiheadAttention is used to calculate the attention weights between nodes. nhead is set to 8 heads. By learning multiple linear mappings, the dynamic spatial dependencies affected by various factors are simulated in different latent subspaces.

[0021] S3: Use the temporal Transformer and the self-attention mechanism to achieve bidirectional temporal dependency modeling across multiple time steps. The flowchart and structure diagram of the method are as follows Figure 2 , attached Figure 4 The specific method of this step is:

[0022] S3.1: Self-attention first calculates the query matrix Q from time t=1 to T t , key matrix K t and the corresponding value matrix V t .

[0023]

[0024] The query, key and value functions in the present invention are a shared linear transformation function. t and K t The dot product of is used to calculate the bidirectional temporal dependency between nodes. The sparse attention function entmax15 is used to normalize the temporal dependency. t Update the node feature and get the new feature of the node as A t .

[0025]

[0026] A t (Q t , K t , V t )=S t V t

[0027] S4: Use the sparse mapping function entmax15 applied to the spatial and temporal Transformer to obtain a spatiotemporal Transformer network with sparse transformation.

[0028] S4.1: Adopt a generalized attention mechanism α-entmax conversion, by introducing an adjustable parameter α, allowing the model to smoothly transition between softmax and sparse max, setting the parameter α = 1.5 to replace the softmax function.

[0029] S5: Design a model framework of spatiotemporal blockchain based on spatiotemporal Transformer, which takes advantage of both dynamic spatial dependency and long-term temporal dependency. The model framework is shown in the attached figure. Figure 1 shown.

[0030] S5.1: The observation trajectory first passes through the fully connected layer to obtain input features, and then enters the sparse spatial Transformer and sparse temporal Transformer. The input of the lth spatiotemporal block is the l-1th spatiotemporal block at time step T. obs-k+1 The sparse spatial transformer and the sparse temporal transformer are stacked to generate the output tensor, while using residual connections to achieve stable training. Extract spatial features.

[0031]

[0032] and Fusion As the input of the subsequent sparse time transformer, a new tensor is finally generated

[0033]

[0034] Finally, we get the output tensor And use this tensor as the input of the l+1th spatiotemporal block. The model allows multiple spatiotemporal blocks to be stacked to improve the model's capabilities.

[0035] S6: Finally, the fully connected layer is used to generate the pedestrian’s future trajectory position based on the pedestrian’s motion state information.

[0036] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention proposes a pedestrian trajectory prediction method based on a sparse spatiotemporal graph Transformer network. The proposed sparse spatial Transformer can effectively dynamically model spatial dependencies, reduce interaction redundancy, and capture pedestrian status information in real time; the proposed sparse temporal Transformer can effectively realize bidirectional time dependency modeling across multiple time steps, and the proposed spatiotemporal block chain model framework simultaneously utilizes dynamic spatial dependencies and long-term time dependencies to effectively improve the accuracy of crowded environments and long-distance trajectory prediction. The present invention can effectively improve the accuracy of pedestrian trajectory prediction in crowded environments and long-distance time steps. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0038] Figure 1 It is a schematic diagram of the overall framework of an embodiment of the present invention.

[0039] Figure 2 Schematic diagram of the structure of the sparse spatiotemporal graph Transformer network of an embodiment of the present invention.

[0040] Figure 3 A schematic diagram of spatial interaction modeling according to an embodiment of the present invention.

[0041] Figure 4 It is a schematic diagram of time-dependency modeling according to an embodiment of the present invention.

[0042] Figure 5 Schematic diagram of experimental effects of different methods in embodiments of the present invention. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. Of course, the specific embodiments described here are only used to explain the present invention and are not used to limit the present invention.

[0044] Example 1

[0045] See also Figures 1 to 4 The present invention provides a technical solution, which is a pedestrian trajectory prediction method based on a sparse spatiotemporal graph Transformer network. The overall process of the method is as shown in the attached Figure 1 The technical solution of the present invention is further described below in conjunction with specific embodiments:

[0046] S1: Obtain the location information of pedestrians from the image frame, perform data preprocessing on the pedestrian trajectory information, obtain the trajectory coordinates of each pedestrian in the current frame, and obtain the trajectory corresponding to all frames of the pedestrian by traversing the image frame.

[0047] S2: Use dynamic spatial Transformer, use self-attention mechanism to dynamically model spatial dependencies, and use multi-head attention mechanism to jointly model multiple models of spatial dependencies. The flowchart and structure diagram of the method are as follows Figure 2 , attached Figure 3 The specific method of this step is:

[0048] S2.1: Dynamic graph convolutional layer.

[0049] The spatial embedding feature projects the input features of each node into a high-dimensional latent subspace by learning a linear mapping, and captures the dynamic spatial dependencies through training and modeling in the high-dimensional latent subspace. First, the input feature X of the sparse spatial Transformer is s After spatial embedding, spatial embedding features are obtained Then embed the features of each time step Projected into a high-dimensional latent subspace. The projection mapping is implemented by a feed-forward neural network consisting of multiple fully connected layers. For each spatial dependency pattern, we embed the feature of each node at each time step Train three latent subspaces, including the query subspace Q s , key subspace K s Sum value subspace V s .

[0050]

[0051]

[0052] in, Q s , K s , V s The weight matrix of .

[0053] Next, use Q s and K s The dot product of the dynamic spatial dependency S between nodes is calculated s .

[0054]

[0055] The sparse attention function 1.5-entmax is used to normalize the spatial dependencies. Therefore, S s Update the node feature and get the new feature of the node as A s, the new feature node passes through the feedforward neural network and learns A s further improve the prediction.

[0056] A s (Q s , K s , V s )=S s V s

[0057] S2.2: Extracting interactions between pedestrians.

[0058] The self-attention mechanism is used for message passing on the graph, and the multi-head attention function MultiheadAttention is used to calculate the attention weights between nodes. nhead is set to 8 heads. By learning multiple linear mappings, the dynamic spatial dependencies affected by various factors are simulated in different latent subspaces.

[0059] S3: Use the temporal Transformer and the self-attention mechanism to achieve bidirectional temporal dependency modeling across multiple time steps. The flowchart and structure diagram of the method are as follows Figure 2 , attached Figure 4 The specific method of this step is:

[0060] S3.1: Self-attention first calculates the query matrix Q from time t=1 to T t , key matrix K t and the corresponding value matrix V t .

[0061]

[0062] The query, key and value functions in the present invention are a shared linear transformation function. t and K t The dot product of is used to calculate the bidirectional temporal dependency between nodes. The sparse attention function entmax15 is used to normalize the temporal dependency. t Update the node feature and get the new feature of the node as A t .

[0063]

[0064] A t (Q t , K t , V t )=S t V t

[0065] S4: Use the sparse mapping function entmax15 applied to the spatial and temporal Transformer to obtain a spatiotemporal Transformer network with sparse transformation.

[0066] S4.1: Adopt a generalized attention mechanism α-entmax conversion, by introducing an adjustable parameter α, allowing the model to smoothly transition between softmax and sparse max, setting the parameter α = 1.5 to replace the softmax function.

[0067] S5: Design a model framework of spatiotemporal blockchain based on spatiotemporal Transformer, which takes advantage of both dynamic spatial dependency and long-term temporal dependency. The model framework is shown in the attached figure. Figure 1 shown.

[0068] S5.1: The observation trajectory first passes through the fully connected layer to obtain input features, and then enters the sparse spatial Transformer and sparse temporal Transformer. The input of the lth spatiotemporal block is the l-1th spatiotemporal block at time step T. obs-k+1 The sparse spatial transformer and the sparse temporal transformer are stacked to generate the output tensor, while using residual connections to achieve stable training. Extract spatial features.

[0069]

[0070] and Fusion As the input of the subsequent sparse time transformer, a new tensor is finally generated

[0071]

[0072] Finally, we get the output tensor And use this tensor as the input of the l+1th spatiotemporal block. The model allows multiple spatiotemporal blocks to be stacked to improve the model's capabilities.

[0073] S6: Finally, the fully connected layer is used to generate the pedestrian’s future trajectory position based on the pedestrian’s motion state information.

[0074] Example 2

[0075] See also Figures 1 to 5The present invention provides a technical solution, a pedestrian trajectory prediction method based on a sparse spatiotemporal graph Transformer network. The embodiment of the present invention provides a comparison of different methods based on a sparse spatiotemporal graph Transformer network. The overall flow chart of this embodiment is as follows Figure 1 , Figure 5 The following is a comparison of the model loss results obtained by different method experiments. The technical solution of the present invention is further described in conjunction with specific embodiments:

[0076] S1: Data preprocessing module, obtains the location information of pedestrians, and performs mean normalization processing on the absolute coordinates of each pedestrian. This step executes the method described in step S1 of Example 1 and will not be elaborated here.

[0077] S2: Use dynamic spatial Transformer, use self-attention mechanism to dynamically model spatial dependencies, and use multi-head attention mechanism to jointly model multiple models of spatial dependencies. s , three different mapping functions are used for calculation, and the specific calculation formula is as follows:

[0078]

[0079] The other steps implement the method described in step S2 of embodiment 1 and will not be elaborated here.

[0080] S3: Use the temporal Transformer and the self-attention mechanism to achieve bidirectional temporal dependency modeling across multiple time steps. Similar to step S2, three different calculation methods are compared for the temporal dependency relationship. The specific calculation formula is as follows:

[0081]

[0082] The other steps execute the method described in step S3 of embodiment 1, which will not be elaborated here.

[0083] S4: Using a 3-layer spatiotemporal Transformer-based spatiotemporal block chain model framework, while utilizing dynamic spatial dependency and long-term temporal dependency. The specific implementation steps are as described in step S5 of Example 1, which will not be elaborated here, wherein the number of spatiotemporal blocks num is set to 3.

[0084] S5: Pedestrian trajectory prediction result output, output the pedestrian trajectory prediction result in the future time period. This step executes the method described in step S6 in Example 1 and will not be elaborated here.

[0085] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A pedestrian trajectory prediction method based on sparse spatiotemporal graph Transformer network, characterized in that: The following steps are involved: S1: Obtain the location information of pedestrians from the image frame, perform data preprocessing on the pedestrian trajectory information, obtain the trajectory coordinates of each pedestrian in the current frame, and obtain the trajectory corresponding to all frames of the pedestrian by traversing the image frame; S2: Use dynamic spatial Transformer, use self-attention mechanism to dynamically model spatial dependencies, and use multi-head attention mechanism to jointly model multiple modes of spatial dependencies; S3: Use the temporal Transformer to model bidirectional temporal dependencies across multiple time steps using the self-attention mechanism; S4: Use the sparse mapping function entmax15 applied to the spatial and temporal Transformer to obtain a spatiotemporal Transformer network with sparse transformation; S5: Design a model framework of spatiotemporal blockchain based on spatiotemporal Transformer, which exploits both dynamic spatial dependencies and long-term temporal dependencies; The step S5 is as follows: the observation trajectory first obtains input features through the fully connected layer, and then enters the sparse spatial Transformer and the sparse temporal Transformer. The input of the lth spatiotemporal block is the l-1th spatiotemporal block at the time step T obs-k+1 The output of the sparse spatial Transformer and the sparse temporal Transformer are stacked to generate the output tensor, while using residual connections to achieve stable training. The sparse spatial Transformer is trained from the input Extract spatial features from and Fusion As the input of the subsequent sparse time transformer, a new tensor is finally generated Get the output tensor And use this tensor as the input of the l+1th spatiotemporal block; S6: Use the fully connected layer to generate the pedestrian’s future trajectory position based on the pedestrian’s motion state information.

2. The pedestrian trajectory prediction method based on sparse spatiotemporal graph Transformer network according to claim 1 is characterized in that: The step S2 specifically includes the following steps: S2.1: Dynamic Graph Convolutional Layer The spatial embedding feature projects the input feature of each node into a high-dimensional latent subspace by learning a linear mapping, and captures the dynamic spatial dependencies through training and modeling in the high-dimensional latent subspace. First, the input feature X of the sparse spatial Transformer is s After spatial embedding, spatial embedding features are obtained Then embed the features of each time step Projected into a high-dimensional latent subspace, the projection mapping is implemented by a feed-forward neural network consisting of multiple fully connected layers. For each spatial dependency pattern, the embedded features of each node at each time step are Train three latent subspaces, including the query subspace Q s , key subspace K s Sum value subspace V s ; in, Q s , K s , V s The weight matrix of Using Q s and K s The dot product of the dynamic spatial dependency S between nodes is calculated s : Use the sparse attention function 1.5-entmax to normalize the spatial dependency and use S s Update the node feature and get the new feature of the node as A s , the new feature node passes through the feedforward neural network and learns A s Improve predictions based on: A s (Q s ,K s ,V s )=S s V s S2.2: Extracting interactions between pedestrians The self-attention mechanism is used for message passing on the graph, and the multi-head attention function MultiheadAttention is used to calculate the attention weights between nodes. The multi-head attention is set to 8 heads. By learning multiple linear mappings, the dynamic spatial dependencies affected by various factors are simulated in different latent subspaces.

3. The pedestrian trajectory prediction method based on sparse spatiotemporal graph Transformer network according to claim 1 is characterized in that: The step S3 is specifically as follows: S3.1: Self-attention first calculates the query matrix Q from time t=1 to T t , key matrix K t and the corresponding value matrix V t ; The query, key, and value functions are a shared linear transformation function, using Q t and K t The dot product of is used to calculate the bidirectional temporal dependency between nodes, and the sparse attention function entmax15 is used to normalize the temporal dependency. t Update the node feature and get the new feature of the node as A t ; A t (Q t ,K t ,V t )=S t V t 。 4. The pedestrian trajectory prediction method based on sparse spatiotemporal graph Transformer network according to claim 1 is characterized in that: The step S4 is specifically as follows: adopting a generalized attention mechanism α-entmax conversion, introducing an adjustable parameter α, allowing the model to smoothly transition between softmax and sparsemax, and setting the parameter α=1.5.

Citation Information

Patent Citations

  • Track prediction method based on space-time diagram and airspace aggregation Transform network

    CN114997067A

  • Pedestrian trajectory prediction method and system based on multi-interaction spatiotemporal graph network

    US11495055B1