Pedestrian trajectory prediction method based on deep learning

By combining social mechanics models, graph convolutional networks, and Transformer networks, spatial and temporal features of pedestrian trajectories are extracted, and Gaussian noise is added, which solves the problem of insufficient prediction accuracy in existing methods and achieves higher prediction accuracy and generalization ability.

CN120913244APending Publication Date: 2025-11-07ANHUI NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510959129.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing pedestrian trajectory prediction methods lack deep integration of spatial and temporal features between pedestrians in complex scenarios, resulting in insufficient prediction accuracy, and the combination of social mechanics models and deep learning methods is insufficient.

Method used

By combining a social mechanics model, graph convolutional networks, and Transformer networks, spatial features between pedestrians are extracted through graph convolutional networks, while temporal features are captured by Transformer networks. Gaussian noise is added during the fusion process to simulate the uncertainty of real-world scenarios, and the future location information of pedestrians is output.

Benefits of technology

It improves the accuracy and generalization ability of pedestrian trajectory prediction, has physical interpretability and strong generalization ability, and is applicable to fields such as intelligent transportation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913244A_ABST
    Figure CN120913244A_ABST
Patent Text Reader

Abstract

The invention discloses a pedestrian trajectory prediction method based on deep learning. The pedestrian trajectory prediction method is characterized by comprising the following steps of S1, receiving input data and performing preprocessing; s2, extracting spatial features from the input data by using a graph convolutional network GCN, wherein the convolutional neural network GCN simulates spatial interaction between pedestrians through an adjacent matrix; s3, extracting time features from the input data by using an encoder layer of a Transform model, wherein the Transform model captures a dependency relationship of pedestrians in a time dimension through a self-attention mechanism; s4, fusing the spatial features and the time features to obtain spatial and temporal features, and extracting spatial and temporal fusion features by using a Transform encoder; and S5, noise is added to the extracted space-time fusion features, then the features are sent to a full connection layer for pedestrian trajectory prediction, and future position information of pedestrians is output. The method improves the prediction precision and robustness, and is suitable for the fields of intelligent transportation and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of pedestrian trajectory prediction, and particularly relates to a pedestrian trajectory prediction method based on deep learning, which relates to the combination of graph convolution network (GCN), Transformer network and social force model, especially the extraction method of pedestrian interaction modeling. BACKGROUND

[0002] Pedestrian trajectory prediction is one of the core technologies in intelligent transportation, robot navigation and public safety, etc., which aims to predict the future moving position and path of pedestrians according to their historical trajectory data. This problem faces many challenges in complex scenarios, such as dynamic interaction between pedestrians, diverse moving patterns, nonlinear motion laws, and behavioral responses to obstacles or targets, etc.

[0003] Traditional trajectory prediction methods are mostly based on rule-based models or statistical methods, such as social force model and Markov chain-based model. These methods can provide certain physical interpretability, but have limited adaptability to complex nonlinear interactions and scenarios. In recent years, deep learning-based trajectory prediction methods have made significant progress, among which long short-term memory network (LSTM), graph convolution network (GCN) and Transformer network are widely used.

[0004] GCN is good at capturing the spatial interaction characteristics between pedestrians, and can effectively model local and global node relationships by constructing a graph structure; Transformer network, with its strong time series modeling capability, performs outstandingly in handling long-time dependency problems. However, existing methods often model spatial and temporal characteristics separately, lacking deep integration of the two, which cannot fully utilize the complex interaction information between pedestrians and the temporal dynamic rules of trajectories, resulting in certain defects in the accuracy of prediction results. In addition, although the social force model has good physical interpretability, its combination with deep learning methods still needs further exploration.

[0005] Therefore, it is of important research value and practical significance to design a trajectory prediction method that combines social force model, graph convolution network and Transformer, which can fully extract spatial and temporal features, and has physical interpretability and strong generalization ability. SUMMARY

[0006] The present application aims to overcome the shortcomings of the prior art and provide a pedestrian trajectory prediction method based on deep learning to improve the accuracy of final pedestrian trajectory prediction. By combining the trajectory prediction method of social force model, graph convolution network and Transformer, it can fully extract spatial and temporal features, and has physical interpretability and strong generalization ability, which has important research value and practical significance.

[0007] To achieve the above object, the technical scheme adopted by the present application is as follows: an interactive extraction method of a multi-feature graph fusion graph convolutional network, which specifically comprises the following steps:

[0008] S1, receiving input data, processing the position coordinates of pedestrians and putting them into two dictionaries pedsPerFrame and trajecPerPedes and performing batch processing;

[0009] S2, extracting spatial features from the input data using a graph convolutional network (GCN), which simulates the spatial interaction between pedestrians through an adjacency matrix;

[0010] S3, extracting temporal features from the input data using the encoder layer of a Transformer model, which captures the dependency relationship of pedestrians in the time dimension through a self-attention mechanism;

[0011] S4, fusing the spatial features and the temporal features to obtain spatio-temporal features, and using a Transformer to extract spatio-temporal fusion features;

[0012] S5, saving a copy of the spatio-temporal fusion graph for the next round of temporal feature extraction, and the temporal feature extraction module can extract spatial features;

[0013] S6, performing pedestrian trajectory prediction based on the spatio-temporal features through a fully connected layer on the data with random Gaussian noise added, and outputting the future position information of the pedestrians.

[0014] In the extraction process of step S1, pedsPerFrame is a dictionary, where each key represents a frame ID (i.e. time step), and each value is a list of IDs of all appearing pedestrians in that frame. The format is as follows: trajecPerPedes is also a dictionary, and each row of the matrix represents information of a time step (frame ID, x-coordinate of pedestrian position, y-coordinate of pedestrian position). The format is as follows:

[0015]

[0016] where each pedestrian ped i has a trajectory composed of T time steps, and each time step contains three values: frame ID f t , x-coordinate and y-coordinate.

[0017] Two dictionaries are further batch processed: in batch processing, the data structure becomes a batch-organized format for model training. Each batch contains the following data: nodes_batch_batch, seq_list_batch, nei_list_batch, nei_num_batch, batch_pednum. They respectively represent the coordinate data of all pedestrians at each time step, whether the position of each pedestrian is valid at each time step, the relative neighbor relationship of each pedestrian with other pedestrians at each time step, the number of neighbors of each pedestrian, and the number of pedestrians at each time step in the current batch.

[0018] The specific method of GCN extracting pedestrian interaction features is as follows: GCN can directly process graph data, where the nodes in the graph represent pedestrians, and the edges represent whether these pedestrians are in a neighbor relationship. The model extracts the spatial features of the graph through the graph convolution layer, including the positions of the pedestrians, the repulsion between pedestrians, and the attraction of the destination to the pedestrians. These features are used for the pedestrian trajectory prediction task.

[0019] The specific method of constructing the graph structure is as follows: suppose there is a graph G = (V, E), where V is the set of pedestrian nodes, and E is the set of edges. Each node where 5 represents the features of the pedestrian node, and each feature is: horizontal coordinate, vertical coordinate, speed, repulsion between pedestrians, and attraction of the target destination.

[0020] The two-dimensional coordinates are obtained by subtracting the position of the first frame from the position of the pedestrian to get the relative coordinate position. The speed is obtained by subtracting the current position from the next position and dividing by the interval time. The formula is as follows:

[0021]

[0022] The repulsion R between pedestrians is calculated based on the surrounding pedestrians of the pedestrian. It can be calculated according to the following formula:

[0023]

[0024] where represents the neighbor set of pedestrian i, which is defined by the neighbor matrix. p i and p j are the positions of pedestrian i and pedestrian j, the numerator ||p i -p j || represents the Euclidean distance between the two pedestrians. The vector represents the unit direction vector of j pointing to i, and A and B are hyperparameters that control the strength and decay range of the force. Finally, the repulsion of the pedestrian is obtained by calculating the modulus of the repulsion, which is taken as one of the physical properties of the pedestrian node.

[0025] The attractiveness of the destination to the pedestrian is defined by the following formula:

[0026]

[0027] a i ||a i ||

[0028] where, is the destination coordinate of pedestrian i (here set as the position where the pedestrian is observed to end), p i is the current position of the pedestrian, C is the attractiveness intensity parameter, which can be self-learned and automatically adjusted in the deep learning network. In order not to affect the prediction accuracy, we assume the length of the attractiveness vector a i ||a i || as the target-oriented component in the node feature.

[0029] After defining the pedestrian node, the neighbor matrix of the pedestrian is constructed: all pedestrians in the same frame constitute a scene, and an NxN matrix Nei is constructed, when the distance between j pedestrian and pedestrian i reaches 5 meters, it is determined to be a neighbor, and the corresponding Nei[i][j] and Nei[j][i] are set to 1

[0030] The specific calculation process of the constructed graph structure input into the GCN network is as follows: further processing in G(V, E) and input into the GCN model, the calculation formula is:

[0031] is the feature vector of node v at the kth layer.

[0032] is the neighbor node set of node v.

[0033] d v and d u are the degrees (i.e. the number of neighbors) of nodes v and u, respectively.

[0034] W (k) is the weight matrix of the kth layer, which controls the feature transformation.

[0035] σ is the activation function, and ReLU function is selected.

[0036] The specific calculation process of the constructed graph structure input into the GCN network is as follows: the input of the model is a tensor with shape (T, N, D), where: T is the sequence length, N is the batch size, and D is the dimension of the input feature.

[0037] TransformerModel adopts stacked encoder layers to process input sequence data through self-attention mechanism and feedforward network. In each layer, the input is converted through multi-head self-attention mechanism and feedforward network, and finally output through residual connection and normalization layer.

[0038] Specifically, the input sequence is calculated by the self-attention mechanism to obtain the attention weight of each position:

[0039]

[0040] Where Q is the query, K is the key, and V is the value. Then, the output is further converted through a linear layer and an activation function.

[0041] The calculation formula of the feedforward network is: FFN(x) = max(0, xW1 + b1)W2 + b2

[0042] Finally, the residual connection is added and normalized: LayerNorm(x + FFN(x))

[0043] Where LayerNorm is a layer normalization operation to ensure the stability of the network.

[0044] The details of the spatio-temporal feature extraction are as follows: After processing the time and space features, the time and space features are fused at each time step through a concatenation operation. First, the output of the last frame of the time Transformer encoder is connected with the output of the space GCN encoder, and then a fully connected layer (fusion_layer) is used for feature compression and fusion. The fused spatio-temporal features are further transmitted to the second Transformer encoder to capture higher-level spatio-temporal dependencies through multi-head attention mechanism. Finally, the output layer (output_layer) is used to obtain the predicted position of the pedestrian at each time step.

[0045] The details of the backup spatio-temporal fusion graph data are as follows: After extraction through the second Transformer, the output is temporarily stored, and in the next round of time feature extraction, the temporarily stored data of the previous N-1 frames are concatenated with the fused data of the last frame for time feature extraction. The extracted data contains both time and space features, which can improve the prediction accuracy of the entire model.

[0046] Details of the final prediction output are as follows: in the process of spatio-temporal feature fusion, in order to enhance the robustness and generalization ability of the model, noise is added to the fused features to simulate the uncertainty in the actual scene. The invention uses Gaussian noise, also known as normal distribution noise. The shape of the noise is (1, 16), which means generating a 16-dimensional noise vector. In order to match the dimension of the spatio-temporal feature, the noise will be repeated in the first dimension, and the number of repetitions is the same as the time step, so as to ensure that the dimension of the noise is consistent with the spatio-temporal feature. In the prediction of each time step, the last spatio-temporal feature is converted into the prediction of the future position of the pedestrian through the output layer (which is a fully connected layer). The spatio-temporal feature after adding noise is mapped to a two-dimensional coordinate space through a linear layer, and the predicted position of each pedestrian is output.

[0047] The advantages of the present application are: it can effectively improve the accuracy of pedestrian trajectory prediction, by combining the social force model, graph convolution network and the trajectory prediction method of Transformer, it can fully extract spatial and temporal features, and has important research value and practical significance. BRIEF DESCRIPTION OF DRAWINGS

[0048] The content expressed by each figure in the specification of the present application and the marks in the figures are briefly described as follows:

[0049] Figure 1 It is the overall model architecture diagram of the present application;

[0050] Figure 2 It is an example diagram of pedestrian spatial interaction feature extraction of the present application;

[0051] Figure 3 It is a schematic diagram of three matrix models designed according to the graph structure;

[0052] Figure 4 It is a schematic diagram of the calculation process principle of the neighbor matrix and the feature vector of the pedestrian node;

[0053] Figure 5 It is a data processing principle block diagram of the prediction method of the present application. DETAILED DESCRIPTION

[0054] The specific embodiments of the present application are further described in detail by comparing the figures and describing the optimal embodiments.

[0055] The application provides a pedestrian trajectory prediction method based on graph convolution network (GCN) and Transformer fusion, which is suitable for complex dynamic scenes. By constructing a pedestrian graph structure, combining a social force model to quantify the interaction relationship, using GCN to extract spatial features, and using Transformer to capture temporal dynamic features, the method finally fuses spatio-temporal features and predicts future trajectories. The method improves the prediction accuracy and robustness and is suitable for intelligent transportation and other fields.

[0056] A pedestrian trajectory spatial feature extraction method based on a multi-feature fusion graph convolution network, the method specifically includes the following steps:

[0057] S1, receiving input data (frame ID, pedestrian ID, X coordinate, Y coordinate in trajectory data), processing the position coordinates of the pedestrians and putting them into the two dictionaries pedsPerFrame and trajecPerPedes, and performing batch processing;

[0058] S2, using a graph convolution network (GCN) to extract spatial features from the input data, the GCN simulates the spatial interaction between pedestrians through an adjacency matrix;

[0059] S3, using the encoder layer of a Transformer model to extract temporal features from the input data, the Transformer model captures the dependency relationship of pedestrians in the time dimension through a self-attention mechanism;

[0060] S4, fusing the spatial features and the temporal features to obtain spatio-temporal features, and using a Transformer decoder to extract spatio-temporal fusion features;

[0061] S5, saving a copy of the spatio-temporal fusion graph backup for the next round of temporal feature extraction, and the temporal feature extraction module can extract spatial features;

[0062] S6, adding random Gaussian noise to the extracted spatio-temporal fusion features to simulate the subjectivity of pedestrians, predicting the pedestrian trajectory based on the spatio-temporal fusion features, and outputting the future position information of the pedestrians.

[0063] Further, the batch processing process in step S1 includes: processing the pedestrian trajectory into two dictionaries pedsPerFrame and trajecPerPedes. Each key of pedsPerFrame represents a frame ID (i.e. time step), and each value is a list of IDs of all appearing pedestrians in the frame. The format is as follows Each row of the trajecPerPedes matrix represents information of a time step (frame ID, x-coordinate of pedestrian position, y-coordinate of pedestrian position). The format is as follows:

[0064]

[0065] where each pedestrian ped i is represented by a trajectory consisting of T time steps, each containing three values: frame ID f t , x coordinate and y coordinate.

[0066] where ped i represents the scene information of pedestrian i, each pedestrian ped i is represented by a trajectory consisting of T time steps, each containing three values: frame ID f t , x coordinate and y coordinate.

[0067] Two dictionaries are further batched: in batch processing, the data structure becomes a batch-organized format for model training. Each batch contains the following data: nodes_batch_batch, seq_list_batch, nei_list_batch, nei_num_batch, batch_pednum. Respectively represent the coordinate data of all pedestrians at each time step, represent whether the position of each pedestrian in each time step is valid, represent the relative neighbor relationship of each pedestrian with other pedestrians in each time step, the number of neighbors of each pedestrian, and the number of pedestrians in each time step in the current batch.

[0068] By calling two dictionaries, an effective data index matrix data_index is generated. This matrix contains three rows of information: frame ID, dataset number, and global frame index. Balanced batch data is generated from the data index. This step is the core of data processing, which includes the following processes:

[0069] Each index in data_index (each frame) is traversed, and the corresponding pedestrian data and trajectory are extracted from pedsPerFrame and trajecPerPedes. The intersection of these two sets ensures that only pedestrians with data in both frames are selected. For each pedestrian, its trajectory segment is obtained from trajecPerPedes. If the trajectory is incomplete or data is missing, the pedestrian is skipped. The trajectory segment of the specified frame is extracted, and the trajectory is checked for completeness and filtered out those pedestrian trajectories with incomplete data or no data in the last observation frame.

[0070] The trajectories of multiple pedestrians in the same frame are combined into a batch. For each batch:

[0071] The trajectory data of each pedestrian is combined into an array with shape (20, N, 2), where 20 is the number of frames and N is the number of pedestrians in the current batch.

[0072] If the number of pedestrians in the current batch exceeds the set around_ped*2, the batch is split into multiple sub-batches; otherwise, the data of multiple scenes is merged into one batch. Specifically, it is divided into three categories: too many pedestrians: if the number of pedestrians in a frame exceeds around_ped*2, the scene is split into two batches. Moderate number of pedestrians: if the number of pedestrians is moderate (greater than or equal to around_ped), the data is directly taken as a batch. Fewer pedestrians: if the number of pedestrians is too few, the data of multiple frames is accumulated until the number of pedestrians in a batch is met.

[0073] The final returned processed results include the following data:

[0074] nodes_batch_batch: All pedestrian coordinate data at each time step, shape seq_length*num_Peds*2, where num_Peds is the total number of pedestrians in the batch, and 2 represents x and y coordinates.

[0075]

[0076] seq_list_batch: A two-dimensional matrix representing whether the pedestrian position is valid at each time step, shape seq_length*num_Peds, where a value of 1 indicates that the pedestrian has valid data at that time step, and a value of 0 indicates that the pedestrian has no valid data at that time step.

[0077]

[0078] nei_list_batch: Neighbor matrix representing the relative neighbor relationship of each pedestrian with other pedestrians at each time step, shape seq_length*num_Peds*num_Peds

[0079]

[0080] nei_num_batch: The number of neighbors of each pedestrian, shape seq_length*num_Peds, indicating the number of neighbors of each pedestrian at each time step.

[0081]

[0082] batch_pednum: The number of pedestrians in the current batch at each time step, shape batch_size, indicating how many pedestrians are in each batch.

[0083]

[0084] As Figure 1 , 2The architecture diagram of the combination of the graph convolution network and the Transformer is shown; first, the trajectory data of the pedestrians are respectively constructed into the structures required by the graph neural network and the Transformer network, the graph structure based on the node and the adjacency matrix is constructed as the input of the graph convolution network GCN through the processed batch data, the spatial feature extraction of the last frame scene is performed through two layers of GCNConv and two layers of Relu, and multiple feature fusion of the pedestrian position, speed, repulsive force between pedestrians, and attraction of the destination is performed therein; second, the continuous trajectory with the time continuous attribute is constructed, the input suitable for the Transformer is output through the fully connected layer, the continuous time feature is extracted through a single layer network of multi-head attention, a normalization layer, a feedforward network layer, and a linear layer. The current scene spatial information extracted by the spatial module and the continuous feature extracted by the time module are fused to obtain the feature with the space-time information, and finally the predicted trajectory is output through the decoding of the Transformer decoder.

[0085] The specific steps of extracting the pedestrian spatial feature in step S2 for the graph convolution are as follows:

[0086] The multiple features in the constructed pedestrian node include the pedestrian position, speed, repulsive force between pedestrians, and attraction of the destination. The pedestrian position uses the relative position, and the relative coordinate position is obtained by subtracting the position of the first frame from the two-dimensional coordinate position of the pedestrian. The formula is as follows:

[0087] r i =p i -p0=(x i -x0,y i -y0)

[0088] Wherein x i and y i are the x and y coordinate positions of the pedestrian i, x0 and y0 are the x and y coordinate positions of the pedestrian i in the 0th frame in the scene.

[0089] The speed at time T is obtained by subtracting the position at time t from the position at time t+1 and dividing by the interval time. The formula is as follows:

[0090]

[0091] Wherein x t , y t , and t t are the x coordinate, y coordinate, and time stamp of the pedestrian at time t.

[0092] The repulsive force R between pedestrians is calculated according to the surrounding pedestrians, and can be calculated according to the following formula:

[0093]

[0094] where represents the neighbor set of pedestrian i, which is defined by the neighbor matrix. p i and p j are the positions of pedestrian i and pedestrian j, respectively, and the molecule ||p i -p j || represents the Euclidean distance between the two pedestrians. The vector represents the unit directional vector of j pointing to i, and A and B are hyperparameters that control the strength of the force and the range of decay. Finally, the repulsive force of the pedestrian is obtained by calculating the modulus of the repulsive force, which is taken as one of the physical properties of the pedestrian node. The attractive force of the destination to the pedestrian is defined as follows:

[0095]

[0096] a i = ||a i ||

[0097] where, is the destination coordinate of pedestrian i (set here as the position where the pedestrian is observed to end), p i is the current position of the pedestrian, and C is the attractive force strength parameter, which can be self-learned and automatically adjusted in the deep learning network. In order not to affect the prediction accuracy, we assume that the modulus of the attractive force vector a i = ||a i || is the target-oriented component in the node feature.

[0098] After defining the pedestrian node features, the neighbor matrix of the pedestrian is constructed: all pedestrians in the same frame form a scene, and a NxN matrix Nei is constructed, which is set to 1 when the distance between j pedestrian and pedestrian i reaches 5 meters, and thus Nei[i][j] and Nei[j][i] are set to 1.

[0099] The above node and edge matrix are constructed into a graph structure G(V, E), which is further processed and input into the GCN model, Figure 2 shows the extraction and calculation process, and the calculation formula is:

[0100]

[0101] is the feature vector of node v at the k+1 layer, is the feature vector of node u at the k layer.

[0102] σ is the activation function, and ReLU function is selected. is the neighbor node set of node v. d v and d uThe degrees of node v and node u (i.e., the number of neighbors of the nodes) respectively. ( k) is the weight matrix of the k-th layer, which controls the feature transformation. Node v and node u represent pedestrians, {v} represents the set of node v, N(v)∪{v} represents the union set of the neighbors of v and v, and represents the set of v and all neighbors of v.

[0103] The input of the Transformer model in S3 is a tensor with shape (T, N, D). Where: T is the sequence length of 20, N is the batch size, which is set to 128, and D is the dimension of the input feature, which is 2. First, the trajectory tensor with shape (20, 128, 2) is converted to a tensor with shape (20, 128, 32) through a fully connected layer FC, which is used as input to the Transformer model.

[0104] The Transformer model adopts a stacked encoder layer to process the input sequence data through a self-attention mechanism (Self-Attention) and a feedforward neural network (Feedforward Network). In each layer, the input is converted through the multi-head self-attention mechanism and the feedforward network, and finally output through the residual connection and the normalization layer.

[0105] Specifically, the input sequence is calculated through the self-attention mechanism to obtain the attention weighting of each position:

[0106]

[0107] Where Q is the query (Query), K is the key (Key), and V is the value (Value). Then, the output is further converted through a linear layer and an activation function.

[0108] The calculation formula of the feedforward network is: FFN(x) = max(0, xW1 + b1)W2 + b2

[0109] Where W1 and W2 represent weights that can be automatically learned and adjusted, and b1 and b2 are constants.

[0110] Finally, the residual connection is added and normalized: LayerNorm(x + FFN(x))

[0111] Where LayerNorm is the normalization operation, x represents the input data, and FFN is the value calculated through the feedforward network, which ensures the stability of the network.

[0112] In S4, after processing the time and space features, the time and space features are fused through a fully connected layer after concatenation operation at each time step. The specific details are as follows: first, the last frame output of the time Transformer encoder is connected with the output of the space GCN encoder, and then a fully connected layer (fusion_layer) is used for feature compression and fusion. The fused spatio-temporal features are further transmitted into a second Transformer encoder to capture higher-level spatio-temporal dependencies through a multi-head attention mechanism. Finally, the output layer (output_layer) is used to obtain the predicted position of each pedestrian at each time step.

[0113] Step S5 is to store the data extracted by the second Transformer after spatio-temporal fusion. These data contain time and space features, and also contain spatial information in the next round of time prediction. The saved data are the fusion results of the time and space features learned by the model. It retains the fine-grained spatio-temporal information of each time step and considers the mutual influence between pedestrians and long-term time dependence. These features can be used for subsequent prediction, especially when generating the final prediction output, the model can use these intermediate features to enhance the prediction ability of future pedestrian trajectories.

[0114] In step S6, in the process of spatio-temporal feature fusion, noise is added to the fused features to simulate the uncertainty in the actual scene and enhance the robustness and generalization ability of the model. The present application uses Gaussian noise, also known as normal distribution noise. The shape of the noise is (1, 16), which means generating a 16-dimensional noise vector. In order to match the dimension of the spatio-temporal features, the noise will be repeated in the first dimension, and the number of repetitions is the same as the time step, so as to ensure that the dimension of the noise is consistent with the spatio-temporal features. In the prediction of each time step, the last spatio-temporal feature is converted into the prediction of the future position of the pedestrian through the output layer (which is a fully connected layer). The spatio-temporal feature after adding noise is mapped to a two-dimensional coordinate space through a linear layer, and the predicted position of each pedestrian is output.

[0115] Obviously, the specific implementation of the present application is not limited by the above method, as long as various non-essential improvements are made by adopting the method concept and technical scheme of the present application, which are within the protection scope of the present application.

Claims

1. A pedestrian trajectory prediction method based on deep learning, characterized in that: The method comprises the following steps: S1, receiving input data and preprocessing; S2, extracting spatial features from the input data using a graph convolution network GCN, the convolutional neural network GCN simulating the spatial interaction between pedestrians through an adjacency matrix; S3, extracting time features from the input data using an encoder layer of a Transformer model, the Transformer model capturing the dependency relationship of pedestrians in the time dimension through a self-attention mechanism; S4, fusing the spatial features and the time features to obtain spatio-temporal features, and extracting the spatio-temporal fusion features using a Transformer encoder; S5, inputting the extracted spatio-temporal fusion features after adding noise into a fully connected layer for pedestrian trajectory prediction, and outputting future position information of the pedestrian. 2.The pedestrian trajectory prediction method based on deep learning according to claim 1, wherein: In step S1, the position coordinates of the pedestrian are processed and put into two dictionaries pedsPerFrame and trajecPerPedes and batch processing is performed; pedsPerFrame is a dictionary, wherein each key represents a frame ID, and each value is a list of IDs of all appearing pedestrians in the frame; the format is as follows: Each row of the matrix of the trajecPerPedes dictionary represents information of a time step: frame ID, x-coordinate of the pedestrian position, and y-coordinate of the pedestrian position, and the format is as follows: where each pedestrian ped i is represented by a trajectory consisting of T time steps, each containing three values: frame ID f t , x coordinate and y coordinate. 3.The pedestrian trajectory prediction method based on deep learning according to claim 2, wherein: In the step S1, the batch processing comprises: changing the data structure of the dictionary into a batch organization format, and each batch contains the following data: nodes_batch_batch, seq_list_batch, nei_list_batch, nei_num_batch, and batch_pednum, which respectively represent coordinate data of all pedestrians at each time step, represent whether the pedestrian position at each time step is valid, represent the relative neighbor relationship of each pedestrian with other pedestrians at each time step, the number of neighbors of each pedestrian, and the number of pedestrians at each time step in the current batch. 4.The pedestrian trajectory prediction method based on deep learning according to claim 1, wherein: In step S2, the convolutional neural network GCN extracts pedestrian interaction features, which comprises: constructing a graph structure of pedestrian trajectories, taking a node to represent a pedestrian, and taking an edge to represent whether the pedestrians are neighbor relationship to construct graph data, extracting spatial features of the graph through the convolutional neural network GCN, including the position of the pedestrian, the repulsion between pedestrians, and the attraction of the destination to the pedestrian. 5.The pedestrian trajectory prediction method based on deep learning according to claim 4, wherein: The construction of the graph structure includes setting a graph G=(V, E), where V is a set of pedestrian nodes, and E is a set of edges; each node wherein 5 represents the characteristics of the pedestrian nodes, and each characteristic is respectively: a horizontal coordinate, a vertical coordinate, a speed, a repulsive force between pedestrians, and an attraction force of a target destination. The two-dimensional coordinates are obtained by subtracting the position of the first frame from the position of the pedestrian to obtain the relative coordinate position; The speed is obtained by subtracting the current position from the next position and dividing by the interval time; The repulsion R between pedestrians is calculated according to the surrounding pedestrians of the pedestrian, and can be calculated according to the following formula: where represents the neighbor set of pedestrian i, which is defined by the neighbor matrix; p i and p j are the positions of pedestrian i and pedestrian j, respectively; the molecule ||p i -p j ||represents the Euclidean distance between the two pedestrians; vector represents the unit directional vector of j pointing to i, and A and B are hyperparameters that control the strength and decay range of the repulsive force; finally, the repulsive force of the pedestrians is obtained by calculating the length of the module, which is used as one of the physical properties of the pedestrian nodes. The attraction of the destination to the pedestrian is defined by the following formula: a i =||a i || wherein, is the destination coordinate of the pedestrian i, p i is the current position of the pedestrian, C is the attraction strength parameter, which is self-learned and automatically adjusted in the deep learning network; the length of the attraction vector a i =||a i || as the target-oriented component in the node feature; After defining the pedestrian node features, the neighbor matrix of the pedestrian is constructed: all pedestrians in the same frame constitute a scene, and an NxN matrix Nei is constructed, and when the distance between the j pedestrian and the pedestrian i reaches a set distance, it is determined to be a neighbor.

6. The pedestrian trajectory prediction method based on deep learning according to claim 4 or 5, characterized in that: Inputting the constructed graph structure data into the GNC network for feature extraction comprises: performing output calculation using the following calculation formula: is the eigenvector of node v at the kth layer; is a set of neighbor nodes of node v; d v and d u are the degrees of node v and node u, respectively; W (k) is the weight matrix of the kth layer, which controls the feature transformation; σ is an activation function, and ReLU function is selected.

7. The pedestrian trajectory prediction method based on deep learning according to claim 6, wherein: The feature extraction of the GNC network further includes that the input of the GCN network model is a tensor with a shape of (T, N, D), where T is a sequence length, N is a batch size, and D is a dimension of input features; the TransformerModel adopts stacked encoder layers to process input sequence data through a self-attention mechanism Self-Attention and a feedforward neural network Feedforward Network; in each layer, the input is converted through the multi-head self-attention mechanism and the feedforward network, and finally output through a residual connection and a normalization layer. 8.The pedestrian trajectory prediction method based on deep learning according to claim 1, wherein: The time feature and the space feature are fused to extract the space-time feature in step S4, which includes that the time feature and the space feature are fused through a splicing operation at each time step, first, the output of the last frame of the time Transformer encoder is connected with the output of the space GCN encoder, and then a full connection layer (fusion_layer) is used for feature compression and fusion to obtain the space-time feature. 9.The pedestrian trajectory prediction method based on deep learning according to claim 1 or 8, wherein: The fused space-time feature is sent to a second Transformer encoder to capture higher-level space-time dependencies through a multi-head attention mechanism, and then the output layer output_layer is used to obtain the predicted position of the pedestrian at each time step. 10.The pedestrian trajectory prediction method based on deep learning according to claim 8, wherein: The feature output extracted through the second Transformer encoder is temporarily stored, and in the process of time feature extraction in the next round, the temporarily stored data of the previous N-1 frames is spliced with the fused data of the last frame for time feature extraction.

Citation Information

Cited By

  • Pedestrian trajectory data compression and recovery method based on key point learning and diffusion

    CN121236505A