Pedestrian trajectory prediction method based on window attention and space diagram interaction network
The method addresses the challenges of long-term dependencies and complex spatial interactions in trajectory prediction by using window attention, hierarchical graph convolution, and multi-scale dilation convolutions, enhancing prediction accuracy and robustness in dynamic environments.
Patent Information
- Application Number
- CN202510341008.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-15
AI Technical Summary
Existing pedestrian trajectory prediction models have difficulties in long-term dependence on modeling and complex spatial interactions, and it is difficult to flexibly extract the attention of key step sizes, resulting in inaccurate prediction results.
The window attention mechanism is used to capture timing features in the time dimension, and the interaction between pedestrians and scenes is modeled in the spatial dimension through a hierarchical heterogeneous graph convolution network, and the future trajectory is generated by combining multi-scale expanded convolution networks. The interaction kernel function is used to quantify social interaction intensity, and an effectiveness mask is introduced to mask missing frames.
It improves the accuracy and robustness of pedestrian trajectory prediction, can better reflect social behavior and nonlinear interaction in real scenes, reduces computational complexity and suppresses noise interference.
Smart Images

Figure CN120318273A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a pedestrian trajectory prediction method based on window attention and spatial graph interaction network. Background Art
[0002] Previous research mainly focused on statistics and physics. With the in-depth research, existing methods have gradually evolved from traditional statistical models to deep learning-based models. The proposal of various models has greatly promoted the development of trajectory prediction technology. Statistics can be further divided into models based on kinematics and social forces; deep learning-based models can be mainly divided into the following categories: models based on recurrent neural networks (cnn) and their variants such as long short-term memory networks (LSTM), generative adversarial networks (GAN), graph convolutional neural networks (GCN), Transformers, and hybrid models such as spatio-temporal graph convolutional neural networks, models combining LSTM and GAN, and so on. The following will be discussed in detail from the following aspects and the limitations of existing methods will be analyzed.
[0003] 1. Trajectory Prediction
[0004] The core of pedestrian trajectory prediction is to predict future trajectories by inputting historical trajectories into the model. Social-LSTM proposed a "social pooling layer" to capture pedestrian features, enabling the sharing of pedestrian state information. Subsequent research also improved Social-LSTM and applied pooling to other models. For example, SFT and SR-LSTM proposed a weighted mechanism and applied it to pedestrian spatio-temporal interaction. Later, it was found that nodes and edges in the graph structure can well fit pedestrian coordinates and interactions between pedestrians. Therefore, models based on graph neural networks have been widely applied to pedestrian trajectory prediction tasks. For example, Social-GAN proposed a model combining generative adversarial network (GAN) and RNN, using an encoder-decoder structure and using a global pooling vector to capture the interaction relationship between pedestrians. In addition, s-attention and PTP-STGCN models both use spatio-temporal graphs to model the interaction between pedestrians and predict future trajectories. SGCN is a graph convolutional sparse model that combines social relationships. It multiplies the spatio-temporal interaction matrix by a sparse mask to obtain a sparse matrix, which is used for trajectory prediction. In recent years, Transformer has become the mainstream method for pedestrian trajectory prediction due to the advantages of its self-attention mechanism. The STAR model proposed the idea of combining temporal transformers and spatial transformers to extract spatio-temporal feature information.
[0005] 2. Social Interaction
[0006] Early pedestrian trajectory prediction was mostly based on physical models. The Social Force Model proposed using attractive and repulsive forces to describe the interactions between pedestrians. This provided inspiration for subsequent pedestrian interaction modeling, especially for the design of interaction kernel functions. On this basis, SoPhie combines scene attention and social attention to generate future trajectories. Among them, social attention is used to extract interaction features, while physical attention is used to extract scene features. Subsequently, a GAN network is used to generate diverse trajectories. AgentFormer proposes to apply intelligent perception attention in the temporal and social dimensions to simulate the potential intentions of pedestrians and distinguish the characteristics between individuals and groups. Social-GAN generates diverse trajectories through a variational autoencoder framework. Trajectron++ captures the social interaction relationships of agents through a graph neural network and generates reasonable trajectories by combining heterogeneous data. Social-STGCNN proposes a spatio-temporal graph convolutional network model to learn the interaction relationships between pedestrians.
[0007] SocialCircle proposes an angle-based pedestrian social interaction model, studies behavioral interactions in layers, and finally demonstrates how pedestrians respond to adjacent pedestrians in real life. Although these methods have made significant progress in modeling social interactions between pedestrians, most of the interaction functions they use, such as Euclidean distance, cannot comprehensively capture the non-linear relationships and highly complex interactions in reality. Therefore, how to design a more physically realistic interaction function has become a research direction in this study.
[0008] 3. Multimodal Trajectory Generation
[0009] Due to the complexity in reality, the multimodal output of pedestrian trajectories is more in line with the actual situation. Social-vae proposes a temporal autoencoder architecture. Although the model has achieved excellent performance, it lacks consideration of spatial interactions. Social-CVAE combines channel attention and self-attention, extracts past spatio-temporal features through the CVAE framework, and then decodes to generate multimodal trajectories, providing more choices for practical applications. The PPNet model obtains pedestrian trajectories with potential intentions through unsupervised learning and then inputs them into the target-conditioned Transformer network architecture to predict the final trajectory probability. MID innovatively proposes a diffusion model Transformer, combines stochastic processes, embeds historical trajectory information and social interaction encoders into the Transformer. The core idea is to perform trajectory prediction through the reverse process of diffusion.
[0010] In summary, these pedestrian trajectory prediction models have certain advantages in their respective fields, but they still face dual challenges such as difficulty in modeling long-term dependencies and complex spatial interactions. How to flexibly extract the attention of key steps while ensuring the prediction accuracy of the model, so that the model can efficiently focus on the most relevant time steps, improve the modeling ability of pedestrian interactions, and more accurately reflect the social behaviors and nonlinear interaction relationships of pedestrians in real scenes is a key problem that needs to be solved in the field of pedestrian trajectory prediction. Summary of the invention
[0011] The purpose of the present invention is to provide a pedestrian trajectory prediction method based on window attention and spatial graph interaction network, which captures temporal features and interaction features in the time dimension and spatial dimension. First, a window attention mechanism is proposed in the time dimension, which mainly multiplies the window mask and the missing frame mask, and then performs attention calculation to adjust the attention receptive field at each moment. In the spatial dimension, a hierarchical heterogeneous graph convolution network is constructed to combine the pedestrian dynamic interaction graph and the scene static semantic graph, and an interaction kernel function based on motion consistency is proposed to model the interaction between individual pedestrians. Finally, a multi-scale dilated convolutional network is used to generate future trajectories, and the multi-scale spatiotemporal features are captured by dilated convolution to enhance the accuracy and robustness of the prediction.
[0012] The inventive concept of the present invention is as follows: comprehensively utilize time and space feature extraction technologies to achieve efficient prediction of pedestrian trajectories through a dilated convolutional neural network. First, on the time scale, instead of using the global attention mechanism of the commonly used Transformer to capture the dependencies of time series, because when the sequence length is too long, the traditional global attention mechanism will cause a sharp increase in computational complexity, and at the same time, data loss, loss of key frames, and noise interference will all affect the prediction results. In addition, the movement of pedestrians within a short time window usually has a certain continuity. Therefore, it is more reasonable to perform local modeling within a short time to avoid irrelevant noise that may be introduced by global modeling. Specifically, the core idea of the fixed-window mask design is to constrain each time step t to only focus on the data within a fixed window size w in its history. Second, to avoid inaccurate prediction results caused by the existence of missing frames, this patent introduces a validity mask Γ∈〖{0,1}〗^T to mask missing frames. Combining the validity mask ensures that invalid frames do not affect the calculation. At the same time, in the spatial dimension, a hierarchical heterogeneous graph convolutional network is designed. Through the cross-graph convolutional operation of the pedestrian dynamic interaction graph G_p and the scene static semantic graph G_s, this network can achieve information exchange between pedestrian nodes and scene elements, thereby more accurately predicting the trajectories of pedestrians. Specifically, a novel interaction function is designed to quantify the social interaction intensity between pedestrians. Subsequently, the pedestrian dynamic interaction graph and the scene static semantic graph are aggregated through cross-graph aggregation convolutional operations. Finally, trajectory prediction is performed through a multi-scale dilated convolutional network, where the dilation factor of the convolutional kernel is increased to handle dependencies over longer time spans, and residual connections are used after each layer of convolution to improve the accuracy and robustness of the prediction. The training of the model is optimized by minimizing the negative log-likelihood loss function. This loss function calculates the error between the predicted trajectory and the true trajectory and updates the model parameters through gradient descent.
[0013] To achieve the above-mentioned inventive purpose, the technical solution adopted by the present invention is specifically as follows: a pedestrian trajectory prediction method based on window attention and spatial graph interaction network, comprising the following steps:
[0014] S1: Obtain the data in the dataset, including the position coordinates of pedestrians at each time, and represent the pedestrian trajectory data as a time series.
[0015] S2: In the time dimension, design a window mask mechanism to adjust the attention receptive field at each moment and effectively capture temporal dependencies;
[0016] S3: In the spatial dimension, construct a hierarchical heterogeneous graph convolutional network, combine the pedestrian dynamic interaction graph and the scene static semantic graph; capture the interaction information between pedestrians - pedestrians and pedestrians - environment;
[0017] S4: Generate future trajectories using a multi-scale dilated convolutional network, capture short-term and long-term motion patterns through dilated convolutions of different scales, and combine residual connections to enhance the stability and accuracy of predictions
[0018] S11: Obtain dataset data, including the position coordinates of pedestrians at each time, represented as
[0019]
[0020] S12: Use positional encoding to temporally encode the trajectory data. The positional encoding generates corresponding encoding vectors for each time step t to represent time information, so as to maintain the transmission of temporal information in the encoder. Each position point is mapped to a high-dimensional space through the positional encoding function as follows:
[0021]
[0022]
[0023] In step S2, the time-dependent features are learned through window attention, which includes the following steps:
[0024] S21: In pedestrian trajectory prediction tasks, the attention mechanism of the transformer is usually used to capture the dependencies in the time series. However, when the sequence length is too long, the traditional global attention mechanism will lead to a sharp increase in computational complexity. At the same time, data loss, loss of key frames, and interference from noise will all affect the prediction results. In addition, the movement of pedestrians within a short-time window usually has a certain continuity. Therefore, it is more reasonable to perform local modeling within a short time to avoid irrelevant noise that may be introduced by global modeling
[0025] Therefore, in the time dimension, this study does not directly apply global attention, but uses a window attention mechanism to reduce complexity and maintain the prediction accuracy. Specifically, the core idea of the fixed window mask design is to constrain each time step t to only focus on the data within a fixed window size w in its history. Specifically, the mask matrix M t (t′) ∈ {0, 1}
[0026] S22: The existence of missing frames in pedestrian trajectory data may lead to inaccurate prediction results. To solve this problem, this patent introduces a validity mask Γ ∈ {0, 1} T to mask the missing frames. Combining the validity mask ensures that invalid frames do not affect the calculation. The calculation method of the mask matrix is as follows:
[0027]
[0028] Fill invalid positions with negative infinity when calculating attention weights.
[0029]
[0030] Mask matrix when there are no missing frames is 1, and the weights remain unchanged at this time. When there are missing frames, the weight matrix A tt′ is infinitesimal. In this way, not only is the computational complexity reduced, but the interference problem of data loss on the model prediction results is also solved, and at the same time, short-term dependence information can be efficiently captured.
[0031] S23: Finally, obtain the time dimension feature F through window attention t as follows:
[0032] F t = WindowAtt(Q, K, V) = Softmax(A tt′ )V t
[0033] Among them, when the attention weight is negative infinity, after the softmax operation, the weight of the masked position is close to zero, ensuring that the missing frames will not affect the prediction. After masking the invalid frames through the mask, the model only calculates the attention weights between the valid time steps within the window, thereby reducing the computational amount and suppressing the influence of invalid interactions. The calculated result is used to generate the feature representation of each time step and further perform trajectory prediction.
[0034] In step S3, construct a hierarchical heterogeneous graph convolutional network, combining the pedestrian dynamic interaction graph and the scene static semantic graph; it includes the following steps:
[0035] S31: Construction of the pedestrian dynamic interaction graph G p of.
[0036] In the model, the pedestrian coordinates are converted into the spatial graph G p =(V p , E p ), which is composed of the node set and the edge set constitutes. Each pedestrian is regarded as a node, and the weight of the edge is calculated from the mutual influence between pedestrians. Specifically, the edge weight between pedestrians is defined through an interaction kernel function, which takes into account factors such as the spatial distance, speed, and relative angle between pedestrians. This interaction function quantifies the social interaction intensity between pedestrians by modeling the motion information between them. The design of the interaction kernel function is as follows:
[0037]
[0038] The interaction intensity is inversely proportional to the physical distance d between pedestrians, that is, the closer the distance, the greater the interaction intensity. Through the inverse square exponential decay, the model can capture the non-linear social behavior between pedestrians. d ij is a threshold. Beyond this distance, the interaction intensity decreases rapidly, which conforms to the law of actual social behavior. The movement speed of pedestrians is an important factor affecting the interaction intensity. Pedestrians with faster speeds have a stronger influence on the surrounding pedestrians. Therefore, the weighted sum of speeds, as an important factor in the interaction intensity, can quantify the degree of influence between pedestrians. When pedestrians move in the same direction, the interaction intensity is relatively large, while when pedestrians move in opposite directions, θ th the angle with the line connecting nodes i and j is negative. After passing through the max function, the numerator takes 0, which conforms to the actual social behavior. The rationality of the function will be explained from three realistic situations below. ij The rationality of the function will be explained from three realistic situations below.
[0039] Approaching each other: In a real scenario, when two pedestrians approach each other face to face, if both of their speeds are relatively fast and the relative angle (between 0° and 90°), the interaction intensity will increase. As the relative angle decreases and the pedestrian speed increases, the interaction intensity will be further enhanced; conversely, when the angle increases or the speed decreases, the interaction intensity will weaken.
[0040] Walking in the same direction: When pedestrians i and j are walking in the same direction, if the relative angle between them is small, the interaction intensity will increase. The model will predict that pedestrian i has a greater influence on pedestrian j. When the distance between pedestrians exceeds a certain threshold, the interaction intensity will weaken, and subsequently, the interaction intensity will be mainly dominated by the angle and speed.
[0041] Walking in opposite directions: When two pedestrians are walking in opposite directions, their interaction intensity is relatively weak. Specifically, when the relative angle is greater than 90°, the interaction intensity is zero, that is, there is no interaction effect between pedestrian i and pedestrian j.
[0042] S32: Construction of the scene static semantic graph G s In the model, the scene coordinates are transformed into the spatial graph G s =(V s , E s ), which is composed of the node set and the edge set Among them, V s is the coordinate position of the observed scene, and the edge set is composed of the adjacency matrix If there is an interaction relationship between node i and node j, Otherwise,
[0043] S33: Cross-graph aggregation convolution mechanism: Pedestrian subgraph Gp and the scene sub - graph G s are two heterogeneous graphs. Through the cross - graph aggregation convolution operation, the specific process is as follows:
[0044]
[0045] Specifically, the adjacency matrix A of the pedestrian sub - graph p and the adjacency matrix A of the scene sub - graph s respectively capture the interactions between pedestrians and the relationships between pedestrians and static elements in the scene. The degree matrix is a diagonal matrix, aiming to prevent the imbalance of information propagation caused by the difference in the number of connections between nodes. is used to balance the weights of neighbor nodes of each node during information propagation, so as to avoid the nodes with larger degrees dominating the information propagation during the process. F (ι) represents the feature matrix of the ι - th layer of the graph convolutional network, where the features of each pedestrian contain their spatial information (such as position, speed, etc.) at the current time step and the historical movement trajectory information. W (ι) controls the propagation of node information in the graph convolution operation. Through cross - graph convolution, the interactions between pedestrians and the relationships between pedestrians and the scene are modeled. At the same time, multi - layer graph convolution is used to iteratively aggregate neighbor information, gradually fusing local interactions and global scene constraints to generate accurate trajectories.
[0046] In the step S4, the time - dimension and space - dimension features are input into a multi - scale dilated convolution network for multi - modal trajectory generation, which includes the following steps:
[0047] S41: The design of the pyramid structure allows the parallel extraction of features at different scales. At each layer, the convolution operation processes the dependencies with longer time spans by increasing the dilation factor of the convolution kernel. The convolution operation at each layer is based on different time scales, enabling the network to adaptively capture long - term and short - term time dependencies. And the dilated convolution network processes the time features in parallel at each scale layer and performs feature fusion. This process ensures that the model's prediction of future time steps is not only based on historical data but also considers the temporal relationships of future predictions.
[0048] In addition, to avoid the problems of information loss and gradient disappearance, residual connections are used after each layer of convolution to improve the prediction accuracy and robustness. The dilation factor is 2 and the convolution kernel is 3*3 for the second layer, and the dilation factor is 4 and the convolution kernel is 5*5 for the second layer. Ensure that each layer of convolution operation captures the dependency relationships at different time scales. The time - dimension and space - dimension features obtained by the model are input into the network for convolution operations. The two - dimensional Gaussian distribution parameters are obtained
[0049] S42: Input the features in the time dimension and space dimension into the multi-scale dilated convolutional network TEP-CNN to generate two-dimensional Gaussian distribution parameters for predicting the future trajectory of pedestrians.
[0050]
[0051] The training of the model is optimized by minimizing the negative log-likelihood loss function. This loss function calculates the error between the predicted trajectory and the true trajectory, and updates the model parameters through gradient descent. Assume the coordinates of pedestrian n in the t time period follow a bivariate Gaussian distribution. The loss function is defined as:
[0052]
[0053] S43: The experiment is verified on the publicly available datasets ETH and UCY. The ETH dataset is a public dataset containing various pedestrian trajectories provided by ETH Zurich, and is usually used to evaluate pedestrian trajectory prediction and crowd behavior analysis. This dataset contains pedestrian trajectory data in different scenarios, and is mainly used for tasks such as multi-objective trajectory prediction and crowd behavior modeling; the UCY dataset is provided by the Aristotle University of Thessaloniki in Greece and contains pedestrian trajectory data in multiple complex environments, especially in dynamic and complex urban environments, including interactions between multiple pedestrians. Predict the future trajectory for the next 4.8 (12 time steps) seconds based on the historical trajectory of the first 3.2 seconds (8 time steps). The model is trained on an NVIDIA 4080 super GPU with 16G of video memory, using the Adam optimizer, an initial learning rate of 3×10 -4 , a batch size of 64, and 300 training epochs. This study uses two metrics, ADE (Average Displacement Error): the average Euclidean distance between the predicted trajectory and the ground truth, and FDE (Final Displacement Error): the Euclidean distance between the predicted position and the end point, to evaluate the accuracy of the predicted trajectory.
[0054]
[0055] Compared with the prior art, the beneficial effects of the present invention are:
[0056] 1. Window attention
[0057] The model pioneeringly embeds window attention into pedestrian trajectory prediction. Using window attention can flexibly extract the attention of key time steps. The model can efficiently focus on the most relevant time steps, while avoiding the impact of noise or data loss on the model performance, reducing the complexity while retaining the temporal dependence.
[0058] 2. Validity mask
[0059] In pedestrian trajectory data, the existence of missing frames may lead to inaccurate prediction results. To solve this problem, we introduce a validity mask Γ ∈ {0, 1} T to mask the missing frames. Combining the validity mask ensures that invalid frames do not affect the calculation. After masking the invalid frames with the mask, the model calculates the attention weights only between the valid time steps within the window, thus reducing the computational amount and suppressing the influence of invalid interactions.
[0060] 3. Interaction-aware kernel function
[0061] The interaction function-aware kernel fuses information such as the cosine of the speed direction and the radial approach rate, and the distance between pedestrians, quantifies the pedestrian interaction risk, and can more accurately reflect the pedestrian social behavior and non-linear interaction relationship in the real scene. This kernel function enables the model to have stronger robustness when dealing with complex interactions between pedestrians. Especially in a high-risk dynamic environment, such as when pedestrians are walking towards each other, it can adjust the predicted trajectory in a timely manner; in this study, the interaction between pedestrians and the environment is modeled as a heterogeneous graph: this structure explicitly associates the spatial constraints between pedestrians, lane lines, and obstacles, and the model can accurately understand the complex relationship between pedestrians and the environment, enabling the effective combination of the geometric information of pedestrian trajectories and road structures, and further improving the prediction ability of the model.
[0062] 4. Multi-scale dilated convolutional network
[0063] The multi-scale dilated convolutional network in the model optimizes the short-term trajectory details and long-term motion trends through the combination of hierarchical dilated convolutions and residual connections. The model can capture the temporal dependencies through multiple levels of convolutional layers, improve the accuracy of time series prediction, and capture the dependencies at different time scales. Each layer of the convolutional network focuses on extracting temporal features with different granularities, so as to better model the time evolution process of pedestrian trajectories; finally, the experiments of this model on the ETH / UCY dataset show that the metrics are improved compared with the existing baselines. The visualization results reveal the advantages of the model in capturing pedestrian social behavior and group interactions. Brief Description of the Drawings
[0064] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention.
[0065] Figure 1 It is the overall model diagram of the present invention.
[0066] Figure 2 It is the schematic diagram of the visualization result of the present invention under the public dataset. Detailed Embodiment
[0067] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. Of course, the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0068] Embodiment 1
[0069] See Figure 1 , the technical solution provided by this embodiment is: a pedestrian trajectory prediction method based on window attention and spatial graph interaction network, including the following steps:
[0070] S1: Obtain the data of the dataset, including the position information of pedestrians at each time, and represent the pedestrian trajectory data as a time series.
[0071] S2: In the time dimension, design a window masking mechanism to adjust the attention receptive field at each moment and effectively capture temporal dependencies;
[0072] S3: In the spatial dimension, construct a hierarchical heterogeneous graph convolutional network to jointly model the pedestrian dynamic interaction graph and the scene static semantic graph;
[0073] S4: Input the features in the time dimension and the spatial dimension into a multi-scale dilated convolutional network for multi-modal trajectory generation.
[0074] Step S1 includes the following steps:
[0075] S11: Obtain the data of the dataset, including the position coordinates of pedestrians at each time, represented as
[0076]
[0077] S12: Use positional encoding to perform temporal encoding on the trajectory data. The positional encoding generates corresponding encoding vectors for each time step t to represent time information, so as to maintain the transmission of temporal information in the encoder. Each position point is mapped to a high-dimensional space through the positional encoding function: Map to a high-dimensional space:
[0078]
[0079]
[0080] In step S2, the time-dependent features are learned through window attention, including the following steps:
[0081] S21: In the pedestrian trajectory prediction task, the attention mechanism of the Transformer is usually used to capture the dependencies of time series. However, when the sequence length is too long, the traditional global attention mechanism will lead to a sharp increase in computational complexity. At the same time, data loss, loss of key frames, and noise interference will all affect the prediction results. In addition, the movement of pedestrians within a short time window usually has a certain continuity. Therefore, it is more reasonable to perform local modeling within a short time to avoid irrelevant noise that may be introduced by global modeling.
[0082] Therefore, in the time dimension, instead of directly applying global attention, this study uses a window attention mechanism to reduce complexity and maintain prediction accuracy. Specifically, the core idea of the fixed-window mask design is to constrain each time step t to only focus on the data within a fixed window size w in its history. Specifically, the mask matrix M t (t′) ∈ {0, 1} is defined as follows:
[0083]
[0084] S22: In pedestrian trajectory data, the existence of missing frames may lead to inaccurate prediction results. To solve this problem, this patent introduces a validity mask Γ ∈ {0, 1} T to mask the missing frames. Combining the validity mask ensures that invalid frames do not affect the calculation. The calculation method of the mask matrix is as follows:
[0085]
[0086] where Γ(t′) = 0 indicates that there is missing data at time step t’, and Γ(t′) = 1 indicates that the data at time step t′ is valid. When calculating the attention weights, negative infinity is filled in the invalid positions.
[0087]
[0088] When there are no missing frames, the mask matrix is 1, and the weights remain unchanged at this time. When there are missing frames, the weight matrix A tt′ is infinitesimal. In this way, not only is the computational complexity reduced, but also the interference problem of data loss on the model prediction results is solved, and at the same time, short-term dependence information can be efficiently captured.
[0089] S23: Finally, the time dimension feature F t is obtained through window attention as follows:
[0090] F t = WindowAtt(Q, K, V) = Softmax(A tt′ )V t
[0091] Among them, when the attention weight is negative infinity, after the softmax operation, the weight at the masked position is close to zero, ensuring that the missing frames will not affect the prediction. After masking the invalid frames, the model calculates the attention weights only between the valid time steps within the window, thereby reducing the computational amount and suppressing the influence of invalid interactions. The calculated results are used to generate the feature representation for each time step and further perform trajectory prediction.
[0092] In step S3, a hierarchical heterogeneous graph convolutional network is constructed, combining the pedestrian dynamic interaction graph and the scene static semantic graph; it includes the following steps:
[0093] S31: Construction of the pedestrian dynamic interaction graph G p Construction.
[0094] In the model, the pedestrian coordinates are converted into the spatial graph G p =(V p , E p ), which is composed of the node set and the edge set . Each pedestrian is regarded as a node, and the weight of the edge is calculated from the mutual influence between pedestrians. Specifically, the edge weight between pedestrians is defined by the interaction kernel function, which takes into account factors such as the spatial distance, speed, and relative angle between pedestrians. This interaction function models the motion information between pedestrians and quantifies the social interaction intensity between pedestrians. The design of the interaction kernel function is as follows:
[0095]
[0096] The interaction intensity is inversely proportional to the physical distance d ij between pedestrians, that is, the closer the distance, the greater the interaction intensity. Through the inverse square exponential decay, the model can capture the non-linear social behavior between pedestrians. d th is a threshold value. Beyond this distance, the interaction intensity decreases rapidly, conforming to the actual social behavior law. The motion speed of pedestrians is an important factor affecting the interaction intensity. Pedestrians with higher speeds have a stronger influence on the surrounding pedestrians. Therefore, the weighted sum of speeds is an important factor in the interaction intensity and can quantify the degree of influence between pedestrians. When pedestrians move in the same direction, the interaction intensity is relatively large, while when pedestrians move in opposite directions, θ ij with the angle of the line connecting nodes i and j is negative. After passing through the max function, the numerator takes 0, which conforms to the actual social behavior. The rationality of the function will be illustrated from three realistic situations below.
[0097] Moving towards each other: In a real - world scenario, when two pedestrians approach each other face - to - face, if their speeds are relatively fast and the relative angle (between 0° and 90°), the interaction intensity will increase. As the relative angle decreases and the pedestrian speed increases, the interaction intensity will be further enhanced; conversely, when the angle increases or the speed decreases, the interaction intensity will weaken.
[0098] Walking in the same direction: When pedestrian i and pedestrian j are walking in the same direction, if the relative angle between them is small, the interaction intensity will increase. The model will predict that pedestrian i has a greater influence on pedestrian j. When the distance between pedestrians exceeds a certain threshold, the interaction intensity will weaken, and subsequently, the interaction intensity will be mainly dominated by the angle and speed.
[0099] Walking away from each other: When two pedestrians are walking in opposite directions, their interaction intensity will be weak. Specifically, when the relative angle is greater than 90°, the interaction intensity is zero, that is, there is no interaction effect between pedestrian i and pedestrian j.
[0100] S32: Construction of the scene static semantic graph G s : In the model, the scene coordinates are converted into the spatial graph G s =(V s , E s ), which is composed of the node set and the edge set . Among them, the edge set is composed of the adjacency matrix . If there is an interaction relationship between node i and node j, Otherwise,
[0101] S33: Cross - graph aggregation convolution mechanism: The pedestrian sub - graph G p and the scene sub - graph G s are two heterogeneous graphs. Through the cross - graph aggregation convolution operation, the specific process is as follows:
[0102]
[0103] Specifically, the adjacency matrix A p of the pedestrian sub - graph and the adjacency matrix A s of the scene sub - graph respectively capture the interaction between pedestrians and the relationship between pedestrians and static elements in the scene. The purpose is to prevent the imbalance of information propagation caused by the difference in the number of connections between nodes. is used to balance the weights of the neighbor nodes of each node during information propagation, so as to avoid the nodes with larger degrees dominating the information propagation during the process. The features of each pedestrian include the spatial information (such as position, speed, etc.) at the current time step and the historical motion trajectory information. W (ι) (ι) Control the propagation of node information in graph convolution operations. Through cross-graph convolution, model the interactions between pedestrians and the relationship between pedestrians and the scene. At the same time, use multi-layer graph convolution to iteratively aggregate neighbor information, gradually fuse local interactions and global scene constraints, and generate accurate trajectories.
[0104] In the step S4, input the time dimension and space dimension features into a multi-scale dilated convolution network for multi-modal trajectory generation, which includes the following steps:
[0105] S41: The design of the pyramid structure allows parallel extraction of features at different scales. At each layer, the convolution operation processes dependencies with longer time spans by increasing the dilation factor of the convolution kernel. The convolution operation at each layer is based on different time scales, enabling the network to adaptively capture long-term and short-term time dependencies. Moreover, the dilated convolution network processes time features in parallel at each scale layer and performs feature fusion. This process ensures that the model's prediction of future time steps is not only based on historical data but also considers the temporal relationship of future predictions.
[0106] In addition, to avoid information loss and gradient vanishing problems, residual connections are used after each layer of convolution to improve the accuracy and robustness of the prediction. The dilation factor is 2 and the convolution kernel is 3*3 for the second layer, and the dilation factor is 4 and the convolution kernel is 5*5 for the second layer. Ensure that each layer of convolution operation captures dependencies at different time scales. Input the time dimension features and space dimension features obtained by the model into the network for convolution operation. Obtain the two-dimensional Gaussian distribution parameters
[0107] S42: Input the features of the time dimension and space dimension into the multi-scale dilated convolution network TEP-CNN to generate two-dimensional Gaussian distribution parameters for predicting the future trajectory of pedestrians.
[0108]
[0109] The training of the model is optimized by minimizing the negative log-likelihood loss function. This loss function calculates the error between the predicted trajectory and the true trajectory and updates the model parameters through gradient descent. Assume the coordinates of pedestrian n in the t time period Follow a bivariate Gaussian distribution. The loss function is defined as:
[0110]
[0111] S43: The experiments were verified on the public datasets ETH and UCY. The ETH dataset is a public dataset containing various pedestrian trajectories provided by ETH Zurich and is commonly used to evaluate pedestrian trajectory prediction and crowd behavior analysis. This dataset contains pedestrian trajectory data in different scenarios and is mainly used for tasks such as multi-objective trajectory prediction and crowd behavior modeling; the UCY dataset is provided by the Aristotle University of Thessaloniki, Greece, and contains pedestrian trajectory data in multiple complex environments, especially in dynamic and complex urban environments, including interactions between multiple pedestrians. Based on the historical trajectories in the first 3.2 seconds (8 time steps), the future trajectories in the next 4.8 (12 time steps) seconds are predicted. The model was trained on an NVIDIA 4080super GPU with 16G of video memory, using the Adam optimizer with an initial learning rate of 3×10 -4 , a batch size of 64, and 300 training epochs. This study used two metrics, ADE (Average Displacement Error): the average Euclidean distance between the predicted trajectory and the ground truth, and FDE (Final Displacement Error): the Euclidean distance between the predicted position and the end point, to evaluate the accuracy of the predicted trajectories.
[0112]
[0113] Example 2
[0114] The experiments were verified on the public datasets ETH and UCY; ETH / UCY contains 5 different scenarios (ETH, HOTEL, UNIV, ZARA1, ZARA2), covering 1536 pedestrian trajectories on campuses, streets, and squares, which is suitable for evaluating the prediction robustness in dense scenarios. Based on the historical trajectories in the first 3.2 seconds (8 time steps), the future trajectories in the next 4.8 (12 time steps) seconds are predicted. The model was trained on an NVIDIA 4080super GPU with 16G of video memory, using the Adam optimizer with an initial learning rate of 3×10 -4 , a batch size of 64, and 300 training epochs. This study used ADE (Average Displacement Error): the average Euclidean distance between the predicted trajectory and the ground truth, and FDE (Final Displacement Error): the Euclidean distance between the predicted position and the end point.
[0115] In this part, through comparative experimental results, the performance of WAGIN and various existing baseline models on the ETH / UCY dataset is analyzed in detail, covering key metrics such as ADE, FDE, the number of parameters, and interaction time. An ablation experiment analysis is also conducted on each component of the model to further elaborate on the contribution of each module to the final performance. First, for the window size ω in window attention, take ω = 9, taking into account both short-term mutation situations and long-term trends. Because the window needs to cover to capture the motion inertia of pedestrians; secondly, too large a window will lead to a doubling of the attention calculation amount and may introduce noise. At the same time, for the interaction function distance d th , according to the social force model, the comfortable distance between pedestrians is about 1 - 2 meters. Therefore, take an interval of 0.2 m between 1 - 2 meters, calculate the values of ade and fde, and finally compare to take the optimal value. The results are shown in Table 1 below.
[0116] Table 1
[0117]
[0118] Table 1 Interaction function d th Average the values of ade and fde on the dataset, and the lower the better the effect. Experiments show that when d th = 1.4, the effect is optimal.
[0119] This model conducts ADE / FDE comparisons with the following state-of-the-art models:
[0120] Social-LSTM: A prediction model based on the LSTM architecture that assigns an independent LSTM network to each pedestrian to learn their motion patterns. Among them, the social interaction between pedestrians is presented through average pooling. At the same time, the hidden states of surrounding pedestrians are shared.
[0121] SGCN: A prediction model based on the GCN architecture that constructs a sparse directed graph to avoid redundancy and high model complexity problems in the graph convolutional network. Subsequently, it is combined with motion trend modeling to effectively reflect the complex interaction relationships between pedestrians.
[0122] S-STGCNN: A prediction model based on the spatio-temporal graph convolutional architecture. Models the pedestrian relationship as a spatio-temporal graph, introduces a kernel function to quantify the degree of spatial interaction between pedestrians, uses the graph convolutional neural network to capture the spatio-temporal relationship in the pedestrian trajectory, and then uses time convolution to extract features.
[0123] STAR: A prediction model based on the GCN and Transformer architectures. Alternately uses spatial transformer and time transformer modules, and then through graph convolutional operations, finally realizes trajectory prediction. It is worth mentioning that it proposes a graph memory module to process features and improves the ability to process time series data.
[0124] Social-VAE: A prediction model based on the transformer and VAE architectures. The core lies in that the data passes through the RNN-based time-varying variational autoencoder it proposed, combines with the encoder based on the attention mechanism to generate random latent variables, and finally uses the clustering method to generate trajectories.
[0125] GroupNet: A prediction model based on the transformer architecture. It models the complex interactions among multiple pedestrians through the multi-scale hypergraph topology module. On the other hand, it models the interaction pattern as three elements through the multi-scale hypergraph information transmission module, and finally performs trajectory prediction through cvae.
[0126] Social-CVAE: A prediction model based on the CVAE architecture. It mainly extracts features from historical trajectory data through channel attention and self-attention, and then generates future trajectories with the help of a bidirectional decoder. The results are shown in Table 2 below:
[0127] Table 2
[0128]
[0129] Comparison of ADE / FDE values of each model in Table 2 under five scenarios in ETH / UCY, where the bold values are the optimal values, the underlined values are the sub-optimal values, and the lower the better the effect.
[0130] This model performs excellently in the dataset scenario and is superior to other baseline methods. Compared with the Socia-CVAE model, the average ADE of the WAGIN model in this embodiment has increased by about 23%, while the average FDE has increased by about 21%. At the same time, compared with the GroupNet model, the average ADE of WAGIN has increased by about 20%, and the average FDE has increased by about 25%. This shows that WAGIN has improved in both the prediction accuracy in dense scenarios and the accuracy of the trajectory end point.
[0131] Compared with other baselines, WAGIN has a large improvement in ADE and FDE under five datasets as well as the final average ADE and average FDE, which proves its effectiveness in multiple scenarios. However, in individual scenarios, this model is slightly lacking compared with the latest baseline, but generally still shows good performance.
[0132] Example 3
[0133] To further analyze the contribution of each component of WAGIN to performance, ablation experiments were conducted in this embodiment. Key modules were gradually removed or replaced, and the changes in the model's performance were observed. The results of the ablation experiments are shown in the figure above. After replacing the window Transformer with the global transformer, both ADE and FDE decreased, proving that the local attention mechanism can better capture emergency motion patterns compared to the traditional global transformer; after replacing the velocity direction cosine interaction kernel with a conventional Euclidean distance kernel function, ADE increased by 14.8%, indicating that the interaction quantization based on motion consistency can more accurately reflect the social intentions between pedestrians than the conventional Euclidean distance. After replacing the multi-scale convolutional neural network structure with an MLP, the final prediction error FDE increased, indicating that the multi-scale dilated convolution has a significant improvement effect on long-term prediction.
[0134] Table 3
[0135]
[0136]
[0137] Table 3 Comparison of ADE / FDE values of this model and the model after removing the corresponding modules in five scenarios of ETH / UCY. Among them, w / o represents the removal operation, WF represents the window attention module, KF represents the interaction function, TEP represents the multi-scale dilated convolution module, the bold is the optimal value, the underlined is the sub-optimal value, and the lower the better.
[0138] Example 4
[0139] In terms of model efficiency, this embodiment also analyzed the number of parameters and inference time of the model. The inference time is the average of multiple single-step inferences, and the lower the inference time, the better. The results are as follows. The number of parameters of the model is significantly lower than that of Social-LSTM and SGCN, and the inference time performance is outstanding, showing its potential in real-time trajectory prediction. The number of parameters of Social-VAE is roughly equal to that of this model, but there is a certain gap in the interaction time. It shows that WAGIN has extremely high inference efficiency while maintaining high accuracy.
[0140] Table 4
[0141]
[0142] Table 4 Comparison of the number of parameters and inference time between the model of this embodiment and other baseline models. The smaller the number of parameters and interaction time, the better.
[0143] Through quantitative analysis and ablation experiments, the model proposed in this embodiment is significantly superior to existing benchmark methods on the ETH / UCY dataset and performs outstandingly on multiple evaluation metrics. In particular, the collaborative optimization of spatio-temporal modeling, the velocity direction cosine interaction kernel, and the multi-scale pyramid convolution structure provide the model with prediction accuracy and robustness beyond the baseline model. Compared with other methods, this model also has a smaller number of parameters and an efficient inference speed, making it suitable for real-time trajectory prediction tasks.
[0144] Example 5
[0145] Qualitative analysis visualization and case study
[0146] Figure 2 It shows a comparison example between the visualization results of trajectory prediction using different methods in the dataset scenario and the ground truth trajectory. One of the twenty trajectories that is closest to the actual situation is selected. Among them, the figure includes historical trajectories (red solid lines), ground truth future trajectories (green dashed lines), predicted trajectories of the WAGIN model (blue dashed lines), and Social-VAE (yellow dashed lines).
[0147] To comprehensively evaluate the model performance, we selected four standard datasets (ETH, HOTEL, UNIV, ZARA). The visualization results include various situations involving multiple pedestrian interaction behaviors. In the ETH scenario, there are pedestrians walking towards each other and passing by each other. When pedestrians are walking side by side towards each other, they are vulnerable to the influence of surrounding pedestrians and exhibit avoidance behaviors. However, when the Social-VAE model's trajectory is dealing with complex scenarios and encounters pedestrians coming from the opposite direction while walking side by side, the trajectory shows unreasonable deviations and even collision behaviors. But the predicted trajectory of our model is basically consistent with the real trajectory. This indicates that the interaction function of our model has advantages in complex situations and is closer to the real trajectory. In the HOTEL scenario, in the Y-shaped intersection, there are pedestrian encounters, but our model can still accurately predict the changes in pedestrian paths with little difference from the real trajectory. In contrast, Social-VAE generates unreasonable trajectories when dealing with this situation, posing a risk of collision. Additionally, in the complex campus scenario of UNIV, there are multiple pedestrians walking, stopping, and frequent interactions. The error between the predicted trajectory of Social-VAE and the real value is relatively large. When dealing with pedestrian interactions, it generates unreasonable zigzag trajectories. Our model does not produce collision behaviors and successfully avoids stationary pedestrians. Although the overall trend is still quite similar to the real trajectory, there is still an error between them. The reason for the error may be that pedestrians who have been stationary for a long time cause deviations in the model's processing. Finally, in the zara scenario, there are interactions between pedestrians and obstacles in the scene. It can be seen that our model is basically consistent with the real trajectory during processing, which also reflects the rationality of the fusion of the dynamic interaction graph of pedestrians and the static semantic graph of the scene. All of the above visualization results show that WAGIN performs excellently in modeling complex social behaviors and can generate high-precision trajectories.
[0148] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A pedestrian trajectory prediction method based on window attention and spatial graph interaction network, characterized in that It includes the following steps: S1: Obtain the data of the dataset, including the position information of pedestrians at each time, and represent the pedestrian trajectory data as a time series; S2: In the time dimension, design a window attention mechanism, use a window mask to constrain the attention receptive field of each time step, and mask out missing frames through a validity mask; S3: In the spatial dimension, construct a hierarchical heterogeneous graph convolutional network, unify the pedestrian dynamic interaction graph and the scene static semantic graph; capture the interaction information between pedestrians - pedestrians and pedestrians - environment; S4: Use a multi-scale dilated convolutional network to generate future trajectories, capture short-term and long-term motion patterns through dilated convolutions of different scales, and combine residual connections to enhance the stability and accuracy of the prediction.
2. The pedestrian trajectory prediction method according to claim 1, wherein The S1 includes the following steps: S11: Obtain the data of the dataset, including the position coordinates of pedestrians at each time, and construct a time series of pedestrian trajectories, expressed as where is the two-dimensional coordinate of the pedestrian at time step t, i represents the individual pedestrian, and T obs is the length of the observation time period; S12: Use position encoding to perform temporal encoding on the trajectory data, calculate the time step information through sine and cosine functions, and map it to a high-dimensional space, defined as follows: Among them, the PE function generates an embedding vector through the sine and cosine functions of the time step. After encoding the trajectory points of all time steps, they are combined to form a trajectory representation matrix X for input to subsequent modules. Among them, pos is the position index of the trajectory time step, and d is the hidden layer dimension of the model.
3. The pedestrian trajectory prediction method according to claim 1, characterized in that The S2 includes the following steps: S21: In the time dimension, use the window attention mechanism to set the window mask matrix M t (t′) ∈ {0, 1}, which is used to constrain that time step t can only focus on the data within the window size ω in its history, and is defined as follows: Among them, t is the current time step, t′ is the past time step, ω is the window size, indicating that at most ω / 2 units of information before and after the unit time step can be focused on; S22: In pedestrian trajectory data, the presence of missing frames leads to inaccurate prediction results. An availability mask Γ ∈ {0, 1} is introduced T to mask the missing frames. Combining with the availability mask, it ensures that invalid frames do not affect the calculation. The mask matrix is calculated as follows: Among them, Γ(t′)=0 indicates that there is missing data at time step t’, and Γ(t′)=1 indicates that the data at time step t′ is valid. Negative infinity padding is performed on invalid positions when calculating the attention weight; where Q t , K t , V t are the query, key, and value matrices, representing the feature representation of pedestrians within the time window; the mask matrix is 1 when there are no missing frames, and in this case the weights remain unchanged. When there are missing frames, the weight matrix A tt′ is infinitesimal; S23: Finally, the time dimension feature F is obtained through window attention t as follows: F t = WindowAtt(Q, K, V) = Softmax(A tt′ )V t Among them, the attention matrix A tt′ is infinitesimal when there are missing frames, and then approaches zero after the Softmax operation. The calculated result is used to generate the feature representation at each time step and further perform trajectory prediction.
4. The pedestrian trajectory prediction method according to claim 1, wherein, In the step S3, construct a hierarchical heterogeneous graph convolutional network, unify the pedestrian dynamic interaction graph and the scene static semantic graph; it includes the following steps: S31: Construction of pedestrian dynamic interaction graph G p : In the model, the pedestrian coordinates are transformed into a spatial graph G p =(V p , E p ), which is composed of a set of nodes and a set of edges . Among them, is the observed coordinate position. Each pedestrian is regarded as a node, and the weight of the edge is calculated from the mutual influence between pedestrians. The edge weight between pedestrians is defined by an interaction kernel function. The interaction kernel function takes into account the spatial distance, speed, and relative angle factors between pedestrians. By modeling the motion information between pedestrians, this interaction function quantifies the social interaction intensity between pedestrians. The design of the interaction kernel function is as follows: where d ij is the Euclidean distance between pedestrian i and pedestrian j; θ ij is the relative motion angle between the velocity direction of pedestrian i and the line connecting the distance to pedestrian j, represents the modulus of the time coordinate difference of pedestrian i from time t - 1 to time t, denoted as Similarly, the interaction intensity is inversely proportional to the physical distance d ij between pedestrians, that is, the closer the distance, the greater the interaction intensity. Through the inverse square exponential decay, the model captures the non-linear social behavior between pedestrians. d th is a threshold value. After exceeding this distance, the interaction intensity rapidly decreases, which conforms to the actual social behavior law; S32: Construction of the scene static semantic graph G s : In the model, the scene coordinates are converted into the spatial graph G s =(V s , E s ), which is composed of the node set and the edge set . Among them, V s is the coordinate position of the observed scene, and the edge set is composed of . If there is an interaction relationship between node i and node j, Otherwise, S33: Cross-graph Aggregation Convolution Mechanism: Pedestrian Subgraph G p and Scene Subgraph G s are two heterogeneous graphs. Through the cross-graph aggregation convolution operation, the specific process is as follows: Among them, σ is the Relu activation function, which is used to introduce non-linear transformation so that the model can learn complex graph features. is the pedestrian adjacency matrix A p and the scene adjacency matrix A s is the concatenation of them, representing the spatial relationship or mutual influence between nodes in the graph. The adjacency matrix A p of the pedestrian sub-graph and the adjacency matrix A s of the scene sub-graph respectively capture the interactions between pedestrians and the relationships between pedestrians and static elements in the scene. is the degree matrix, which is used for graph normalization operations; the degree matrix is a diagonal matrix, aiming to prevent the imbalance of information propagation caused by the difference in the number of connections between nodes. is the inverse square root of the degree matrix, which is used to balance the weights of neighbor nodes of each node during information propagation; F (ι) represents the feature matrix of the ι-th layer of the graph convolutional network, where the features of each pedestrian contain the spatial information at the current time step and the historical movement trajectory information. W (ι) is the learnable weight matrix of the ι-th layer, which controls the propagation of node information in the graph convolution operation.
5. The pedestrian trajectory prediction method according to claim 4, wherein The S4 includes the following steps: S41: The design of the pyramid structure allows parallel extraction of features at different scales. Design a convolutional network with a pyramid structure, and capture dependencies at different time scales through convolutional kernels with different dilation factors; the dilation factor of the second layer is 2, the convolutional kernel is 3*3, the dilation factor of the second layer is 4, the convolutional kernel is 5*5, and residual connections are introduced after each layer of convolutional operation to avoid information loss and gradient vanishing problems; S42: Input the features in the time dimension and the spatial dimension into the multi-scale dilated convolutional network TEP-CNN to generate two-dimensional Gaussian distribution parameters for predicting the future trajectories of pedestrians; Among them, among them, is the mean of the distribution, is the standard deviation of the distribution, is the correlation coefficient between x and y, W T is the learnable weight of the TEP-CNN network; the training of the model is optimized by minimizing the negative log-likelihood loss function, which calculates the error between the predicted trajectory and the true trajectory, and updates the model parameters through gradient descent; assume the coordinates of any pedestrian n at time t follow a bivariate Gaussian distribution, and the loss function is defined as: S43: The experiment is verified on the publicly available datasets ETH and UCY, and two metrics, the Euclidean distance between the predicted position and the end point, are used to evaluate the accuracy of the predicted trajectories; Among them, ADE is the average displacement error, which is used to measure the average Euclidean distance between the predicted trajectory and the true trajectory over the entire prediction time period. FDE is the final displacement error, which is used to measure the error between the final position predicted by the model and the true final position. N is the total number of samples of the predicted trajectory, and Tpred is the total number of prediction time steps, that is, the number of future time steps that the model needs to predict. is the position coordinate of pedestrian n predicted by the model at time step t. is the position coordinate of pedestrian n at time step t. is the Euclidean distance between the predicted trajectory point and the true trajectory point.
Citation Information
Cited By
Mining trackless rubber-tyred vehicle control system and method capable of adaptively judging obstacles
CN120802951A
Personnel trajectory drawing method and system based on AI algorithm
CN120932297A
A personnel trajectory drawing method and system based on an AI algorithm
CN120932297B
High-speed rail hub passenger microscopic trajectory prediction and crowd macroscopic risk identification method
CN121190869A
A high-speed rail hub passenger micro trajectory prediction and crowd macro risk identification method
CN121190869B