Short-term prediction method for rail passenger flow based on dynamic multi-graph fusion spatiotemporal deep learning

Through dynamic multi-graph integration of space-time deep learning methods, the time-changing dynamic graph and the introduction of causal time multi-head attention mechanisms are solved, and the problems of dynamic changes and multi-factor influence between stations in traditional methods are achieved, and more accurate short-term passenger flow forecasts and better management support are achieved.

CN119416976BActive Publication Date: 2025-08-19BEIJING UNIV OF CIVIL ENG & ARCHITECTURE +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411582878.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-07
Publication Date
2025-08-19
Estimated Expiration
2044-11-07

AI Technical Summary

Technical Problem

Traditional graph convolutional networks cannot effectively capture the dynamic changes and multi-factor influence of spatial relationships between stations in urban rail transit passenger flow prediction, resulting in a decrease in prediction accuracy.

Method used

The dynamic multi-graph fusion space-time deep learning method is adopted, and the time-varying dynamic graph is constructed through the Tucker decomposition method, combining multiple predefined graph convolution and causal time multi-head attention mechanism, a fusion graph convolution module is constructed to perform multi-layer spatiotemporal feature extraction.

Benefits of technology

The model's ability to extract space-time features is improved, more accurate short-term passenger flow prediction results are provided, the interpretability of dynamic spatial relationships is enhanced, and the effective management of urban rail transit systems is supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119416976B_ABST
    Figure CN119416976B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for short-term rail passenger flow prediction based on dynamic multi-graph fusion spatiotemporal deep learning. The method first includes convolving historical passenger flow feature data to form input data; then constructing a fusion graph convolution module, which includes multiple predefined graph convolutions and dynamic graph convolutions, and introducing causal time into the fusion graph convolution module to form a spatiotemporal learner; then using the spatiotemporal learner to process the input data; finally, the output of the spatiotemporal learner is connected through a jump connection layer, and the jump connection layer introduces the output of the spatiotemporal learner into the output layer, and the prediction result is obtained through the output layer. The present invention uses a combination of predefined graph convolution and dynamic graph convolution as an adjacency matrix, which can effectively learn the complex spatial characteristics of passenger flow, and introduces a causal time multi-head attention mechanism to perform multi-layer spatiotemporal feature extraction, thereby improving the model's ability to extract spatiotemporal features and providing a more accurate feature representation for short-term passenger flow prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of short-term rail passenger flow prediction, and in particular to a short-term rail passenger flow prediction method based on dynamic multi-graph fusion spatiotemporal deep learning. Background Art

[0002] With the rapid expansion of the Urban Rail Transit (URT) network and the continuous increase in passenger volume, the temporal and spatial distribution of passengers has become increasingly complex. Short-term passenger flow forecasts provide precise data, enabling operators to adjust strategies in real time to ensure smooth system operation. Such forecasts not only help passengers plan their journeys more efficiently, thereby reducing costs, but also enhance overall system management. Therefore, accurate short-term passenger flow forecasts are essential for the effective management and operation of urban rail transit systems.

[0003] Currently, graph convolutional networks (GCNs) are commonly used in urban rail transit passenger flow forecasting. However, this approach requires complex domain expertise and can suffer from incomplete information utilization. Alternatively, methods adaptively construct spatial adjacency matrices based on data. These methods utilize shared parameters across layers to construct a matrix and update it using methods such as gradient descent. While GCNs have achieved promising results in urban rail transit forecasting, the following issues remain.

[0004] 1. The dynamic spatiotemporal nature of urban rail transit. Due to rapidly changing factors such as commuting patterns and traffic congestion, the interactions between stations or nodes within the network are dynamic. For example, the traffic volume or density between two stations during peak hours (such as the morning rush hour) may be significantly different from that during off-peak hours (such as the afternoon). Traditional models often fail to capture the dynamic and evolving nature of these inter-station spatial relationships, resulting in reduced prediction accuracy.

[0005] 2. The variables influencing urban rail transit passenger flow are multifactorial and complex. The physical structure of the urban rail transit network, historical traffic patterns, and other environmental and temporal factors such as weather, holidays, or special events are all interrelated. These factors are often represented using predefined diagrams. However, each predefined diagram only captures a fragment of the overall narrative, bypassing the interactions between these factors. For example, a diagram that emphasizes the physical structure of the urban rail network may ignore the influence of historical passenger flow patterns and environmental variables. This one-sided dynamic perspective further affects the predictive accuracy of traditional models. Summary of the Invention

[0006] The purpose of the present invention is to provide a short-term prediction method for rail passenger flow based on dynamic multi-graph fusion spatiotemporal deep learning to address the above-mentioned problems, and overcome the limitations of static matrices in traditional methods. In the prior art, traditional graph neural network methods usually use static adjacency matrices in urban rail transit predictions, which ignores the dynamic changes in inter-station correlations during the day. The present invention uses the Tucker decomposition method to construct a time-varying dynamic graph, which can accurately capture the spatial correlation of passenger flow in different time intervals, effectively overcoming this limitation; enriching the extraction of multivariate spatial correlations of passenger flow, previous research methods were relatively single in extracting passenger flow spatial correlations. The present invention constructs three predefined graphs, namely the physical structure graph, the passenger flow similarity graph, and the OD correlation graph, and fuses them with the dynamic graph to construct a fused graph convolution module, which greatly enriches the multivariate spatial correlations of passenger flow and provides more comprehensive information for more accurate short-term passenger flow predictions; and improves the spatiotemporal feature extraction capability, as the prior art has deficiencies in spatiotemporal feature extraction. This paper introduces a causal temporal multi-head attention mechanism, which together with the fusion graph convolution module forms a spatiotemporal learner, which can perform multi-layer spatiotemporal feature extraction, thereby improving the model's ability to extract spatiotemporal features and providing more accurate feature representation for short-term passenger flow prediction.

[0007] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is as follows:

[0008] According to one aspect of the present invention, a method for short-term rail passenger flow prediction based on dynamic multi-graph fusion spatiotemporal deep learning is provided, comprising the following steps:

[0009] S1. Convolve the historical passenger flow feature data to form input data;

[0010] S2. Construct a fused graph convolution module, wherein the fused graph convolution module includes multiple predefined graph convolutions and dynamic graph convolutions, and introduces a causal temporal multi-head attention mechanism into the fused graph convolution module to form a spatiotemporal learner;

[0011] S3, connecting the plurality of spatiotemporal learners in series, and using each of the spatiotemporal learners to process the input data respectively;

[0012] S4. The outputs of all the spatiotemporal learners are connected through a skip connection layer. The skip connection layer introduces the outputs of the spatiotemporal learners into an output layer, and a prediction result is obtained through the output layer.

[0013] Preferably, in step S1, the historical passenger flow characteristic data is represented by a historical passenger flow sequence, specifically:

[0014]

[0015] Among them, X t represents the historical passenger flow sequence; represents the inflow data of the nth station in the tth time interval; N represents the total number of stations.

[0016] Preferably, in step S2, the output of the fusion graph convolution module is expressed as follows:

[0017]

[0018] in, is the output of the fused graph convolution; is the gate value; H p is the output of multiple predefined graph convolutions; H d is the output of dynamic graph convolution; ⊙ is the Hadamard product.

[0019] Preferably, the gate value is represented by the following formula:

[0020]

[0021] Among them, H p is the output of multiple predefined graph convolutions; H d is the output of dynamic graph convolution; is the gate value; is the bias parameter.

[0022] Preferably, the output H of the multiple predefined graph convolutions p It can be expressed by the following formula:

[0023] H p =Z p ||Z s ||Z c

[0024] Among them, Z p is the convolution output of the physical structure graph; Z s is the output of the passenger flow similarity graph convolution; Z c is the output of OD related graph convolution; ‖ is the tensor connection.

[0025] Preferably, the physical structure graph convolution output Z p , the output Z of the passenger flow similarity graph convolution s , OD related graph convolution Z c They are expressed by the following formulas:

[0026]

[0027] Among them, GCN(·) is the graph convolution operation; G p is the physical structure diagram; G s is the passenger flow similarity graph; G c is the OD correlation diagram; is the output of the causal temporal multi-head attention mechanism.

[0028] Preferably, in step S2, the dynamic space graph can be G d Indicates that at a specific time t, the dynamic space graph can be expressed as:

[0029]

[0030] Among them, S represents the site set; represents the spatial connectivity at time t; represents the corresponding weight between node i and node j.

[0031] Preferably, in step S4, the output of the spatiotemporal learner can be expressed as:

[0032]

[0033] in, is the output of the fused graph convolution, H l+1 is the output of the spatiotemporal learner at this layer, H l It is the output of the spatiotemporal learner in the previous layer.

[0034] Preferably, in step S4, the prediction result obtained by the output layer can be expressed by the following formula:

[0035]

[0036] Where L is the number of spatiotemporal learners, H l is the output of the previous spatiotemporal learner.

[0037] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0038] The present invention first convolves historical passenger flow feature data for dimensional alignment; then, a spatiotemporal learner is constructed by fusing graph convolution and causal temporal mechanism. In the process of fusing graph convolution, a combination of predefined graph convolution and dynamic spatial graph convolution is used as the adjacency matrix, which can effectively learn the complex spatial characteristics of passenger flow. The causal temporal multi-head attention mechanism is introduced, and together with the fused graph convolution module, a spatiotemporal learner is formed, which can perform multi-layer spatiotemporal feature extraction, thereby improving the model's ability to extract spatiotemporal features and providing more accurate feature representation for short-term passenger flow prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic flow diagram of the present invention;

[0040] Figure 2 It is a schematic diagram of the dynamic graph construction process of the present invention;

[0041] Figure 3 It is a schematic diagram of the structure of the fusion graph convolution module of the present invention;

[0042] Figure 4 Schematic diagram of visualization of CT-MSA of the present invention. DETAILED DESCRIPTION

[0043] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention is further described below with reference to the accompanying drawings and by way of preferred embodiments. However, it should be noted that many of the details listed in this specification are merely provided to help the reader gain a thorough understanding of one or more aspects of the present invention, and these aspects of the present invention can be practiced even without these specific details.

[0044] See also Figure 1 The present invention provides a short-term rail passenger flow prediction method based on dynamic multi-graph fusion spatiotemporal deep learning. The technical solution is as follows:

[0045] A short-term rail passenger flow prediction method based on dynamic multi-graph fusion spatiotemporal deep learning includes the following steps:

[0046] S1. Convolve the historical passenger flow feature data to form input data.

[0047] Specifically, a historical passenger flow sequence is extracted. The historical passenger flow sequence includes data from the Automatic Fare Collection (AFC) system, which includes the passenger's card number, departure station, arrival station, and respective timestamps. After extracting the historical passenger flow sequence from the AFC system data, the historical passenger flow sequence is structured into a matrix to form passenger flow features configured at different time intervals. In this embodiment, the time intervals are set at intervals of 10, 15, and 60 minutes. The sequence can be represented by the following formula:

[0048]

[0049] Among them, X t represents the historical passenger flow sequence; represents the inflow data of the nth station in the tth time interval; N represents the total number of stations.

[0050] After obtaining the passenger flow sequence, 1x1 convolution is performed on the passenger flow sequences in different time periods. The convolution results can be used as input data for subsequent processing.

[0051] S2. Construct a fused graph convolution module, wherein the fused graph convolution module includes multiple predefined graph convolutions and dynamic graph convolutions, and introduces a causal temporal multi-head attention mechanism into the fused graph convolution module to form a spatiotemporal learner.

[0052] First, a dynamic spatial graph is constructed;

[0053] Existing methods usually rely on predefined dependencies, such as distance and functional similarity, or use static graphs. This results in the inability to explain the heterogeneity of different time intervals. Therefore, in this embodiment, a dynamic spatial graph construction method for dynamic graph convolution is used to tailor the spatial representation of potential attributes to specific time intervals to capture the dynamic spatial dependencies in passenger flow data. Specifically, based on the concept of periodicity, it is assumed that inputs from the same time period on different days can be used to construct a shared dynamic graph. The dynamic graph can be represented as G d , at a specific time t, the dynamic space graph can be expressed as:

[0054]

[0055] Among them, S represents the site set; represents the spatial connectivity at time t; Represents the corresponding weight between node i and node j. The method of constructing a shared dynamic graph reduces the number of separate constructions of N t A dynamic graph approach. Using tensors Indicates building N t To determine the appropriate dynamic graph at a specific time t, use Calculate the corresponding index and select As a dynamic graph at time t. Using Tucker decomposition to reorganize the adjacency tensor can effectively reduce the number of required parameters. Allocate three matrices and a core tensor: the time slot embedding matrix Source node embedding matrix Target node embedding matrix And a core tensor E k ∈R d×d×d The adjacency tensor can be determined by the above elements:

[0056] A=Softmax(ReLU(E k ×1E t ×2E s ×3E e ))

[0057] Among them, N t 、N s 、N e and d represent the time slot, original node, target node, and embedding dimension, respectively. The performance and stability of dynamic graph convolution can be improved by normalizing and activating the adjacent tensor through the Softmax function and ReLU function. Figure 2 This paper demonstrates the construction of dynamic graphs using Tucker decomposition. The core tensor and factor matrix are learned from data to capture the latent spatial correlations at different time intervals.

[0058] Then, in order to solve the problem that relying solely on dynamic graphs cannot fully explore hidden spatial dependencies and lacks sufficient interpretability (although the dynamic spatial graph generated by the dynamic graph builder can reflect the dynamic correlation of passenger flow characteristics of different stations within a day, the dynamic graph is essentially a supplement to the uncertain relationship between nodes). In order to more deeply explore these hidden spatial dependencies and enhance the explanatory power of the model, in this embodiment, Figure 3 As shown in Figure 1, a fused graph convolution module is constructed and a causal temporal multi-head attention mechanism is introduced into the fused graph convolution module to form a spatial learner, thereby adding more prior knowledge. Specifically, the fused graph convolution module consists of a dynamic graph convolution module and multiple predefined graph convolution modules. The dynamic graph convolution module uses the constructed dynamic adjacency matrix A to capture and integrate the dynamic spatial information in the rail passenger flow network. This process can be expressed mathematically as follows:

[0059]

[0060] Among them, GCN(·) represents the graph convolution operation, with the dynamic adjacency matrix a and the output of the causal time MSA module as its input.

[0061] For GCN(·), the urban rail network is represented as a graph G = (V, E), where V is a set of nodes (stations) and E is an edge set, representing the abstract connection relationship between different stations. The adjacency matrix of the graph can be expressed as A∈R n×n , where n represents the number of nodes. The element W located in the i-th row and j-th column of the matrix ij Represents the weight coefficient, which represents the strength of the relationship between stations. By using a graph convolutional network (GCN) to operate on different adjacency matrices, different spatial dependencies between stations are mined.

[0062] The GCN operation can be expressed as follows:

[0063]

[0064] in, is the node representation matrix of the lth layer; H 0 =X; is the weight matrix to be learned in the lth layer; A = A + I is the adjacency matrix with self-connection added, where I is the identity matrix; D is the degree matrix of A.

[0065] In the multiple predefined graph convolution modules, using professional knowledge in fields such as topology and similarity, a variety of graphs can be designed to illustrate the various spatial connections between nodes (especially stations). The passenger flow change patterns of stations with similar functions are also similar. At the same time, the size of the OD (origin-destination) passenger flow between stations can also reflect the correlation between the two stations. Therefore, based on prior knowledge, three types of graphs representing different spatial relationships are constructed. These three types of graphs are physical adjacency matrix (physical graph of URT network), functional similarity matrix (functional similarity graph) and OD correlation matrix (origin-destination (OD) correlation graph). They are named G p =(S,E p ,W p ), G s =(S,E s ,W s ) and G c =(S,E c ,W c ). Where, S={s1,s2,…,s n} represents the station set; n represents the number of URT stations; e ij ∈E k (k=p,s,c) represents G k In order to capture the multifaceted spatial interdependencies in the entire URT network, we define the edges between stations i and j in E k (k=p,s,c) Create multiple weight matrices W k The following describes the three types of graphs:

[0066] 1) Physical diagram:

[0067] Figure G p is constructed directly from the physical layout of the subway system under study. p If nodes i and j in the graph are connected to each other in the real world, an edge is established to connect nodes i and j. In order to assign weights to these edges, we first construct a physical connection matrix P with dimension R N×N If there is an edge between nodes i and j, the element P(i,j) is set to 1, otherwise P(i,j) is set to 0. Finally, the edge weight W is obtained by linear normalization of each row. p Specifically, the element W p The calculation method of (i,j) is as follows:

[0068]

[0069] 2) Functional similarity graph:

[0070] While the spatial layout of stations is crucial, it is equally crucial to assess their functional coherence. Despite being geographically distant, the connections between some stations can be attributed to their similar roles, such as commuting or commercial centers. Therefore, this functional similarity must be incorporated when capturing spatial interdependencies. To further explore this issue, Represented as site s i Historical passenger attributes of , where C represents the total passenger characteristics and TS represents the time step in interval t. i and s j The similarity weight between can be expressed as:

[0071]

[0072] According to the given similarity weight matrix W s (i,j)∈R N×N , select sites with high similarity weight to construct edge E s A threshold S(i,j) is pre-set to identify these sites as follows:

[0073]

[0074] In the modified similarity weight matrix After that, each row is normalized to facilitate training. The calculation formula of the normalized matrix is:

[0075]

[0076] 3) OD correlation diagram:

[0077] The OD (origin-destination) information of a target station is an important indicator for measuring the closeness of the connection between that station and other stations. To capture the closeness between stations, the OD features between different stations are extracted from the original AFC data to form an OD correlation graph. The OD correlation coefficient from station i to station j is defined as follows, denoted as W_c(i,j):

[0078]

[0079] Where count(i,j) represents the total number of passengers from station i to station j; N represents the total number of stations.

[0080] Determine the station with a high allocation ratio by comparing W_c(i,j) and the preset allocation ratio W_OD:

[0081]

[0082] In order to further improve the graph, the rows and columns are normalized to obtain the OD correlation weight matrix W_c(i,j):

[0083]

[0084] At each layer of the ST-Feature Learner, the output of the causal temporal attention mechanism is input, and the convolution results of the three different types of graphs mentioned above are obtained. They are then concatenated and fused into tensors to finally obtain the output of multiple predefined graph convolutions, which are expressed as follows:

[0085]

[0086] H p =Z p ||Z s ||Z e

[0087] Among them, Z p is the convolution output of the physical structure graph; Z s is the output of the passenger flow similarity graph convolution; Z c is the output of OD related graph convolution; || is the tensor connection; H p The output of multiple predefined graph convolutions.

[0088] The high-order spatial information representations extracted from two types of multiple predefined graph convolutions and dynamic graph convolutions are respectively and On this basis, the gate control value is further calculated as follows:

[0089]

[0090] Among them, H d is the output of dynamic graph convolution; is the gate value; is the bias parameter.

[0091] Therefore, the fused graph convolution output of the l-th ST feature learner layer is calculated as follows:

[0092]

[0093] in, is the output of the fused graph convolution, H l+1 is the output of the ST feature learner, and ⊙ is the Hadamard product.

[0094] The above fusion graph convolution module can derive the spatial dependency of passenger flow. In addition to spatial dependency, the passenger flow of a rail station is also affected by its historical passenger flow characteristics. Therefore, the Causal Temporal MSA (CT-MSA) is introduced. Figure 4As shown in Figure 2, its implementation is similar to that of standard MSA. However, the difference is that CT-MSA makes two major modifications by utilizing the domain knowledge of time series. They are:

[0095] Local Window: Because adjacent time steps typically exhibit stronger correlation than distant time steps, MSA is implemented within non-overlapping windows to encapsulate local interactions between time steps, resulting in a computational cost of O(TWC), effectively reducing TW compared to standard MSA (W is the window size). At the same time, to maintain the wide receptive field of standard MSA, the window size is gradually increased in different ST feature learners.

[0096] Temporal Causality: Because the passenger flow characteristics of the current step are irrelevant to their future state, we incorporate causality into MSA based on WaveNet to ensure that the model does not violate the temporal order of the input data. Causality is implemented through a masked attention mechanism. To facilitate location-aware MSA, a learnable absolute position encoding is introduced into the input of CT-MSA.

[0097] In terms of extracting passenger flow time features, the multi-head attention mechanism, compared with traditional recurrent neural networks, can process information of each time step in parallel, effectively handle the dependencies of various time spans, and improve the prediction accuracy of the model.

[0098] In MSA, each element of the input sequence is mapped into three vectors, namely query (Q), key (K), and value (V), through a separate linear transformation. The attention score is then calculated by taking the dot product of Q and K, and then a soft maximum operation is performed to ensure that the scores form a valid probability distribution. These scores are then used to calculate the weighted sum of the V vector, thereby providing a time-related representation for each specific time period in the prediction model. In the multi-head design, each "head" corresponds to an independent attention calculation process, so that temporal information of different dimensions can be extracted simultaneously, thereby enriching the expressive power of the model.

[0099] The form of MSA operation can be expressed as follows:

[0100] MSA(Q,K,V)=Concat(head1,…,head n )W o

[0101] The calculation formula for each head is:

[0102] head i =Attention(QW Qi ,KW Ki ,VW Vi )

[0103] The attention function is defined as:

[0104] Attention(Q,K,V)=softmax(QK T )V

[0105] By introducing the Causal Temporal MSA (CT-MSA), both historical passenger flow data and causal relationships are considered, thereby enhancing the processing of time series characteristics and being able to more accurately learn the temporal dynamic changes of rail passenger flow.

[0106] S3. Connect multiple spatiotemporal learners in series and use each spatiotemporal learner to process the input data separately.

[0107] Specifically, multiple spatiotemporal learners are connected in series, the output of the previous spatiotemporal learner is connected to the input of the next spatiotemporal learner, and the last spatiotemporal learner is connected to the output module. The causal temporal MSA module and the fusion graph convolution module of each learner process the input, and the residual connection ensures effective gradient backpropagation.

[0108] S4. The outputs of all the spatiotemporal learners are connected through a jump connection module, and the output of the jump connection module is introduced into an output module to obtain a prediction result through the output module.

[0109] The jump connection layer (jump connection module) can effectively avoid gradient vanishing by integrating information from different layers (different spatiotemporal learners), thereby improving the prediction accuracy of the model. Specifically, the output of each spatiotemporal learner is first summed up as a tensor, and then a 1×L i The standard convolution transforms the information connected to the output module into a sequence of length 1, where L i Represents the input sequence length of each connection layer.

[0110] The output module consists of two 1×1 standard convolutional layers to adjust the input channel dimension to achieve the desired output dimension. Specifically, if the task is to predict one step into the future, the required output dimension is 1. However, if the prediction spans Q consecutive steps, the output dimension is Q.

[0111] By finding a function (denoted by F), the function is used to predict the future passenger flow data within a specified time period. The prediction is based on a series of historical passenger flow data, represented by X = {X t-T ,…,X t-2 ,X t-1}, where the variable T represents the historical time step. Other factors, such as the predefined URT network graph (denoted as Gp, Gs, and Gc) and the dynamic space graph (denoted as Gd), are all integral to the forecast. Therefore, the mathematical expression of the function F can be expressed as:

[0112] F(X,G a ,G s ,G c ,G d )→Y t

[0113] Among them, Y t represents the predicted passenger flow at the next time step t.

[0114] The prediction of future passenger flow can be expressed as follows:

[0115]

[0116] Y t It is the combined result of the individual outputs of all L layers in the model.

[0117] The prediction process of the present invention is as follows: first, a 1x1 convolution is performed on the historical passenger flow feature data to align the dimensions. Then, a spatiotemporal learner is composed of a causal temporal MSA module and a fusion graph convolution module. Multiple spatiotemporal learners are connected in series to form the DMSTGCN model. The causal temporal MSA module and the fusion graph convolution module of each learner process the input, and the residual connection ensures effective gradient backpropagation. Then, the outputs of all learners are connected through jumps and input into the output layer to obtain the final result. In the process of fusion graph convolution, a combination of predefined and dynamic spatial graphs is used as the adjacency matrix, which can effectively learn the complex spatial characteristics of passenger flow.

[0118] The present invention overcomes the limitations of static matrices in traditional methods. Traditional graph neural network methods usually use static adjacency matrices in urban rail transit predictions, which ignores the dynamic changes in inter-station correlations throughout the day. The present invention uses the Tucker decomposition method to construct a time-varying dynamic graph, which can accurately capture the spatial correlation of passenger flow in different time intervals, effectively overcoming this limitation; and overcomes the singleness of existing methods in extracting passenger flow spatial correlation, enriching the multivariate spatial correlation extraction of passenger flow. By constructing three predefined graphs, namely the physical structure graph, the passenger flow similarity graph and the OD correlation graph, and fusing them with the dynamic graph to construct a fused graph convolution module, the multivariate spatial correlation of passenger flow is greatly enriched, providing more comprehensive information for more accurate short-term passenger flow predictions. The causal temporal multi-head attention mechanism is introduced, and together with the fused graph convolution module, a spatiotemporal learner is formed. The spatiotemporal learner can perform multi-layer spatiotemporal feature extraction, thereby improving the model's ability to extract spatiotemporal features and providing more accurate feature representation for short-term passenger flow predictions.

[0119] The beneficial effects of the present invention are as follows: by constructing a time-varying dynamic graph and a predefined graph that integrates multi-dimensional spatial correlations, and introducing a causal temporal multi-head attention mechanism to extract multi-layer spatiotemporal features, the short-term passenger flow of urban rail transit can be predicted more accurately. A large number of experimental results conducted on three real-world data sets have confirmed the superior performance of this model, providing a more reliable decision-making basis for urban traffic management and crowd control; the interpretability of dynamic spatial relationships, the present invention not only provides accurate short-term passenger flow prediction results, but also shows the intuitive results of the dynamic spatial graph, further proving its effectiveness. This interpretability enables traffic managers and relevant decision makers to better understand the dynamic changes in inter-station correlations, so as to take more targeted measures for traffic management and crowd control; the dynamic multi-graph fusion spatiotemporal graph convolutional network (DMSTGCN) proposed in the present invention is a new model architecture that integrates Tucker decomposition, predefined graph fusion and causal temporal multi-head attention mechanism. This architecture provides new ideas and methods for short-term passenger flow prediction of urban rail transit, and is expected to promote technological development in this field.

[0120] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A short-term rail passenger flow prediction method based on dynamic multi-graph fusion spatiotemporal deep learning, characterized by: The following steps are involved: S1. Convolve the historical passenger flow feature data to form input data; S2. Construct a fusion graph convolution module, wherein the fusion graph convolution module includes multiple predefined graph convolutions and dynamic graph convolutions. The multiple predefined graph convolutions are constructed by physical structure graph convolution, passenger flow similarity graph convolution, and OD correlation graph convolution. A causal temporal multi-head attention mechanism is introduced into the fusion graph convolution module to form a spatiotemporal learner. S3, connecting the plurality of spatiotemporal learners in series, and using each of the spatiotemporal learners to process the input data respectively; S4. The outputs of all the spatiotemporal learners are connected through a skip connection layer. The skip connection layer introduces the outputs of the spatiotemporal learners into an output layer, and a prediction result is obtained through the output layer.

2. The method for short-term rail passenger flow prediction based on dynamic multi-graph fusion spatiotemporal deep learning according to claim 1 is characterized by: In step S1, the historical passenger flow feature data is represented by a historical passenger flow sequence, specifically: Among them, X t represents the historical passenger flow sequence; represents the inflow data of the nth station in the tth time interval; N represents the total number of stations.

3. The method for short-term rail passenger flow prediction based on dynamic multi-graph fusion spatiotemporal deep learning according to claim 1 is characterized by: In step S2, the output of the fused graph convolution module is expressed as follows: in, is the output of the fused graph convolution; is the gate value; H p is the output of multiple predefined graph convolutions; H d is the output of dynamic graph convolution; ⊙ is the Hadamard product.

4. The method for short-term rail passenger flow prediction based on dynamic multi-graph fusion spatiotemporal deep learning according to claim 3 is characterized by: The gate value is expressed by the following formula: Among them, H p is the output of multiple predefined graph convolutions; H d is the output of dynamic graph convolution; is the gate value; is the bias parameter.

5. The method for short-term rail passenger flow prediction based on dynamic multi-graph fusion spatiotemporal deep learning according to claim 3 is characterized by: The output H of the multiple predefined graph convolutions p It can be expressed by the following formula: H p =Z p ‖Z s ‖Z c Among them, Z p is the convolution output of the physical structure graph; Z s is the output of the passenger flow similarity graph convolution; Z c is the output of OD related graph convolution; ‖ is the tensor connection.

6. The method for short-term rail passenger flow prediction based on dynamic multi-graph fusion spatiotemporal deep learning according to claim 5 is characterized by: The physical structure graph convolution output Z p , the output Z of the passenger flow similarity graph convolution s , OD related graph convolution Z c They are expressed by the following formulas: Among them, GCN(·) is the graph convolution operation; G p is the physical structure diagram; G s is the passenger flow similarity graph; G c is the OD correlation diagram; is the output of the causal temporal multi-head attention mechanism.

7. The method for short-term rail passenger flow prediction based on dynamic multi-graph fusion spatiotemporal deep learning according to claim 1 is characterized by: In step S2, the dynamic space graph can be used G d Indicates that at a specific time t, the dynamic space graph can be expressed as: Among them, S represents the site set; represents the spatial connectivity at time t; represents the corresponding weight between node i and node j.

8. The method for short-term rail passenger flow prediction based on dynamic multi-graph fusion spatiotemporal deep learning according to claim 3 is characterized by: In step S4, the output of the spatiotemporal learner can be expressed as: in, is the output of the fused graph convolution, H l+1 is the output of the spatiotemporal learner at this layer, H l It is the output of the spatiotemporal learner in the previous layer.

9. The method for short-term rail passenger flow prediction based on dynamic multi-graph fusion spatiotemporal deep learning according to claim 4 is characterized by: In step S4, the prediction result obtained by the output layer can be expressed by the following formula: Where L is the number of spatiotemporal learners, H l is the output of the previous spatiotemporal learner.

Citation Information

Patent Citations

  • Space-time Transform traffic flow prediction method based on dynamic correlation

    CN116543554A

  • Traffic scheduling method and device

    CN116629500A

  • Space-time adaptive dynamic graph convolutional network traffic flow prediction method

    CN118629226A