Holiday subway passenger flow prediction method and system based on spatiotemporal dynamic graph clustering

By using a spatiotemporal dynamic graph clustering method, combined with social media data and deep learning algorithms, subway network stations are dynamically clustered, solving the problem of insufficient accuracy in passenger flow prediction during holidays and achieving more efficient subway operation management.

CN118863984BActive Publication Date: 2026-07-21ZHEJIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2024-07-04
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing subway passenger flow forecasting methods struggle to accurately capture dynamic changes and complex station relationships during holidays, resulting in insufficient forecast accuracy and failing to meet the actual needs of subway operations.

Method used

A spatiotemporal dynamic graph clustering method is adopted, which combines social media data and deep learning algorithms. Through topology adaptive graph convolutional networks and multi-head self-attention mechanisms, subway network stations are dynamically clustered to build a passenger flow prediction model and capture passenger flow patterns and trends during holidays.

Benefits of technology

It improves the accuracy and adaptability of subway passenger flow forecasting during holidays, helps subway management departments to allocate resources rationally, alleviate passenger congestion, and provide a more comfortable passenger experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118863984B_ABST
    Figure CN118863984B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of passenger flow prediction, and in particular to a holiday subway passenger flow prediction method and system based on space-time dynamic graph clustering, which comprises obtaining holiday characteristic variable information; according to the station passenger flow characteristic variable information, the dynamic graph clustering algorithm is used to realize the subway network station group characteristics, and the stations with similar travel modes are clustered; based on the obtained holiday characteristic variable information, the results of the dynamic graph clustering algorithm are used to construct a passenger flow prediction model in different station clustering clusters, and the holiday passenger flow of the subway station is predicted. The present application adopts the GTN method combining the Transformer architecture and the graph neural network to predict the subway holiday full-network entry passenger flow. This method extends the self-attention mechanism to the subway travel network graph nodes and edges, pays weighted attention to the connection relationship between different stations, more accurately grasps the correlation between passenger flow modes of different stations, and improves the prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of passenger flow prediction technology, specifically a method and system for predicting subway passenger flow during holidays based on spatiotemporal dynamic graph clustering. Background Technology

[0002] With the acceleration of urbanization and the increase in transportation demand, subway networks have developed and expanded rapidly in many cities. As a fast and efficient mode of public transportation, the subway has become an important choice for urban residents. With the continuous extension of subway lines and the expansion of coverage, subway passenger flow has also shown a rapid growth trend, making subway passenger flow forecasting increasingly important, especially during holidays such as New Year's Day, Dragon Boat Festival, Mid-Autumn Festival, and National Day, when passenger flow fluctuations vary across different holidays.

[0003] Each holiday has its unique characteristics and influencing factors, which lead to different changes in passenger flow. Spring Festival is the most important traditional festival in China. During Spring Festival, the flow of people between cities is very active. The passenger flow from the city center to the surrounding rural areas increases significantly, while the passenger flow in the city center decreases. In addition, shopping and tourism activities also reach their peak during Spring Festival, and the passenger flow in commercial areas and tourist attractions also increases significantly.

[0004] The National Day holiday typically lasts seven days. During this time, people use the holiday for travel and visiting relatives and friends. Tourist attractions and popular destinations experience a surge in visitor traffic, while office and commercial areas typically see a decrease. Compared to the relatively stable normal period, which is constrained by commuting times, school dismissal times, and commercial activities, the mechanisms and patterns of passenger flow during holidays are completely different. Passenger flow fluctuates dramatically during this period, posing a significant challenge to passenger flow forecasting. By accurately predicting the changing trends of subway passenger flow during holidays, subway management departments can rationally arrange station staffing, better meet passenger demand, alleviate congestion, and provide passengers with a safer and more comfortable travel experience.

[0005] There is a close correlation between subway holiday travel passenger flow and social media data. Social media data can provide real-time and diverse information, helping to understand and respond to passenger flow during holidays. With the development of the internet, social media platforms have gradually become important channels for people to communicate, share, and obtain information. Weibo is one of China's largest social media platforms, similar to Twitter internationally. Its real-time nature, short texts, trending topics, and social interaction have made it popular among users, becoming one of the important platforms for Chinese internet users to share information, obtain news, and express opinions. Social media data can provide real-time, large-scale user behavior information. By monitoring user activity and topic discussions on social media, information such as people's travel intentions, travel plans, and activity participation during holidays can be obtained. This data can reflect people's travel needs and behavioral trends, thus providing real-time reference for passenger flow forecasting and helping to predict changes in passenger flow in specific time periods and regions. Social media data can provide geographic location information and user distribution, providing a geographic distribution reference for passenger flow forecasting, helping to determine passenger flow around popular attractions, business districts, or transportation hubs, and thus guiding transportation organization and resource scheduling.

[0006] Dynamic clustering of data is necessary and advantageous for predicting subway passenger flow during holidays. Dynamic clustering can dynamically identify typical characteristics and passenger flow trends between different stations based on real-time updates and trends in passenger flow data, providing a foundation for predictive modeling. This is particularly important for passenger flow changes during special circumstances such as holidays. Clustering algorithms group similar objects into the same cluster by calculating the similarity or distance between them, while assigning dissimilar objects to different clusters. Subway systems consist of numerous stations and complex line networks, with clear spatial relationships between nodes that change with passenger flow over time. Traditional static clustering methods based on feature similarity may not be able to fully capture these complex dynamic relationships and connections. Dynamic graph clustering methods, however, can capture the topology of the subway network and the connections between nodes in real time, dynamically revealing geographical, functional, and behavioral similarities within the network. This allows for a more accurate characterization of passenger flow patterns and interactions between stations, as stations within the same cluster typically share similar passenger flow trends and passenger numbers.

[0007] During holidays, passenger flow varies across different areas and stations at different times. For example, passenger flow around tourist attractions decreases the day before a holiday, while it increases in commercial centers. Conversely, on the first day of a holiday, passenger flow around tourist attractions increases, while it decreases in commercial centers. General-purpose predictive models may not be well-suited to adapt to the dynamic demand changes at various stations during holidays. Dynamic graph clustering analysis can identify different passenger flow patterns and behavioral characteristics, allowing for the selection of specific model algorithms for different holiday scenarios. This provides a foundation for customized predictive models, helping operators better predict and adjust holiday operating plans, and providing more efficient and comfortable subway services.

[0008] With the rapid development of the Internet, the Internet of Things, and digital technologies, massive amounts of data are constantly being generated, collected, and stored. The rise of big data has provided favorable conditions for the popularity of deep learning algorithms. Deep learning performs pattern recognition and feature extraction through multi-layered neural network structures. It relies on large-scale data to train and optimize models, and performs exceptionally well in clustering and prediction tasks. In clustering, deep learning algorithms can automatically cluster data by learning feature representations and similarities in real time. Compared to traditional clustering algorithms, deep learning algorithms can dynamically learn higher-level, more abstract feature representations, thereby better capturing the potential structures and patterns in irregular passenger flow data during holidays. For example, using autoencoders or variational autoencoders, unsupervised clustering can be achieved on large-scale datasets. In prediction, deep learning algorithms can learn from large amounts of multi-dimensional historical feature data to build complex models to predict future trends and outcomes.

[0009] During normal times, subway passenger flow usually follows certain patterns, such as the concentration of people during rush hour on specific time periods and subway lines. Traditional methods can predict this regular passenger flow relatively well because they are easier to model and predict. However, passenger flow during holidays often exhibits irregular, sudden, and highly volatile characteristics. Summary of the Invention

[0010] To overcome the shortcomings of existing technologies and solve the aforementioned technical problems, this invention proposes a method and system for predicting subway passenger flow during holidays based on spatiotemporal dynamic graph clustering.

[0011] On the one hand, this invention proposes a method for predicting subway passenger flow during holidays based on spatiotemporal dynamic graph clustering. The method includes the following steps:

[0012] S10: Obtain information on holiday-related variable characteristics;

[0013] S20: Based on the passenger flow characteristic variable information of the stations, the dynamic graph clustering algorithm is used to realize the clustering characteristics of subway network stations and cluster stations with similar travel patterns.

[0014] S30: Based on the obtained holiday characteristic variable information, and using the results of dynamic graph clustering algorithm, a passenger flow prediction model is constructed in different station clusters to predict the holiday passenger flow of subway stations.

[0015] Specifically, S10 includes:

[0016] S101: Obtain the number of blog posts on social media platforms to acquire passenger travel data. The number of blog posts includes the current number of blog posts and the number of blog posts at historical time steps.

[0017] S102: Define the number of blog posts in time period t as b t The number of blog posts at historical time steps is used to help predict passenger traffic for the next period; among which:

[0018] B t+1 =[b t-s ,b t-s+1 ,…,b t ]

[0019] Where s is the historical time step;

[0020] S103: Introduce holiday coding features to predict changes in subway passenger flow during holidays.

[0021] Specifically, the dynamic graph clustering algorithm in S20 includes:

[0022] S201: A topology-adaptive method is used for graph embedding to generate a spatial topology that adapts to the input data;

[0023] S202: The solver of the lower-level clustering task is embedded as a differentiable layer and incorporated into the training process of the deep learning model;

[0024] S203: Establish a topology-adaptive graph convolutional network and construct a dynamic travel network topology graph based on passenger travel data. in ε represents the nodes of the graph, i.e., stations, and ε represents the edges of the graph, i.e., the connections between stations. The characteristics of the edges in the graph are described by the amount of travel between the start and end points.

[0025] S204: Using an adjacency matrix Describes the spatial connections between stations, where N is the number of stations;

[0026] S205: Graph node feature matrix x t This includes the station's longitude and latitude, daily passenger flow demand for entering and exiting the station, and hourly passenger flow demand for entering and exiting the station.

[0027] Specifically, the establishment of a topology-adaptive graph convolutional network in S203 includes:

[0028] S2031: Create a filter in the Fourier domain using a graph convolutional network model to perform convolution operations on graph nodes;

[0029] S2032: Captures spatial features between nodes by capturing their first-order neighbor relationships, and constructs a graph convolutional network model by stacking multiple convolutional layers, as shown in the following formula:

[0030]

[0031] in, It is the normalized adjacency matrix of the graph. It is an adjacency matrix, I N It is the identity matrix. It is the degree matrix of the travel network l represents the l-th hidden layer. This represents the f-th graph filter. These are the polynomial coefficients of a graph filter; graph convolution is the product of a matrix and a vector. It is the f-th output feature map, b f It is a learnable bias term. It is the output of layer l, and σ is the activation function such as ReLU;

[0032] S2033: Formula Improved to This allows the filters in a topology-adaptive graph convolutional network to dynamically adapt to the graph's topology during the convolution process, where k is the filter size.

[0033] The following formula is used to calculate the weight of information transmission between nodes;

[0034]

[0035] Where p represents a path of length k from node j to i, and the node sequence of p is...

[0036] S2034: Module degree is used as a loss function for error backpropagation to optimize parameters and iteratively improve graph representation learning and cluster assignment. Module degree is denoted by Q and calculated using the following formula:

[0037]

[0038] Where m is the total number of edges in the network; c is the number of cluster categories; A ij It is the adjacency matrix of the network, indicating whether there is a connection between node i and node j; d i It is the degree of node i; if node i is assigned to class c, r icIt is 1 if it is true, otherwise it is 0.

[0039] Specifically, the passenger flow prediction model built in S30 includes:

[0040] S301: Establish a multi-head self-attention mechanism to obtain relevant information in different subspaces. The multi-head attention calculation formula is as follows:

[0041]

[0042] in, Let d represent the input vector. c The hidden layer size is the c-th attention head. For each attention head, it is first set with different trainable parameters. Each Convert to query vector key vector Sum value vector

[0043] S3022: Encode the edge features and add them to the key vector as additional information for each layer. The multi-head attention mechanism for the edge from node j to node i is calculated by the following formula:

[0044] e c,ij =W c,e e ij +b c,e #(14)

[0045]

[0046] in, For the characteristics of a node, e ij Features of the edges;

[0047] S3023: Based on graph multi-head attention, message aggregation from distant node j to source node i is calculated by the following formula:

[0048]

[0049] Where || represents the concatenation operation of C attention heads, and the multi-head attention matrix replaces the original normalized adjacency matrix as the transition matrix for message passing;

[0050] On the other hand, this invention proposes a holiday subway passenger flow prediction system based on spatiotemporal dynamic graph clustering, which includes:

[0051] The acquisition module is used to obtain information on holiday-related variable characteristics.

[0052] The clustering module is used to cluster stations with similar travel patterns by using a dynamic graph clustering algorithm based on station passenger flow characteristic variables.

[0053] The prediction module is used to construct a passenger flow prediction model in different station clusters based on the obtained holiday characteristic variable information and the results of dynamic graph clustering algorithm, so as to predict the holiday passenger flow of subway stations.

[0054] The beneficial effects of this invention are as follows:

[0055] 1. This invention employs the GTN method, which combines the Transformer architecture and graph neural networks, to predict passenger flow entering subway stations across the entire network during holidays. This method extends the self-attention mechanism to nodes and edges in the subway travel network graph, giving weighted attention to different stations and the connections between them, thus more accurately grasping the correlation of passenger flow patterns between different stations and improving prediction accuracy.

[0056] 2. This invention considers multi-dimensional data such as social media data and holiday coding. Social media data can provide information on people's activities, interests, and behaviors during holidays; holiday coding data provides information such as holiday types and durations, enabling a more comprehensive understanding of people's travel destination preferences during holidays and better capturing the periodic changes and special patterns of passenger flow.

[0057] 3. This invention proposes an end-to-end dynamic graph clustering algorithm, TAkmeans, for dynamic clustering of subway travel networks. Unlike a simple combination of graph convolutional networks and clustering techniques, it iteratively optimizes neural network parameters through graph representation and clustering, integrating the two into a unified framework and achieving fully automated training. Attached Figure Description

[0058] The invention will now be further described with reference to the accompanying drawings.

[0059] Figure 1 This is a flowchart of the training process to maximize accuracy in this invention;

[0060] Figure 2 This is a flowchart of the training process to maximize decision quality in this invention;

[0061] Figure 3 This is a partial topology diagram of the Beijing subway line provided in this invention;

[0062] Figure 4 This is a graph showing the fluctuation of passenger flow in and out of a commercial area station during the National Day holiday in 2023 and during a normal week after the National Day holiday, as provided in this invention.

[0063] Figure 5 This refers to the number of blog posts published on different days within a week, as provided in the experiment conducted by this invention.

[0064] Figure 6This is a correlation coefficient matrix diagram of passenger flow and blog post volume in the experiment provided by this invention;

[0065] Figure 7 This is a diagram showing the evolution of local graph clustering of the dynamic travel network of Beijing Metro over time, as provided in this invention.

[0066] Figure 8 This is a comparison chart of the actual passenger flow and GTN algorithm prediction values ​​for different functional area stations during a normal week and the National Day holiday, provided by the present invention.

[0067] Figure 9 This is a flowchart of the method in this invention. Detailed Implementation

[0068] This invention provides a method for predicting subway passenger flow during holidays based on spatiotemporal dynamic graph clustering. It adopts the GTN method, which combines the Transformer architecture and graph neural network, to predict the passenger flow entering the entire subway network during holidays. This method extends the self-attention mechanism to the nodes and edges of the subway travel network graph, and gives weighted attention to different stations and the connection relationships between stations, so as to more accurately grasp the correlation of passenger flow patterns between different stations and improve the prediction accuracy.

[0069] After introducing the basic principles of the present invention, various non-limiting embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0070] Example 1:

[0071] like Figure 9 As shown in the figure, this invention provides a method for predicting subway passenger flow during holidays based on spatiotemporal dynamic graph clustering. The method includes the following steps:

[0072] S10: Obtain information on holiday-related variable characteristics;

[0073] S20: Based on the passenger flow characteristic variable information of the stations, the dynamic graph clustering algorithm is used to realize the clustering characteristics of subway network stations and cluster stations with similar travel patterns.

[0074] S30: Based on the obtained holiday characteristic variable information, and using the results of dynamic graph clustering algorithm, a passenger flow prediction model is constructed in different station clusters to predict the holiday passenger flow of subway stations.

[0075] Specifically, firstly, regarding holiday-related variable information, social media data provides a way to indirectly reflect changes in subway passenger flow during holidays compared to weekdays. During holidays, people tend to share their travel plans, activity schedules, and attraction recommendations on social media.

[0076] The quantity and content of these social media campaigns reflect people's interest in and demand for travel during the holiday season. Blog postings are a key indicator of social media activity, reflecting user engagement, activity levels, and travel information.

[0077] Monitoring blog post volume on social media platforms can yield valuable insights. For example, an increase in blog posts may indicate that more people plan to travel by subway during holidays, or that there is more traffic flowing to certain popular destinations. Conversely, a decrease in blog posts may suggest that people are choosing alternative modes of transportation or opting to stay home during holidays. By incorporating blog post volume into subway passenger flow prediction models, this invention can better understand and predict passenger flow during holidays.

[0078] This predictive model, which comprehensively considers social media data, can help subway operators more accurately predict holiday passenger flow and adjust capacity and operational strategies accordingly, alleviating the pressure caused by surges in passenger traffic. This invention statistically counts the number of blog posts using "Beijing subway station" as a keyword across different time periods, defining the number of blog posts in time period t as b. t The number of blog posts at historical time steps is used to help predict passenger traffic in the next period.

[0079] B t+1 =[b t-s ,b t-s+1 ,…,b t ]

[0080] Where s is the historical time step.

[0081] To accurately predict subway passenger flow changes during holidays, it's necessary to consider not only the passenger flow on the holiday day itself, but also the day before, the day after, and even longer timeframes. Passenger flow patterns can differ significantly between different days of a holiday. For example, the day before a holiday is typically when people are preparing for the holiday. People need to go to stores to buy groceries and gifts, or to the city center to purchase special items.

[0082] Therefore, the subway may face slightly increased passenger flow on this day. On holidays, people tend to stay home to celebrate rather than go out frequently, meaning the subway system may experience relatively stable passenger flow. However, the day after a holiday presents a completely different picture. People begin to visit relatives and friends, tour attractions, or travel. On this day, the subway system may face higher passenger flow peaks, requiring reasonable scheduling and arrangements in advance to ensure a smooth travel experience for passengers. To incorporate the impact of these date types into the passenger flow prediction model, integer encoding can be used to add holiday characteristics as a discrete feature column to the prediction model.

[0083] The model can learn passenger flow patterns and trends for different date types based on this feature, enabling more accurate predictions of passenger flow changes during holidays. Different date types can be represented by integers such as -1, 0, 1, 2, 3, and 4. For example, -1 represents the day before a holiday, 0 represents a weekday (not a holiday), 1 represents the first day of a holiday, 2 represents the second day of a holiday, 3 represents the third day of a holiday, and so on. By introducing holiday coding features, the model can learn that passenger flow is likely to be higher the day before a holiday and lower on weekdays (not holidays). This feature helps the model more accurately predict changes in subway passenger flow during holidays.

[0084] Secondly, this embodiment discloses a novel method, TAkmeans, that combines node feature representation and clustering algorithms to dynamically mine the spatiotemporal travel network structure of subways through graph learning and optimization. It surpasses the simple combination of graph convolutional networks and clustering techniques. The simple combination first trains the model using standard loss, then inserts the obtained node representations into the clustering algorithm; the two stages are completely separate, as shown below. Figure 1 As shown. This training method, which minimizes the standard loss function (e.g., cross-entropy), is not designed for a specific clustering task, thus limiting clustering performance. The graph clustering method introduced in this section employs a topology-adaptive approach for graph embedding, generating a spatial topology that adapts to the input data. It embeds the solver of the lower-level clustering task as a differentiable layer into the training process of the deep learning model. It iteratively optimizes the neural network parameters through graph representation and clustering, using the quality of downstream problem resolution as the loss function, integrating both into a unified framework to achieve fully automated training, such as... Figure 2 As shown.

[0085] Graph Convolutional Networks (GCNs) overcome the limitation of convolutional neural networks (CNNs) in processing Euclidean-structured data. They consider neighboring nodes and update the features of the current node by aggregating information from neighboring nodes. This information propagation mechanism enables GCNs to effectively handle non-Euclidean-structured data, demonstrating excellent performance in extracting spatial dependencies in subway networks. However, GCNs use fixed-size convolutional kernels for smooth shifting, pooling, and non-linear activation functions, and assume that the topology of the input graph is fixed. In contrast, the subway system is a dynamic transportation network, influenced by various factors such as peak / off-peak hours, weekdays / weekends, and holidays. The topology of the subway's spatiotemporal travel network is dynamically changing. To dynamically mine the changes in the subway's spatiotemporal travel network and discover different travel network structures, this section proposes a representation learning model called Topology Adaptive Graph Convolutional Networks (TAGCNs). TAGCNs combine adaptive neighbor sampling and adaptive weight aggregation mechanisms to adapt to graph data with different topological structures.

[0086] Constructing a dynamic travel network topology based on passenger travel data in ε represents the nodes of the graph, i.e., the stations, and ε represents the edges of the graph, i.e., the connections between stations. Adjacency matrix. This describes the spatial connections between stations, where N is the number of stations. The node feature matrix x of the graph... t It consists of the station's longitude and latitude, daily passenger flow demand for entering and exiting the station, and hourly passenger flow demand for entering and exiting the station. The features of the edges in the graph are described by the travel volume between the origin and destination. The GCN model creates a filter in the Fourier domain, performs convolution operations on the graph nodes, captures the spatial features between nodes by capturing the first-order neighbor relationships, and constructs the GCN model by stacking multiple convolutional layers, as shown in Equation 2-5.

[0087]

[0088] in, It is the normalized adjacency matrix of the graph. It is an adjacency matrix, I N It is the identity matrix. It is the degree matrix of the travel network l represents the l-th hidden layer. This represents the f-th graph filter. These are the polynomial coefficients of a graph filter; graph convolution is the product of a matrix and a vector. It is the f-th output feature map, b fIt is a learnable bias term. σ is the output of layer l, and σ is the activation function such as ReLU.

[0089] TAGCN is a variant of GCN. First, TAGCN dynamically selects each node's neighbor nodes through adaptive neighbor sampling. Figure 3 This is a partial topology diagram of the Beijing subway system. Edge connections represent passenger entry and exit between the origin and destination (OD) points within a certain time period, and the OD passenger flow is the edge weight. Taking station 1 (Nanluoguxiang Station) as an example, the GCN, which only uses a Chebyshev polynomial with k=1 to approximate the convolution kernel, simply considers fixed first-order neighbor stations in the network, i.e., only stations 2 and 7. In contrast, TAGCN employs a diverse set of k-localized filters to implement graph convolution in the vertex domain. These filters have variable sizes, ranging from 1 to K, allowing each filter to be flexibly adjusted according to the characteristics of the graph. That is, TAGCN will sequentially consider first-order neighbor stations 2 and 7, second-order stations 3, 4, and 6, and so on. This flexibility is achieved through the polynomial coefficients of the filters. This is achieved through [the following]. Therefore, TAGCN improves Equation 3 of GCN to Equation 6, enabling the TAGCN filter to dynamically adapt to the graph topology during convolution.

[0090]

[0091] Where k is the size of the filter.

[0092] Secondly, TAGCN adjusts the information transfer weights between nodes through adaptive weight aggregation. Traditional GCNs use fixed weight aggregation methods, which cannot adapt to the relationships between different nodes. TAGCN calculates the importance of each node's neighboring nodes and dynamically adjusts the weights based on importance, enabling each node to better integrate information from its neighbors. The calculation method for the information transfer weights between nodes is shown in Formula 7. For example, Figure 3 There are two paths with a length of 3 (i.e., k=3) from station 1 to station 6: (1,2,7,6) and (1,2,3,6), represented by blue and purple shading respectively. Therefore, according to Formula 7, we can calculate... So, what is the output of layer f in TAGCN? The result is obtained by modifying Formula 4 into Formula 8.

[0093]

[0094] Where p represents a path of length k from node j to c, and the node sequence of p is...

[0095] After extracting node and edge features using the TAGCN layer, the structural and attribute information of the original travel network is represented in a reduced-dimensionality manner. Subsequently, the output of the hidden layer is passed as input to the clustering module, where this invention uses a differentiable K-means clustering algorithm. K-means clustering aims to divide data points into different clusters by assigning them to the nearest centroid. The core idea of ​​this algorithm is iterative optimization to find the optimal cluster centers, thereby achieving data clustering and grouping. However, the k-means algorithm itself is a non-differentiable process because it involves a hard threshold function and discrete cluster assignment steps. To achieve differentiable k-means clustering, this invention employs the Gumbel Softmax approximation method, using the Gumbel distribution to approximate the discrete cluster assignment process and using the Softmax function to associate the probability distribution with the data points. Specifically, it uses the Gumbel-Softmax distribution to assign a probability distribution to each data point, representing the probability that the data point belongs to each cluster. Then, the centroids of the clusters are updated by maximizing the expected distance between the data points and the cluster centroids. This makes the entire process differentiable, and optimization methods such as gradient descent can be used to optimize centroid and cluster assignment. Backpropagation is allowed during optimization, enabling joint training with other neural network models.

[0096] This invention proposes a graph clustering algorithm, TAkmeans, that combines feature representation and clustering. It uses TAGCN to extract and represent the dynamic spatiotemporal travel network structure and features, and then employs a differentiable K-means clustering algorithm to group data based on similarity and assign them to corresponding clusters. The final clustering results are used as a loss function, with modularity as the loss function, for error backpropagation to optimize parameters and iteratively improve graph representation learning and cluster assignment. This framework fully integrates graph representation learning and optimization, enabling customized training. Modularity is an indicator used to measure the quality of graph network structure partitioning. It measures the difference between the actual graph network structure and the stochastic expected network structure. If a graph network structure has a high modularity value, it means that the connections between nodes in the structure are more confined within subgraphs, while the connections between subgraphs are fewer, indicating a better network structure partitioning. Typically, a Q value between 0.3 and 0.7 is considered to indicate strong network modularity, calculated using Equation 9.

[0097]

[0098] Where m is the total number of edges in the network; c is the number of cluster categories; A ij It is the adjacency matrix of the network, indicating whether there is a connection between node i and node j; d i It is the degree of node i; if node i is assigned to class c, r ic It is 1 if it is true, otherwise it is 0.

[0099] Finally, this invention discloses a prediction algorithm, GTN, combining Transformer and Graph Neural Networks. By extending the self-attention mechanism to nodes and edges in the subway network to capture local and global dependencies, it better predicts subway passenger flow during holidays. In a recurrent neural network, each neuron in the hidden layer receives the state of the previous neuron and the current input, calculates the new state and output, and then passes it to the next neuron. This limits the size of its input and output, which is usually fixed in a given task, and its sequential computational nature makes parallel computation impossible, resulting in slow training speed and difficulty in handling long-distance dependencies. To address these issues, Vaswani et al. proposed the Transformer model in 2017, a sequence-to-sequence model based on a self-attention mechanism. The main components of the Transformer model include an encoder and a decoder. The encoder consists of multiple identical layers stacked together, each layer comprising two sub-layers: a multi-head self-attention mechanism and a feedforward neural network. The multi-head self-attention mechanism allows the model to pay attention to all positions in the input sequence at each time step and weights the sum of the values ​​by calculating attention weights. The feedforward neural network is a fully connected feedforward network used to perform a nonlinear transformation on the output of the self-attention mechanism. The decoder also consists of multiple identical layers stacked together, each layer comprising three sub-layers: a multi-head self-attention mechanism, an encoder-decoder attention mechanism, and a feedforward neural network. The encoder-decoder attention mechanism allows the decoder to pay attention to the encoder's hidden representations and generate context vectors based on the attention weights. This enables the decoder to acquire information from the encoder to help generate the correct output sequence.

[0100] The Transformer class, by introducing multi-head self-attention and positional encoding, considers all positions in the input sequence simultaneously and allows for parallel computation. This enables the Transformer model to better capture long-range dependencies in the sequence and achieves faster training speeds. Positional encoding embeds positional information from the sequence into the model by adding positional information to the embedding vector of the input sequence, allowing the model to perceive the relative order of different positions in the sequence. The multi-head self-attention mechanism is essentially composed of multiple independent, parallel attention mechanisms, which, through this concatenation, can obtain relevant information across different subspaces. Equations 10-13 are the calculation formulas for multi-head attention.

[0101]

[0102] in, Let d represent the input vector. cThe hidden layer size is the c-th attention head. For each attention head, it is first set with different trainable parameters. Each Convert to query vector key vector Sum value vector

[0103] GTN combines the Transformer architecture with graph neural networks, utilizing a self-attention mechanism to weight different parts of the graph data during representation learning. It determines the importance of each node or edge by calculating the similarity between the node or edge and its neighbors, and then weights and sums the features of its neighbors based on importance. This mechanism allows it to better capture the correlation between node and edge features and their surroundings when learning them, thereby improving the understanding and representation of graph structures. A significant advantage of GTN compared to traditional graph neural networks is its ability to automatically learn useful connections and node representations from raw graph data, reducing reliance on manually designed features, enhancing the model's generalization and adaptability, providing broader applicability to various types of graph data, and achieving state-of-the-art performance across different domains and tasks. Unlike processing sequential data, GTN operates on the graph's adjacency matrix and feature matrix, encoding the graph's structural information into the self-attention mechanism and calculating a weighted sum of neighbor node representations. Specifically, it encodes edge features and adds them to the key vector as additional information for each layer; the multi-head attention mechanism for the edge from node j to node i is calculated by Equation 15.

[0104] e c,ij =W c,e e ij +b c,e

[0105]

[0106] in, For the characteristics of a node, e ij Features of the edges.

[0107] After obtaining the multi-head attention of the graph, the message aggregation from the distant node j to the source node i is calculated by Equation 16.

[0108]

[0109] Here, || represents the concatenation operation of C attention heads, and the multi-head attention matrix replaces the original normalized adjacency matrix as the transition matrix for message passing.

[0110] Example 2: This invention provides a holiday subway passenger flow prediction system based on spatiotemporal dynamic graph clustering. The system includes:

[0111] The acquisition module is used to obtain information on holiday-related variable characteristics.

[0112] The clustering module is used to cluster stations with similar travel patterns by using a dynamic graph clustering algorithm based on station passenger flow characteristic variables.

[0113] The prediction module is used to construct a passenger flow prediction model in different station clusters based on the obtained holiday characteristic variable information and the results of dynamic graph clustering algorithm, so as to predict the holiday passenger flow of subway stations.

[0114] To demonstrate the accuracy of the holiday subway passenger flow prediction method based on spatiotemporal dynamic graph clustering provided in this invention, the following experiments were conducted:

[0115] This study analyzes passenger flow patterns and trends during holidays using Automatic Fare Collection (AFC) card swipe data and social media data from Beijing's urban rail transit system in September and October 2023. AFC card swipe data refers to passenger card swipe information recorded by the rail transit system, including the swipe time and station. Analysis of AFC card swipe data reveals differences in subway travel patterns during holidays compared to weekdays. Figure 4 This represents the fluctuations in passenger flow during the National Day holiday and the normal week following the holiday at a certain commercial area station in 2023.

[0116] It can be seen that the overall travel patterns during holidays differ significantly from those on normal weeks. On a typical weekday, the station in this commercial area saw 19,892 passengers entering the station; however, on a holiday weekday, the number reached 38,111, an increase of 18,219 passengers compared to a normal day. Therefore, accurately predicting passenger flow, especially during holidays, remains a necessary and challenging task. Social media data is collected from user-posted content related to subway travel on social media platforms (mainly Weibo), such as travel plans, experience sharing, and attraction recommendations. Analyzing social media data can provide information about travel purposes during holidays, visitor traffic at popular attractions, and passenger experiences. Figure 5 The chart shows the number of blog posts published on different days of the week, with the horizontal axis representing the time period and the vertical axis representing the number of posts. The number of posts published in each time period of the day is cumulatively stacked on the bar chart. It can be seen that people tend to publish blogs related to subway travel at the beginning of the holiday, between 10:00 and 12:00 noon. For example, 17 blogs were published between 11:30 and 12:00 on Tuesday, while only 2 were published between 20:30 and 21:00.

[0117] Through the quantitative description and analysis of social media data and subway inbound passenger flow data, the correlation between the number of Weibo posts and the subway inbound passenger flow can be explored. Taking 30 minutes as the time granularity, count the number of Weibo posts and the subway inbound passenger flow within different unit times, and draw a correlation coefficient matrix diagram as shown in Figure 6 shown.

[0118] The correlation coefficient matrix diagram can visually express the correlation between the number of blog posts and the inbound flow. The numbers in the diagram represent the Pearson correlation coefficient. The closer the value is to 1, the stronger the positive correlation between the two variables. For example, the correlation coefficient between the inbound passenger flow on October 3 and the blog release volume on October 2 is 0.52. The larger the origin in the diagram, the stronger the correlation. The positive correlation is red, and the negative correlation is blue; the redder the origin, the closer the correlation coefficient is to 1, and the stronger the positive correlation. The * in the origin is a significance marker. When the significance level p <= 0.001, it is marked as ***; when the secondary correlation is 0.001 < p <= 0.01, it is marked as **; when p <= 0.05, it is marked as *. It can be seen from the diagram that there is a significant positive correlation between the number of blog posts and the inbound passenger flow at the significance level of 0.05.

[0119] Among them, the comparative analysis of the dynamic graph clustering algorithm is as follows:

[0120] The dynamic graph clustering method TAkmeans first obtains node features through representation learning, and then performs clustering analysis on these node features to obtain the clustering results in the subway network. The present invention conducts a comparative analysis on the proposed TAkmeans in terms of the graph neural network algorithm for representation learning and the clustering task based on the Beijing subway network. In terms of the selection of the graph neural network algorithm for representation learning, compare algorithms such as GCN and Graph Attention Network (GAT) to select the most suitable one. In terms of the selection of the algorithm for the clustering task, the present invention will compare the traditional K-means algorithm and some other commonly used clustering algorithms, such as K-means++ and Mean Shift. By comparing their performance in the subway network clustering task, the most suitable clustering algorithm for this study can be determined, and its differences from other algorithms can be evaluated.

[0121] Among them, GCN can adapt to the spatial correlation between stations in the subway network and use information such as the geographical location of adjacent stations and the inbound and outbound passenger flow to enrich the feature representation, and has greater flexibility in dealing with non-Euclidean features and graph structure relationships;

[0122] GAT is a graph neural network model based on the attention mechanism. It effectively captures the important relationships between subway stations by learning the attention weights between stations and their neighbor stations in the subway network, thereby improving the performance of the graph neural network;

[0123] Simplifying Graph Convolutional Networks (SGCN) are models for processing spatiotemporal graph structure data. Combining GCN and spatiotemporal modeling methods, it effectively captures spatiotemporal relationships and evolutionary patterns between nodes, making it suitable for spatiotemporal data analysis and prediction tasks. This is particularly significant for the dynamic identification of communities in subway networks where network structure and passenger flow at stations are constantly changing node characteristics.

[0124] Graph Sampling and Aggregated Convolution (GraphSAGE) updates node representations by sampling and aggregating features from neighboring nodes, effectively capturing both local and global information in graph data. GraphSAGE reduces computational complexity while preserving node contextual information, making it well-suited for processing rail networks containing numerous subway stations and lines.

[0125] The Feature Statistics Convolution (FeaStConv) algorithm updates graph data features by modeling statistical information (such as maximum, minimum, mean, and variance) of node features, including passenger flow and location. It can more comprehensively capture the distribution information of node features, effectively analyzing and understanding the similarities and differences between different stations, and their roles in the entire subway network.

[0126] Local Edge Convolution (LEConv) determines the degree of association between nodes and edges by calculating the similarity between node features and edge features, and uses this similarity to update node representations. This helps to understand the connectivity and importance between different stations in a metro network, providing valuable insights for the dynamic partitioning of the metro network.

[0127] Residual Gated Graph Convolution (ResGatedConv) is a graph convolutional layer that combines residual connections and gating mechanisms. It adds the original features to the updated features through residual connections to preserve the original information, and controls the flow of information and the importance of features through gating mechanisms, thereby better capturing the correlation and feature distribution between stations in the subway network and making more reasonable classifications of the subway network.

[0128] Table 1 shows the comparison results of the proposed TAGCN algorithm with other algorithms in terms of modularity values ​​during weekdays and holidays. Modularity is an indicator of the quality of graph network structure partitioning. It measures the difference between the actual graph network structure and the stochastic expected network structure. The closer the modularity value is to 1, the more nodes in the graph network structure are connected within subgraphs, and the fewer connections are between subgraphs, indicating a better network structure partitioning. On weekdays, GCN achieved a modularity value of 0.683 on the subway network. SGCN, however, only achieved 0.519, and by using GAT, ResGatedConv, and LEConv layers, it achieved modularity values ​​of 0.687, 0.698, and 0.710, respectively. However, the modularity value using the TAGCN layer was higher, reaching 0.769. During holidays, the modularity values ​​of GCN and GAT layers on the subway network were 0.669 and 0.693, respectively, while TAGCN achieved a higher modularity value of 0.770. These results demonstrate that the TAGCN algorithm for data representation performs best in metro networks and achieves a higher modularity value compared to the benchmark algorithm.

[0129] Table 1. Comparison of modularity between TAGCN and other benchmark algorithms

[0130]

[0131] Note: The clustering algorithms used in the table for subsequent clustering tasks all employ the differentiable K-means algorithm.

[0132] Regarding algorithm selection for clustering tasks, this invention compares the performance of the classic K-means algorithm with several commonly used clustering algorithms. Table 2 compares the modularity results of different clustering algorithms (including differentiable K-means, K-means++, and MeanShift) on subway travel networks. K-means++ is an improved k-means clustering algorithm that uses a probability-weighted approach when selecting initial centroids, making data points farther from the selected centroids more likely to become the next centroid, thereby improving the algorithm's convergence speed. Mean Shift is a density-based nonparametric clustering algorithm used to discover cluster structures in data. Its main idea is to locate cluster centers by iteratively finding the direction of the maximum density gradient of data points. Differentiability of K-means++ and Mean Shift can also be achieved using the Gumbel Softmax approximation method. However, since K-means++ itself is an iterative, locally optimized process, it may get trapped in local optima during the optimization process.

[0133] The computational complexity of the Mean Shift algorithm is directly proportional to the size and dimensionality of the dataset. For large-scale datasets and high-dimensional data, the execution time of the algorithm increases significantly, resulting in low efficiency. As shown in Table 2, for the subway travel dataset, the differentiable K-means clustering algorithm yields the best clustering results for most feature extraction algorithms. For example, for the LEConv feature extraction algorithm, the modularity of the differentiable K-means clustering algorithm is 0.710 during weekdays, while the modularity of K-means++ and Mean Shift clustering algorithms are 0.618 and 0.360, respectively, significantly lower than K-means. The same applies during holidays; the K-means algorithm with a modularity of 0.609 performs better than K-means++ and Mean Shift clustering algorithms with modularity values ​​of 0.589 and 0.375, respectively. Furthermore, using TAGCN for feature extraction and differentiable K-means, K-means++, and Mean Shift clustering algorithms for community partitioning, the modularity values ​​obtained during normal days were 0.769, 0.756, and 0.579, respectively, while the modularity values ​​during holidays were 0.770, 0.718, and 0.557. This result demonstrates that the TAGCN algorithm achieves higher modularity compared to other graph-based feature extraction algorithms (such as LEConv and GAT) and clustering algorithms (such as K-means++ and Mean Shift). It is also worth noting that the Mean Shift algorithm requires more training time.

[0134] Table 2 Comparison of Module Degree of Clustering Algorithms

[0135]

[0136]

[0137] Figure 7 This is a clustering result of a partial map of the Beijing Metro dynamic travel network over time. The color of a station represents its category, with stations of the same category represented by the same color. The connection between stations is determined by whether there is passenger flow during that time period, and the weight (thickness) and color of the connecting edges reflect the size of the passenger flow. Figure 7 (a) shows the local clustering results for the period from 7:00 to 7:30 AM. It can be seen that the Takmeans algorithm clusters sites with strong connections into category A (red), while sites with weaker connections to category A are classified into other categories (yellow, purple, etc.). As time progresses, from 12:00 to 12:30 PM, the connections between some sites and category A sites gradually strengthen, thus some sites may move from other categories to category A, such as... Figure 7As shown in (b), this shift is highlighted by the green arrow. During the 18:00-18:30 period, there was a significant decrease in OD (Original Demand) traffic between some stations, causing some stations to fall out of their original category A and be reclassified into other categories, such as... Figure 7 As shown in (c), this transformation is highlighted by the green arrow. As time progresses, during the 21:00-21:30 period, the classification of stations changes according to the strength of the connectivity between them, as the OD (Original Demand) volume between stations increases or decreases. Figure 7 As shown in (d).

[0138] The specific analysis is as follows:

[0139] To investigate the impact of various feature variables on passenger flow prediction, this invention developed several models for experimentation. First, there's the M-FL model, which uses only historical time steps, historical passenger flow data for the same period, and the station's latitude and longitude as graph node features. To explore the impact of clustering on improving passenger flow prediction accuracy, the M-FLC model was constructed, which incorporates station category features obtained using the TAkmeans graph clustering algorithm, in addition to the passenger flow time series features and station location features. To evaluate the impact of social media data and holiday integer encoding on holiday passenger flow, this invention created the M-FLCSH model, which incorporates historical time step blog post volume and holiday features.

[0140] To evaluate the effectiveness of the model integrating clustering, social media data, and integer-encoded holiday features for predicting passenger flow on weekdays and holidays, this invention used real Beijing Metro AFC data from the month of National Day and the month before National Day for experiments. The proposed prediction model GTN was compared with various baseline models for travel demand prediction, including Random Forest (RF), Gradient Boosting Decision Trees (GBDT), GCN, GAT, SGCN, and TAGCN. The model's accuracy was measured using the coefficient of determination (R²). 2 The mean absolute error (MAE) and root mean square error (RMSE) are used for evaluation.

[0141]

[0142] Where y t and These represent the actual and predicted passenger flow at time t. Let t be the average number of passengers actually entering the station at time t, T be the final end time, and N be the total number of time points experienced from the start to the end.

[0143] Table 3 shows the performance of the M-FL, M-FLC, and M-FLCSH models in predicting passenger flow on holidays and weekdays. The results indicate that GTN performs better than other prediction algorithms on both weekdays and holidays. For example, during weekdays, RF, SGCN, and TAGCN show lower R² values ​​on the M-FLCSH model. 2 The R values ​​were 0.71, 0.46, and 0.72, respectively; the MAE values ​​were 92, 138, and 94, respectively; and the RMSE values ​​were 258, 335, and 253, respectively. Meanwhile, the R value of GTN on the M-FLCSH model was... 2 The R² values ​​for MAE, RMSE, and RF were 0.75, 82, and 220, respectively, representing improvements of 4.17%, 10.86%, and 13.04% compared to the best baseline model. During the holiday period, RF, SGCN, and TAGCN showed improved R² values ​​on the M-FLCSH model. 2 The R values ​​were 0.71, 0.40, and 0.73, respectively; the MAE values ​​were 96, 122, and 89, respectively; and the RMSE values ​​were 259, 306, and 202, respectively. Meanwhile, the R value of GTN on the M-FLCSH model was... 2 The MAE and RMSE were 0.75, 75, and 195, respectively. Furthermore, adding clustering, social media data, and integer-encoded holiday features can improve prediction accuracy. During normal days, GTN achieved a higher R² value on M-FL. 2 The MAE and RMSE values ​​were 0.69, 103, and 251, respectively; while in the M-FLC model with added clustering features, R... 2 The MAE and RMSE were 0.72, 119, and 245 respectively, indicating a certain improvement in prediction accuracy. This suggests that clustering is beneficial for improving prediction accuracy. On the M-FLCSH model, R... 2 The MAE and RMSE values ​​are 0.75, 82, and 220 respectively. Compared to M-FL, incorporating clustering, social media data, and holiday integer encoding features can improve R... 2 The MAE increased by 8.70%, the RMSE increased by 20.39%, and the RMSE increased by 12.35%.

[0144] Figure 8 The chart shows a comparison between the actual passenger flow and the GTN algorithm prediction values ​​for different functional area stations during a typical week and the National Day holiday.

[0145] Figure 8 (a)(b) show the comparison of passenger flow forecasts for residential area stations during weekdays and holidays. It can be seen that passenger flow at residential area stations decreases significantly during holidays. Although the forecast and actual passenger flow trends are generally consistent, there is still a lot of room for improvement. Figure 8(c)(d) show the passenger flow forecast results for commercial area stations during weekdays and holidays. Due to factors such as commuting to work, the population flow in commercial areas is relatively high on weekdays, but less on weekends. In the first few days of holidays, as people shift more from commuting to leisure and tourism, the passenger flow decreases significantly. The GTN algorithm captures this large fluctuation pattern of passenger flow during holidays very well, and the forecast is relatively accurate. Figure 8 (e)(f) show the comparison results of passenger flow predictions for scenic spot stations during weekdays and holidays. It can be seen from the figure that the passenger flow of scenic spot stations increases significantly during holidays, from hundreds to thousands. Moreover, the passenger flow during holidays no longer shows a clear difference between weekdays and rest days. However, the GTN algorithm proposed in this invention can adapt well to this fluctuation and make relatively accurate predictions.

[0146] Table 3 shows the performance of various models that integrate different features in predicting passenger flow to the station.

[0147]

[0148]

[0149] Note: This table uses a clustering-then-prediction framework to predict passenger flow.

[0150] The above embodiments illustrate and describe the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for predicting subway passenger flow during holidays based on spatiotemporal dynamic graph clustering, characterized in that, The method includes the following steps: S10: Obtain information on holiday-related variable characteristics; S20: Based on the passenger flow characteristic variable information of the stations, the dynamic graph clustering algorithm is used to realize the clustering characteristics of subway network stations and cluster stations with similar travel patterns. S30: Based on the obtained holiday characteristic variable information, and using the results of the dynamic graph clustering algorithm, a passenger flow prediction model is constructed in different station clusters to predict the holiday passenger flow of subway stations. S10 includes: S101: Obtain the number of blog posts on social media platforms to acquire passenger travel data. The number of blog posts includes the current number of blog posts and the number of blog posts at historical time steps. S102: Definition The number of blog posts during the period was The number of blog posts at historical time steps is used to help predict passenger traffic for the next period; among which: in, For historical time steps; S103: Introducing holiday coding features to predict changes in subway passenger flow during holidays; The dynamic graph clustering algorithms in S20 include: S201: A topology-adaptive method is used for graph embedding to generate a spatial topology that adapts to the input data; S202: The solver of the lower-level clustering task is embedded as a differentiable layer and incorporated into the training process of the deep learning model; S203: Establish a topology-adaptive graph convolutional network and construct a dynamic travel network topology graph based on passenger travel data. ,in The nodes of the graph, i.e., the stations, The edges of the graph represent the connections between stations. The characteristics of the edges in the graph are described by the travel volume between the origin and the destination. S204: Using an adjacency matrix Describe the spatial connections between sites. Number of stations; S205: Node feature matrix of the graph This includes the station's longitude and latitude, daily passenger flow demand for entering and exiting the station, and hourly passenger flow demand for entering and exiting the station.

2. The method for predicting subway passenger flow during holidays based on spatiotemporal dynamic graph clustering as described in claim 1, characterized in that, The creation of a topology-adaptive graph convolutional network in S203 includes: S2031: Create a filter in the Fourier domain using a graph convolutional network model to perform convolution operations on graph nodes; S2032: Captures spatial features between nodes by capturing their first-order neighbor relationships, and constructs a graph convolutional network model by stacking multiple convolutional layers, as shown in the following formula: in, middle, It is an adjacency matrix. It is the identity matrix. It is the degree matrix of the travel network , Representing the Hidden layers Indicates the first Individual graph filters, These are the polynomial coefficients of a graph filter; graph convolution is the product of a matrix and a vector. , It is the first Each output feature map It is a learnable bias term. yes The output of the layer, It is the activation function ReLU; S2033: Formula Improved to This allows the filters in a topology-adaptive graph convolutional network to dynamically adapt to the graph's topology during the convolution process, where... This represents the size of the filter, i.e., the path length. The following formula is used to calculate the weight of information transmission between nodes; in, Represents a node arrive Path length is A certain path, The node sequence is ; S2034: Modularity is used as a loss function for error backpropagation to optimize parameters and iteratively improve graph representation learning and cluster assignment. Modularity is expressed as... And calculate using the following formula: in, It is the total number of edges in the network; It is the number of cluster categories; It is the adjacency matrix of the network, representing the nodes. and nodes Is there a connection between them? It is a node The degree; if the node Assigned to category , It is 1 if it is true, otherwise it is 0.

3. The method for predicting subway passenger flow during holidays based on spatiotemporal dynamic graph clustering as described in claim 2, characterized in that, The passenger flow prediction model constructed in S30 includes: S301: Establish a multi-head self-attention mechanism to obtain relevant information in different subspaces. The multi-head attention calculation formula is as follows: in, Represents the input vector. It is the first The size of the hidden layer for each attention head is determined by first setting different trainable parameters. Each Convert to query vector key vector Sum value vector ; S3022: Encode the edge features and add them to the key vector as additional information for each layer, from the node. To the node The multi-head attention mechanism of the edge is calculated by the following formula: in, Features of the edges; S3023: Based on graph multi-head attention, starting from distant nodes To the source node The message aggregation is calculated using the following formula: in, express The concatenation operation of multiple attention heads replaces the original normalized adjacency matrix with a multi-head attention matrix, which becomes the transition matrix for message passing.

4. A holiday subway passenger flow prediction system based on spatiotemporal dynamic graph clustering, applicable to the holiday subway passenger flow prediction method based on spatiotemporal dynamic graph clustering as described in any one of claims 1-3, characterized in that, The system includes: an acquisition module for acquiring holiday characteristic variable information; a clustering module for using a dynamic graph clustering algorithm to realize the clustering characteristics of subway network stations based on station passenger flow characteristic variable information, and clustering stations with similar travel patterns; and a prediction module for constructing a passenger flow prediction model in different station clusters based on the acquired holiday characteristic variable information and the results of the dynamic graph clustering algorithm, and predicting the holiday passenger flow of subway stations.