Urban rail transit network passenger flow distribution calculation failure early warning method
Through the relationship graph convolution neural network model and knowledge graph technology, the relationship between urban rail transit network structure and passenger flow distribution calculation failure is analyzed, and the problem of large differences in the calculation results of existing systems under complex networks is solved, and early warning of passenger flow distribution calculation failure and model effectiveness is achieved.
Patent Information
- Application Number
- CN202510109167.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
Under the complex network and diversified travel behavior, the difference between the calculation results and the actual passenger flow distribution of the existing urban rail transit network passenger flow distribution gradually expands, resulting in increased operation and management difficulties.
The relationship graph convolutional neural network model is used to construct a knowledge graph for passenger flow distribution calculation failure calculation by building a knowledge graph for urban rail transit networks, analyze the relationship between network structure and passenger flow distribution calculation failure, design an automated early warning method and visual display.
It has realized the early warning of the failure of passenger flow distribution calculation of urban rail transit networks, helping operation managers to identify key structures, improve quantitative verification of model effectiveness, and reduce the cost of passenger flow surveys and model improvement.
Smart Images

Figure CN120069049A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent transportation, and particularly relates to a method for warning of the failure of calculating the passenger flow distribution in an urban rail transit network. Background Art
[0002] At present, in the "one-ticket transfer" mode of rail transit network operation, it is impossible to directly obtain the passenger travel path only relying on the Automatic Fare Collection (AFC) data of rail transit. Moreover, the more complex the urban rail transit network is, the more prominent this situation becomes. In order to be able to calculate the spatio-temporal distribution of passenger flow in the rail transit network more accurately, a clearing model based on passenger flow distribution has emerged. With the continuous increase in the complexity of the scale and structure of the rail transit network, the diversity of train operation modes, and the differences in passenger travel behaviors, the difference between the calculation results of the existing clearing system mainly based on AFC data and the actual passenger flow distribution is constantly expanding, which brings difficulties to the effective preparation of operation plans by urban rail transit operation management departments and the precise organization of passenger flow at stations.
[0003] Existing research can determine the actual passenger travel path through the investigation of the passenger flow in the urban rail transit network or a data-driven estimation model, so as to make a quantitative evaluation of the calculation results of the passenger flow distribution model. By increasing the consideration factors of the comprehensive impedance calculation of the passenger travel path and improving the probability selection model structure, the accuracy of the model for calculating the passenger flow distribution can be improved.
[0004] However, in order to verify the accuracy of the passenger flow distribution calculation, regularly conducting large-scale passenger flow surveys across the network not only has great implementation difficulties, but also lacks survey means, and requires high time and labor costs. If it is possible to first warn of the effectiveness of the calculation of the passenger flow distribution in the urban rail transit network, and then arrange passenger flow surveys and model improvements according to the warning results, it can not only effectively reduce the cost of model verification, but also improve the refined level of operation management. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for warning of the failure of calculating the passenger flow distribution in an urban rail transit network. Based on the research on the failure mechanism of the calculation of the urban rail transit passenger flow distribution, analyze and summarize the relationship between the urban rail transit network structure (lines or stations) and the failure of the passenger flow distribution calculation, research and propose a method for warning of the failure of calculating the passenger flow distribution in the urban rail transit network, and design a corresponding algorithm to complete the automatic implementation and visual display of the failure warning method.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A method for warning of the failure of calculating the passenger flow distribution in an urban rail transit network, comprising:
[0008] S1, construct a knowledge graph of the failure of the passenger flow distribution calculation for urban rail transit networks;
[0009] S2, based on the knowledge graph of the failure of the passenger flow distribution calculation, construct a heterogeneous graph required for the relational graph convolutional neural network model, construct a feature matrix for all entity nodes in the heterogeneous graph, determine the relationship to be predicted, and add feature labels to the specified relationship and the nodes at both ends as the input items for subsequent model training; the heterogeneous graph refers to a graph formed by abstracting the topological network of the knowledge graph with real numbers into a digital topological matrix;
[0010] S3, during the training process of the neural network model, split the heterogeneous graph into two different subgraphs by setting masks, one subgraph is used as the model training set, and the other is used as the model validation set; by setting different masks and measures of different mask bits, multiple different subgraphs are generated for training to make up for the data sparsity problem of the knowledge graph. After multiple rounds of model iteration, the finally trained relational graph convolutional neural network model is obtained;
[0011] S4, use the finally trained relational graph convolutional neural network model to score the entities that may have a specific relationship, and the high or low score indicates the probability that the above specific relationship may exist between the entities; further combine the actual impact on passenger flow and operation management measures to determine whether the knowledge inference information of this specific relationship is output as the early warning information of the failure of the passenger flow distribution calculation.
[0012] Optionally, before step S1, it further includes: extracting knowledge according to the actual physical network of rail transit and the list of failure causes to obtain the triple data of the knowledge graph; completing the batch import, entity creation and relationship pairing of the triple data in the database through computer programming to form a knowledge graph of the failure mechanism of the passenger flow distribution calculation.
[0013] Optionally, the steps of extracting knowledge according to the actual physical network of rail transit and the list of failure causes to obtain knowledge graph triple data, and completing batch import, entity creation, and relationship pairing of the triple data in the database through computer programming to form a knowledge graph of the failure mechanism of passenger flow distribution calculation are specifically as follows: S01, extraction of knowledge triples of the failure mechanism: The failure of passenger flow distribution calculation is jointly determined by the model calculation selection probability, the actual path selection ratio, the model-derived selected train number, the actual travel selected train number, the effective path set, and the actual travel path of passengers. A causal and inclusion relationship is formed between these six and the failure of passenger flow distribution calculation. Extract and organize the above relationships into triple form; for the model calculation selection probability, the actual path selection ratio, the model-derived selected train number, the actual travel selected train number, the effective path set, and the actual travel path of passengers, there are usually clear data for rail transit operators, and the actual selection probability can be obtained through passenger flow surveys or data-driven methods. S02, extraction of knowledge triples of the network structure: The actual network structure knowledge of urban rail transit includes the multi-layer composition structure of the network, lines, and stations, the basic attributes of the departure interval, line type, and station type of the lines, and the transfer relationships between the urban rail transit network and suburban lines, railways, aviation, and trams. Extract and organize the above relationships into triple form; S03, construction of the knowledge graph of the failure mechanism of passenger flow distribution calculation: Input the triple information into the Neo4J tool, and based on the Neo4J tool, add urban rail network structure information in the knowledge graph, such as the characteristics of rail transit lines and stations, and organize historical passenger flow distribution calculation failure scenarios and related operation situations and add them to the failure graph together. Use the Neo4J tool to complete the database construction and visualization display of the knowledge graph of the failure mechanism of passenger flow distribution calculation, and represent different categories of entities and relationships through different colors and sizes.
[0014] Further, step S2 specifically includes: S21, constructing a heterogeneous graph required for the relational graph convolutional neural network model based on the passenger flow distribution calculation failure knowledge graph: The knowledge in the knowledge graph is stored in the form of triples of "entity-relationship-entity" or "entity-attribute-attribute value". When constructing the corresponding heterogeneous graph, it is necessary to add an entity category and an entity number to each entity. Entities of the same type should have the same entity category, and entities under the same category are numbered incrementally starting from 0; S22, constructing a feature matrix for all entity nodes in the heterogeneous graph: The nodes in the heterogeneous graph carry features of name, category, and number, and a multi-dimensional vector is used to describe the features of the nodes; S23, determining the relationship to be predicted, and adding feature labels to the specified relationship and its two end nodes: Knowledge graph link prediction is to infer the possibility of a specific relationship between other entities after learning the feature expression of a specific relationship. This relationship is in the form of a triple of "specific category entity 1-specific relationship-specific category entity 2"; In this step, it is necessary to add a feature label composed of a multi-dimensional vector to all the relationships to be predicted and their two end entities in the heterogeneous graph.
[0015] Optionally, the passenger flow distribution calculation failure mechanism knowledge graph is described as G=(V, E, R), where V represents the set of entity nodes in the knowledge graph vi∈V, E represents the set of edges with relationship labels (vi, r, vj)∈E, and R is the set of relationship types r∈R.
[0016] Optionally, the neural network model adopts a relational graph convolutional neural network composed of multiple R-GCN layers to learn the relationship features of the nodes and their domains in the local graph of the heterogeneous graph, and gradually extends to the learning of the feature representation of the large-scale relational network through the transitivity of relationship-nodes. The encoder composed of the neural network model needs to perform feature learning and update on each node in the relational graph, and multiple R-GCN layers need to be stacked to allow learning of the hidden relationships between nodes across multiple relationships. The propagation model for calculating the forward propagation update of the feature representation of each node in the relational graph is as follows:
[0017]
[0018] where refers to all neighbor nodes of the node with index i connected by the relationship r, is a normalization constant. is the representation matrix of node i at the l-th layer of the model, can be represented by one-hot encoding or other features; is the feature matrix corresponding to each relationship, The representation vector of the node that summarizes the l-th layer into i is passed to the representation vector of the node in the (l+1)-th layer; σ(·) represents any activation function, and the role of the activation function is to aggregate the neighbor nodes of the node and the corresponding relationship information, and at the same time pass the information of the node itself, so that each layer of the neural network model can learn the structural features of the node.
[0019] Optionally, in step S4, during the actual training of the model, for each relationship (s, r, o) ∈ e existing in the knowledge graph, that is, the positive sample, randomly break the head entity s or the tail entity o of the relationship to generate w negative samples, and use the cross-entropy loss to distinguish between positive and negative samples, so that the scoring function gives high scores to positive samples and low scores to negative samples. The calculation formula of the cross-entropy loss function is:
[0020]
[0021] where Γ is the set of positive and negative samples, when (s, r, o) is a positive sample, y takes the value of 1, otherwise it takes 0, l(·) is the sigmoid function, which belongs to an activation function; in this loss function, ylogl(f(s, r, o)) is used to optimize the discrimination of positive samples, and (1 - y)log(1 - l(f(s, r, o))) is used to optimize the discrimination ability of negative samples; |ε| represents the number of elements in this set, and the entire denominator (1 + w)|ε| represents normalizing all samples.
[0022] Due to the above technical solutions, the beneficial effects of the present invention include: The present invention designs a passenger flow distribution calculation failure warning method based on a relational graph convolutional neural network model, which can learn system structure features and historical operation experience from the urban rail network failure mechanism knowledge graph, identify the key structures in the network that cause the failure of the passenger flow distribution calculation model, and assist operation managers to quantitatively verify the effectiveness of the model by comprehensively considering factors such as passenger flow volume. Description of the Drawings
[0023] Figure 1 is the implementation flowchart of the present invention.
[0024] Figure 2 is the system structure diagram of the present invention.
[0025] Figure 3 is the passenger flow distribution calculation failure mechanism knowledge graph (partial) of the present invention.
[0026] Figure 4 is the schematic diagram of node feature representation calculation and update of the present invention.
[0027] Figure 5 is the example road network diagram of the embodiment of the present invention.
[0028] Figure 6 is the average cross-entropy loss of the prediction group and the reference group in the calculation example of the present invention.
[0029] Figure 7 is a schematic diagram of the prediction results of the prediction group and the reference group in the calculation example of the present invention. Specific Embodiments
[0030] In order to make the objectives, technical solutions and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0031] For example, a method for early warning of calculation failure of passenger flow distribution in an urban rail transit network includes:
[0032] S1, constructing a knowledge graph for calculating the failure of passenger flow distribution in an urban rail transit network;
[0033] S2, constructing a heterogeneous graph required for a relational graph convolutional neural network model based on the knowledge graph for calculating the failure of passenger flow distribution, constructing a feature matrix for all entity nodes in the heterogeneous graph, determining the relationship to be predicted, and adding feature labels to the specified relationship and its two end nodes as input items for subsequent model training; wherein, the heterogeneous graph refers to a graph formed by abstracting the topological network of the knowledge graph with real numbers into a digital topological matrix. The specific steps of step S2 are as follows:
[0034] S21, constructing a heterogeneous graph required for a relational graph convolutional neural network model based on the knowledge graph for calculating the failure of passenger flow distribution: The knowledge in the knowledge graph is stored in the form of triples of "entity-relationship-entity" or "entity-attribute-attribute value". When constructing the corresponding heterogeneous graph, it is necessary to add an entity category and an entity number to each entity. Entities of the same type should have the same entity category, and entities under the same category are numbered incrementally starting from 0.
[0035] S22, constructing a feature matrix for all entity nodes in the heterogeneous graph: The nodes in the heterogeneous graph have features of name, category and number, and a multi-dimensional vector is used to describe the features of the nodes;
[0036] S23, determining the relationship to be predicted, and adding feature labels to the specified relationship and its two end nodes: Link prediction in the knowledge graph is to infer the possibility of a specific relationship existing between other entities after learning the feature expression of a specific relationship. This relationship is in the form of a triple of "specific category entity 1-specific relationship-specific category entity 2"; in this step, it is necessary to add a feature label composed of a multi-dimensional vector to all relationships to be predicted and their two end entities in the heterogeneous graph.
[0037] S3. During the training process of the neural network model, the heterogeneous graph is split into two different subgraphs by setting masks. One subgraph is used as the model training set, and the other is used as the model validation set. By setting different masks and measures of different mask bits, multiple different subgraphs are generated for training to make up for the data sparsity problem of the knowledge graph. After multiple rounds of model iteration, the finally trained relational graph convolutional neural network model is obtained.
[0038] S4. Use the finally trained relational graph convolutional neural network model to score the entities that may have a specific relationship. The level of the score indicates the probability that the entities may have the above specific relationship. Further combined with the actual impact on passenger flow and operation management measures, it is determined whether the knowledge inference information (the possibility of a specific relationship between entities) of this specific relationship is output as the early warning information for the failure of passenger flow distribution calculation. The entities include network structure entities and failure cause entities. The network structure entity is a multi-layer structure of network, line and station, and the failure cause entity is the failure of passenger flow distribution calculation. The specific relationship refers to analyzing the possibility of failure factors between network structure entities with the failure of passenger flow distribution calculation as the phenomenon.
[0039] Optionally, before step S1, it further includes: extracting knowledge graph triple data according to the actual physical network of rail transit and the failure cause list; completing the batch import, entity creation and relationship pairing of the triple data in the database through computer programming to form a knowledge graph of the failure mechanism of passenger flow distribution calculation. Specifically as follows:
[0040] S01. Extraction of failure mechanism knowledge triples: The failure of passenger flow distribution calculation is jointly determined by the model calculation selection probability, the actual path selection ratio, the model calculation selected train number, the actual travel selected train number, the effective path set and the actual travel path of passengers. A causal and inclusion relationship is formed between these six and the failure of passenger flow distribution calculation. Extract and organize the above relationships into triple form; for the model calculation selection probability, the actual path selection ratio, the model calculation selected train number, the actual travel selected train number, the effective path set and the actual travel path of passengers, there are usually clear data for the rail transit operator, and the actual selection probability can be obtained through passenger flow surveys or data-driven methods.
[0041] S02. Extraction of network structure knowledge triples: The actual network structure knowledge of urban rail transit includes the multi-layer composition structure of network, line and station, the basic attributes of the departure interval, line type, station type of the line, and the transfer relationship between the urban rail transit network and suburban lines, railways, aviation, and tramways. Extract and organize the above relationships into triple form;
[0042] S03. Construct a knowledge graph of the failure mechanism of passenger flow distribution calculation: Input the triple information into the Neo4J tool. Based on the Neo4J tool, add the urban rail network structure information to the knowledge graph, such as the characteristics of rail transit lines and stations, and organize the historical passenger flow distribution calculation failure scenarios and related operation situations and add them to the failure graph together. Use the Neo4J tool to complete the database construction and visualization display of the knowledge graph of the failure mechanism of passenger flow distribution calculation, and represent different categories of entities and relationships through different colors and sizes.
[0043] Furthermore, the knowledge graph of the failure mechanism of passenger flow distribution calculation is described as G=(V, E, R), where V represents the set of entity nodes in the knowledge graph, vi∈V, E represents the set of edges with relationship labels, (vi, r, vj)∈E, and R is the set of relationship types, r∈R.
[0044] Optionally, the neural network model adopts a relational graph convolutional neural network composed of multiple R-GCN layers to learn the relationship features of nodes and their domains in the heterogeneous graph local graph. Through the transitivity of relationship-nodes, it gradually expands to the learning of large-scale relational network feature representations. The encoder composed of the neural network model needs to perform feature learning and update on each node in the relational graph, and multiple R-GCN layers need to be stacked to allow learning of hidden relationships between nodes across multiple relationships. The propagation model for calculating the forward propagation update of the feature representations of each node in the relational graph is as follows:
[0045]
[0046] where refers to all neighbor nodes of the node with index i connected by the relationship r, is a normalization constant. is the representation matrix of node i at the l-th layer of the model, which can be represented by one-hot encoding or other features; is the feature matrix corresponding to each relationship, used to transfer the representation vector of the node with index i at the l-th layer to the representation vector of the node at the (l + 1)-th layer; σ(·) represents any activation function. The role of the activation function is to aggregate the neighbor nodes of the node and the corresponding relationship information, and at the same time transfer the information of the node itself, so that each layer of the neural network model can learn the structural features of the node.
[0047] Optionally, in step S4, during the actual model training process, for each relationship (s, r, p) ∈ e existing in the knowledge graph, that is, the positive sample, w negative samples are randomly generated by breaking the relationship of the subject entity s or the object entity o. The cross-entropy loss is used to distinguish between positive and negative samples, so that the scoring function gives high scores to positive samples and low scores to negative samples. The calculation formula of the cross-entropy loss function is:
[0048]
[0049] where Γ is the set of positive and negative samples. When (s, r, p) is a positive sample, y takes the value of 1, otherwise it takes 0. l(·) is the sigmoid function, which belongs to an activation function. In this loss function, ylogl(f(s, r, o)) is used to optimize the discrimination of positive samples, and (1 - y)log(1 - l(f(s, r, o))) is used to optimize the discrimination ability of negative samples. |ε| represents the number of elements in this set, and the entire denominator (1 + w)|ε| represents the normalization process for all samples.
[0050] As Figure 2 shown, the system consists of an input module, an encoder module, a decoder module, and a cross-entropy calculation module. These software modules are all independent and can be installed on one machine according to the actual situation, or installed on multiple machines. In addition, it can also be integrated into the urban rail transit operation management platform as a subsystem of the urban rail transit operation management auxiliary decision support system.
[0051] Now, in combination with Figure 2 briefly describe the functions of the above modules, and focus on combining Figure 1 to illustrate the implementation process and results of the method of the present invention.
[0052] 1. Input module
[0053] 1.1 Knowledge graph of the failure mechanism of passenger flow distribution calculation
[0054] The knowledge graph of the failure mechanism of passenger flow distribution calculation in this module can be described as G = (V, E, R), where V represents the set of entity nodes vi ∈ V in the knowledge graph, E represents the set of edges with relationship labels (vi, r, vj) ∈ E, and R is the set of relationship types r ∈ R. Knowledge extraction is carried out according to the actual physical network of rail transit and the failure cause list to obtain the knowledge graph triples. Through computer programming, the batch import, entity creation, and relationship pairing of the triple data in the Neo4J database are completed to form the knowledge graph of the failure mechanism of passenger flow distribution calculation ( Figure 3 ). Neo4j is a high-performance graph database. It is an embedded, disk-based, Java persistence engine with full transaction characteristics, but it stores structured data on the network.
[0055] Table 1 Triple of the Structural Relationship between Rail Transit Lines and Stations
[0056]
[0057] Table 2 Triple of the Departure Interval Duration Attribute of Rail Transit Lines
[0058]
[0059]
[0060] Table 3 Triple of the Failure Cause Mechanism
[0061]
[0062] 1.2 Input Data Preprocessing
[0063] When building the failure mechanism knowledge graph, in order to represent the interaction relationship between the network structures of each layer of rail transit and the calculation of passenger flow distribution for failure causes as detailed as possible, the node types and relationship types of the knowledge graph triples are refined. First, the node types and relationship types are merged to make the specified triples more comprehensively cover the target nodes; then the nodes are numbered in order by type to more efficiently identify the unique entity under the specified type from hundreds of entities in the knowledge graph; finally, the proof sample triples are constructed for subsequent model training.
[0064] Table 4 Triple Structure of Data Preprocessing
[0065]
[0066]
[0067] 2. Encoder Module
[0068] This module is a relational graph convolutional neural network composed of multiple R-GCN layers. Its working principle is to learn the relationship features of nodes and their domains in the local graph, and gradually expand to the learning of large-scale relational network feature representations through the transitivity of relationships - nodes. The encoder composed of the neural network model needs to perform feature learning and update for each node in the relational graph, and multiple R-GCN layers need to be stacked to allow learning of the hidden relationships between nodes across multiple relationships. The propagation model for calculating the forward propagation update of the feature representations of each node in the relational graph is as follows:
[0069]
[0070] Among them refers to all neighbor nodes of the node with index i connected by relationship r, is a normalization constant. is the representation matrix of node i at the l-th layer of the model, which can be represented by one-hot encoding or other features. is the feature matrix corresponding to each relationship, used to transfer the representation vector of the node that shrinks to i at the l-th layer to the representation vector of this node at the (l + 1)-th layer. σ(·) represents any activation function. The role of the activation function is to aggregate the neighbor nodes of the node and the corresponding relationship information, and at the same time transfer the information of the node itself, so that each layer of the neural network model can learn the structural features of the node.
[0071] As Figure 4 shown, the red circle represents the node to be calculated in each layer of the R-GCN network, the blue circle represents the neighbor nodes adjacent to this entity, and the direction of the arrow represents the relationship between the neighbor node and this entity. Node 1 represents the node type with multiple input nodes and multiple output nodes, that is, the many-to-many node; Node N represents the node type with a single input node and multiple output nodes, that is, the one-to-many / many-to-one node; the self-loop represents the node type where the node itself forms input and output with itself. The R-GCN model can learn the feature representations of the above several types of nodes:
[0072] Step 1, collect the activation vector representations from the adjacent nodes represented by the blue rectangle;
[0073] Step 2, perform matrix transformation on the input edges and output edges of each relationship type respectively to obtain the vector representations represented by the green rectangle;
[0074] Step 3, normalize and aggregate the vector representations obtained by the matrix transformation;
[0075] Step 4, pass the result obtained by aggregating the vector representations through the ReLU activation function.
[0076] 3. Decoder module
[0077] DistMult is a machine learning model used to complete entity relationship prediction in the knowledge graph. Its core idea is to represent the semantic association between them by mapping entities and relationships to a low-dimensional continuous vector space, and it has characteristics such as simplicity, efficiency, and interpretability. At the same time, DistMult is also a tool for modeling the relationship between two input vectors, which can learn the non-linear relationship between the input vectors and can capture the interaction between the input vectors, which is very consistent with the purpose of the failure warning task of the present invention. Its formal expression is:
[0078]
[0079] Among them, The embedding representation vector of the subject node representing the relationship r The embedding representation vector of the object node representing the relationship r, R r ∈R d×d is a diagonal matrix associated with the relationship r
[0080] 4. Cross-entropy loss calculation module
[0081] During the actual training process of the model, for each relationship (s, r, o) ∈ e existing in the knowledge graph, that is, the positive sample, randomly break the subject entity s or the object entity o of the relationship to generate w negative samples, and distinguish positive and negative samples through cross-entropy loss, so that the scoring function gives high scores to positive samples and low scores to negative samples. The calculation formula of the cross-entropy loss function is as follows:
[0082]
[0083] where Γ is the set of positive and negative samples. When (s, r, o) is a positive sample, y takes the value of 1, otherwise it takes 0. l(·) is the sigmoid function, which belongs to an activation function. In this loss function, ylogl(f(s, r, o)) is used to optimize the discrimination of positive samples, and (1 - y)log(1 - l(f(s, r, o))) is used to optimize the discrimination ability of negative samples. |ε| represents the number of elements in this set, and the entire denominator (1 + w)|ε| represents normalizing all samples. The symbol meanings can be referred to in Section 4 of the reference Modeling Relational Data with Graph Convolutional Networks
[0084] 5. Verification of the failure warning method
[0085] Taking the triple of the influence relationship between the Shanghai urban rail network structure and the causes of the failure of passenger flow distribution calculation as the historical operation experience knowledge input, following the processes of each module of the passenger flow distribution calculation failure warning model, predict and output the probability of the influence relationship between the urban rail network structure and the failure result of passenger flow distribution calculation, that is, according to the size of the station / OD passenger flow, judge the actual influence of this warning information on the passenger flow distribution calculation to determine whether to push actual warnings
[0086] Its implementation process can be specifically illustrated by the following example
[0087] The basic road network of the example takes the Shanghai urban rail transit in 2024 as an example (as Figure 5 shown), the road network includes 20 rail transit lines, including Rail Transit Line 1 to Rail Transit Line 18, Pujiang Line and Maglev Line, with a total of 508 stations, of which 83 are transfer stations, the operating mileage is 831 kilometers in total, and the highest daily passenger volume reaches 13.294 million person-times
[0088] As an interchange station of Line 1 and Line 5, Xinzhuang Station of the rail transit has the characteristics of large daily passenger flow and large morning peak passenger flow, which also has a certain impact on the path reliability. At the same time, the path reliability will cause the failure of the spatial dimension of the passenger flow distribution calculation, and the large passenger flow will cause the failure of the time dimension. After learning the above characteristics, the model can predict the possibility of the failure of the passenger flow distribution calculation at Xinzhuang Station; on the contrary, Huaning Road Station is neither a special type of station nor a large passenger flow station, and can be used as a reference group to observe the effectiveness of the model in learning network characteristics. Figure 6 )
[0089] Table 5 Results of Failure Early Warning of Calculation Examples
[0090]
[0091] The smaller the cross-entropy loss value, the higher the possibility of the existence of the relationship. As Figure 7 shown, Xinzhuang Station has the characteristics of interchange station, large daily passenger flow and large morning peak passenger flow, and all three station characteristics are related to the failure of the spatio-temporal dimension of the passenger flow distribution calculation. The model believes that the possibility of Xinzhuang Station affecting the failure of the passenger flow distribution calculation is relatively high, that is, it is difficult to accurately deduce the passenger flow distribution of the passenger flow with Xinzhuang Station as Station O in the urban rail network; on the contrary, as an ordinary station, Huaning Road Station has a relatively low possibility of affecting the failure of the passenger flow distribution calculation. The experimental results verify the effectiveness of the present invention.
[0092] The above description of the embodiments is to enable those of ordinary skill in the art to understand and use the present invention. Obviously, those skilled in the art can easily make various modifications to these embodiments and apply the general principles described herein to other embodiments without creative efforts. Therefore, the present invention is not limited to the above embodiments, and all improvements and modifications made by those skilled in the art without departing from the scope of the present invention according to the disclosure of the present invention should be within the protection scope of the present invention.
Claims
1. A method for early warning of failure of passenger flow distribution calculation in urban rail transit network, characterized in that: include: S1, constructing a knowledge graph of failure in calculating passenger flow distribution in urban rail transit networks; S2, based on the passenger flow distribution calculation failure knowledge graph, constructs the heterogeneous graph required by the relationship graph convolutional neural network model, constructs a feature matrix for all entity nodes in the heterogeneous graph, determines the relationship to be predicted, and adds feature labels to the specified relationship and its two end nodes as input items for subsequent model training; S3, during the neural network model training process, the heterogeneous graph is split into two different subgraphs by setting a mask, one subgraph is used as the model training set, and the other is used as the model verification set; by setting different masks and different mask bit measures, multiple different subgraphs are generated for training to make up for the data sparsity problem of the knowledge graph. After multiple rounds of model iterations, a trained relational graph convolutional neural network model is finally obtained; S4, use the trained relationship graph convolutional neural network model to score entities that may have specific relationships. The higher the score, the higher the probability that the above-mentioned specific relationship may exist between the entities; further combined with the actual impact on passenger flow and operational management measures, determine whether the knowledge reasoning information of this specific relationship should be output as a warning information for passenger flow distribution calculation failure.
2. The urban rail transit network passenger flow distribution calculation failure warning method according to claim 1 is characterized in that: Before step S1, the method further includes: Extract knowledge based on the actual physical network of rail transit and the list of failure causes to obtain knowledge graph triple data; Through computer programming, batch import of triple data in the database, entity creation and relationship matching are completed to form a knowledge graph of the failure mechanism of passenger flow distribution calculation.
3. The urban rail transit network passenger flow distribution calculation failure warning method according to claim 2 is characterized in that: The knowledge is extracted according to the actual physical network of rail transit and the list of failure causes to obtain knowledge graph triple data; The steps of batch importing triple data into the database, creating entities and pairing relationships through computer programming to form a knowledge graph of the failure mechanism of passenger flow distribution calculation include: S01, extraction of failure mechanism knowledge triples: The failure of passenger flow distribution calculation is jointly determined by the model calculation selection probability, the actual route selection ratio, the model estimation selection number, the actual travel selection number, the effective route set and the actual travel route of the passengers. These six factors form a causal and inclusion relationship with the failure of passenger flow distribution calculation. The above relationships are extracted and organized into triples; S02, network structure knowledge triple extraction: The actual network structure knowledge of urban rail transit includes the multi-layer structure of network, line, and station, the departure interval of the line, the basic attributes of the line type, and the station type, as well as the transfer relationship between the urban rail transit network and the urban line, railway, aviation, and tram. The above relationships are extracted and organized into triples; S03, construct the knowledge graph of failure mechanism of passenger flow distribution calculation: input the triple information into the Neo4J tool, use the Neo4J tool to complete the database construction and visualization display of the knowledge graph of failure mechanism of passenger flow distribution calculation, and use different colors and sizes to represent different categories of entities and relationships.
4. The urban rail transit network passenger flow distribution calculation failure warning method according to claim 1 is characterized in that: Step S2 specifically includes: S21, based on the knowledge graph of passenger flow distribution calculation failure, the heterogeneous graph required for the relationship graph convolutional neural network model is constructed: the knowledge in the knowledge graph is stored in the form of "entity-relationship-entity" or "entity-attribute-attribute value" triples. When constructing the heterogeneous graph, it is necessary to add entity categories and entity numbers for each entity. Entities of the same type should have the same entity category, and entities under the same category are numbered incrementally starting from 0; S22, construct a feature matrix for all entity nodes in the heterogeneous graph: the nodes in the heterogeneous graph have features of name, category and number, and a multi-dimensional vector is used to describe the features of the nodes; S23, determine the relationship to be predicted, and add feature labels to the specified relationship and its two end nodes: Knowledge graph link prediction is to infer the possibility of the existence of specific relationships between other entities after learning the feature expression of specific relationships. This relationship is in the form of a triple of "specific category entity 1-specific relationship-specific category entity 2"; in this step, it is necessary to add a feature label consisting of a multidimensional vector to all the relationships to be predicted and the entities at both ends of the heterogeneous graph.
5. The urban rail transit network passenger flow distribution calculation failure warning method according to claim 1 is characterized in that: The knowledge graph of the failure mechanism of passenger flow distribution calculation is described as G=(V,E,R), where V represents the set of entity nodes vi∈V in the knowledge graph, E represents the set of edges with relationship labels (vi,r,vj)∈E, and R is the set of relationship types r∈R.
6. The urban rail transit network passenger flow distribution calculation failure warning method according to claim 1 is characterized in that: The neural network model uses a relational graph convolutional neural network composed of multiple R-GCN layers to learn the relational features of nodes and their fields in the local graph of the heterogeneous graph, and gradually expands to the learning of large-scale relational network feature representation through the transitivity of the relationship-node. The encoder composed of the neural network model needs to learn and update the features of each node in the relationship graph. It needs to be stacked with multiple R-GCN layers to allow the hidden relationships between nodes to be learned across multiple relationships. The propagation model used to calculate the forward transfer update of the feature representation of each node in the relationship graph is as follows: in It refers to all neighbor nodes of the node with index i connected by relationship r. is a normalizing constant. is the representation matrix of node i in the lth layer of the model, Can be represented by one-hot encoding or other features; is the feature matrix corresponding to each relationship, It is used to transfer the representation vector of the node i in the lth layer to the representation vector of the node in the (l+1)th layer; σ(·) represents any activation function. The role of the activation function is to aggregate the neighbor nodes of the node and the corresponding relationship information, and at the same time transfer the information of the node itself, so that each layer of the neural network model can learn the structural characteristics of the node.
7. The urban rail transit network passenger flow distribution calculation failure warning method according to claim 1 is characterized in that: In step S4, during the actual model training process, for each relationship (s, r, o) ∈ e in the knowledge graph, i.e., a positive sample, w negative samples are generated by randomly destroying the subject entity s or the object entity o of the relationship. The positive samples and negative samples are distinguished by the cross entropy loss, so that the scoring function gives high scores to positive samples and low scores to negative samples. The calculation formula of the cross entropy loss function is: Where Γ is a set of positive and negative samples, y takes the value of 1 when (s, r, o) is a positive sample and 0 otherwise, l(·) is a sigmoid function, which is an activation function. In this loss function, ylogl(f(s, r, o)) is used to optimize the discrimination of positive samples and (1-y)log(1-l(f(s, r, o))) is used to optimize the discrimination of negative samples. |ε| represents the number of elements in the set, and the denominator (1+w)|ε| represents the normalization of all samples.