A method and system for identifying a station group with high passenger flow correlation degree in a complex railway network
Patent Information
- Application Number
- CN202611341485.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-09-01
- Publication Date
- 2026-09-29
AI Technical Summary
[0007]本发明的目的在于克服现有技术中无法从网络角度挖掘车站间隐藏客流关联模式、缺乏动态自适应分割机制以及分析结果可解释性不足的缺陷,提供一种复杂铁路网络高客流关联度车站群识别方法及系统
1.本发明提供一种复杂铁路网络高客流关联度车站群识别方法,通过将目标铁路网络中的车站作为节点构建客流网络图,基于列车运行数据计算任意两个车站之间的客流特征量作为连接关系的权重,再根据权重进行车站群识别并施加基于预设最大规模的规模约束,能够从网络拓扑角度动态挖掘车站间隐藏的客流关联模式,自动生成符合规模要求的精细车站群划分结果,克服了现有静态分析方法缺乏自适应分割机制的缺陷,提高了分析结果的可解释性与实用性。
Smart Images

Figure CN122840441A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of railway transportation planning and intelligent transportation technology, and in particular to a method and system for identifying station groups with high passenger flow correlation in complex railway networks. Background Technology
[0002] Currently, the mainstream methods for analyzing passenger flow in railway networks mainly include spatial visualization analysis based on geographic information systems, passenger flow forecasting based on time series data, and traditional statistical analysis methods. These methods can provide basic information on passenger flow distribution at a macro level, but they still have the following shortcomings in practical applications:
[0003] First, existing methods typically do not systematically analyze the passenger flow relationship structure between stations from a network topology perspective, making it difficult to discover the group relationship patterns hidden in the passenger flow between stations, such as the high-frequency, high-passenger-flow interconnection relationship formed between multiple stations.
[0004] Second, most analytical methods are static assessments, which usually only divide the railway network at a specific time or for a specific line at a one-time time. They lack dynamic and adaptive station group detection and segmentation mechanisms, resulting in weak generalization ability of the methods and difficulty in adapting to railway networks of different sizes and with different passenger flow characteristics.
[0005] Third, the interpretability of the analysis results is limited, often only outputting statistical reports or visualizations, lacking further quantification of structured indicators such as passenger flow correlation strength and station node importance, making it difficult to provide direct and clear support for transportation decisions such as train operation plan optimization and transfer hub planning.
[0006] In recent years, complex network theory has been gradually introduced into the field of traffic analysis. However, existing studies have mostly focused on urban road traffic networks or community discovery based on single passenger flow indicators (such as passenger volume or number of trains). There is still a lack of systematic solutions for the comprehensive utilization of multi-dimensional passenger flow characteristics of railway networks (such as passenger volume, number of trains, and average number of trains per train), as well as for how to impose scale constraints and fine segmentation on detected station groups. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of existing technologies, such as the inability to mine hidden passenger flow correlation patterns between stations from a network perspective, the lack of dynamic adaptive segmentation mechanisms, and insufficient interpretability of analysis results, and to provide a method and system for identifying station groups with high passenger flow correlation in complex railway networks.
[0008] In a first aspect, the present invention provides a method for identifying high-passenger-flow-related station groups in complex railway networks, comprising: S1: Using stations in the target railway network as nodes, construct a passenger flow network graph, where a connection is established between any two stations if there are train operation records between them; S2: Based on train operation data, calculate the passenger flow characteristics between any two stations, and use the passenger flow characteristics as the weight of the corresponding connection relationship in the passenger flow network graph; S3: Based on the weights, identify station groups in the passenger flow network map to obtain preliminary station group identification results; S4: Based on the preset maximum station group size, the preliminary identification results of the station group are constrained by size, and the final identification results of the station group are output.
[0009] Preferably, the train operation data includes: train number, origin station, destination station, train dispatch volume, passenger arrival and departure times between stations, distance between stations, and station coordinates.
[0010] Preferably, the passenger flow characteristics include: the total number of trains departing (or arriving), the total number of trains, and the average number of trains departing (or arriving) per train, with any two stations as the origin and destination.
[0011] Preferably, the station group identification includes: S31: Initialize each station node as an independent station group; S32: Calculate the modularity gain when each station node is assigned to each adjacent station group, and assign each station node to the adjacent station group that maximizes the modularity gain; if all modularity gains are less than or equal to zero, keep the original station group affiliation of the corresponding station node; repeat S32 until the first convergence condition is met. S33: Map each station group obtained after S32 to a new station node. The weight between any two new station nodes is the sum of all weights between the corresponding two original station groups. Construct a simplified network graph based on the new station nodes and weights. S34: For the simplified network graph, repeat S31 to S33 until the second convergence condition is met, output the final simplified network graph, and take each station group in the final simplified network graph as the preliminary identification result of the station group.
[0012] Furthermore, the first convergence condition includes: the change in modularity after two consecutive executions of S32 is less than the first convergence threshold; the second convergence condition includes: the change in modularity of the simplified network graph obtained in two consecutive executions is less than the second convergence threshold; wherein, the modularity is calculated based on the weights.
[0013] Preferably, the size constraint includes: S41: Determine whether the size of each station group exceeds the maximum station group size; if not, retain the station group. S42: If the number of nodes exceeds the limit, check whether all nodes in the station group are connected within the station group. If they are connected, identify the station group and generate multiple sub-station groups. If there are nodes in the station group that are not connected within the station group, split the station group into multiple sets of nodes with connected internal nodes and treat each set of nodes as a sub-station group. S43: For each sub-station group generated by S42, repeat S41 to S42 until the size of all sub-station groups is less than or equal to the size of the maximum station group. S44: If, after repeating S41 to S42 a preset number of times, there is still a sub-station group whose size exceeds the maximum station group size, then the sub-station group is forcibly split.
[0014] Furthermore, the forced segmentation includes: S441: Calculate the importance score of each node in the sub-station group. The importance score is obtained by weighted summation based on the degree ratio of the node within the sub-station group, the weight ratio of the edge where the node is located, and the betweenness centrality of the node. S442: Select the node with the highest score from the nodes of the sub-station group as the seed node according to the importance score from high to low. S443: Using the seed node as the core, expand the sub-station group within the sub-station group through breadth-first search to form a new sub-station group with interconnected internal nodes, until the number of nodes contained in the new sub-station group reaches the maximum station group size; S444: Remove the nodes that have formed a new sub-station group from the sub-station group, and repeat S442 to S443 for the remaining nodes until all nodes in the sub-station group have been assigned to the new sub-station group.
[0015] Preferably, the method further includes performing the following steps on the final identification result of the station group: Based on the final identification results of the station groups, the internal connection strength of each station group is calculated, and the N station groups with the highest internal connection strength are selected as the key analysis objects; where N is a preset positive integer. Calculate the weighted degree centrality of each station node within the key analysis object and output it in descending order.
[0016] In a second aspect, the present invention provides a system for identifying station clusters with high passenger flow correlation in complex railway networks, comprising: The network construction module is used to construct a passenger flow network graph by using stations in the target railway network as nodes. If there are train operation records between any two stations, a connection relationship is established. The weight calculation module is used to calculate the passenger flow characteristics between any two stations based on train operation data, and use the passenger flow characteristics as the weight of the corresponding connection relationship in the passenger flow network graph. The station group identification module is used to identify station groups in the passenger flow network map according to the weights, and obtain preliminary station group identification results. The scale constraint module, based on the preset maximum station group size, applies scale constraints to the preliminary identification results of the station group and outputs the final identification results of the station group.
[0017] In a third aspect, the present invention provides a system for identifying station groups with high passenger flow correlation in complex railway networks, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the program to implement a method for identifying station groups with high passenger flow correlation in complex railway networks as described in any of the first aspects.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention provides a method for identifying station groups with high passenger flow correlation in complex railway networks. By constructing a passenger flow network graph using stations in the target railway network as nodes, calculating the passenger flow characteristics between any two stations based on train operation data as the weight of the connection relationship, and then identifying station groups according to the weights and applying a scale constraint based on a preset maximum size, it can dynamically mine hidden passenger flow correlation patterns between stations from the perspective of network topology, automatically generate fine station group division results that meet the scale requirements, overcome the shortcomings of existing static analysis methods that lack adaptive segmentation mechanisms, and improve the interpretability and practicality of the analysis results.
[0019] 2. This invention provides a station group identification system with high passenger flow correlation in complex railway networks. The system uses a network construction module to treat stations as nodes and establish connections between them. A weight calculation module assigns weights based on passenger flow characteristics. The station group identification module and the scale constraint module sequentially complete the preliminary identification and scale constraint output of the final identification results. This system achieves fully automated processing from train operation data to station group division, making it easy for transportation planners to use directly and improving the efficiency of passenger flow analysis and decision support capabilities.
[0020] 3. This invention provides a station group identification system for complex railway networks with high passenger flow correlation. By storing the station group identification method in the form of a computer program in a memory and executing it with a processor, it realizes the effective integration of passenger flow network graph construction, weight calculation, station group identification and scale constraint algorithm into the actual railway operation hardware platform. This reduces the application threshold of complex network analysis methods in engineering sites and improves the real-time response capability and operational stability of passenger flow analysis in large-scale railway networks. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating the overall process of identifying a high-passenger-flow-related station group in a complex railway network, as described in this embodiment. Figure 2 This is a flowchart illustrating the specific implementation of a method for identifying high-passenger-flow-related station groups in a complex railway network, as described in this embodiment. Figure 3 This is a flowchart of the station group identification method in the embodiment; Figure 4 This is a flowchart of the scale constraint method in the embodiment; Figure 5 This is a flowchart of the forced segmentation method in the embodiment; Figure 6 This is a diagram showing the station group identification results based on inter-station transmission volume in the embodiment; Figure 7 This is a diagram showing the station group identification results based on inter-station train services in the embodiment. Detailed Implementation
[0022] The present invention will now be described in further detail with reference to specific embodiments. However, this should not be construed as limiting the scope of the present invention to the following embodiments; all technologies implemented based on the content of the present invention fall within the scope of the present invention.
[0023] Unless otherwise specified, the terms "upper," "lower," "left," "right," "center," "inner," and "outer," etc., used in the description of specific embodiments of the present invention to indicate orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings, or the orientation or positional relationship in which the product / equipment / device is usually placed during use. These terms are merely for the purpose of facilitating the description of the present invention or simplifying the description in specific embodiments, and for enabling those skilled in the art to quickly understand the solution, and do not indicate or imply that a particular device / component / element must have a specific orientation, or be constructed and operated in a specific positional relationship. Therefore, they should not be construed as limitations on the present invention.
[0024] Furthermore, the use of terms such as "horizontal," "vertical," "suspended," "parallel," and "coaxial" does not imply that the corresponding device / component / element must be absolutely horizontal, vertical, suspended, parallel, or coaxial. Slight tilt or deviation is permissible, as long as it does not affect the normal function of the relevant component. For example, "horizontal" simply means that its direction is more horizontal relative to "vertical," not that the structure must be perfectly horizontal; a slight tilt is acceptable. "Coaxial" means that two components are arranged as coaxially as possible, allowing them to move coaxially or approximately coaxially when their relative positions change. Alternatively, it can be simplified to mean that the corresponding device / component / element, when arranged in "horizontal," "vertical," "suspended," "parallel," or "coaxial" directions, can have an error / deviation of ±10% relative to the corresponding direction, more preferably within ±8%, more preferably within ±6%, more preferably within ±5%, and more preferably within ±4%. For example, the deviation in the "coaxial" direction is controlled within 0.2-1mm, preferably within 0.2-0.5mm. As long as the corresponding device / component / element is within the error / deviation range, it can still achieve its function in the solution of the present invention.
[0025] Furthermore, the use of terms such as "first," "second," and "third" in terminology is merely for distinguishing descriptions of identical or similar components and should not be interpreted as emphasizing or implying the relative importance of a particular component.
[0026] Furthermore, in the description of the embodiments of the present invention, "several", "more than", and "a number of" represent at least two. The number can be any number, such as two, three, four, five, six, seven, eight, or nine, and can even exceed nine.
[0027] Furthermore, in the description of the technical solution of this invention, unless otherwise explicitly specified / limited / restricted, the terms "set up," "install," "connect," "link," "provided with," "laid out," and "arranged" should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to connection methods commonly used in the art, such as welding, riveting, bolting, and threaded connections. Such connections can be mechanical, electrical, or communication connections; they can be direct connections or indirect connections through an intermediate medium; and they can refer to the internal communication between two components.
[0028] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0029] The method and system for identifying station groups with high passenger flow correlation in complex railway networks provided by this invention can be applied to the fields of railway transportation planning and intelligent transportation technology, and is particularly suitable for passenger flow analysis and train operation scheme optimization in large-scale railway networks such as intercity railways and urban rail transit. This method abstracts railway stations as network nodes, constructs a multi-dimensional passenger flow network graph based on train operation data, detects station groups using the Louvain algorithm, and introduces adaptive scale constraints and forced segmentation mechanisms. Finally, it outputs station groups with high passenger flow correlation and a ranking of station importance.
[0030] Figure 1 This is a flowchart illustrating the overall process of a method for identifying high-passenger-flow-related station groups in complex railway networks, as provided by the present invention. Figure 2 The specific implementation flowchart shows that the method includes the following steps: In step S1, stations in the target railway network are used as nodes to construct a passenger flow network graph, wherein a connection relationship is established between any two stations if there are train operation records between them.
[0031] For example, the target railway network can be all railway stations within a certain area, such as stations A, B, C, and D. For any two stations, if there is at least one train operation record (i.e., a train actually runs between these two stations), a connection is established in the graph. For example, if there is a train operation between station A and station B, a connection is established between A and B; if there is no train operation record between station A and station C, no connection is established. The passenger flow network graph constructed in this way is an undirected graph, where nodes represent stations, and node attributes include station name, administrative division, coordinates, etc., and the connections between nodes represent actual train operation links between stations.
[0032] In step S2, based on train operation data, the passenger flow characteristic quantity between any two stations is calculated, and the passenger flow characteristic quantity is used as the weight of the corresponding connection relationship in the passenger flow network graph.
[0033] For example, train operation data specifically includes information such as train number, origin station, destination station, and train dispatch volume. Optionally, train operation data may also include passenger arrival and departure times between stations, distances between stations, station coordinates, etc., which can be used to assist in analysis or recording. In addition, train operation data can be preprocessed, including data cleaning, missing value handling, and outlier removal, to ensure data quality.
[0034] In one possible implementation, for each pair of stations, the total number of departures (or arrivals), the total number of trains, and the average number of departures (or arrivals) per train for all train records with these two stations as origin-destination (OD) are calculated.
[0035] It should be noted that in this invention, passenger flow characteristics can be expressed using either dispatch volume (including total dispatch volume and average dispatch volume per train) or arrival volume (including total arrival volume and average arrival volume per train). Since the total dispatch volume of the entire railway passenger network equals the total arrival volume, and the total number of trains between any origin-destination (OD) pair is exactly the same for both dispatch and arrival directions, the total dispatch volume and total arrival volume are balanced, and the average dispatch volume per train is also equal to the average arrival volume per train. Therefore, using either dispatch volume or arrival volume to characterize passenger flow is completely equivalent and does not affect the subsequent calculation results of station correlation. To maintain a consistent calculation method, this embodiment and the following steps all use dispatch volume as the basis for calculating passenger flow characteristics.
[0036] Specifically, total number of transmissions ,in For the number of times the vehicle is connected, For the first Number of train departures; total number of train trips Average number of transmissions per column For example, if there are 10 trains traveling from station A to station B, with a total passenger volume of 5000, and 8 trains traveling from station B to station A, with a total passenger volume of 4000, then the combined total passenger volume is 9000, the total number of trains is 18, and the average passenger volume per train is 9000 / 18 = 500. These three values are used as connection weights for three different dimensions, thus constructing a first passenger flow network graph with the total passenger volume as the weight. Second passenger flow network diagram with total number of trips as weight Third passenger flow network diagram with average number of passengers per train as weight. The technical advantage of this step is that it can comprehensively characterize the strength of passenger flow correlation between stations from three perspectives: total dispatch volume, traffic frequency, and passenger carrying capacity per vehicle.
[0037] In step S3, station group identification is performed on the passenger flow network map according to the weights to obtain preliminary station group identification results. For example, the station group identification can employ the Louvain algorithm, performing the station group identification process on the first, second, and third passenger flow network maps respectively. The specific implementation of this step is as follows... Figure 3 As shown.
[0038] In step S31, each station node is initialized as an independent station group. For example, station A is grouped as a separate group, station B as a separate group, and so on.
[0039] In step S32, the modularity gain of each station node is calculated when it is assigned to each of the adjacent station groups, and each station node is assigned to the adjacent station group that maximizes the modularity gain. If all modularity gains are less than or equal to zero, the original station group affiliation of the corresponding station node is maintained. S32 is repeated until the first convergence condition is met.
[0040] For example, the first convergence condition can be that the change in modularity after two consecutive executions of step S32 is less than a first convergence threshold. This first convergence threshold is a positive number close to 0. By repeatedly moving nodes, the modularity is gradually increased until a local optimum is reached. Here, the first convergence threshold is a quantitative expression of the convergence condition that "the modularity no longer increases." Using "change less than the threshold" instead of "unchanged" avoids the engineering problem of "completely unchanged" due to floating-point precision, and also covers various practical application scenarios. The modularity is an indicator used to measure the quality of station group partitioning, calculated based on connection weights. The modularity gain can be calculated based on parameters such as the sum of the weights of the edges between the current node and adjacent groups, and the node degree.
[0041] Specifically, modularity The calculation formula is:
[0042] in, For nodes and Connection weights between nodes; It is half the sum of all connection weights; For nodes The weighted degree; For nodes The weighted degree; Represents a node The group of stations to which it belongs; Represents a node The group of stations to which it belongs; For the Kronecker function, when the node and The value is 1 if the stations belong to the same station group, and 0 otherwise. The connection weight in the above formula can be the transmission volume, the number of trains, or the average transmission volume per train, depending on the network. Calculate the corresponding .
[0043] Modularity gain The calculation method is as follows:
[0044] in, This is the sum of connection weights between nodes within the station group, including only internal connection relationships; The sum of all connection weights connected to nodes within the station group, including internal and external connections; Station node The degree; Station node The sum of connection weights to nodes within the station group; This is half the sum of all connection weights in the graph. The connection weights in the above formula can be the transmission volume, the number of vehicles, or the average transmission volume per column, depending on the network. Calculate the corresponding .
[0045] In step S33, each station group obtained after step S32 is mapped to a new station node. The weight between any two new station nodes is the sum of all weights between the corresponding two original station groups. A simplified network graph is constructed based on the new station nodes and weights. For example, after step S32, three station groups are formed: station group 1 contains stations A and B, station group 2 contains stations C and D, and station group 3 contains station E. Each station group is mapped to a new station node. The weight between the new station node 1 (i.e., station group 1) and the new station node 2 (i.e., station group 2) is equal to the sum of the weights of all edges between the original A and B and C and D, that is, the weights of edges AC, AD, BC, and BD are added together. The number of nodes in the simplified network graph after construction is reduced.
[0046] In step S34, steps S31 to S33 are repeated for the simplified network graph until the second convergence condition is met. The second convergence condition can be that the change in modularity of the simplified network graph obtained in two consecutive iterations is less than a second convergence threshold, which is a positive number close to 0. Using "change less than the threshold" instead of "unchanged" avoids the engineering problem of achieving "completely unchanged" due to floating-point precision, and can cover various practical application scenarios. Finally, the last simplified network graph is output, and the various station groups in this simplified network graph are used as the preliminary identification results of the station groups. This multi-level node merging and network aggregation method can discover station group structures at different scales, improving the accuracy of identification.
[0047] In step S4, based on a preset maximum station group size, the preliminary station group identification results are subject to size constraints, and the final station group identification results are output. Maximum station group size This is a preset positive integer, for example, set to 10, indicating that each station group contains a maximum of 10 stations. For example, this size constraint is applied to the preliminary station group identification results of the first, second, and third passenger flow network maps, respectively. Figure 4 As shown, the size constraint specifically includes the following sub-steps: In step S41, it is determined whether the size of each station group exceeds the maximum station group size. If it does not exceed the maximum station group size, the station group is retained. If it does exceed the maximum station group size, the process proceeds to step S42: checking whether all nodes within the station group are connected. If they are connected, the station group is identified (i.e., the identification process in step S3 is applied again), generating multiple sub-station groups. If there are nodes within the station group that are not connected, the station group is split into multiple sets of nodes with connected internal nodes, and each set of nodes is considered a sub-station group. For example, a station group containing 12 nodes but not all of them are connected internally may be split into two sub-station groups of 6 nodes each, with all nodes connected to each other. In step S43, for each sub-station group generated in step S42, steps S41 to S42 are repeated until the size of all sub-station groups is less than or equal to the maximum station group size. In step S44, if, after repeating steps S41 to S42 a preset number of times, there are still sub-station groups whose size exceeds the maximum station group size, then these sub-station groups are forcibly split. The aforementioned scale constraints ensure that the final output of the station cluster is of a reasonable size and internally connected, making it easy for transportation planners to use directly.
[0048] It should be noted that the preset number of iterations can be set according to the actual application scenario. In engineering practice, the size of the station group usually stabilizes after several iterations of the size constraint processing. At this point, continuing to apply the Louvain algorithm can no longer effectively reduce the size of the station group. Considering that the Louvain algorithm itself has a certain degree of randomness, but its results after multiple runs are usually relatively stable, a preset number of iterations is set as a termination condition, for example, the preset number of iterations can be set to 3 or 5. When S41 to S42 are repeated to reach this preset number of iterations, if there are still sub-station groups with a size exceeding the maximum station group size, it indicates that the Louvain algorithm alone cannot meet the size constraint, and at this time, the forced segmentation process is initiated. This method ensures processing efficiency and can handle all sub-station groups that exceed the size limit.
[0049] In one possible implementation, such as Figure 5 As shown, the forced segmentation includes the following steps S441 to S444.
[0050] Step S441: Calculate the importance score of each node in the sub-station group. The importance score is obtained by weighted summation of the degree ratio of the node within the sub-station group, the weight ratio of the edge where the node is located, and the betweenness centrality of the node.
[0051] Specifically, node importance score The calculation method is as follows:
[0052] in, Let i be the degree of station node i within the station group; Let i be the total degree of node i; For node betweenness centrality; For a group of stations; For the set of all station nodes; Station node and The connection weight between them can be the amount of data sent, the number of vehicles sent, or the average number of data sent per column. Let be the weighting coefficient, satisfying .
[0053] Step S442: Select the node with the highest score from the nodes of the sub-station group as the seed node according to the importance score from high to low.
[0054] Step S443: Using the seed node as the core, expand the sub-station group within the sub-station group using breadth-first search to form a new sub-station group with interconnected internal nodes, until the number of nodes in the new sub-station group reaches the maximum station group size. During the breadth-first search process, neighboring nodes with higher importance scores are added first.
[0055] Step S444: Remove the nodes from the newly formed sub-station group from the sub-station group, and repeat steps S442 to S443 for the remaining nodes until all nodes in the sub-station group are assigned to the new sub-station group.
[0056] By using the forced partitioning described above, we can ensure that even when the Louvain algorithm fails to meet the size constraints, we can still obtain a reasonable partition that meets the size constraints and is centered on important nodes.
[0057] After outputting the final identification results of the station groups, optionally, further analysis can be performed to provide more intuitive decision support. In one possible implementation, based on the final identification results of the station groups from the first, second, and third passenger flow network diagrams, the internal connectivity strength of each station group is calculated, and the top N station groups with the highest internal connectivity strength are selected as the focus of analysis, where N is a preset positive integer. Internal connectivity strength is defined as the ratio of the sum of the actual connection weights within a station group to the sum of the theoretical maximum connection weights in the complete graph state. The specific calculation formula is as follows:
[0058] in, For internal connection strength; For a group of stations; Station node and The connection weight between them can be the amount of data sent, the number of vehicles sent, or the average number of data sent per column. For station clusters The number of nodes in the group; the closer this value is to 1, the denser the connections within the group.
[0059] Simultaneously, the weighted degree centrality of each station node within the key analysis object is calculated and output in descending order. Weighted degree centrality refers to the sum of the connection weights of all nodes connected to that node, i.e. ,in Station node and The connection weight between them can be the number of transmissions, the number of vehicles, or the average number of transmissions per column.
[0060] Based on the calculations of internal connectivity strength and weighted degree centrality, the results can be organized and output to form an analytical report that can be directly used for decision-making. In one possible implementation, the following four parts are output for the first, second, and third passenger flow network diagrams respectively: First, a list of station clusters identified based on the data of this dimension, for example, listing all detected station clusters and their constituent station nodes one by one. Second, a list of station clusters sorted by internal connectivity strength, arranged in descending order, with station clusters having higher internal connectivity strength ranking higher in the list. Third, a list of station nodes sorted by weighted degree centrality, i.e., for each key analysis object under this dimension, the station nodes are sorted and output in descending order of weighted degree centrality. Fourth, optimization suggestions are generated based on the above output results: for station clusters with high internal connectivity strength, they are designated as key research areas for optimizing train operation schemes, such as increasing train frequency or adjusting timetables between these station clusters; for stations located at the center of station clusters in multiple dimensions, they are planned as transfer hubs to improve the overall transfer efficiency of the network. Through the above results, this invention can directly provide structured and quantitative support information for railway transportation planning and operation decisions.
[0061] To better illustrate the beneficial effects of this method, the following example uses actual monthly operating data from the Jiangsu intercity railway network. This case includes 136 stations. Based on the train operation records between each station, 5952 pairs of inter-station origin-destination (OD) relationships were constructed, and a passenger flow network diagram was built based on this. The maximum station group size was preset to 50. Subsequently, station group identification and size constraint processing were performed based on two dimensions: inter-station passenger volume and inter-station train frequency.
[0062] When identifying stations based on inter-station transmission volume, they are ranked by internal connectivity strength. The top five station groups each have 46 stations (e.g., ...). Figure 6As shown), 24, 43, 24, and 19 stations respectively. The top-ranked station group mainly includes stations on the Xuzhou-Yancheng Passenger Dedicated Line, the Lianyungang-Zhenjiang Railway, and the Shanghai-Nanjing Intercity Railway, indicating that these intercity railways have the closest connections in terms of passenger volume, and the demand for passenger flow between stations is relatively high. When identifying stations based on the number of trains operating between them, the top five station groups are ranked by internal connectivity strength, with each group having 30 stations (e.g., 24, 43, 24, and 19 stations respectively). Figure 7 (As shown), 39, 41, 10, and 27 stations respectively. The station group ranked first mainly includes stations on the Lianyungang-Zhenjiang Railway, Qingdao-Yancheng Railway, and Yancheng-Tongzhou High-speed Railway, indicating that these intercity railways are most closely connected in terms of train frequency, with a large number of trains running between the stations.
[0063] Normally, the number of trains between stations is strongly correlated with the volume of passenger traffic between stations. However, after performing cluster analysis on the intercity railway network in Jiangsu Province from different dimensions, a discrepancy was found between the volume of passenger traffic between stations and the number of trains. This further indicates that there are problems with the operation plan of the intercity railway network in Jiangsu Province during this period, which need to be improved. This fully demonstrates that the technical solution provided by this method can reveal the passenger flow correlation between stations from multiple dimensions, providing a clear quantitative basis for optimizing train operation plans.
[0064] Based on the same inventive concept, this invention also provides a system for identifying station groups with high passenger flow correlation in complex railway networks. The functional modules of this system correspond one-to-one with the steps of the method described above. The network construction module executes step S1, constructing a passenger flow network graph using stations in the target railway network as nodes, and establishing connections between stations with existing train operation records. The weight calculation module executes step S2, calculating passenger flow characteristics between any two stations based on train operation data and using these characteristics as weights for the corresponding connections. The station group identification module executes step S3, identifying station groups in the passenger flow network graph based on the weights to obtain preliminary identification results. The scale constraint module executes step S4, applying scale constraints based on a preset maximum station group size and outputting the final identification results. The specific implementation of each module can employ software programs, application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs).
[0065] Based on the same inventive concept, the present invention also provides another system for identifying station groups with high passenger flow correlation in complex railway networks, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in any of the above method embodiments.
[0066] For example, a processor may include one or more processing units, such as a neural network processing unit (NPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a digital signal processor (DSP), a baseband processor, etc. The different processing units may be independent devices or integrated into one or more processors. The controller can generate operation control signals based on the instruction opcode and timing signals to control instruction fetching and execution.
[0067] The memory can be used to store executable program code, including instructions. Internal memory may include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function, etc. The data storage area may store data created during the use of the electronic device. Furthermore, internal memory may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFs), etc. The processor executes various functional applications and data processing of the electronic device by running instructions stored in the internal memory and / or instructions stored in memory located within the processor.
[0068] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for identifying station groups with high passenger flow correlation in complex railway networks, characterized in that, include: S1: Using stations in the target railway network as nodes, construct a passenger flow network graph, where a connection is established between any two stations if there are train operation records between them; S2: Based on train operation data, calculate the passenger flow characteristics between any two stations, and use the passenger flow characteristics as the weight of the corresponding connection relationship in the passenger flow network graph; S3: Based on the weights, identify station groups in the passenger flow network map to obtain preliminary station group identification results; S4: Based on the preset maximum station group size, the preliminary identification results of the station group are constrained by size, and the final identification results of the station group are output.
2. The method for identifying high-passenger-flow-related station groups in a complex railway network according to claim 1, characterized in that, The train operation data includes: train number, origin station, destination station, train dispatch volume, passenger arrival and departure times between stations, distance between stations, and station coordinates.
3. The method for identifying high-passenger-flow-related station groups in a complex railway network according to claim 1, characterized in that, The passenger flow characteristics include: the total number of trains, the total number of trains, and the average number of trains per train, for all trains with any two stations as origin and destination.
4. The method for identifying high-passenger-flow-related station groups in a complex railway network according to claim 1, characterized in that, The station group identification includes: S31: Initialize each station node as an independent station group; S32: Calculate the modularity gain when each station node is assigned to each adjacent station group, and assign each station node to the adjacent station group that maximizes the modularity gain; if all modularity gains are less than or equal to zero, keep the original station group affiliation of the corresponding station node; repeat S32 until the first convergence condition is met. S33: Map each station group obtained after S32 to a new station node. The weight between any two new station nodes is the sum of all weights between the corresponding two original station groups. Construct a simplified network graph based on the new station nodes and weights. S34: For the simplified network graph, repeat S31 to S33 until the second convergence condition is met, output the final simplified network graph, and take each station group in the final simplified network graph as the preliminary identification result of the station group.
5. The method for identifying high-passenger-flow-related station groups in a complex railway network according to claim 4, characterized in that, The first convergence condition includes: the change in modularity after two consecutive executions of S32 is less than the first convergence threshold; the second convergence condition includes: the change in modularity of the simplified network graph obtained in two consecutive executions is less than the second convergence threshold; wherein, the modularity is calculated based on the weights.
6. The method for identifying high-passenger-flow-related station groups in a complex railway network according to claim 1, characterized in that, The size constraints include: S41: Determine whether the size of each station group exceeds the maximum station group size; if not, retain the station group. S42: If the number of nodes exceeds the limit, check whether all nodes in the station group are connected within the station group. If they are connected, identify the station group and generate multiple sub-station groups. If there are nodes in the station group that are not connected within the station group, split the station group into multiple sets of nodes with connected internal nodes and treat each set of nodes as a sub-station group. S43: For each sub-station group generated by S42, repeat S41 to S42 until the size of all sub-station groups is less than or equal to the size of the maximum station group. S44: If, after repeating S41 to S42 a preset number of times, there is still a sub-station group whose size exceeds the maximum station group size, then the sub-station group is forcibly split.
7. The method for identifying high-passenger-flow-related station groups in a complex railway network according to claim 6, characterized in that, The forced segmentation includes: S441: Calculate the importance score of each node in the sub-station group. The importance score is obtained by weighted summation based on the degree ratio of the node within the sub-station group, the weight ratio of the edge where the node is located, and the betweenness centrality of the node. S442: Select the node with the highest score from the nodes of the sub-station group as the seed node according to the importance score from high to low. S443: Using the seed node as the core, expand the sub-station group within the sub-station group through breadth-first search to form a new sub-station group with interconnected internal nodes, until the number of nodes contained in the new sub-station group reaches the maximum station group size; S444: Remove the nodes that have formed a new sub-station group from the sub-station group, and repeat S442 to S443 for the remaining nodes until all nodes in the sub-station group have been assigned to the new sub-station group.
8. The method for identifying high-passenger-flow-related station groups in a complex railway network according to claim 1, characterized in that, The process also includes performing the following steps on the final identification results of the station group: Based on the final identification results of the station groups, the internal connection strength of each station group is calculated, and the N station groups with the highest internal connection strength are selected as the key analysis objects; where N is a preset positive integer. Calculate the weighted degree centrality of each station node within the key analysis object and output it in descending order.
9. A system for identifying station groups with high passenger flow correlation in complex railway networks, characterized in that, include: The network construction module is used to construct a passenger flow network graph by using stations in the target railway network as nodes. If there are train operation records between any two stations, a connection relationship is established. The weight calculation module is used to calculate the passenger flow characteristics between any two stations based on train operation data, and use the passenger flow characteristics as the weight of the corresponding connection relationship in the passenger flow network graph. The station group identification module is used to identify station groups in the passenger flow network map according to the weights, and obtain preliminary station group identification results. The scale constraint module, based on the preset maximum station group size, applies scale constraints to the preliminary identification results of the station group and outputs the final identification results of the station group.
10. A system for identifying high-passenger-flow-related station groups in a complex railway network, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements a method for identifying a group of stations with high passenger flow correlation in a complex railway network as described in any one of claims 1 to 8.