A method and system for assessing passenger flow risk at subway stations based on commuting attributes
By using a commuter-attribute-based passenger flow risk assessment method for subway stations, and leveraging subway card swiping data and multi-algorithm analysis, the problem of not considering differences in station functional attributes in existing technologies is solved, enabling accurate assessment and dynamic management of passenger flow risks at subway stations.
Patent Information
- Application Number
- CN202511048193.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-29
AI Technical Summary
Existing subway passenger flow risk assessment methods do not fully consider the differences in station functional attributes, resulting in a discrepancy between the assessment results and the actual risk distribution. Furthermore, they lack objectivity and accuracy, making it difficult to identify high-risk stations.
The subway station passenger flow risk assessment method based on commuting attributes extracts passenger flow data from subway card swiping data, combines entropy weight method and TOPSIS method for objective weighting, constructs a differentiated risk assessment index system, and uses k-means clustering analysis to classify risk levels.
It enables accurate classification and risk assessment of commuter and non-commuter stations, provides a scientific basis for risk management, and improves the efficiency of handling large passenger flows and the accuracy of resource allocation.
Smart Images

Figure CN120542943B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban rail transit operation and management technology, specifically to a method and system for assessing passenger flow risk at subway stations based on commuting attributes. Background Technology
[0002] With the acceleration of urbanization in my country, urban rail transit is playing an increasingly prominent role in commuting. As an important component of the urban public transportation system, rail transit, with its advantages of large carrying capacity and high transportation efficiency, has become a key means of alleviating traffic pressure in large cities; its punctuality, efficiency, and convenience also make it the preferred mode of transportation for many residents.
[0003] However, the continuous expansion of urban areas and the rapid agglomeration of population have led to a sharp increase in demand for rail transit commuting services. Especially during weekday morning and evening rush hours, station passenger flow has exploded, with some hub stations exceeding their design capacity during peak hours, revealing the public's strong demand for rail transit commuting services. This increased demand presents both development opportunities and serious challenges to station service quality. The tidal nature of commuter flow and the spatial-temporal mismatch between station service capacity are becoming increasingly prominent, becoming a key bottleneck restricting the improvement of rail transit system efficiency.
[0004] Existing subway passenger flow risk assessment methods often employ single indicators or generic models, failing to adequately consider the differences in station functional attributes (such as the differences in passenger flow characteristics between commuter and non-commuter stations), leading to discrepancies between assessment results and actual risk distribution. For example, traditional methods typically use daily passenger volume as the primary indicator, but neglect key factors such as the clustering effect during morning and evening rush hours and the elasticity of changes during holidays. Furthermore, the weighting of indicators in existing technologies relies heavily on subjective experience, lacking objectivity; risk classification methods are also relatively simple, making it difficult to accurately identify high-risk stations. Summary of the Invention
[0005] This invention provides a method and system for assessing passenger flow risk at subway stations based on commuting attributes, in order to solve at least one of the above-mentioned technical problems.
[0006] The technical solution of this invention to solve the above-mentioned technical problems is as follows: A method for assessing passenger flow risk at subway stations based on commuting attributes, comprising:
[0007] S1, based on different dates and different time granularities, statistically analyzes the inbound and outbound passenger flow of each station in each time window from the subway card swiping data, and obtains inbound passenger flow dataset and outbound passenger flow dataset;
[0008] S2, based on the inbound passenger flow dataset and the outbound passenger flow dataset, calculate multiple commuting characteristic indicators for each station based on morning and evening peak hours, and use the entropy weight method combined with the TOPSIS method to objectively assign weights and perform multi-attribute decision processing on each commuting characteristic indicator for each station to obtain the commuting index for each station, and determine the commuting attribute of each station by combining the surrounding land use attributes; wherein, the commuting attribute includes commuting category and non-commuting category;
[0009] S3, construct a commuter passenger flow risk assessment index system for commuter stations and a non-commuter passenger flow risk assessment index system for non-commuter stations; based on the commuter attributes of each station, construct a commuter passenger flow risk assessment matrix and a non-commuter passenger flow risk assessment matrix using the commuter passenger flow risk assessment index system and the non-commuter passenger flow risk assessment index system respectively.
[0010] S4. Using the entropy weight method combined with the TOPSIS method, objective weighting and multi-attribute decision processing are performed on the commuter passenger flow risk assessment matrix and the non-commuter passenger flow risk assessment matrix respectively, to obtain the passenger flow risk score of each commuter station and the passenger flow risk score of each non-commuter station.
[0011] S5. The k-means clustering analysis method is used to perform cluster analysis on the passenger flow risk scores of each commuter station and each non-commuter station to obtain the passenger flow risk level of each commuter station and each non-commuter station.
[0012] Based on the above-mentioned method for assessing passenger flow risk at subway stations based on commuting attributes, this invention also provides a system for assessing passenger flow risk at subway stations based on commuting attributes.
[0013] A subway station passenger flow risk assessment system based on commuting attributes includes:
[0014] The passenger flow statistics module is used to calculate the inbound and outbound passenger flow of each station within each time window based on different dates and time granularities from the subway card swiping data, and obtain the inbound passenger flow dataset and the outbound passenger flow dataset.
[0015] The station commuting attribute judgment module is used to calculate multiple commuting characteristic indicators for each station based on the inbound passenger flow dataset and the outbound passenger flow dataset, and to perform objective weighting and multi-attribute decision processing on each commuting characteristic indicator of each station using the entropy weight method combined with the TOPSIS method to obtain the commuting index of each station, and to judge the commuting attribute of each station in combination with the surrounding land use attributes; wherein, the commuting attribute includes commuting category and non-commuting category;
[0016] The risk indicator construction module is used to construct a commuter passenger flow risk assessment indicator system for commuter stations and a non-commuter passenger flow risk assessment indicator system for non-commuter stations. Based on the commuting attributes of each station, the module constructs a commuter passenger flow risk assessment matrix and a non-commuter passenger flow risk assessment matrix using the commuter passenger flow risk assessment indicator system and the non-commuter passenger flow risk assessment indicator system, respectively.
[0017] The risk score calculation module is used to objectively assign weights and perform multi-attribute decision processing on the commuter passenger flow risk assessment matrix and the non-commuter passenger flow risk assessment matrix using the entropy weight method combined with the TOPSIS method, respectively, to obtain the passenger flow risk score of each commuter station and the passenger flow risk score of each non-commuter station.
[0018] The risk level classification module uses k-means clustering analysis to perform cluster analysis on the passenger flow risk scores of each commuter station and each non-commuter station, respectively, to obtain the passenger flow risk level of each commuter station and each non-commuter station.
[0019] The beneficial effects of this invention are as follows: This invention provides a subway station passenger flow risk assessment method and system based on commuting attributes. First, it extracts multi-time-granularity passenger flow data through data preprocessing, combining indicators such as the proportion of morning and evening peak hours and the ratio of inbound to outbound passenger flow to accurately distinguish between commuting and non-commuting stations. Second, it constructs differentiated risk assessment indicator systems for the two types of stations: commuting stations focus on weekday peak pressure, while non-commuting stations introduce flexible passenger flow indicators during holidays. Furthermore, it quantifies station risk scores through objective weighting using the entropy weight method and ranking using the TOPSIS method. Finally, it classifies risk levels based on k-means clustering to achieve hierarchical management. This invention integrates the advantages of multiple algorithms, solving the shortcomings of traditional models such as poor scenario adaptability and coarse classification, providing subway operation departments with a scientific and dynamic basis for risk management, and effectively improving the efficiency of handling large passenger flows and the accuracy of resource allocation. Attached Figure Description
[0020] Figure 1 This is a flowchart of a subway station passenger flow risk assessment method based on commuting attributes according to the present invention;
[0021] Figure 2 This is a structural block diagram of a subway station passenger flow risk assessment system based on commuting attributes, according to the present invention. Detailed Implementation
[0022] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0023] like Figure 1As shown, a method for assessing passenger flow risk at subway stations based on commuting attributes includes the following S1~S5:
[0024] S1, based on different dates and different time granularities, statistically analyzes the inbound and outbound passenger flow of each station in each time window from the subway card swiping data, and obtains inbound passenger flow dataset and outbound passenger flow dataset;
[0025] S2, based on the inbound passenger flow dataset and the outbound passenger flow dataset, calculate multiple commuting characteristic indicators for each station based on morning and evening peak hours, and use the entropy weight method combined with the TOPSIS method to objectively assign weights and perform multi-attribute decision processing on each commuting characteristic indicator for each station to obtain the commuting index for each station, and determine the commuting attribute of each station by combining the surrounding land use attributes; wherein, the commuting attribute includes commuting category and non-commuting category;
[0026] S3, construct a commuter passenger flow risk assessment index system for commuter stations and a non-commuter passenger flow risk assessment index system for non-commuter stations; based on the commuter attributes of each station, construct a commuter passenger flow risk assessment matrix and a non-commuter passenger flow risk assessment matrix using the commuter passenger flow risk assessment index system and the non-commuter passenger flow risk assessment index system respectively.
[0027] S4. Using the entropy weight method combined with the TOPSIS method, objective weighting and multi-attribute decision processing are performed on the commuter passenger flow risk assessment matrix and the non-commuter passenger flow risk assessment matrix respectively, to obtain the passenger flow risk score of each commuter station and the passenger flow risk score of each non-commuter station.
[0028] S5. The k-means clustering analysis method is used to perform cluster analysis on the passenger flow risk scores of each commuter station and each non-commuter station to obtain the passenger flow risk level of each commuter station and each non-commuter station.
[0029] The following is a detailed explanation of each step:
[0030] Before step S1, the following preparatory work is required: Collect raw card swiping data (Automatic Fare Collection System Data, or AFC data) from each subway station to obtain the raw subway card swiping dataset for each station. Each raw subway card swiping data entry includes information such as card number, transaction sequence number, transaction time, transaction route, transaction station, transaction gate, transaction type, card type, entry transaction time, transaction route, and entry transaction station. Clean the raw subway card swiping dataset to remove abnormal records, such as duplicate test card swiping and data with time logic errors, to obtain the final subway card swiping dataset.
[0031] Based on the subway card swipe dataset, and combined with the card swipe ID to identify the origin and destination points of travel, a travel origin and destination point dataset is obtained, which identifies the passenger's entry and exit stations and the lines they belong to. It should be noted that the same passenger may enter and exit the station multiple times within a day, so the data of the nearest entry and exit station needs to be grouped together. Based on the travel origin and destination station dataset, according to station information and entry and exit times, and based on different time granularities (i.e., time intervals, such as 5 minutes, 15 minutes, 30 minutes, 1 hour), the passenger flow entering and exiting each station can be statistically analyzed to obtain the passenger flow entering and exiting each station within each time window. For example, taking a time granularity of 15 minutes as an example, firstly, based on the specific operating time of the subway, the card swiping data between 0:00 and 6:00 is determined to be staff card swiping data, and the effective travel data focuses on the period from 6:00 to 24:00; then, with 15-minute intervals, it can be divided into 72 time windows, with each time window being 6:00-6:15, 6:15-6:30, ..., 23:45-24:00; finally, the passenger flow entering and exiting the station on different dates in each time window is statistically analyzed.
[0032] In S1, the inbound passenger flow dataset and the outbound passenger flow dataset are respectively represented as:
[0033] , ;
[0034] in, This represents the inbound passenger flow dataset. This represents the outbound passenger flow dataset. It belongs to a set of dates consisting of multiple different dates. It belongs to a set of time granularities consisting of multiple different time granularities; and They represent the first time. The date and the The inbound and outbound passenger flow matrices at different time granularities, and and The expressions are as follows:
[0035] , ;
[0036] in, It belongs to a collection of sites consisting of multiple sites. It belongs to a set of time windows consisting of multiple time windows. This represents the total number of stations in the station set (where transfer stations are counted separately by line; for example, if a station belongs to two subway lines, it should be counted as two stations when counting the total number of stations). This indicates the total number of time windows in the time window set; and They represent the first time. The date and the The first time granularity The site at the Passenger flow entering and exiting the station within a specific time window.
[0037] Using different dates and time granularities to count the inbound and outbound passenger flow at each station can provide more refined input for subsequent analysis; for example, weekday data can focus on morning and evening peak hours, while holiday data reflects flexible travel characteristics.
[0038] In S2, considering the significant differences between commuter stations and non-commuter stations in terms of passenger flow characteristics such as time distribution and personnel composition, as well as functional positioning, it is necessary to determine whether a station is a commuter station or a non-commuter station before conducting a passenger flow risk assessment in order to conduct a targeted assessment of station risks and customize preventive measures.
[0039] Specifically, the entropy weight method combined with the TOPSIS method is used to process various commuting characteristic indicators for each station, including the following S21~S24:
[0040] S21, construct a commuting feature index matrix based on the various commuting feature indicators of each station; the commuting feature index matrix is represented as:
[0041] ;
[0042] in, It belongs to a set of commuting characteristic indicators, which consists of multiple commuting characteristic indicators. This represents the total number of commuting feature indicators in the commuting feature indicator set; Indicates the first The date and the A matrix of commuting characteristic indicators at different time granularities. Indicates the first The date and the The first time granularity The first site Commuting characteristic indicators.
[0043] S22, Standardize the commuting feature index matrix to obtain a standardized commuting feature index matrix; the standardized commuting feature index matrix is expressed as follows:
[0044] ;
[0045] in, Indicates the first The date and the A standardized commuting characteristic index matrix at a time granularity. Indicates the first The date and the The first time granularity The first site Standardized values of commuting characteristic indicators, and In addition, the data in the matrix generally needs to be positiveized before standardization, but since the values of the commuting characteristic indicators are all positive, this invention does not require positiveization.
[0046] S23, calculate the relative entropy probability of each commuting feature index based on the standardized commuting feature index matrix; the formula for calculating the relative entropy probability of each commuting feature index is as follows:
[0047] ;
[0048] in, Indicates the first The date and the The first time granularity The first site The relative entropy probability of each commuting characteristic indicator;
[0049] Based on the relative entropy probability of each commuting characteristic indicator, the information entropy of each commuting characteristic indicator is calculated; the formula for calculating the information entropy of each commuting characteristic indicator is as follows:
[0050] ;
[0051] in, Indicates the first Information entropy of commuting characteristic indicators;
[0052] The credit validity value of each commuting feature indicator is calculated based on its information entropy; the formula for calculating the credit validity value of each commuting feature indicator is as follows:
[0053] ;
[0054] in, Indicates the first Credit validity values of each commuting characteristic indicator;
[0055] The entropy weight of each commuting feature indicator is calculated based on its effective credit value; the formula for calculating the entropy weight of each commuting feature indicator is as follows:
[0056] ;
[0057] in, Indicates the first Entropy weights of commuting characteristic indicators.
[0058] S23 is the process of determining the weights of each indicator using the entropy weight method.
[0059] S24, construct a commuting feature index weighting matrix based on the entropy weights of each commuting feature index and the standardized commuting feature index matrix; the commuting feature index weighting matrix is expressed as:
[0060] ;
[0061] in, Indicates the first The date and the A weighted matrix of commuting characteristic indicators at various time granularities. Indicates the first The date and the The first time granularity The first site The weighted values of each commuting characteristic indicator, and ;
[0062] The optimal and worst vectors of the commuting feature indicators are constructed based on the weighted matrix of the commuting feature indicators; the optimal and worst vectors of the commuting feature indicators are expressed as follows:
[0063] ;
[0064] ;
[0065] in, and These represent the optimal and worst vectors of commuting characteristic indicators, respectively. The optimal vector representing commuting characteristic indicators is related to the first... The element corresponding to the weighted value of the commuting characteristic index, and used to characterize the The date and the The first time granularity for all stations The maximum value of the weighted average of the commuting characteristic indicators; The worst vector representing commuting characteristic indicators is related to the first... The element corresponding to the weighted value of the commuting characteristic index, and used to characterize the The date and the The first time granularity for all stations The minimum weighted value of each commuting characteristic indicator;
[0066] Calculate the Euclidean distance between each station and the optimal and worst vectors of the commuting feature indicators; the formula for calculating the Euclidean distance between each station and the optimal and worst vectors of the commuting feature indicators is as follows:
[0067] ;
[0068] ;
[0069] in, Indicates the first The date and the The first time granularity The Euclidean distance (also known as the positive ideal solution) between each station and the optimal vector of commuting characteristic indicators. Indicates the first The date and the The first time granularity The Euclidean distance (also known as the negative ideal solution) between each station and the worst vector of commuting characteristic indicators.
[0070] The commuting index for each station is calculated based on the Euclidean distance between each station and the optimal and worst vectors of the commuting characteristic indicators. The formula for calculating the commuting index for each station is as follows:
[0071] ;
[0072] in, Indicates the first The date and the The first time granularity Commuting index for each station.
[0073] S24 is the process of calculating the commuting index of each station using the TOPSIS method. The commuting index is also called the TOPSIS score.
[0074] After obtaining the commuting index of each station, a comprehensive judgment is made in combination with the surrounding land use attributes (such as the distribution of residential areas and commercial areas), and finally the station attributes are divided into commuting and non-commuting categories, so as to achieve accurate classification based on passenger flow characteristics.
[0075] In this embodiment, the plurality of commuting characteristic indicators include commuting characteristic indicators for characterizing the proportion of passenger flow entering and exiting the station during morning and evening peak hours to the total daily passenger flow, commuting characteristic indicators for characterizing the ratio of passenger flow entering to exiting the station during the morning peak hours, and commuting characteristic indicators for characterizing the ratio of passenger flow exiting to entering the station during the evening peak hours.
[0076] The morning peak is from 7:30 to 8:30, and the evening peak is from 17:30 to 18:30. Of course, the time periods for the morning and evening peaks can be adjusted appropriately to meet actual needs in order to adapt to different regions or other conditions.
[0077] In S3, in order to accurately assess the passenger flow risk of stations with different attributes, and to fully consider the significant differences in passenger flow characteristics between commuter stations and non-commuter stations, a corresponding passenger flow risk evaluation index system is constructed to measure the passenger flow risk status of each station from multiple dimensions and in all aspects, so as to provide a strong basis for the subsequent formulation of targeted risk prevention and control measures.
[0078] In this embodiment, the evaluation indicators in the commuter passenger flow risk assessment indicator system include:
[0079] a1: Daily passenger flow in and out of the station on weekdays (10,000 people / day). This indicator reflects the overall passenger flow on weekdays and the daily carrying capacity requirements of the station.
[0080] a2: Passenger flow in and out of the station during the morning peak (10,000 people / hour). This indicator focuses on the passenger flow intensity during the commuting peak hours and measures the operational pressure of the station during the busiest time.
[0081] a3: Weekend daily passenger flow (10,000 people / day). This indicator shows the passenger flow at the station during the weekend and provides insight into the passenger flow level on non-working days.
[0082] a4: Weekend passenger flow elasticity ratio (a3 / a1). This indicator assesses the range of passenger flow changes between weekdays and weekends by comparing passenger flow on weekends and weekdays, and judges the stability and fluctuation of passenger flow.
[0083] In this embodiment, the evaluation indicators in the non-commuter passenger flow risk assessment indicator system include:
[0084] b1: Daily passenger flow in and out of the station on weekdays (10,000 people / day). This indicator reflects the overall passenger flow of non-commuter stations on weekdays.
[0085] b2: Weekend daily passenger flow (10,000 people / day). This indicator shows the passenger flow situation of non-commuter stations on weekends.
[0086] b3: Weekend passenger flow elasticity ratio (b2 / b1), this indicator assesses the difference in passenger flow changes between weekends and weekdays for non-commuter stations;
[0087] b4: Daily passenger flow in and out of the station during holidays (10,000 people / day). This indicator highlights the passenger flow during holidays. For non-commuter stations, holidays are often the peak period for passenger flow.
[0088] b5: Holiday passenger flow elasticity ratio (b4 / b1). This indicator measures the change in passenger flow between holidays and weekdays, and helps to understand the passenger flow fluctuation characteristics of non-commuter stations during holidays.
[0089] Furthermore, based on the commuting attributes of each station, the commuting passenger flow risk assessment matrix and the non-commuting passenger flow risk assessment matrix, constructed using the aforementioned commuting passenger flow risk assessment index system and the aforementioned non-commuting passenger flow risk assessment index system, are respectively expressed as follows:
[0090] , ;
[0091] in, It belongs to a collection of commuter stations, consisting of multiple commuter stations. It belongs to a collection of non-commuter sites consisting of multiple non-commuter sites. This belongs to the aforementioned commuter passenger flow risk assessment index system. This belongs to the aforementioned non-commuter passenger flow risk assessment index system. This represents the total number of commuter stations in the commuter station set. This represents the total number of non-commuter stations in the set of non-commuter stations. This indicates the total number of evaluation indicators in the commuter passenger flow risk assessment indicator system. This indicates the total number of evaluation indicators in the non-commuter passenger flow risk assessment indicator system; In the first The date and the A risk assessment matrix for commuter passenger flow at a time granularity. In the first The date and the Risk assessment matrix for non-commuter passenger flow at different time granularities. Indicates the first The date and the The first time granularity The first commuter station One evaluation indicator, Indicates the first The date and the The first time granularity The first non-commuter station One evaluation indicator.
[0092] In S4, the process of objectively weighting and multi-attribute decision processing of the commuter passenger flow risk assessment matrix and the non-commuter passenger flow risk assessment matrix using the entropy weight method combined with the TOPSIS method is the same as the process described in S21 to S24 above, and will not be repeated here. It is only necessary to replace the processing objects (commuter characteristic indicators) in S21 to S24 above with evaluation indicators. The final index is the passenger flow risk score.
[0093] Overall, the entropy weight method-TOPSIS method combines objective weighting with multi-attribute decision-making, effectively avoiding subjective bias.
[0094] Specifically, S5 is:
[0095] S51, using the Hopkins statistic to test the clustering trend of passenger flow risk scores for each commuter station and each non-commuter station, respectively, and obtain the clustering attributes of each commuter station and each non-commuter station.
[0096] S52, the silhouette coefficient method is used to analyze the clustering attributes of each commuter station and each non-commuter station respectively, so as to determine the optimal number of clusters for each commuter station and each non-commuter station.
[0097] S53, risk levels are divided according to the optimal number of clusters for each commuter station and the optimal number of clusters for each non-commuter station, and the passenger flow risk levels for each commuter station and each non-commuter station are obtained accordingly.
[0098] Specifically, referring to the methods in S21-S24, after obtaining the Euclidean distances between each station and the optimal and worst vectors of the evaluation indicators in S4, the Euclidean distances between each station and the optimal vectors of the evaluation indicators are used as the x-axis values, and the corresponding Euclidean distances between each station and the worst vectors of the evaluation indicators are used as the y-axis values. This yields the coordinate points corresponding to each station. The set of coordinate points corresponding to all stations constitutes a coordinate point set, denoted as the coordinate point set of all commuter stations. Let the set of coordinates of all commuter stations be denoted as . .
[0099] In step S51, a clustering trend test is performed, and the Hopkins statistic is used to verify the clusterability of the data. The Hopkins statistic is selected to determine the spatial randomness of the data. The Hopkins statistic is a spatial statistic used to test the spatial randomness of spatially distributed variables, thereby determining whether the data can be clustered. The calculation steps for the Hopkins statistic are as follows:
[0100] First, uniformly from the set of coordinate points Extraction sample points , ... For each coordinate point ( ), find exist The nearest neighbor in, and make for Rather than The distance between the nearest neighbors in the array, i.e. Then, evenly from Extraction sample points , ... For each point coordinate point ( ), find exist The nearest neighbor in, and make for Rather than The distance between the nearest neighbors in the array, i.e. Finally, calculate the set of coordinate points. Johns Hopkins statistics The calculation formula is as follows:
[0101] .
[0102] If the extracted coordinate points are close to a random distribution, The value approaches 0.5. If the clustering trend is obvious, the distance between randomly generated coordinate points should be much greater than the distance between actual coordinate points, i.e. The value approaches 1.
[0103] Similarly, for a set of coordinate points Alternatively, the coordinate point set can be calculated using the method described above. The Hopkins statistics will not be elaborated here.
[0104] In step S52, different clustering results will be formed when the number of clusters is different. To evaluate which result is better in an unsupervised learning environment, the dispersion and compactness of the clusters can be analyzed to assess the quality of the clustering. This invention uses the silhouette coefficient method to determine the optimal number of clusters. For example ( (same), calculation for Each coordinate point Calculate the average distance between it and all other coordinate points within the same cluster, denoted as . Then for each coordinate point Find the nearest non-self cluster (called the nearest neighbor cluster) and calculate its coordinates. The average distance to all points within the cluster is denoted as . For each coordinate point The formula for calculating the profile coefficient is:
[0105] ;
[0106] Profile coefficient The value of is in the range of [-1, 1]. A value closer to 1 indicates that the coordinate point is closely clustered with its own cluster and well separated from its nearest neighbor cluster. A value closer to 0 indicates that the coordinate point is located on the boundary between two clusters, resulting in mediocre clustering. A value closer to -1 indicates that the coordinate point may have been incorrectly assigned to the wrong cluster. The optimal clustering number is then obtained. , .
[0107] In step S53, risk levels are classified according to the optimal number of clusters. This is based on the coordinate point set of commuter stations. Optimal number of clusters For example (the processing steps are the same for non-commuter sites), first randomly select... ( = ) coordinate points Using the initial cluster centroid as the initial point, calculate the distance between each coordinate point and each cluster centroid, and assign it to the nearest cluster. Find the nearest cluster using an iterative method. The cluster partitioning scheme minimizes the cost function corresponding to the clustering results.
[0108] Define the cost function:
[0109] ;
[0110] in, Indicates the category to which the coordinate point belongs; It is the first The centroid of the cluster; Represents coordinate points With cluster center The square of the Euclidean distance between them.
[0111] make =0,1,2,… represents the iteration step number. During the iteration process, assume the current… If the minimum value is not reached, then first fix the cluster center { Adjust each coordinate point Category Let Functions are reduced, then fixed { }, Adjust cluster center {},make Decrease. These two processes alternate in a cycle. Monotonically decreasing, until When decreasing to the minimum value, { }and{ The convergence also occurs simultaneously. The final result is the commuter station classification data, i.e., the dataset. … Based on the tiered results, management can optimize strategies for handling large crowds.
[0112] The method of the present invention will be illustrated below with specific examples:
[0113] This example discloses a method for evaluating 266 subway stations (transfer stations are counted as one) in a city. First, raw subway card swipe data is acquired and preprocessed. This data is then cleaned by combining card swipe IDs and entry / exit times, and the number of passengers entering and exiting each station is statistically analyzed at different time granularities. Next, based on indicators such as the proportion of passenger flow during morning and evening peak hours, and considering the land use attributes surrounding the stations, the commuting attributes of each station are accurately determined, identifying 60 commuter stations and 206 non-commuter stations. Then, a differentiated passenger flow risk assessment index system is established for stations with different attributes. The objective weights of each index are determined using the entropy weight method, and the TOPSIS method is used to rank the passenger flow risk scores of the two types of stations. Finally, k-means clustering is used to classify the passenger flow risk of commuter and non-commuter stations into three levels (high, relatively high, and relatively low). The aim is to identify the commuting attributes of stations, and through the ranking and classification of passenger flow risks based on different attributes, to promptly identify and screen high-risk stations, and to develop targeted countermeasures based on these classifications. Specifically:
[0114] Obtain raw subway card swiping data; the raw subway card swiping data includes transaction number, entry / exit transaction time, line where the entry / exit transaction occurred, station where the entry / exit transaction occurred, gate number where the entry / exit transaction occurred, card type, etc.
[0115] The original subway card swiping data was cleaned by combining card swipe ID, entry and exit time, etc. Taking into account the concentration of passenger flow at stations during weekday peak hours and the elastic changes in residents' travel on weekends and holidays, the origin and destination of passenger travel were identified, and the passenger flow entering and exiting at each station was statistically analyzed at time granular intervals of 5 minutes, 15 minutes, 30 minutes, and 1 hour.
[0116] For ease of explanation, this specific example uses the average passenger flow in and out of the station on a weekday of a certain week, the average passenger flow in and out of the station on the same weekend, and the passenger flow in and out of the station during the National Day holiday.
[0117] Considering the significant differences between commuter and non-commuter stations in terms of passenger flow characteristics such as time distribution and demographics, as well as their functional positioning, it is necessary to determine whether a station is a commuter station before conducting a passenger flow risk assessment in order to effectively assess station risks and customize preventative measures. Commuting characteristic indicators include:
[0118] e1: The percentage of passenger flow entering and exiting the station during morning and evening peak hours out of the total daily passenger flow;
[0119] e2: The ratio of passenger flow entering the station to passenger flow exiting the station during the morning peak;
[0120] e3: The ratio of passenger flow exiting the station to passenger flow entering the station during the evening peak.
[0121] It should be added that the calculation time for the whole day is from 6:00 to 24:00, the calculation time for the morning peak period is from 7:30 to 8:30, and the calculation time for the evening peak period is from 17:30 to 18:30.
[0122] After data cleaning in the previous step, the commuting characteristic indicators were processed using the entropy weight method and the TOPSIS method. First, positiveization and standardization were performed to obtain the objective weights of each evaluation indicator, which were 0.09, 0.55, and 0.36, respectively. Then, the commuting index of each station was ranked based on the TOPSIS method. Finally, considering factors such as the surrounding land conditions, it was determined whether a station was a commuting station. Of the final 266 stations, 60 were commuting stations and 206 were non-commuting stations. Table 1 below shows the commuting characteristic indicators and commuting classification results for some stations.
[0123] Table 1: Partial Station Indicator Table and Commuting Classification Results Table
[0124]
[0125] Step 3: Establish differentiated passenger flow risk assessment indicators.
[0126] By designing differentiated indicators, we can avoid a "one-size-fits-all" approach to evaluation and improve the model's adaptability to different scenarios.
[0127] For commuter stations, the selected passenger flow risk assessment index system includes the following evaluation indicators: a1 All-day passenger flow in and out of all stations on weekdays (10,000 people / day), a2 Passenger flow in and out of stations during morning peak hours (10,000 people / hour), a3 All-day passenger flow in and out of stations on weekends (10,000 people / day), and a4 Weekend passenger flow elasticity ratio (a3 / a1).
[0128] For non-commuter stations, the selected passenger flow risk assessment index system includes the following evaluation indicators: b1. Daily passenger flow of all stations across the network on weekdays (10,000 people / day); b2. Daily passenger flow of all stations on weekends (10,000 people / day); b3. Weekend passenger flow flexibility ratio (b2 / b1); b4. Daily passenger flow of all stations on holidays (10,000 people / day); b5. Holiday passenger flow flexibility ratio (b4 / b1). Table 2 below shows the differentiated passenger flow risk assessment index system.
[0129] Table 2: Differentiated Passenger Flow Risk Assessment Index System
[0130]
[0131] Step 4: Determine the passenger flow risk ranking using the entropy weight method-TOPSIS method.
[0132] For commuter stations, the objective weights of each evaluation indicator were obtained using the entropy weight method, which were 0.24, 0.24, 0.26, and 0.26 respectively. For non-commuter stations, the objective weights of each evaluation indicator were obtained using the entropy weight method, which were 0.22, 0.28, 0.07, 0.36, and 0.08 respectively. Then, the TOPSIS method was used to rank the passenger flow risk scores of the two types of stations, with higher scores indicating greater risk. Table 3 shows the standardized results and risk levels of some commuter station indicators, and Table 4 shows the standardized results and risk levels of some non-commuter station indicators. As can be seen from Tables 3 and 4, among commuter stations, station 11 has the highest passenger flow risk; among non-commuter stations, station 14 has the highest passenger flow risk.
[0133] Table 3: Standardization Results and Risk Levels of Some Commuter Stations
[0134]
[0135] Table 4: Standardization Results and Risk Levels of Some Non-Commuter Stations
[0136]
[0137] Step 5: Passenger flow risk classification based on k-means clustering.
[0138] Combining the entropy weight method with the TOPSI method, clustering techniques can be used to classify station passenger flow risks and determine risk levels. This facilitates the subsequent development of tiered operational management plans and the allocation of operational resources. Common clustering methods include k-means clustering, hierarchical clustering, and DBSCAN density clustering.
[0139] This example uses k-means clustering. Before performing K-means clustering, the data needs to be pre-analyzed to determine if it is suitable for cluster analysis. This example uses the Hopkins H-squared normality index and uniformity index for illustration. These two indices are used to assess the clustering tendency of a dataset by measuring the probability that a given dataset is generated from a uniform data distribution. If clustering exists in the dataset, H will be close to 1; when H is higher than 0.75, it indicates that there is a clustering tendency in the dataset at a 90% confidence level. Then, the optimal k-value can be calculated using the silhouette coefficient mean.
[0140] For commuter stations, the Hopkins statistical H-normality index is 0.8, indicating a clustering trend. The optimal k-value is 2, and the mean silhouette coefficient is 0.6. This suggests that this data set is suitable for cluster analysis. K-means clustering was used for classification, the process of which is not detailed here. The classification results are shown in Table 5 below. For commuter stations, stations with a cluster classification result of 1 have higher passenger flow risk, while stations with a cluster classification result of 3 have relatively lower passenger flow risk.
[0141] For non-commuter stations, the Hopkins statistical H uniformity index is 0.9, indicating a clustering trend. The optimal k value is 3, and the mean silhouette coefficient is 0.8. This indicates that this data set is suitable for cluster analysis, and the classification results are shown in Table 5. Stations with a cluster classification result of 0 have high passenger flow risk, stations with a cluster classification result of 1 have relatively high passenger flow risk, and stations with a cluster classification result of 2 have relatively low passenger flow risk.
[0142] Table 5: Clustering Classification Table of Some Commuter and Non-Commuter Stations
[0143]
[0144] Based on the above-mentioned method for assessing passenger flow risk at subway stations based on commuting attributes, this invention also provides a system for assessing passenger flow risk at subway stations based on commuting attributes.
[0145] like Figure 2 As shown, a subway station passenger flow risk assessment system based on commuting attributes includes:
[0146] The passenger flow statistics module is used to calculate the inbound and outbound passenger flow of each station within each time window based on different dates and time granularities from the subway card swiping data, and obtain the inbound passenger flow dataset and the outbound passenger flow dataset.
[0147] The station commuting attribute judgment module is used to calculate multiple commuting characteristic indicators for each station based on the inbound passenger flow dataset and the outbound passenger flow dataset, and to perform objective weighting and multi-attribute decision processing on each commuting characteristic indicator of each station using the entropy weight method combined with the TOPSIS method to obtain the commuting index of each station, and to judge the commuting attribute of each station in combination with the surrounding land use attributes; wherein, the commuting attribute includes commuting category and non-commuting category;
[0148] The risk indicator construction module is used to construct a commuter passenger flow risk assessment indicator system for commuter stations and a non-commuter passenger flow risk assessment indicator system for non-commuter stations. Based on the commuting attributes of each station, the module constructs a commuter passenger flow risk assessment matrix and a non-commuter passenger flow risk assessment matrix using the commuter passenger flow risk assessment indicator system and the non-commuter passenger flow risk assessment indicator system, respectively.
[0149] The risk score calculation module is used to objectively assign weights and perform multi-attribute decision processing on the commuter passenger flow risk assessment matrix and the non-commuter passenger flow risk assessment matrix using the entropy weight method combined with the TOPSIS method, respectively, to obtain the passenger flow risk score of each commuter station and the passenger flow risk score of each non-commuter station.
[0150] The risk level classification module uses k-means clustering analysis to perform cluster analysis on the passenger flow risk scores of each commuter station and each non-commuter station, respectively, to obtain the passenger flow risk level of each commuter station and each non-commuter station.
[0151] The specific functions of each module in the subway station passenger flow risk assessment system based on commuting attributes of this invention are described in the specific steps of the subway station passenger flow risk assessment method based on commuting attributes of this invention, and will not be repeated here.
[0152] This invention presents a method and system for assessing passenger flow risk at subway stations based on commuting attributes. First, it extracts multi-time-granularity passenger flow data through data preprocessing, combining indicators such as the proportion of morning and evening peak hours and the entry / exit ratio to accurately distinguish between commuting and non-commuting stations. Second, it constructs differentiated risk assessment indicator systems for the two types of stations: commuting stations focus on weekday peak pressure, while non-commuting stations incorporate flexible passenger flow indicators during holidays. Further, it quantifies station risk scores through objective weighting using the entropy weight method and ranking using the TOPSIS method. Finally, it classifies risk levels based on k-means clustering to achieve hierarchical management. This invention integrates the advantages of multiple algorithms, overcoming the shortcomings of traditional models such as poor scenario adaptability and coarse classification, providing subway operation departments with a scientific and dynamic basis for risk management, and effectively improving the efficiency of handling large passenger flows and the accuracy of resource allocation.
[0153] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for assessing passenger flow risk at subway stations based on commuting attributes, characterized in that, include: S1, based on different dates and different time granularities, statistically analyzes the inbound and outbound passenger flow of each station in each time window from the subway card swiping data, and obtains inbound passenger flow dataset and outbound passenger flow dataset; S2, based on the inbound passenger flow dataset and the outbound passenger flow dataset, calculate multiple commuting characteristic indicators for each station based on morning and evening peak hours, and use the entropy weight method combined with the TOPSIS method to objectively assign weights and perform multi-attribute decision processing on each commuting characteristic indicator for each station to obtain the commuting index for each station, and determine the commuting attribute of each station by combining the surrounding land use attributes; wherein, the commuting attribute includes commuting category and non-commuting category; S3, construct a commuter passenger flow risk assessment index system for commuter stations and a non-commuter passenger flow risk assessment index system for non-commuter stations; based on the commuter attributes of each station, construct a commuter passenger flow risk assessment matrix and a non-commuter passenger flow risk assessment matrix using the commuter passenger flow risk assessment index system and the non-commuter passenger flow risk assessment index system respectively. S4. Using the entropy weight method combined with the TOPSIS method, objective weighting and multi-attribute decision processing are performed on the commuter passenger flow risk assessment matrix and the non-commuter passenger flow risk assessment matrix respectively, to obtain the passenger flow risk score of each commuter station and the passenger flow risk score of each non-commuter station. S5. The k-means clustering analysis method is used to perform clustering analysis on the passenger flow risk scores of each commuter station and each non-commuter station to obtain the passenger flow risk level of each commuter station and each non-commuter station. The multiple commuting characteristic indicators include commuting characteristic indicators that characterize the proportion of passenger flow entering and exiting the station during the morning and evening peak hours to the total daily passenger flow, commuting characteristic indicators that characterize the ratio of passenger flow entering to exiting the station during the morning peak hour, and commuting characteristic indicators that characterize the ratio of passenger flow exiting to entering the station during the evening peak hour. In S3, the evaluation indicators in the commuter passenger flow risk assessment indicator system include: Weekday passenger flow, morning peak passenger flow, weekend passenger flow, and weekend passenger flow elasticity ratio; The evaluation indicators in the non-commuter passenger flow risk assessment indicator system include: the total number of passengers entering and exiting the station on weekdays, the total number of passengers entering and exiting the station on weekends, the flexible ratio of passenger flow on weekends, the total number of passengers entering and exiting the station on holidays, and the flexible ratio of passenger flow on holidays. Among them, the weekend passenger flow elasticity ratio is the ratio of the total passenger flow in and out of the station on weekends to the total passenger flow in and out of the station on weekdays, and the holiday passenger flow elasticity ratio is the ratio of the total passenger flow in and out of the station on holidays to the total passenger flow in and out of the station on weekdays.
2. The subway station passenger flow risk assessment method based on commuting attributes according to claim 1, characterized in that, In S1, the inbound passenger flow dataset and the outbound passenger flow dataset are respectively represented as: , ; in, This represents the inbound passenger flow dataset. This represents the outbound passenger flow dataset. It belongs to a set of dates consisting of multiple different dates. It belongs to a set of time granularities consisting of multiple different time granularities; and They represent the first time. The date and the The inbound and outbound passenger flow matrices at different time granularities, and and The expressions are as follows: , ; in, It belongs to a collection of sites consisting of multiple sites. It belongs to a set of time windows consisting of multiple time windows. This indicates the total number of sites in the site collection. This indicates the total number of time windows in the time window set; and They represent the first time. The date and the The first time granularity The site at the Passenger flow entering and exiting the station within a specific time window.
3. The subway station passenger flow risk assessment method based on commuting attributes according to claim 1, characterized in that, In S2, the entropy weight method combined with the TOPSIS method is used to process various commuting characteristic indicators of each station, specifically including: S21, Construct a commuting feature index matrix based on the various commuting feature indicators of each station; S22, Standardize the commuting feature index matrix to obtain a standardized commuting feature index matrix; S23, calculate the relative entropy probability of each commuting feature index based on the standardized commuting feature index matrix; calculate the information entropy of each commuting feature index based on the relative entropy probability of each commuting feature index; calculate the credit validity value of each commuting feature index based on the information entropy of each commuting feature index; calculate the entropy weight of each commuting feature index based on the credit validity value of each commuting feature index. S24. Construct a commuting feature index weighting matrix based on the entropy weights of each commuting feature index and the standardized commuting feature index matrix; construct the optimal and worst vectors of the commuting feature indexes according to the commuting feature index weighting matrix; calculate the Euclidean distance between each station and the optimal and worst vectors of the commuting feature indexes respectively; calculate the commuting index of each station according to the Euclidean distance between each station and the optimal and worst vectors of the commuting feature indexes respectively.
4. The subway station passenger flow risk assessment method based on commuting attributes according to claim 3, characterized in that, In step S21, the commuting characteristic index matrix is represented as follows: ; in, It belongs to a set of dates consisting of multiple different dates. It belongs to a set of time granularities composed of multiple different time granularities. It belongs to a collection of sites consisting of multiple sites. It belongs to a set of commuting characteristic indicators, which consists of multiple commuting characteristic indicators. This indicates the total number of sites in the site collection. This represents the total number of commuting feature indicators in the commuting feature indicator set; Indicates the first The date and the A matrix of commuting characteristic indicators at different time granularities. Indicates the first The date and the The first time granularity The first site One commuting characteristic indicator; In S22, the standardized commuting characteristic index matrix is represented as follows: ; in, Indicates the first The date and the A standardized commuting characteristic index matrix at a time granularity. Indicates the first The date and the The first time granularity The first site Standardized values of commuting characteristic indicators, and .
5. The subway station passenger flow risk assessment method based on commuting attributes according to claim 4, characterized in that, In S23, The formulas for calculating the relative entropy probability of each commuting characteristic indicator are as follows: ; in, Indicates the first The date and the The first time granularity The first site The relative entropy probability of each commuting characteristic indicator; The formula for calculating the information entropy of each commuting characteristic indicator is as follows: ; in, Indicates the first Information entropy of commuting characteristic indicators; The formulas for calculating the credit validity value of each commuting characteristic indicator are as follows: ; in, Indicates the first Credit validity values of each commuting characteristic indicator; The formula for calculating the entropy weight of each commuting characteristic indicator is as follows: ; in, Indicates the first Entropy weights of commuting characteristic indicators.
6. The subway station passenger flow risk assessment method based on commuting attributes according to claim 5, characterized in that, In S24, The weighted matrix of the commuting feature indicators is represented as follows: ; in, Indicates the first The date and the A weighted matrix of commuting characteristic indicators at various time granularities. Indicates the first The date and the The first time granularity The first site The weighted values of each commuting characteristic indicator, and ; The optimal and worst vectors of commuting characteristic indicators are represented as follows: ; ; ; ; in, and These represent the optimal and worst vectors of commuting characteristic indicators, respectively. The optimal vector representing commuting characteristic indicators is related to the first... The element corresponding to the weighted value of the commuting characteristic index, and used to characterize the The date and the The first time granularity for all stations The maximum value of the weighted average of the commuting characteristic indicators; The worst vector representing commuting characteristic indicators is related to the first... The element corresponding to the weighted value of the commuting characteristic index, and used to characterize the The date and the The first time granularity for all stations The minimum weighted value of each commuting characteristic indicator; The formulas for calculating the Euclidean distance between each station and the optimal and worst vectors of the commuting characteristic indicators are as follows: ; ; in, Indicates the first The date and the The first time granularity The Euclidean distance between each station and the optimal vector of commuting characteristic indicators. Indicates the first The date and the The first time granularity Euclidean distance between each station and the worst vector of commuting characteristic indicators; The formula for calculating the commuting index for each station is as follows: ; in, Indicates the first The date and the The first time granularity Commuting index for each station.
7. The subway station passenger flow risk assessment method based on commuting attributes according to claim 1, characterized in that, Specifically, S5 is: S51, using the Hopkins statistic to test the clustering trend of passenger flow risk scores for each commuter station and each non-commuter station, respectively, and obtain the clustering attributes of each commuter station and each non-commuter station. S52, the silhouette coefficient method is used to analyze the clustering attributes of each commuter station and each non-commuter station respectively, so as to determine the optimal number of clusters for each commuter station and each non-commuter station. S53, risk levels are divided according to the optimal number of clusters for each commuter station and the optimal number of clusters for each non-commuter station, and the passenger flow risk levels for each commuter station and each non-commuter station are obtained accordingly.
8. A subway station passenger flow risk assessment system based on commuting attributes, characterized in that, The method for assessing passenger flow risk at subway stations based on commuting attributes as described in any one of claims 1 to 7 includes: The passenger flow statistics module is used to calculate the inbound and outbound passenger flow of each station within each time window based on different dates and time granularities from the subway card swiping data, and obtain the inbound passenger flow dataset and the outbound passenger flow dataset. The station commuting attribute judgment module is used to calculate multiple commuting characteristic indicators for each station based on the inbound passenger flow dataset and the outbound passenger flow dataset, and to perform objective weighting and multi-attribute decision processing on each commuting characteristic indicator of each station using the entropy weight method combined with the TOPSIS method to obtain the commuting index of each station, and to judge the commuting attribute of each station in combination with the surrounding land use attributes; wherein, the commuting attribute includes commuting category and non-commuting category; The risk indicator construction module is used to construct a commuter passenger flow risk assessment indicator system for commuter stations and a non-commuter passenger flow risk assessment indicator system for non-commuter stations. Based on the commuting attributes of each station, the module constructs a commuter passenger flow risk assessment matrix and a non-commuter passenger flow risk assessment matrix using the commuter passenger flow risk assessment indicator system and the non-commuter passenger flow risk assessment indicator system, respectively. The risk score calculation module is used to objectively assign weights and perform multi-attribute decision processing on the commuter passenger flow risk assessment matrix and the non-commuter passenger flow risk assessment matrix using the entropy weight method combined with the TOPSIS method, respectively, to obtain the passenger flow risk score of each commuter station and the passenger flow risk score of each non-commuter station. The risk level classification module uses k-means clustering analysis to perform cluster analysis on the passenger flow risk scores of each commuter station and each non-commuter station, respectively, to obtain the passenger flow risk level of each commuter station and each non-commuter station.
Citation Information
Patent Citations
Traffic hub passage risk index discrimination method and system
CN113988721A