A method for identifying the causes of low passenger flow at intercity railway stations based on classification and multiple indicators
By establishing a passenger flow potential indicator system for intercity railway stations and adopting the natural break point method, CRITIC weighting method and k-means clustering method, the reasons for low passenger flow at intercity railway stations are identified, the problem of intercity railway passenger flow polarization is solved, and a quantitative analysis of passenger flow improvement strategies is achieved.
Patent Information
- Application Number
- CN202410262890.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-07
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-03-07
AI Technical Summary
Intercity railway passenger flow polarization is serious, with some lines facing low passenger flow and severe operating losses. Existing research lacks quantitative analysis of the reasons for low passenger flow at intercity railway stations, making it difficult to formulate effective strategies to increase passenger flow.
Based on intercity railway operation data and mobile phone signaling data, a station demand and line competitiveness indicator system was established. The natural break point method was used to screen low passenger flow stations. The CRITIC weighting method and k-means clustering method were combined to identify the reasons for low passenger flow at stations.
It provides a quantitative analysis of the reasons for low passenger flow, helps formulate targeted strategies to increase passenger flow, and improves the operational efficiency of intercity railways.
Smart Images

Figure CN118037128B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of public transportation facility planning, and in particular relates to a method for identifying causes of low passenger flow at intercity railway stations based on classification and multiple indicators. Background Art
[0002] With the rapid advancement of urbanization in my country, connections between cities within urban agglomerations are continuously strengthening. The operation of numerous intercity railways has further accelerated the flow of technology, information, talent, and other factors between cities within these agglomerations. Intercity railways are fast, convenient, and high-density passenger-dedicated lines specifically serving adjacent cities or urban agglomerations, with passenger trains designed for speeds of 200 km / h or less. While providing efficient intercity transportation, intercity railways also face fierce competition from other rail transit modes (high-speed rail, intercity subways) and road transportation (buses, cars, etc.) along the same corridor. Passenger flow on currently operating lines shows significant polarization of intercity rail traffic. While a few lines enjoy good passenger traffic performance, many lines face low passenger volume and significant operating losses. Strategies to increase passenger flow from planning and operational perspectives are urgently needed to boost intercity rail passenger volume.
[0003] Existing research on rail transit passenger flow has focused more on urban rail transit passenger flow patterns and less on intercity rail. The passenger flow characteristics of the two differ significantly. Intercity rail generally serves intercity passengers. Compared to urban rail transit systems, intercity rail travels longer distances, has larger station spacing, and serves a wider range of stations. Passenger volume is more closely related to station connections and the service quality of the line itself. Improving intercity rail passenger flow requires consideration at both the station and line levels. First, the causes of low passenger flow must be analyzed so that specific strategies can be implemented for improving passenger flow at different stations. Existing research has provided limited quantitative analysis of the causes of low passenger flow at intercity rail stations, and research on intercity rail passenger flow patterns is relatively weak. However, with the recent increase in the number of operating intercity rail lines and the resulting wealth of operational data, further in-depth research on intercity rail passenger flow is possible. Summary of the Invention
[0004] To address the above issues, the present invention proposes a method for identifying the causes of low passenger flow at intercity railway stations based on classification and multiple indicators. A station passenger flow potential indicator system is established from the perspectives of station demand and line competitiveness. A cluster analysis method is used to classify and identify the causes of low passenger flow at different stations, which can pave the way for the next step of research on strategies to increase passenger flow at intercity railways.
[0005] The technical solutions of the present invention are as follows:
[0006] A method for identifying the causes of low passenger flow at intercity railway stations based on classification and multiple indicators includes the following steps:
[0007] S1: Based on existing intercity railway operation data and mobile phone signaling data, we collected the average daily entry and exit data of each station, calculated the daily travel volume around each station, and used the natural break point method to screen low passenger flow stations;
[0008] S2: For stations with a certain travel volume but low passenger flow around the station, an intercity railway station passenger flow potential index model is established from the perspectives of station demand and line competitiveness. The index model is weighted using the CRITIC weighting method, and the comprehensive evaluation index of station demand and line competitiveness is calculated for each station.
[0009] S3: Use the k-means clustering method to classify all stations into four categories, and determine the reasons for low passenger flow at each station based on the classification results.
[0010] Preferably, step S1 is specifically as follows:
[0011] S11: For a set of intercity railway stations N, obtain the daily average entry and exit data for each station, and calculate the daily travel volume around each station based on mobile phone signaling data. The station perimeter is a 3000-meter radius around the station.
[0012] S12: Using the natural break point method, the stations are divided into four categories based on the average daily entry and exit volume: high, medium-high, medium-low, and low, corresponding to the station sets. and The stations are divided into four categories based on the daily travel volume around the stations: high, medium-high, medium-low and low. and The details are as follows:
[0013] Step 1: Determine the number of categories k1 = 4;
[0014] Step 2: Arrange the daily average entry and exit volume of the station or the daily travel volume around the station from small to large, select k1-1 discontinuity points to divide the data into k1 categories, with a total of classification schemes, where |N| is the number of elements in the intercity railway station set N, that is, the total number of intercity railway stations;
[0015] Step 3: Traverse each classification scheme and calculate the sum of the total squared differences of its classes SDCM, that is
[0016]
[0017] In the formula, k1 is the number of groups, m is i is the number of elements in the i-th group, v ij is the jth element in the i-th group, is the average value of each element in the i-th group;
[0018] Step 4: Select the classification scheme with the smallest SDCM value, which is the best classification scheme obtained by the natural break point method;
[0019] Step 5: Based on the classification results of the two sets of data, the average daily station entry and exit volume and the daily travel volume around the station, select two types of stations: low demand, low passenger flow and high and medium demand, low passenger flow; among them, the low demand, low passenger flow station set is The collection of high-demand and low-passenger-flow stations is
[0020] Preferably, step S2 is specifically as follows:
[0021] S21: For the set of stations N2 with high-demand and low-passenger-flow selected in step S12, define an intercity railway station passenger flow potential index model; the index model includes a station demand index C1 and a line competitiveness index C2;
[0022] The station demand index C1={x 11 ,x 12 ,x 13 ,x 14}, including the number of trips generated by cities along the line around the station x 11 , the number of trips generated along the intercity line corridor around the station x 12 、Total number of employed people covered by public transportation within 45 minutes x 13 Total employment of the covered population within 30 minutes of driving 14 ;
[0023] The line competitiveness index C2 = {x 21 ,x 22 ,x 23 ,x 24}, including intercity railway design speed x 21 、Number of trains arriving and departing from the station throughout the day x 22 , fare comparison coefficient x 23 、Number of connected subway lines x 24 ;
[0024] S22: Use the CRITIC weighting method to assign weights to the indicator model. The specific steps are as follows:
[0025] Step 1: Perform dimensionless processing on the values of each indicator, and use the maximum / minimum value normalization method. For the indicator x (x∈C1∪C2), if it is a positive indicator, the positive indicator includes x 11 、x 12 、x 13 、x 14 、x 21 、x 22 、x24 , then the index value x(i) of the i-th sample is normalized to:
[0026]
[0027] In the formula, x(i)′ is the normalized index, max(x) and min(x) represent the maximum and minimum values of the index x in each sample respectively; if the index x is a negative index, the negative index is x 23 , then the index value x(i) of the i-th sample is normalized to:
[0028]
[0029] Step 2: Calculate the coefficient of variation of each indicator; the coefficient of variation is defined as the ratio of the standard deviation of each sample value to the mean. For indicator x (x∈C1∪C2), the coefficient of variation calculation formula is:
[0030]
[0031] Where, is the coefficient of variation of indicator x, σ x is the standard deviation of each sample value of indicator x, is the average value of each sample value of indicator x;
[0032] Step 3: Calculate the independence coefficient of each indicator in the station demand index C1 and the line competitiveness index C2 respectively; for the indicator x,y∈{x 11 ,x 12 ,x 13 ,x 14} or x,y∈{x 21 ,x 22 ,x 23 ,x 24} and x and y are different indicators, the correlation coefficient between them is:
[0033]
[0034] Where r xy is the correlation coefficient between indicator x and indicator y, n is the number of station samples, ∑xy, ∑x, ∑y, ∑x 2 ,∑y 2 They represent the sum of the products of the sample values of indicators x and y, the sum of the sample values of indicator x, the sum of the sample values of indicator y, the sum of the squares of the sample values of indicator x, and the sum of the squares of the sample values of indicator y respectively;
[0035] The indicator x(x∈{x 11 ,x 12 ,x 13 ,x14})'s independence coefficient for
[0036]
[0037] The index x(x∈{x 21 ,x 22 ,x 23 ,x 24})'s independence coefficient for
[0038]
[0039] Step 4: Calculate the comprehensive evaluation index c1 of station demand and the comprehensive evaluation index c2 of line competitiveness;
[0040] Calculate the objective weight coefficient W of indicator x x for:
[0041]
[0042] The comprehensive evaluation index c1 of station demand is calculated as:
[0043] c1=W 11 x 11 +W 12 x 12 +W 13 x 13 +W 14 x 14
[0044] Where W 11 、W 12 、W 13 、W 14 They are respectively indicators x 11 、x 12 、x 13 and x 14 The objective weight coefficient of
[0045] The comprehensive evaluation index c2 of route competitiveness is calculated as:
[0046] c2=W 21 x 21 +W 22 x 22 +W 23 x 23 +W 24 x 24
[0047] Where W 21 、W 22 、W 23 、W24 They are respectively indicators x 21 、x 22 、x 23 、x 24 The objective weight coefficient.
[0048] Preferably, in step S3, the k-means clustering method is applied to divide the elements in the set N2 of high-demand and low-passenger-flow stations into k2=4 categories according to the values of the comprehensive evaluation index c1 of station demand and the comprehensive evaluation index c2 of line competitiveness, and the reasons for the low passenger flow of each station are determined according to the classification results. The specific steps are as follows:
[0049] S31: Construct indicator feature set X:
[0050]
[0051] Where n is the number of elements in set N2, c 1i and c 2i are the values of c1 and c2 indicators of the i-th station, 1≤i≤n, x i Indicates that the i-th station is located at c 1i and c 2i The vector composed of x i =(c 1i ,c 2i );
[0052] S32: Randomly select k2 station samples from X as initial cluster centers and place them into the cluster center set C;
[0053] S33: Calculate the shortest Euclidean distance between the remaining stations and each cluster center, and classify the remaining stations:
[0054]
[0055] Where, D(x i ) is the sample x i The shortest Euclidean distance between the k2 cluster centers, x k is the cluster center and k=1,2,…,k2; establish the corresponding classification set for k2 cluster centers The sample x i Divide into the category with the shortest distance to its cluster center, that is:
[0056]
[0057] S34: For classification sets Recalculate its cluster center x k , update the cluster center set C:
[0058]
[0059] S35: Return to step S33, for sample x i Reclassify and obtain each classification set Determine whether the elements in each classification set have changed after reclassification. If so, proceed to step S34 and enter the next clustering iteration; if not, terminate the iteration and obtain the final clustering result.
[0060] S36: Determine the reasons for low passenger flow at the station based on the four classification results.
[0061] Preferably, step S36 is specifically as follows:
[0062] ① For stations with high station demand indicators and high line competitiveness indicators, the main reason for low passenger flow is insufficient competitiveness compared with adjacent high-passenger flow stations on the same line, or the rail travel habits of residents around the station need to be further cultivated;
[0063] ② For stations with high station demand indicators and low line competitiveness indicators, the main reason for low passenger flow is the low service level of intercity railways;
[0064] ③ For stations with low station demand indicators and high line competitiveness indicators, the main reason for low passenger flow is low demand along the intercity railway direction or low level of transportation connection services, resulting in low total travel demand attracted by the station;
[0065] ④ For stations with low station demand indicators and line competitiveness indicators, the reason for the low passenger flow is that the travel demand along the intercity railway around the station and the intercity railway service level are both low.
[0066] The present invention proposes a method for identifying the causes of low passenger flow at intercity railway stations based on classification and multiple indicators. Based on existing intercity railway operation data and mobile phone signaling data reflecting the size of travel demand around the stations, a passenger flow potential indicator system for intercity railway stations is established. Natural break point method, CRITIC weighting method, k-means clustering and other methods are used to classify and identify the causes of low passenger flow at intercity railway stations.
[0067] Compared with existing technologies, this invention offers the advantage of proposing a method for determining and identifying the causes of intercity railway stations experiencing low passenger flow. This method fills a gap in existing research and provides a quantitative analytical basis for research into increasing passenger flow at intercity railway stations. Based on this invention, targeted passenger flow improvement strategies can be implemented for intercity railway stations experiencing low passenger flow to improve the passenger flow efficiency of intercity railway operations. Alternatively, further research can be conducted based on this foundation, further refining the causes of low passenger flow through enhanced comparative analysis of various indicators.
[0068] This method is applicable to urban agglomerations with multiple intercity railway lines in operation, ensuring a sufficient number of stations to serve as samples for the classification study. This method is portable and allows for the selection of appropriate indicator systems and indicator weighting methods based on the characteristics of the study area and data availability. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is a scatter plot of station classification.
[0070] Figure 2 This is a schematic diagram of the station clustering results. DETAILED DESCRIPTION
[0071] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0072] In order to address the defects and shortcomings of the above-mentioned technologies, the present invention proposes a method for identifying the causes of low passenger flow at intercity railway stations based on classification and multiple indicators. Based on the existing operation data and mobile phone signaling data of intercity railways, the method first collects the full-day entry and exit data of each station, calculates the full-day travel generation around the station, and uses the natural break point method to screen low passenger flow stations; then, for stations with a certain travel scale but low passenger flow around the station, an intercity railway station passenger flow potential index system is established from the two perspectives of station demand and line competitiveness, the CRITIC weighting method is used to assign weights to the index system, and the comprehensive evaluation index of station demand and the comprehensive evaluation index of line competitiveness of each station are calculated; finally, the k-means clustering method is used to divide all stations into four categories, and the reasons for low passenger flow at each station are identified according to the classification results.
[0073] Step 1: Screen intercity railway stations with low passenger flow
[0074] For the set N of intercity railway stations, the average daily entry and exit volume data of all stations is obtained, and the daily travel volume around each station is calculated based on the mobile phone signaling data (the present invention defines the area around the station as a coverage area with a radius of 3,000 meters, the same below).
[0075] Using the natural break point method, the stations are divided into four categories according to the daily entry and exit volume: high, medium-high, medium-low and low, corresponding to the station set and The stations are divided into four categories based on the daily travel volume within 3000 meters: high, medium-high, medium-low and low, corresponding to the station collections. and
[0076] The natural breaks method is a classification method for discrete data sets. Its principle is to divide the research objects into several groups and determine the best classification by iteratively comparing the sum of the squared differences between the mean and the observed values of the elements in each group. The specific steps used in the present invention are:
[0077] Step 1: Determine the number of categories k1 = 4;
[0078] Step 2: Arrange the data set (daily station entry and exit volume or daily travel volume around the station) from small to large, select k1-1 discontinuity points to divide the data into k1 categories, with a total of classification schemes (|N| is the number of elements in set N, i.e., the total number of intercity railway stations);
[0079] Step 3: Traverse each classification scheme and calculate the sum of the total square difference SDCM (sum of deviation from the class mean), that is,
[0080]
[0081] In the formula, k1 is the number of groups, m is i is the number of elements in the i-th group, v ij is the jth element in the i-th group, is the average value of each element in the i-th group;
[0082] Step 4: Select the classification scheme with the smallest SDCM value, which is the best classification scheme obtained by the natural break point method.
[0083] Based on the classification results of the two sets of data, daily station entry and exit volume and daily travel volume, we screened out two types of stations: low demand, low passenger flow, and high and medium demand, low passenger flow. Among them, the set of low demand, low passenger flow stations is The main reason for the low passenger flow of this type of station is that the travel demand around the station is low and the population employment needs to be further concentrated; the collection of stations with high demand and low passenger flow is There is already a high demand for travel around this type of station, but the passenger flow at the station is relatively low, and the reasons for the low passenger flow need further analysis.
[0084] Step 2: Establish an indicator system for passenger flow potential of intercity railway stations
[0085] For the set of stations N2 with high- to medium-demand and low-passenger-flow identified in step 1, a passenger flow potential index system for intercity railway stations is defined. The index system is divided into two categories: station demand index C1 and line competitiveness index C2.
[0086] The station demand index C1 is an indicator that reflects the travel demand structure and transportation connection coverage. For intercity railway stations with high surrounding travel volume but low passenger flow, the reasons for the low passenger flow may be that the line service direction does not match the main travel direction of the station surrounding area, the travel demand in the line corridor is low, and the transportation connection efficiency is low. Station demand index C1 = {x 11 ,x 12,x 13 ,x 14}, including the number of trips generated by cities along the line around the station x 11 , the number of trips generated along the intercity line corridor around the station x 12 、Total number of employed people covered by public transportation within 45 minutes x 13 Total employment of the covered population within 30 minutes of driving 14 .
[0087] The line competitiveness index C2 reflects the competitiveness of the intercity railway line to which the station belongs. The competitiveness of the intercity railway itself with high-speed rail, subway, bus, car and other modes affects the travel choices of residents along the line, and thus affects the passenger flow of the station. Line competitiveness index C2 = {x 21 ,x 22 ,x 23 ,x 24}, including intercity railway design speed x 21 、Number of trains arriving and departing from the station throughout the day x 22 , fare comparison coefficient x 23 、Number of connected subway lines x 24 .
[0088] The definitions of each variable are shown in Table 1.
[0089] Table 1 Selection of passenger flow potential indicators for intercity railway stations
[0090]
[0091] Step 3: Empower the passenger flow potential indicator system
[0092] To comprehensively reflect the demand around a station and the competitiveness of its associated lines, it is necessary to combine and weight the indicators in each of the C1 and C2 categories. This paper uses the CRITIC weighting method (CRiteria Importance Through Intercriteria Correlation) to calculate the weights of each indicator, thereby obtaining a comprehensive evaluation value for the station demand index and the line competitiveness index. The specific steps are:
[0093] Step 1: Perform dimensionless processing on the values of each indicator. Using the maximum / minimum value normalization method, for the indicator x (x∈C1∪C2), if it is a positive indicator (that is, the larger the indicator value, the better the station passenger flow efficiency, positive indicators include x 11 、x 12 、x 13 、x 14 、x 21 、x 22 、x 24), then the index value x(i) of the i-th sample is normalized to:
[0094]
[0095] Where x(i)′ is the normalized index, max(x) and min(x) represent the maximum and minimum values of the index x in each sample, respectively.
[0096] If the indicator x is a negative indicator (that is, the smaller the indicator value, the better the station passenger flow efficiency, x 23 is a negative indicator), then the index value x(i) of the i-th sample is normalized as follows:
[0097]
[0098] Step 2: Calculate the coefficient of variation of each indicator. The coefficient of variation is defined as the ratio of the standard deviation of each sample value to the mean. For indicator x (x∈C1∪C2), the coefficient of variation calculation formula is:
[0099]
[0100] Where, is the coefficient of variation of indicator x, σ x is the standard deviation of each sample value of indicator x, is the average value of each sample of indicator x.
[0101] The coefficient of variation reflects the size of the difference in values between samples of the indicator. The larger the coefficient of variation, the greater the amount of information the indicator can reflect, and the greater the weight of the indicator.
[0102] Step 3: Calculate the independence coefficient of each indicator in the C1 and C2 classifications respectively. 11 ,x 12 ,x 13 ,x 14} or x,y∈{x 21 ,x 22 ,x 23 ,x 24} and x and y are different indicators, the correlation coefficient between them is:
[0103]
[0104] Where r xy is the correlation coefficient between indicator x and indicator y, n is the number of station samples, ∑xy, ∑x, ∑y, ∑x 2 ,∑y 2They respectively represent the sum of the products of the sample values of indicators x and y, the sum of the sample values of indicator x, the sum of the sample values of indicator y, the sum of the squares of the sample values of indicator x, and the sum of the squares of the sample values of indicator y.
[0105] The indicator x(x∈{x 11 ,x 12 ,x 13 ,x 14})'s independence coefficient for
[0106]
[0107] The indicator x(x∈{x 21 ,x 22 ,x 23 ,x 24})'s independence coefficient for
[0108]
[0109] Step 4: Calculate the comprehensive evaluation index of station demand index and line competitiveness index. The objective weight coefficient W of index x x for:
[0110]
[0111] The comprehensive evaluation index c1 of station demand is:
[0112] c1=W 11 x 11 +W 12 x 12 +W 13 x 13 +W 14 x 14
[0113] Where W 11 、W 12 、W 13 、W 14 They are respectively indicators x 11 、x 12 、x 13 and x 14 The objective weight coefficient
[0114] The comprehensive evaluation index c2 of line competitiveness is:
[0115] c2=W 21 x 21 +W 22 x 22 +W 23 x 13+W 24 x 24
[0116] Where W 21 、W 22 、W 23 、W 24 They are respectively indicators x 21 、x 22 、x 23 、x 24 The objective weight coefficient.
[0117] Step 4: Cluster analysis to identify the reasons for low passenger flow at the station
[0118] The k-means clustering method is used to divide the elements in the set N2 of high-demand and low-passenger-flow stations into k2 = 4 categories based on the values of the comprehensive evaluation index c1 of station demand and the comprehensive evaluation index c2 of line competitiveness, corresponding to the four reasons for low passenger flow. The specific steps are as follows:
[0119] Step 1: Construct indicator feature set X:
[0120]
[0121] Where n is the number of elements in set N2, c 1i and c 2i are the values of c1 and c2 indicators of the i-th station, 1≤i≤n, x i Indicates that the i-th station is located at c 1i and c 2i The vector composed of x i =(c 1i ,c 2i )
[0122] Step 2: Randomly select k2 station samples from X as initial cluster centers and place them in the cluster center set C.
[0123] Step 3: Calculate the shortest Euclidean distance between the remaining stations and each cluster center, and classify the remaining stations:
[0124]
[0125] Where, D(x i ) is the sample x i The shortest Euclidean distance between the k2 cluster centers, x k is the cluster center and k=1,2,…,k2.
[0126] Establish corresponding classification sets for k2 cluster centers The sample x i Divide into the category with the shortest distance to its cluster center, that is:
[0127]
[0128] Step 4: For the classification set Recalculate its cluster center x k , update the cluster center set C:
[0129]
[0130] Step5: Return to step 3 and perform the calculation on the sample x. i Reclassify and obtain each classification set Determine whether the elements in each classification set have changed after reclassification. If so, continue to Step 4 and enter the next clustering iteration; if not, terminate the iteration and obtain the final clustering result.
[0131] Step 6: Determine the reasons for low passenger flow at the station based on the four classification results:
[0132] ① For stations with high station demand and line competitiveness indicators, there is already a certain base of intercity travel demand in the vicinity of the station, and the intercity railway's own service level is competitive. The main reason for low passenger flow is insufficient competitiveness compared with adjacent high-passenger flow stations on the same line, or the rail travel habits of residents around the station need to be further cultivated. Such stations have high potential for increasing passenger flow;
[0133] ② For stations with high station demand indicators and low line competitiveness indicators, there is already a certain travel demand base in the vicinity of the station, but the intercity railway itself lacks competitiveness. The main reason for the low passenger flow is the low service level of the intercity railway, and residents around the station tend to choose other intercity modes of travel;
[0134] ③ For stations with low station demand indicators and high line competitiveness indicators, although there is a certain scale of travel volume around the station, the demand along the intercity railway direction is low, or the level of transportation connection service is low, the total travel demand attracted by the station is small, resulting in low passenger flow;
[0135] ④ For stations with low station demand indicators and line competitiveness indicators, the reason for the low passenger flow is that the travel demand along the intercity railway around the station and the intercity railway service level are both low. Specific embodiments
[0137] A case study was conducted on a city cluster with 6 intercity railways in operation. There are a total of 82 intercity railway stations in operation on each line (numbered 1 to 82). The daily passenger flow data of each station was collected, and the passenger travel demand within 3,000 meters of the station was calculated based on the mobile phone signaling data. The natural break point method was used to divide each station into four levels: high, medium-high, medium-low, and low according to the daily entry and exit volume and the surrounding full-day travel volume. The station classification results are shown in Table 2. 31 stations with high and medium demand and low passenger flow, 33 stations with low demand and low passenger flow, and 17 other stations were classified. The station classification scatter plot is shown in Table 2. Figure 1 shown.
[0138] Table 2 Station classification results
[0139]
[0140]
[0141] For the 36 stations with high to medium demand and low passenger flow, the reasons for their low passenger flow were further identified. The station demand index and line competitiveness index values of each station were calculated and obtained, and normalized. The CRITIC weighting method was used to assign weights to the indicators. The station demand comprehensive evaluation index and line competitiveness comprehensive evaluation index values of each station were calculated. The k-means clustering method was used to classify the stations. The classification results and the reasons for low passenger flow are shown in Table 3. The cluster scatter plot is shown in Figure 2 shown.
[0142] Table 3 Station clustering results
[0143]
[0144] The present invention proposes a method for identifying the causes of low passenger flow at intercity railway stations based on classification and multiple indicators. Based on existing intercity railway operation data and mobile phone signaling data reflecting the size of travel demand around the stations, a passenger flow potential indicator system for intercity railway stations is established. Natural break point method, CRITIC weighting method, k-means clustering and other methods are used to classify and identify the causes of low passenger flow at intercity railway stations.
[0145] Compared with existing technologies, this invention offers the advantage of proposing a method for determining and identifying the causes of intercity railway stations experiencing low passenger flow. This method fills a gap in existing research and provides a quantitative analytical basis for research into increasing passenger flow at intercity railway stations. Based on this invention, targeted passenger flow improvement strategies can be implemented for intercity railway stations experiencing low passenger flow to improve the passenger flow efficiency of intercity railway operations. Alternatively, further research can be conducted based on this foundation, further refining the causes of low passenger flow through enhanced comparative analysis of various indicators.
[0146] This method is applicable to urban agglomerations with multiple intercity railway lines, ensuring a sufficient number of stations to serve as samples for the classification study. The method is portable and can be used to select appropriate indicator systems and indicator weighting methods based on the characteristics of the study area and data availability.
[0147] For those skilled in the art, several modifications and improvements can be made to the embodiments of the present invention without departing from the inventive concept of the present application, and all of these modifications and improvements fall within the scope of protection of the present application.
Claims
1. A method for identifying the causes of low passenger flow at intercity railway stations based on classification and multiple indicators, characterized in that: The steps include: S1: Based on existing intercity railway operation data and mobile phone signaling data, we collected the average daily entry and exit data of each station, calculated the daily travel volume around each station, and used the natural break point method to screen low passenger flow stations; Step S1 is specifically as follows: S11: For a set of intercity railway stations N, obtain the daily average entry and exit data for each station, and calculate the daily travel volume around each station based on mobile phone signaling data. The station perimeter is a 3000-meter radius around the station. S12: Using the natural break point method, the stations are divided into four categories based on the average daily entry and exit volume: high, medium-high, medium-low, and low, corresponding to the station sets. and The stations are divided into four categories based on the daily travel volume around the stations: high, medium-high, medium-low and low. and The details are as follows: Step 1: Determine the number of categories k1 = 4; Step 2: Arrange the daily average entry and exit volume of the station or the daily travel volume around the station from small to large, select k1-1 discontinuity points to divide the data into k1 categories, with a total of classification schemes, where |N| is the number of elements in the intercity railway station set N, that is, the total number of intercity railway stations; Step 3: Traverse each classification scheme and calculate the sum of the total squared differences of its classes SDCM, that is In the formula, k1 is the number of groups, m is i is the number of elements in the i-th group, v ij is the jth element in the i-th group, is the average value of each element in the i-th group; Step 4: Select the classification scheme with the smallest SDCM value, which is the best classification scheme obtained by the natural break point method; Step 5: Based on the classification results of the two sets of data, the average daily station entry and exit volume and the daily travel volume around the station, select two types of stations: low demand, low passenger flow and high and medium demand, low passenger flow; among them, the low demand, low passenger flow station set is The collection of high-demand and low-passenger-flow stations is S2: For stations with a certain travel volume but low passenger flow around the station, an intercity railway station passenger flow potential index model is established from the perspectives of station demand and line competitiveness. The index model is weighted using the CRITIC weighting method, and the comprehensive evaluation index of station demand and line competitiveness is calculated for each station. S3: Use the k-means clustering method to classify all stations into four categories, and determine the reasons for low passenger flow at each station based on the classification results.
2. The method for identifying the causes of low passenger flow at intercity railway stations based on classification and multiple indicators according to claim 1 is characterized in that: Step S2 is specifically as follows: S21: For the set of stations N2 with high-demand and low-passenger-flow selected in step S12, define an intercity railway station passenger flow potential index model; the index model includes a station demand index C1 and a line competitiveness index C2; The station demand index C1={x 11 ,x 12 ,x 13 ,x 14 }, including the number of trips generated by cities along the line around the station x 11 , the number of trips generated along the intercity line corridor around the station x 12 、Total number of employed people covered by public transportation within 45 minutes x 13 Total employment of the covered population within 30 minutes of driving 14 ; The line competitiveness index C2 = {x 21 ,x 22 ,x 23 ,x 24 }, including intercity railway design speed x 21 、Number of trains arriving and departing from the station throughout the day x 22 , fare comparison coefficient x 23 、Number of connected subway lines x 24 ; S22: Use the CRITIC weighting method to assign weights to the indicator model. The specific steps are as follows: Step 1: Perform dimensionless processing on the values of each indicator, and use the maximum / minimum value normalization method. For the indicator x (x∈C1∪C2), if it is a positive indicator, the positive indicator includes x 11 、x 12 、x 13 、x 14 、x 21 、x 22 、x 24 , then the index value x(i) of the i-th sample is normalized to: In the formula, x(i)′ is the normalized index, max(x) and min(x) represent the maximum and minimum values of the index x in each sample respectively; If the indicator x is a negative indicator, the negative indicator is x 23 , then the index value x(i) of the i-th sample is normalized to: Step 2: Calculate the coefficient of variation of each indicator; the coefficient of variation is defined as the ratio of the standard deviation of each sample value to the mean. For indicator x (x∈C1∪C2), the coefficient of variation calculation formula is: Where, is the coefficient of variation of indicator x, σ x is the standard deviation of each sample value of indicator x, is the average value of each sample value of indicator x; Step 3: Calculate the independence coefficient of each indicator in the station demand index C1 and the line competitiveness index C2 respectively; for the indicator x,y∈{x 11 ,x 12 ,x 13 ,x 14 } or x,y∈{x 21 ,x 22 ,x 23 ,x 24 } and x and y are different indicators, the correlation coefficient between them is: Where r xy is the correlation coefficient between indicator x and indicator y, n is the number of station samples, ∑xy, ∑x, ∑y, ∑x 2 ,∑y 2 They represent the sum of the products of the sample values of indicators x and y, the sum of the sample values of indicator x, the sum of the sample values of indicator y, the sum of the squares of the sample values of indicator x, and the sum of the squares of the sample values of indicator y respectively; The indicator x(x∈{x 11 ,x 12 ,x 13 ,x 14 })'s independence coefficient for The index x(x∈{x 21 ,x 22 ,x 23 ,x 24 })'s independence coefficient for Step 4: Calculate the comprehensive evaluation index c1 of station demand and the comprehensive evaluation index c2 of line competitiveness; Calculate the objective weight coefficient W of indicator x x for: The comprehensive evaluation index c1 of station demand is calculated as: c1=W 11 x 11 +W 12 x 12 +W 13 x 13 +W 14 x 14 Where W 11 、W 12 、W 13 、W 14 They are respectively indicators x 11 、x 12 、x 13 and x 14 The objective weight coefficient of The comprehensive evaluation index c2 of route competitiveness is calculated as: c2=W 21 x 21 +W 22 x 22 +W 23 x 23 +W 24 x 24 Where W 21 、W 22 、W 23 、W 24 They are respectively indicators x 21 、x 22 、x 23 、x 24 The objective weight coefficient.
3. The method for identifying the causes of low passenger flow at intercity railway stations based on classification and multiple indicators according to claim 1 is characterized in that: In step S3, the k-means clustering method is applied to divide the elements in the set N2 of high-demand and low-passenger-flow stations into k2=4 categories based on the values of the comprehensive evaluation index c1 of station demand and the comprehensive evaluation index c2 of line competitiveness. The reasons for the low passenger flow of each station are determined based on the classification results. The specific steps are as follows: S31: Construct indicator feature set X: Where n is the number of elements in set N2, c 1i and c 2i are the values of c1 and c2 indicators of the i-th station, 1≤i≤n, x i Indicates that the i-th station is located at c 1i and c 2i The vector composed of x i =(c 1i ,c 2i ); S32: Randomly select k2 station samples from X as initial cluster centers and place them into the cluster center set C; S33: Calculate the shortest Euclidean distance between the remaining stations and each cluster center, and classify the remaining stations: Where, D(x i ) is the sample x i The shortest Euclidean distance between the k2 cluster centers, x k is the cluster center and k=1,2,…,k2; Establish corresponding classification sets for k2 cluster centers The sample x i Divide into the category with the shortest distance to its cluster center, that is: S34: For classification sets Recalculate its cluster center x k , update the cluster center set C: S35: Return to step S33, for sample x i Reclassify and obtain each classification set Determine whether the elements in each classification set have changed after reclassification compared to before classification. If so, proceed to step S34 and enter the next clustering iteration; if not, terminate the iteration and obtain the final clustering result; S36: Determine the reasons for low passenger flow at the station based on the four classification results.
4. The method for identifying the causes of low passenger flow at intercity railway stations based on classification and multiple indicators according to claim 3 is characterized in that: Step S36 is specifically as follows: ① For stations with high station demand indicators and high line competitiveness indicators, the main reason for low passenger flow is insufficient competitiveness compared with adjacent high-passenger flow stations on the same line, or the rail travel habits of residents around the station need to be further cultivated; ② For stations with high station demand indicators and low line competitiveness indicators, the main reason for low passenger flow is the low service level of intercity railways; ③ For stations with low station demand indicators and high line competitiveness indicators, the main reason for low passenger flow is low demand along the intercity railway direction or low level of transportation connection services, resulting in low total travel demand attracted by the station; ④ For stations with low station demand indicators and line competitiveness indicators, the reason for the low passenger flow is that the travel demand along the intercity railway around the station and the intercity railway service level are both low.
Citation Information
Patent Citations
Urban rail transit transport capacity and passenger flow adaptability evaluation method, system and equipment
CN117236790A