Urban rail transit station classification method and system based on multi-feature fusion

By integrating multiple features with an improved k-means algorithm, combined with the time-varying passenger flow and static topological features of rail transit stations, the limitations of existing station classification methods are overcome, achieving more accurate station classification and improving operational efficiency.

CN120687900APending Publication Date: 2025-09-23SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510782427.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The existing urban rail transit station classification method relies on a single indicator, lacks theoretical basis, is difficult to adapt to the dynamic expansion of the line network, and cannot fully characterize the differentiated effects of built environment factors. In addition, the existing classification method based on clustering model is easily affected by noise and has limited classification accuracy.

Method used

A multi-feature fusion method is adopted, combined with the time-varying passenger flow characteristics and static topological characteristics of rail transit stations, feature dimensionality reduction is performed through principal component analysis and station clustering is performed using the k-means algorithm. The initial cluster center selection method is improved and the initial cluster center is obtained through multiple iterations.

Benefits of technology

It has achieved accurate classification of stations, improved the operational efficiency of urban rail transit, and promoted differentiated management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687900A_ABST
    Figure CN120687900A_ABST
Patent Text Reader

Abstract

The invention discloses an urban rail transit station classification method and system based on multi-feature fusion. The method comprises the following steps: constructing a network topological relation; constructing a passenger flow characteristic index system of the station, and calculating passenger flow characteristic parameters; calculating static characteristic parameters of the site, and standardizing the static characteristic parameters; according to the passenger flow characteristic parameters and the static characteristic parameters, carrying out fusion and dimension reduction processing on the station characteristic parameters to obtain a fused station characteristic data set; determining an initial clustering center according to the fused site feature data set, and carrying out site clustering based on a k-means algorithm; calculating the sum of squares of errors of clustering categories corresponding to different k values and centers of the clustering categories, and determining an optimal clustering number; and according to a cluster division result corresponding to the optimal clustering number, outputting categories corresponding to all stations. According to the method, the time-varying passenger flow characteristics and the static topology characteristics of the rail transit stations are comprehensively considered, the stations are classified through fusion of the two types of characteristic parameters, and accurate classification of the stations is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of urban rail transit and relates to rail transit station classification technology, and specifically to an urban rail transit station classification method and system based on multi-feature fusion. Background Art

[0002] Urban rail transit stations are key nodes in urban transportation networks, and their classification plays a crucial role in improving transportation operations and management. Different types of stations, such as transfer stations, terminal stations, and business district stations, carry varying passenger flow characteristics and transportation demands. By rationally classifying stations and accurately identifying the passenger flow characteristics of each type, targeted rail transit station operations and management strategies can be developed, effectively reducing operating costs and improving overall operational efficiency.

[0003] Existing urban rail transit station classification methods mainly rely on static indicators such as station location, surrounding land use characteristics, and connection methods for classification. This type of method has obvious limitations. First, the empirical threshold setting of traditional methods is too subjective and lacks theoretical basis, making it difficult to adapt to the changes in classification dimensions brought about by the dynamic expansion of the line network. Second, a single indicator cannot fully characterize the differentiated effects of built environment factors (such as commercial facility density and residential population distribution) on station functions. Third, it ignores the chaotic characteristics and spatial interaction effects of passenger flow time series, resulting in the failure to identify commuter-consumption mixed stations. These defects greatly limit the applicability of the classification results.

[0004] In recent years, the widespread deployment of automatic fare collection (AFC) systems has provided abundant data foundation for station classification research, driving a shift in station classification methods toward a data-driven intelligent classification paradigm. In the context of big data, existing technologies have begun to utilize advanced machine learning methods and deep learning architectures for station classification research. These methods leverage the passenger flow big data provided by AFC systems to mine the time-varying characteristics of passenger flow at stations, thereby achieving accurate station classification. For example, Gaussian mixture models, random forest algorithms, decision tree algorithms, hierarchical clustering algorithms, and support vector machines have been widely used in station classification research.

[0005] Despite significant progress in site classification research using big data, existing research still lacks systematic modeling of multidimensional site characteristics, resulting in a lack of a site classification method that comprehensively considers multiple characteristic factors. Furthermore, existing clustering-based classification methods are susceptible to noise when processing complex datasets, and classification accuracy is limited by the initial cluster centers, reducing the reliability of the classification results. Summary of the Invention

[0006] Purpose of the invention: In order to overcome the deficiencies in the prior art, a method and system for classifying urban rail transit stations based on multi-feature fusion is provided, which comprehensively considers the time-varying passenger flow characteristics and static topological characteristics of rail transit stations, classifies stations by fusing the two types of feature parameters, and achieves accurate classification of stations, which has positive significance for promoting differentiated management of rail transit stations and improving operational efficiency.

[0007] Technical solution: To achieve the above objectives, the present invention provides a method for classifying urban rail transit stations based on multi-feature fusion, comprising the following steps:

[0008] S1: Construct the line network topology relationship based on the obtained urban rail transit operation information;

[0009] S2: Based on the obtained passenger flow data of rail transit stations, a passenger flow characteristic index system of the station is constructed and passenger flow characteristic parameters are calculated;

[0010] S3: Calculate the static characteristic parameters of the site based on the network topology information and standardize the static characteristic parameters;

[0011] S4: Based on the passenger flow characteristic parameters and static characteristic parameters, the station characteristic parameters are fused and dimensionally reduced based on the principal component analysis method to obtain the fused station characteristic dataset;

[0012] S5: Determine the initial cluster center based on the fused site feature dataset and perform site clustering based on the k-means algorithm;

[0013] S6: Calculate the sum of squared errors between cluster categories and their centers corresponding to different k values, and determine the optimal number of clusters based on the principle of minimizing the sum of squared errors.

[0014] S7: Output the categories corresponding to all stations based on the cluster division results corresponding to the optimal number of clusters.

[0015] Furthermore, the rail transit operation information in step S1 includes a rail transit operation diagram, line length, station adjacency, station spacing, and operation time. A line network topology relationship matrix is ​​constructed based on the rail transit operation information. The line network topology relationship matrix includes station ID, upstream and downstream stations, and adjacent station spacing.

[0016] Furthermore, the passenger flow data at the rail transit station in step S2 includes the card swiping data of passengers entering each station provided by the automatic ticket vending and checking system. The calculation of the passenger flow characteristic parameters includes:

[0017] A1: Select multiple characteristic parameters at different levels to construct a station passenger flow characteristic indicator system. The first-level indicator is set as the station passenger flow distribution characteristics, and the second-level indicators are the passenger flow time imbalance coefficient, passenger flow entropy value, and daily average hourly passenger flow. The second-level indicator passenger flow time imbalance coefficient also includes third-level indicators: peak hour coefficient and weekday coefficient.

[0018] A2: Quantify the distribution characteristics of passenger flow into each station by calculating the secondary indicators of passenger flow time imbalance coefficient, passenger flow entropy value and flow sequence.

[0019] Furthermore, the step A2 specifically includes:

[0020] A2-1: Based on the multi-day passenger flow data of N rail transit stations, let a station be s, s∈S, where S is the set of all stations in the line network. Calculate the hourly average passenger flow on weekdays and non-workdays using the following formula:

[0021]

[0022] in, Is it a working day / non-working day T i The average passenger flow entering the station during the period, T i is the i-th hour from the start of the operation time, dset is the set of days, among which there are 5 dsets for working days from Monday to Friday, and 2 dsets for non-working days, namely Saturday and Sunday. is the passenger flow of station s within the Ti time granularity during the dset period, and n is the length of the dset set;

[0023] A2-2: Based on the hourly average passenger flow at station s, the average passenger flow entering the station during peak hours is calculated as: The average daily passenger flow into the station is The peak hour coefficient x1 is calculated by the following formula:

[0024]

[0025] According to the above formula, the peak hour coefficients of working days and non-working days are calculated respectively, and the peak hour coefficient of working days is set as The peak hour factor for non-working days is

[0026] Assume that the average passenger flow of station s on weekdays is Passenger flow on non-working days The working day coefficient x2 is calculated by the following formula:

[0027]

[0028] A2-3: Calculate the passenger flow entropy value x3 based on the station's all-day passenger flow. This value is used to measure the degree of balance in the station's passenger flow distribution throughout the day. The formula is as follows:

[0029]

[0030] Among them, x3 is the entropy value of the passenger flow entering the station, Q t is the passenger flow entering the station during period t, T is the set of t, including the morning peak, flat peak, and evening peak periods. The passenger flow entropy values ​​of working days and non-working days are calculated according to the above formula, which is calculated as

[0031] A2-4: Calculate the average daily hourly passenger flow x 4 based on the average passenger flow during the working and non-working hours. The formula is as follows:

[0032]

[0033] Where H is the number of operating hours per day, It's T i The average passenger flow entering the station during the period; the average hourly passenger flow on working days and non-working days is calculated according to the above formula, among which the average hourly passenger flow on working days is The average hourly passenger flow on non-working days is

[0034] Furthermore, the static characteristic parameters in step S3 include two types of indicators: node association and proximity centrality of the site, and the calculation method includes:

[0035] According to the rail transit network topology matrix, the number of stations connected to station s is counted, which is the node association degree of the station, calculated as x5;

[0036] Calculate the closeness centrality of each station in the network, which is x6, according to the formula:

[0037]

[0038] Where x6(s) represents the closeness centrality of site s, d(s,v) is the shortest path length between sites s and v, and N represents the total number of sites.

[0039] Furthermore, the standardization of the static characteristic parameters in step S3 includes:

[0040] All station passenger flow characteristics and static feature parameters are combined into a feature matrix X, that is, Any column of the feature matrix X is a certain type of feature of N sites. The data in each column is standardized as follows:

[0041]

[0042] Where x ij is the normalized parameter, x ij is the jth feature of site i, μ j , σ j are the mean and standard deviation of the j-th feature respectively.

[0043] Furthermore, the step S4 specifically includes:

[0044] B1: Set the standardized feature matrix to X ' , calculate X based on principal component analysis ' The covariance matrix of is obtained, and the eigenvalue used to represent the variance size in the direction of each principal component and the eigenvector used to represent the direction of each principal component are obtained. The eigenvalue is set to Λ and the corresponding eigenvector is set to V;

[0045] B2: Sort by the size of the eigenvalues ​​and select the first m principal components whose cumulative contribution rate reaches a certain threshold, m < ρ, where the threshold G is set. m ≥85%, calculate the cumulative contribution rate G according to the formula i :

[0046]

[0047] Project the original data onto the principal component direction to obtain the data after dimensionality reduction and feature fusion. The projection formula is:

[0048] Y=X′V m

[0049] Among them, Y is the matrix after feature fusion, V m is the matrix consisting of the first m eigenvectors.

[0050] Furthermore, the step S5 specifically includes:

[0051] C1: For the fused feature dataset Y, where the matrix size is N×m, select any data point u1, that is, a row of the matrix, as the first cluster center, and set j=1;

[0052] C2: For the unselected data points in the dataset Y, calculate the distance between each point and the nearest center point in the selected cluster center set U, where U = {u1,u2,...,u j}, the formula is as follows:

[0053]

[0054] Among them, d(y,u i ) represents the data point y and the selected center point ui The distance between them, 1 ≤ i ≤ j, using the Euclidean distance as the metric, according to the formula:

[0055]

[0056] where M is the feature dimension, y m and u im are the values of y and u i at the m-th feature dimension respectively;

[0057] C3: Calculate the probability p(y) that the data point y is selected as the center. The formula is as follows:

[0058]

[0059] where D(y) 2 is the square of the distance between the data point y and the nearest center, and y' is other points in the data set Y;

[0060] Randomly select the next center point from the subset of the data set Y that does not contain the center points U according to the selection probability p(y). According to the formula:

[0061] u j+1 = randsrc(1, 1, [Y - U, p])

[0062] where u j+1 is the next selected cluster center, the randsrc function is used to generate random numbers of the specified size and distribution, and p is the vector composed of the selection probabilities of all non-center data points;

[0063] C4: When j < k, let j = j + 1, and repeat steps C2 - C3 to repeatedly select new cluster centers until k center points are selected. That is, when j = k, stop this process, and set the final initial center set as {u1, u2,..., u k};

[0064] C5: Use the standard k-Means algorithm for iteration, assign the data point y to the nearest center u i to form clusters, and update the center point of each cluster to the mean value of the data points within its cluster. The formula is:

[0065]

[0066] where C i represents the cluster corresponding to the center u i |C i | represents the number of points contained in this cluster, until the iteration proceeds until the center points no longer change significantly or reach the preset number of iterations.

[0067] Furthermore, the step S6 specifically includes:

[0068] D1: For the number of clusters k, calculate the sum of squared errors within the cluster results. The formula is as follows:

[0069]

[0070] Where c is C i The sample points in Indicates C i The mean of all samples in ;

[0071] D2: Let k range from 1 to N, where N is the total number of sites. Iteratively calculate SSE. When the absolute change in SSE |SSE(k)-SSE(k-1)| is less than the given threshold, the algorithm stops and this k value is used as the optimal number of clusters.

[0072] Based on the above method, the present invention also provides an urban rail transit station classification system based on multi-feature fusion, comprising:

[0073] Data acquisition module, used to obtain urban rail transit operation information and passenger flow data at rail transit stations;

[0074] Characteristic parameter calculation module, used to calculate and obtain passenger flow characteristic parameters and static characteristic parameters;

[0075] The feature fusion module is used to fuse and reduce the dimensionality of site feature parameters to obtain the fused site feature dataset;

[0076] The site clustering module is used to determine the initial cluster centers and perform site clustering based on the k-means algorithm;

[0077] An optimal cluster number determination module, used to determine the optimal cluster number;

[0078] Category output module, used to output the categories corresponding to all stations.

[0079] Beneficial effects: Compared with the existing technology, the present invention aims to integrate the multi-dimensional characteristics of the station based on the dynamic time-varying passenger flow data provided by the AFC system and the extracted static characteristic parameters of the line network, improve the initial cluster center selection method, and use multiple iterations to obtain the initial cluster center to achieve accurate classification of the station. It has positive significance for promoting differentiated management of urban rail transit stations and improving operational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 Schematic diagram of the process of the present invention;

[0081] Figure 2 Schematic diagram of the specific steps of step S1 of the method of the present invention;

[0082] Figure 3 Schematic diagram of the specific steps of step S2 of the method of the present invention;

[0083] Figure 4 Schematic diagram of the specific steps of step S22 of the method of the present invention;

[0084] Figure 5 Schematic diagram of the specific steps of step S3 of the method of the present invention;

[0085] Figure 6 It is the overall process framework diagram of the method of the present invention;

[0086] Figure 7 This is a graph showing changes in passenger flow at the site within a day;

[0087] Figure 8 It is the rail transit network topology map;

[0088] Figure 9 Display diagram of the extracted initial cluster centers;

[0089] Figure 10 is the change graph of the sum of squared errors;

[0090] Figure 11 A three-dimensional display diagram for site classification. DETAILED DESCRIPTION

[0091] The present invention is further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0092] Example 1:

[0093] like Figure 1 and Figure 6 As shown, this embodiment provides a method for classifying urban rail transit stations based on multi-feature fusion, including the following steps:

[0094] S1: Construct the line network topology relationship based on the obtained urban rail transit operation information;

[0095] like Figure 2 As shown, step S1 specifically includes:

[0096] S11: Obtain rail transit operation information, including rail transit operation diagram, line length, station adjacency, station spacing, operation hours, and rail transit passenger flow information, including passenger entry card swiping data at each station provided by the automatic ticket vending and checking system;

[0097] S12: Construct a line network topology relationship matrix based on rail transit operation information. The line network topology relationship matrix includes station IDs, upstream and downstream stations, and the distance between adjacent stations.

[0098] S2: Based on the obtained passenger flow data of rail transit stations, a passenger flow characteristic index system of the station is constructed and passenger flow characteristic parameters are calculated;

[0099] like Figure 3 As shown, the passenger flow data of the rail transit station in step S2 includes the card swiping data of passengers entering each station provided by the automatic ticket vending and checking system. The calculation of the passenger flow characteristic parameters includes:

[0100] S21: Select multiple characteristic parameters at different levels to construct a station passenger flow characteristic index system, wherein the first-level index is set as the station passenger flow distribution characteristics, the second-level index is set as the passenger flow time imbalance coefficient, the passenger flow entropy value, and the daily average hourly passenger flow. The second-level index passenger flow time imbalance coefficient also includes the third-level indicators: peak hour coefficient and weekday coefficient;

[0101] S22: Quantify the distribution characteristics of passenger flow into each station by calculating the secondary indicators passenger flow time imbalance coefficient, passenger flow entropy value and flow sequence.

[0102] like Figure 4 As shown, step S22 specifically includes:

[0103] S22-1: Based on the multi-day passenger flow data of N rail transit stations, set a station as s, s∈S, where S is the set of all stations in the line network, and calculate the hourly average passenger flow on weekdays and non-workdays. The formula is as follows:

[0104]

[0105] in, Is it a working day / non-working day T i The average passenger flow entering the station during the period, T i is the i-th hour from the start of the operation time, dset is the set of days, among which there are 5 dsets for working days from Monday to Friday, and 2 dsets for non-working days, namely Saturday and Sunday. is the passenger flow of station s within the Ti time granularity during the dset period, and n is the length of the dset set;

[0106] S22-2: Based on the hourly average passenger flow at station s, the average passenger flow entering the station during peak hours is calculated as The average daily passenger flow into the station is The peak hour coefficient x1 is calculated by the following formula:

[0107]

[0108] According to the above formula, the peak hour coefficients of working days and non-working days are calculated respectively, and the peak hour coefficient of working days is set as The peak hour factor for non-working days is

[0109] Assume that the average passenger flow of station s on weekdays is Passenger flow on non-working days The working day coefficient x2 is calculated by the following formula:

[0110]

[0111] S22-3: Calculate the passenger flow entropy value x3 based on the station's all-day passenger flow. This value is used to measure the degree of balance in the station's passenger flow distribution throughout the day. The formula is as follows:

[0112]

[0113] Among them, x3 is the entropy value of the passenger flow entering the station, Q t is the passenger flow entering the station during period t, T is the set of t, including the morning peak, flat peak, and evening peak periods. The passenger flow entropy values ​​of working days and non-working days are calculated according to the above formula, which is calculated as

[0114] S22-4: Calculate the average daily hourly passenger flow x 4 based on the average passenger flow during the working and non-working hours. The formula is as follows:

[0115]

[0116] Where H is the number of operating hours per day, It's T i The average passenger flow entering the station during the period; the average hourly passenger flow on working days and non-working days is calculated according to the above formula, among which the average hourly passenger flow on working days is The average hourly passenger flow on non-working days is

[0117] S3: Calculate the static characteristic parameters of the site based on the network topology information and standardize the static characteristic parameters. The static characteristic parameters include two types of indicators: node association and proximity centrality.

[0118] like Figure 5 As shown, step S3 specifically includes:

[0119] S31: According to the rail transit network topology matrix, count the number of stations connected to station s, which is the node association degree of the station, calculated as x5;

[0120] S32: Calculate the closeness centrality of each station in the network, which is x6, according to the formula:

[0121]

[0122] Where x6(s) represents the closeness centrality of site s, d(s,v) is the shortest path length between sites s and v, and N represents the total number of sites.

[0123] S33: Combine all station passenger flow characteristics and static feature parameters into a feature matrix X, that is, Any column of the feature matrix X is a certain type of feature of N sites. The data in each column is standardized as follows:

[0124]

[0125] Where x ij is the normalized parameter, x ij is the jth feature of site i, μ j , σ j are the mean and standard deviation of the j-th feature respectively.

[0126] S4: Based on the passenger flow characteristic parameters and static characteristic parameters, the station characteristic parameters are fused and dimensionally reduced based on the principal component analysis method to obtain the fused station characteristic dataset;

[0127] Step S4 specifically includes:

[0128] S41: Set the standardized feature matrix to X ' , calculate X based on principal component analysis ' The covariance matrix of is obtained, and the eigenvalue used to represent the variance size in the direction of each principal component and the eigenvector used to represent the direction of each principal component are obtained. The eigenvalue is set to Λ and the corresponding eigenvector is set to V;

[0129] S42: Sort by the size of the eigenvalues ​​and select the first m principal components whose cumulative contribution rate reaches a certain threshold, m < ρ, where the threshold G is set. m ≥85%, calculate the cumulative contribution rate G according to the formula i :

[0130]

[0131] Project the original data onto the principal component direction to obtain the data after dimensionality reduction and feature fusion. The projection formula is:

[0132] Y=X′V m

[0133] Among them, Y is the matrix after feature fusion, Vm is the matrix consisting of the first m eigenvectors.

[0134] S5: Determine the initial cluster center based on the fused site feature dataset and perform site clustering based on the k-means algorithm;

[0135] S5 as a whole is an innovation. The initial cluster centers in the standard k-means algorithm are completely random, which can cause the algorithm to converge to a local optimum or require multiple iterations. With this improvement, the initial centers are determined with a certain probability, resulting in more stable and reliable clustering results. Steps S52 and S53 are innovative, generating probabilities for data points to be selected as centers and selecting cluster centers based on these probabilities.

[0136] Step S5 specifically includes:

[0137] S51: For the fused feature dataset Y, where the matrix size is N×m, select any data point u1, that is, a row of the matrix as the first cluster center, and set j=1;

[0138] S52: For the unselected data points in the data set Y, calculate the distance between each point and the nearest center point in the selected cluster center set U, where U = {u1, u2, ..., u j}, the formula is as follows:

[0139]

[0140] Among them, d(y,u i ) represents the data point y and the selected center point u i The distance between them, 1≤i≤j, uses Euclidean distance as a metric, according to the formula:

[0141]

[0142] Where M is the feature dimension, y m and u im y and u respectively i The value in the mth feature dimension;

[0143] S53: Calculate the probability p(y) that the data point y is selected as the center. The formula is as follows:

[0144]

[0145] Among them, D(y) 2 is the square of the distance between the data point y and the nearest center, and y' is the other points in the data set Y;

[0146] From the subset of data set Y that does not contain the center point U, randomly select the next center point according to the probability of selection p(y), according to the formula:

[0147] u j+1 = randsrc(1, 1, [Y - U, p])

[0148] where u j+1 is the next selected clustering center, the randsrc function is used to generate random numbers of a specified size and distribution, and p is a vector composed of the selection probabilities of all non - center data points;

[0149] S54: When j < k, let j = j + 1, and repeat steps S52 - S53 to repeatedly select new clustering centers until k center points are selected. That is, when j = k, stop this process, and set the final initial center set as {u1, u2,..., u k};

[0150] S55: Use the standard k - Means algorithm to iterate, assign the data point y to the nearest center u i to form clusters, and update the center point of each cluster to the mean value of the data points within its cluster. The formula is:

[0151]

[0152] where C i represents the cluster corresponding to the center u i |C i | represents the number of points contained in this cluster, until the iteration proceeds until the center points no longer change significantly or reach the preset number of iterations.

[0153] S6: Calculate the sum of squared errors of the clustering categories corresponding to different k values and their centers, and determine the optimal number of clusters based on the principle of the minimum sum of squared errors;

[0154] Step S6 specifically includes:

[0155] D1: For the number of clusters k, calculate the within - cluster sum of squared errors of the clustering result. The formula is as follows:

[0156]

[0157] where c is the sample point in C i and represents the mean value of all samples in C i ;

[0158] D2: Let k range from 1 to N, where N is the total number of stations. Iteratively calculate SSE. When the absolute change amount |SSE(k) - SSE(k - 1)| of SSE is less than the given threshold, the algorithm stops, and this k value is used as the optimal number of clusters.

[0159] S7: Output the categories corresponding to all stations based on the cluster division results corresponding to the optimal number of clusters.

[0160] Example 2:

[0161] Based on the method of Example 1, this embodiment provides an urban rail transit station classification system based on multi-feature fusion, including:

[0162] Data acquisition module, used to obtain urban rail transit operation information and passenger flow data at rail transit stations;

[0163] Characteristic parameter calculation module, used to calculate and obtain passenger flow characteristic parameters and static characteristic parameters;

[0164] The feature fusion module is used to fuse and reduce the dimensionality of site feature parameters to obtain the fused site feature dataset;

[0165] The site clustering module is used to determine the initial cluster centers and perform site clustering based on the k-means algorithm;

[0166] An optimal cluster number determination module, used to determine the optimal cluster number;

[0167] Category output module, used to output the categories corresponding to all stations.

[0168] This embodiment also provides a computer storage medium that stores a computer program that can implement the method described above when a processor executes the computer program. The computer-readable medium can be considered to be tangible and non-transitory. Non-limiting examples of non-transitory tangible computer-readable media include non-volatile memory circuits (such as flash memory circuits, erasable programmable read-only memory circuits, or mask read-only memory circuits), volatile memory circuits (such as static random access memory circuits or dynamic random access memory circuits), magnetic storage media (such as analog or digital tapes or hard drives), and optical storage media (such as CDs, DVDs, or Blu-ray discs). The computer program includes processor-executable instructions stored on at least one non-transitory tangible computer-readable medium. The computer program may also include or rely on stored data. The computer program may include a basic input / output system (BIOS) that interacts with the hardware of a special-purpose computer, device drivers that interact with specific devices of the special-purpose computer, one or more operating systems, user applications, background services, background applications, etc.

[0169] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0170] Example 3:

[0171] In this embodiment, simulation experiments are conducted to verify the effects of the present invention. The specific experiments and data analysis are as follows:

[0172] Reference Figure 6 This embodiment provides a method for classifying urban rail transit stations based on multi-feature fusion, including the following steps:

[0173] S1: Construct the line network topology relationship based on the obtained urban rail transit operation information;

[0174] S2: Based on the obtained passenger flow data of rail transit stations, a passenger flow characteristic index system of the station is constructed and passenger flow characteristic parameters are calculated;

[0175] S3: Calculate the static characteristic parameters of the site based on the network topology information and standardize the static characteristic parameters;

[0176] S4: Based on the passenger flow characteristic parameters and static characteristic parameters, the station characteristic parameters are fused and dimensionally reduced based on the principal component analysis method to obtain the fused station characteristic dataset;

[0177] S5: Determine the initial cluster center based on the fused site feature dataset and perform site clustering based on the k-means algorithm;

[0178] S6: Calculate the sum of squared errors between cluster categories and their centers corresponding to different k values, and determine the optimal number of clusters based on the principle of minimizing the sum of squared errors.

[0179] S7: Output the categories corresponding to all stations based on the cluster division results corresponding to the optimal number of clusters.

[0180] In step S1 of this embodiment, the passenger flow of a certain station of a certain city's rail transit changes within a day as follows: Figure 7 As shown in Figure 2, the rail transit network topology is as follows: Figure 8 shown. Figure 7 In the figure, the time aggregation interval of passenger flow is 5 minutes, and the two curves represent the passenger flow variation patterns of a certain station on weekdays and weekends respectively. Figure 8This is a city's line network topology diagram used in the embodiment, reflecting the station adjacency relationship (not representing the actual spatial distribution of the stations), including 2 lines and 58 stations (including 1 transfer station). The station numbering in the figure is consistent with that in Tables 1 and 2.

[0181] In this embodiment, the station passenger flow characteristic parameters calculated in step S2 and the station static characteristic parameters extracted in step S3 are shown in Table 1. This table gives all the characteristic parameters of some stations, station numbers and Figure 8 Each row represents the values ​​of the seven characteristic parameters of the site before standardization.

[0182] Table 1 Characteristic parameters of rail transit stations (partial)

[0183]

[0184]

[0185] The distribution of the characteristic parameters of each site after fusion in step S4 of this embodiment in the three-dimensional coordinate system, and the initial cluster center extracted in step S5 are as follows: Figure 9 shown. Figure 9 In the figure, the three axes represent the first three principal components (i.e., the fused features) whose cumulative contribution reaches the threshold (85%). Their cumulative contribution G3 reaches 87.6%. It can be seen that feature fusion reduces the feature data dimension (from 9 dimensions to 3) while retaining the key information. The boxes in the figure indicate the four initial cluster centers determined in step S5.

[0186] The change of the sum of squared errors (SSE) in step S6 of this embodiment is as follows: Figure 10 The figure shows the variation trend of SSE when the number of clusters k ranges from 1 to 10. It can be seen that when k>4, the variation of SSE decreases significantly and is lower than the given threshold (20). Therefore, the optimal number of clusters in this embodiment is 4.

[0187] The three-dimensional display of the site classification in step S7 of this embodiment is as follows Figure 11 As shown in the figure, the distribution of characteristic parameters corresponding to the four types of sites in three-dimensional space is represented by different colors. The site classification results are shown in Table 2, which shows the number of sites and site numbers in each category. The "type identifier" is only used to intuitively describe the site category. It is derived from the classification results by analyzing the common characteristics of each category of sites (such as location or surrounding land use).

[0188] Table 2 Site classification results

[0189]

Claims

1. A method for classifying urban rail transit stations based on multi-feature fusion, characterized in that: The steps include: S1: Construct the line network topology relationship based on the obtained urban rail transit operation information; S2: Based on the obtained passenger flow data of rail transit stations, a passenger flow characteristic index system of the station is constructed and passenger flow characteristic parameters are calculated; S3: Calculate the static characteristic parameters of the site based on the network topology information and standardize the static characteristic parameters; S4: Based on the passenger flow characteristic parameters and static characteristic parameters, the station characteristic parameters are fused and dimensionally reduced based on the principal component analysis method to obtain the fused station characteristic dataset; S5: Determine the initial cluster center based on the fused site feature dataset and perform site clustering based on the k-means algorithm; S6: Calculate the sum of squared errors between cluster categories and their centers corresponding to different k values, and determine the optimal number of clusters based on the principle of minimizing the sum of squared errors. S7: Output the categories corresponding to all stations based on the cluster division results corresponding to the optimal number of clusters.

2. The urban rail transit station classification method based on multi-feature fusion according to claim 1 is characterized in that: The rail transit operation information in step S1 includes a rail transit operation diagram, line length, station adjacency, station spacing, and operation time. A line network topology relationship matrix is ​​constructed based on the rail transit operation information. The line network topology relationship matrix includes station IDs, upstream and downstream stations, and adjacent station spacing.

3. The urban rail transit station classification method based on multi-feature fusion according to claim 1 is characterized in that: The passenger flow data at the rail transit station in step S2 includes the card swiping data of passengers entering each station provided by the automatic ticket vending and checking system. The calculation of the passenger flow characteristic parameters includes: A1: Select multiple characteristic parameters at different levels to construct a station passenger flow characteristic indicator system. The first-level indicator is set as the station passenger flow distribution characteristics, and the second-level indicators are the passenger flow time imbalance coefficient, passenger flow entropy value, and daily average hourly passenger flow. The second-level indicator passenger flow time imbalance coefficient also includes third-level indicators: peak hour coefficient and weekday coefficient. A2: Quantify the distribution characteristics of passenger flow into each station by calculating the secondary indicators of passenger flow time imbalance coefficient, passenger flow entropy value and flow sequence.

4. The urban rail transit station classification method based on multi-feature fusion according to claim 3 is characterized in that: The step A2 specifically includes: A2-1: Based on the multi-day passenger flow data of N rail transit stations, let a station be s, s∈S, where S is the set of all stations in the line network. Calculate the hourly average passenger flow on weekdays and non-workdays using the following formula: in, Is it a working day / non-working day T i The average passenger flow entering the station during the period, T i is the i-th hour from the start of the operation time, dset is the set of days, among which there are 5 dsets for working days from Monday to Friday, and 2 dsets for non-working days, namely Saturday and Sunday. is the passenger flow of station s within the Ti time granularity during the dset period, and n is the length of the dset set; A2-2: Based on the hourly average passenger flow at station s, the average passenger flow entering the station during peak hours is calculated as: The average daily passenger flow into the station is The peak hour coefficient x1 is calculated by the following formula: According to the above formula, the peak hour coefficients of working days and non-working days are calculated respectively, and the peak hour coefficient of working days is set as The peak hour factor for non-working days is Assume that the average passenger flow of station s on weekdays is Passenger flow on non-working days The working day coefficient x2 is calculated by the following formula: A2-3: Calculate the passenger flow entropy value x3 based on the station's all-day passenger flow. This value is used to measure the degree of balance in the station's passenger flow distribution throughout the day. The formula is as follows: Among them, x3 is the entropy value of the passenger flow entering the station, Q t is the passenger flow entering the station during period t, T is the set of t, including the morning peak, flat peak, and evening peak periods. The passenger flow entropy values ​​of working days and non-working days are calculated according to the above formula, which is calculated as A2-4: Calculate the average daily hourly passenger flow x 4 based on the average passenger flow during the working and non-working hours. The formula is as follows: Where H is the number of operating hours per day, It's T i The average passenger flow entering the station during the period; the average hourly passenger flow on working days and non-working days is calculated according to the above formula, among which the average hourly passenger flow on working days is The average hourly passenger flow on non-working days is 5. The urban rail transit station classification method based on multi-feature fusion according to claim 1 is characterized in that: The static characteristic parameters in step S3 include two types of indicators: node association and proximity centrality of the site. The calculation method includes: According to the rail transit network topology matrix, the number of stations connected to station s is counted, which is the node association degree of the station, calculated as x5; Calculate the closeness centrality of each station in the network, which is x6, according to the formula: Where x6(s) represents the closeness centrality of site s, d(s,v) is the shortest path length between sites s and v, and N represents the total number of sites.

6. The urban rail transit station classification method based on multi-feature fusion according to claim 5 is characterized in that: The step S3 of normalizing the static characteristic parameters includes: All station passenger flow characteristics and static feature parameters are combined into a feature matrix X, that is, Any column of the feature matrix X is a certain type of feature of N sites. The data in each column is standardized as follows: Where x ij is the normalized parameter, x ij is the jth feature of site i, μ j , σ j are the mean and standard deviation of the j-th feature respectively.

7. The urban rail transit station classification method based on multi-feature fusion according to claim 6 is characterized in that: The step S4 specifically includes: B1: Set the standardized feature matrix to X ' , calculate X based on principal component analysis ' The covariance matrix of is obtained, and the eigenvalue used to represent the variance size in the direction of each principal component and the eigenvector used to represent the direction of each principal component are obtained. The eigenvalue is set to Λ and the corresponding eigenvector is set to V; B2: Sort by the size of the eigenvalues ​​and select the first m principal components whose cumulative contribution rate reaches the threshold, m < ρ, where the threshold G is set. m ≥85%, calculate the cumulative contribution rate G according to the formula i : Project the original data onto the principal component direction to obtain the data after dimensionality reduction and feature fusion. The projection formula is: Y=XV m Among them, Y is the matrix after feature fusion, V m is the matrix consisting of the first m eigenvectors.

8. The urban rail transit station classification method based on multi-feature fusion according to claim 7 is characterized in that: The step S5 specifically includes: C1: For the fused feature dataset Y, where the matrix size is N×m, select any data point u1, that is, a row of the matrix, as the first cluster center, and set j=1; C2: For the unselected data points in the dataset Y, calculate the distance between each point and the nearest center point in the selected cluster center set U, where U = {u1,u2,...,u j }, the formula is as follows: Among them, d(y,u i ) represents the data point y and the selected center point u i The distance between them, 1≤i≤j, uses Euclidean distance as a metric, according to the formula: Where M is the feature dimension, y m and u im y and u respectively i The value in the mth feature dimension; C3: Calculate the probability p(y) that data point y is selected as the center. The formula is as follows: Among them, D(y) 2 is the square of the distance between the data point y and the nearest center, and y' is the other points in the data set Y; From the subset of data set Y that does not contain the center point U, randomly select the next center point according to the probability of selection p(y), according to the formula: u j+1 =randsrc(1,1,[Y-U,p]) Among them, u j+1 To select the next cluster center, the randsrc function is used to generate random numbers of specified size and distribution, and p is a vector consisting of the probabilities of selection of all non-center data points; C4: When j < k, let j = j + 1, and repeat steps C2 - C3 to repeatedly select new cluster centers until k center points are selected. That is, when j = k, stop this process, and set the final initial center set as {u1, u2,..., u k}; C5: Use the standard k-Means algorithm to iterate and assign the data point y to its nearest center u i Form clusters and update the center point of each cluster to the mean of the data points in the cluster. The formula is: Among them, C i Represents the center u i The corresponding cluster, |C i | represents the number of points contained in the cluster, and the iteration continues until the center points no longer change significantly or the preset number of iterations is reached.

9. The urban rail transit station classification method based on multi-feature fusion according to claim 8 is characterized in that: The step S6 specifically includes: D1: For the number of clusters k, calculate the sum of squared errors within the cluster results. The formula is as follows: Where c is C i The sample points in Indicates C i The mean of all samples in ; D2: Let k range from 1 to N, where N is the total number of sites. Iteratively calculate SSE. When the absolute change in SSE |SSE(k)-SSE(k-1)| is less than the given threshold, the algorithm stops and this k value is used as the optimal number of clusters.

10. An urban rail transit station classification system based on multi-feature fusion, characterized in that: include: Data acquisition module, used to obtain urban rail transit operation information and passenger flow data at rail transit stations; Characteristic parameter calculation module, used to calculate and obtain passenger flow characteristic parameters and static characteristic parameters; The feature fusion module is used to fuse and reduce the dimensionality of site feature parameters to obtain the fused site feature dataset; The site clustering module is used to determine the initial cluster centers and perform site clustering based on the k-means algorithm; An optimal cluster number determination module, used to determine the optimal cluster number; Category output module, used to output the categories corresponding to all stations.