Rail transit station classification method and device based on multi-source data
By integrating multi-source data and building a multi-dimensional characteristic index system, the K-means clustering method is used to classify rail transit stations, which solves the problems of strong operational subjectivity and poor results consistency in the existing technology, and realizes the detailed division and scientific evaluation of the functions of rail transit stations.
Patent Information
- Application Number
- CN202411857145.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-06
AI Technical Summary
The existing technology has strong subjectivity in the classification of rail transit stations, poor results consistency, and lacks a classification system that comprehensively considers multiple functions and unique positioning.
A multi-source data-based method is adopted to integrate urban land use, rail transit operation, geographical information and mobile signaling data, and a characteristic index system including five dimensions: passenger flow scale, residents' needs, area development, site connection and surrounding facilities, and accurately classify rail transit stations through the K-means clustering method.
It has achieved detailed classification and scientific evaluation of the functions of urban rail transit stations, improved the accuracy and practicality of classification, and can more comprehensively reflect the diverse functions and unique positioning of rail stations.
Smart Images

Figure CN119939411A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of rail transit technology, and in particular to a rail transit station classification method based on multi-source data and a rail transit station classification device based on multi-source data. Background Art
[0002] The passenger flow of rail transit stations is affected by many factors, including the geographical location of the station, surrounding land use, population density, distribution of commercial facilities, etc. Traditional passenger flow analysis and prediction methods often rely on historical data and fail to fully consider the dynamically changing urban environment and diverse travel needs. Therefore, research methods based on multi-source data have gradually attracted attention, which can comprehensively consider different types of data and provide a more comprehensive perspective for passenger flow analysis.
[0003] In recent years, with the development of big data technology, city managers can obtain a large amount of real-time data, including traffic flow, Internet data, mobile user data, etc. These multi-source data provide new possibilities for analyzing the characteristics of rail transit stations and their relationship with passenger flow. Through in-depth mining and analysis of these data, we can not only identify the key factors affecting passenger flow, but also better understand passengers' travel behavior, thereby providing a scientific basis for the planning and management of urban rail transit.
[0004] Related prior arts include: Invention Patent 1 "A Classification Method and Device for Rail Transit Stations", which obtains multiple rail transit stations to be classified and the station attribute information corresponding to each rail transit station, and clusters multiple rail transit stations according to the station attribute information corresponding to each rail transit station to obtain several cluster sets. Invention Patent 2 "A Classification Method for Rail Transit Stations Based on Big Data" constructs Thiessen polygons around rail transit stations to determine the influence range of the stations, and obtains five types of interest point data within the research range from the network map through data mining technology, and finally determines the classification of rail transit stations by calculating the dominance index and uniformity index of the interest points.
[0005] Disadvantages of invention patent 1:
[0006] The invention proposes to cluster rail stations according to the station attribute information corresponding to each station, but this method does not clearly define the specific content of the rail transit station attribute information to be extracted, resulting in strong subjectivity in the operation and poor consistency in the results; at the same time, the lack of a clear classification system makes the process of determining the category attributes of the clustering results and mapping them to a single station lack objective standards and verification mechanisms.
[0007] Disadvantages of invention patent 2:
[0008] The invention proposes to collect data of five types of points of interest, including companies, banks, residential areas, restaurants and schools within the spatial influence range of rail transit stations, and to determine the type of rail stations by calculating the dominance index and uniformity index of the points of interest, and finally divide the stations into four types: residential, commercial, office and mixed. The invention clarifies the attribute information of rail transit stations to be extracted and proposes a classification system. However, the limitation of the invention is that it only focuses on the data of points of interest and ignores other elements around the rail stations that may be equally important or more distinctive. For example, key factors such as population positions, cultural facilities, leisure and entertainment venues, convenience of public transportation transfers, and walkability are not taken into consideration. These factors often have a significant impact on the functional characteristics, passenger flow attraction and quality of life of residents of rail stations. Therefore, this method is not comprehensive enough in reflecting the specific characteristics and comprehensive attributes of rail stations, and cannot accurately capture the multiple functions and unique positioning of the stations. In order to improve the accuracy and practicality of classification, future research should consider incorporating more dimensions of information to build a more comprehensive and detailed rail transit station classification system. Summary of the invention
[0009] In view of the above problems, the present invention is proposed to provide a rail transit station classification method based on multi-source data and a corresponding rail transit station classification device based on multi-source data to overcome the above problems or at least partially solve the above problems.
[0010] The present invention discloses a rail transit station classification method based on multi-source data, the method comprising:
[0011] Acquire multiple rail transit stations to be classified, determine the influence range of each rail transit station, and acquire station attribute information corresponding to each rail transit station based on the influence range of each rail transit station;
[0012] Construct a characteristic indicator system containing multiple dimensions, and based on the multiple characteristic indicators contained in each dimension, extract the characteristics of each rail transit station from the station attribute information corresponding to each rail transit station;
[0013] The characteristic data of all rail transit stations are constructed into a rail transit original characteristic data matrix, and the original rail transit characteristic data matrix is standardized to obtain a rail transit standard characteristic data matrix;
[0014] Performing dimensionality reduction processing on the rail transit standard characteristic data matrix to obtain a rail transit standard reduced dimensionality characteristic data matrix;
[0015] Clustering all rail transit stations according to the rail transit standard dimension reduction feature data matrix to obtain multiple cluster sets;
[0016] The dominant feature analysis and differential feature comparison of multiple cluster sets are performed to determine the category and functional positioning of each cluster set, and the category of the rail transit stations in the cluster set is determined according to the category of each cluster set.
[0017] Optionally, the site attribute information includes urban land use and planning data, rail transit operation data, geographic information data and mobile phone signaling data.
[0018] Optionally, the multiple dimensions include passenger flow scale dimension, resident demand dimension, district development dimension, station connection dimension and surrounding facilities dimension; the multiple characteristic indicators of the passenger flow scale dimension include average daily gathering and distribution volume on weekdays, gathering and distribution volume during peak hours on weekdays, average daily gathering and distribution volume on holidays, and average daily transfer volume on weekdays; the multiple characteristic indicators of the resident demand dimension include population job density, average daily travel scale, and commuting travel ratio; the multiple characteristic indicators of the district development dimension include land development mix, key development projects, and average housing prices; the multiple characteristic indicators of the station connection dimension include the number of bus connection lines, slow traffic accessibility, road network density, and parking facility density; the multiple characteristic indicators of the station connection dimension include commercial facility coverage, public service facility coverage, and educational facility coverage.
[0019] Optionally, determine the impact range of each rail transit station, including:
[0020] According to the geographical location of the site, the city is divided into multiple hierarchical areas, including the central urban area, the sub-city center, the key development area, and the peripheral area;
[0021] According to the multiple level areas, corresponding influence radius is set for the rail transit stations in each level area; among which, the influence radius of the station in the central urban area is smaller than the influence radius of the station in the urban sub-center or the key development area, and the influence radius of the station in the urban sub-center or the key development area is smaller than the influence radius of the station in the peripheral area.
[0022] Optionally, the characteristic data of all rail transit stations are constructed into a rail transit original characteristic data matrix, and the rail transit original characteristic data matrix is standardized to obtain a rail transit standard characteristic data matrix, including:
[0023] The characteristic data of all rail transit stations are constructed into a rail transit original characteristic data matrix;
[0024] Calculate the mean value of the feature data of each feature in the rail transit original feature data matrix and the standard deviation of each feature data;
[0025] According to the feature data mean of each feature and the standard deviation of each feature data, each feature data is converted into a standard normal distribution to obtain the rail transit standard feature data matrix; the conversion formula is as follows:
[0026]
[0027] Among them, z ij represents the standardized value, x ij Represents the original feature data, μ i represents the mean of the feature data, σ i Represents the standard deviation of the feature data.
[0028] Optionally, the rail transit standard feature data matrix is subjected to dimensionality reduction processing to obtain a rail transit standard dimensionality reduction feature data matrix, including:
[0029] Calculate the covariance matrix of the rail transit standard characteristic data matrix;
[0030] Calculating the eigenvalues and eigenvectors of the covariance matrix;
[0031] Calculate the variance share of each principal component according to the size of the eigenvalue, sum the variance shares of all principal components to obtain the cumulative explained variance and draw a trend graph, and select the top k principal components with the highest contribution rate according to the trend graph;
[0032] The rail transit standard feature data matrix is projected onto the feature vectors corresponding to the first k principal components with the highest contribution rates to obtain the rail transit standard reduced-dimensional feature data matrix.
[0033] Optionally, all rail transit stations are clustered according to the rail transit standard dimension reduction feature data matrix to obtain multiple cluster sets, including:
[0034] S1, according to the rail transit standard dimension reduction feature data matrix, calculate the intra-cluster error sum of squares WCSS under different cluster numbers, draw a WCSS curve and select the position where the error reduction speed is significantly slowed down as the optimal value to obtain the optimal cluster number;
[0035] S2, randomly selecting the optimal number of clustering feature data in the rail transit standard dimensionality reduction feature data matrix as the initial cluster center;
[0036] S3, calculating the Euclidean distance between each feature data in the rail transit standard dimensionality reduction feature data matrix and all cluster centers, and assigning each feature data to the cluster center with the closest distance;
[0037] S4, for each cluster, calculating the average value of all feature data in the cluster, and taking the average value as the new cluster center;
[0038] S5, repeatedly executing S3 and S4 until the change of the cluster center reaches a preset convergence condition or reaches a preset maximum number of iterations, and obtaining the optimal clustering number cluster sets.
[0039] Optionally, a dominant feature analysis and a differential feature comparison are performed on multiple cluster sets to determine the category and functional positioning of each cluster set, including:
[0040] Performing feature analysis on the plurality of cluster sets based on the passenger flow scale dimension, the resident demand dimension, the area development dimension, the station connection dimension, and the surrounding facilities dimension to determine the dominant features of each cluster set;
[0041] Compare the characteristics of rail transit stations in multiple cluster sets to determine the difference characteristics between each cluster set and other cluster sets;
[0042] According to the difference characteristics between each cluster set and other cluster sets, and the dominant characteristics of each cluster set, the category and functional positioning of each cluster set are determined.
[0043] Optionally, the categories of the rail transit stations include core area comprehensive stations, residential function-dominated stations, business and office-dominated stations, transfer hub stations, suburban comprehensive stations, leisure and tourism function stations, and general stations.
[0044] The present invention also discloses a rail transit station classification device based on multi-source data, the device comprising:
[0045] A station influence range determination module is used to obtain multiple rail transit stations to be classified, determine the influence range of each rail transit station, and obtain station attribute information corresponding to each rail transit station based on the influence range of each rail transit station;
[0046] The rail transit station feature extraction module is used to construct a feature index system containing multiple dimensions, and based on the multiple feature indicators contained in each dimension, extract the features of each rail transit station from the station attribute information corresponding to each rail transit station;
[0047] A rail transit station feature standardization module is used to construct the feature data of all rail transit stations into a rail transit original feature data matrix, and perform standardization processing on the rail transit original feature data matrix to obtain a rail transit standard feature data matrix;
[0048] A rail transit station feature dimensionality reduction module, used to perform dimensionality reduction processing on the rail transit standard feature data matrix to obtain a rail transit standard dimensionality reduction feature data matrix;
[0049] A rail transit station clustering module, used to cluster all rail transit stations according to the rail transit standard dimension reduction feature data matrix to obtain multiple cluster sets;
[0050] The rail transit station category determination module is used to perform dominant feature analysis and differential feature comparison on multiple cluster sets, determine the category and functional positioning of each cluster set, and determine the category of the rail transit stations in the cluster set according to the category of each cluster set.
[0051] The present invention includes the following advantages:
[0052] The rail transit station classification method based on multi-source data of the present invention integrates multi-source information such as urban land use and planning data, rail transit operation data, geographic information data and mobile phone signaling data, and constructs a characteristic index system including passenger flow scale, resident demand, area development, station connection and surrounding facilities in five dimensions, and extracts 17 characteristic indicators to comprehensively calibrate the characteristics of rail stations. Subsequently, these characteristic indicators are analyzed using the K-means clustering method, and urban rail transit stations are accurately classified into seven categories: core area comprehensive stations, residential function-dominated stations, business office-dominated stations, transfer hub stations, suburban comprehensive stations, leisure and tourism function stations and general stations, thereby realizing the detailed division and scientific evaluation of the functions of urban rail stations. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a flowchart of the steps of a rail transit station classification method based on multi-source data provided by an embodiment of the present invention;
[0054] Figure 2 It is a schematic diagram of the classification characteristic index system of urban rail transit stations provided by the present invention;
[0055] Figure 3 is a flow chart of the K-means clustering algorithm provided by the present invention;
[0056] Figure 4 It is a schematic diagram of the visualization result of classifying the sample data of the present invention using the K-means clustering algorithm;
[0057] Figure 5 It is a schematic diagram of visualization results of clustering result analysis and site classification of the example data of the present invention;
[0058] Figure 6 It is a structural block diagram of a rail transit station classification device based on multi-source data provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0059] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0060] Reference Figure 1 , shows a flowchart of a method for classifying rail transit stations based on multi-source data provided in an embodiment of the present invention, which may specifically include the following steps:
[0061] Step 101, obtaining a plurality of rail transit stations to be classified, determining the influence range of each rail transit station, and obtaining station attribute information corresponding to each rail transit station based on the influence range of each rail transit station;
[0062] Step 102, constructing a feature index system including multiple dimensions, and extracting the features of each rail transit station from the station attribute information corresponding to each rail transit station based on the multiple feature indexes included in each dimension;
[0063] Step 103, constructing the characteristic data of all rail transit stations into a rail transit original characteristic data matrix, and performing standardization processing on the rail transit original characteristic data matrix to obtain a rail transit standard characteristic data matrix;
[0064] Step 104, performing dimensionality reduction processing on the rail transit standard feature data matrix to obtain a rail transit standard reduced dimensionality feature data matrix;
[0065] Step 105, obtaining a dimension reduction feature data matrix according to the rail transit standard, clustering all rail transit stations to obtain multiple cluster sets;
[0066] Step 106 , performing dominant feature analysis and differential feature comparison on multiple cluster sets, determining the category and functional positioning of each cluster set, and determining the category of the rail transit station in the cluster set according to the category of each cluster set.
[0067] The rail transit station classification method based on multi-source data of the present invention is implemented. First, a plurality of rail transit stations to be classified are obtained from the city. Then, the geographical location of the rail transit stations is used to determine the influence range of each rail transit station, and based on these influence ranges, the station attribute information corresponding to each station is obtained. This information covers multiple aspects such as urban land use, rail transit operation, geographic information, and mobile phone signaling.
[0068] Specifically, we collect multi-source data including urban land use and planning data, rail transit operation data, geographic information data, and mobile phone signaling data. We clean the collected data, remove noise and redundant data, and process missing values to ensure the integrity and reliability of the data. This study mainly requires the following types of data:
[0069] 1) Land use and planning data
[0070] Data including land use nature, development intensity, key construction projects, population job distribution, etc. can be obtained and vectorized from the public data of urban planning departments.
[0071] 2) Rail transit operation data
[0072] The AFC card swiping data of each station on the urban rail transit network within a certain period of time includes passenger ID, entry time, exit time, entry and exit stations, etc., provided by the urban rail transit operating company.
[0073] 3) Geographic Information Data
[0074] It includes the geographic coordinates of urban public transportation stations, road network vector data, public service facility POI coordinate data, and slow traffic reach vector data, which can be obtained through the open source map website API and processed again in conjunction with ArcGIS.
[0075] 4) Mobile phone signaling data
[0076] This includes information such as users’ travel trajectory, travel time, travel frequency, travel mode and length of stay, provided by telecom operators.
[0077] At the same time, the present invention constructs a feature index system containing five dimensions. For each dimension, multiple feature indexes are further designed to comprehensively calibrate the characteristics of each rail transit station. The goal of constructing the feature index system and extracting features is to extract the most representative and discriminative features from the original data to help the clustering analysis model better understand the inherent laws of the data, thereby achieving more accurate prediction or classification performance.
[0078] Specifically, the present invention extracts 17 characteristic indicators from five dimensions, namely, station passenger flow, resident demand, area development, station connection and surrounding facilities, to calibrate the characteristics of rail stations. The specific characteristic system and indicators are as follows: Figure 2 shown.
[0079] 1) Passenger flow scale
[0080] The indicators of passenger flow scale reflect the distribution volume and transfer situation of rail transit stations at different time periods, which can help evaluate the passenger flow situation and transportation hub function of the station. Among them:
[0081] (1) Average daily distribution volume on weekdays
[0082] This indicator represents the average passenger flow of a railway station on weekdays, reflecting the daily passenger flow scale of the station and is an important indicator for judging the passenger flow load of the station. The data is obtained through the statistics of the AFC card swiping data of the station, and the total number of people entering and leaving the station in a specific time period is counted, and the average value of multiple days is taken. The calculation formula of this indicator is as follows:
[0083]
[0084] In the formula is the calculated value of the average daily distribution index of the ith station on working days, is the number of working days in the data statistical period of the i-th site, represents the period of the i-th station The total site distribution volume between .
[0085] (2) Peak hours on weekdays
[0086] This indicator is used to reflect the passenger flow distribution during peak hours on weekdays, and can evaluate the congestion and passenger flow pressure of the station during peak hours. The peak hour distribution data is obtained through the station AFC card swiping data, by counting the number of people entering and leaving the station during the morning peak (such as 7:00-9:00) and evening peak (such as 17:00-19:00) on weekdays and taking the maximum value of the distribution volume during the morning and evening peaks. The calculation formula of this indicator is as follows:
[0087]
[0088] In the formula is the calculated value of the peak-hour distribution index of the i-th station during the working day, is the number of working days in the data statistical period of the i-th station, P i peak represents the cycle of the i-th station Total peak-hour site collection and distribution volume.
[0089] (3) Average daily gathering and distribution volume during holidays
[0090] This indicator is used to reflect the passenger flow performance of the station during special periods such as holidays. The data is obtained through the statistics of AFC card swiping data at the station, by counting the total number of people entering and leaving the station during statutory holidays and taking the average value. The calculation formula of this indicator is as follows:
[0091]
[0092] In the formula is the calculated value of the average daily gathering and distribution volume index of the i-th station on holidays, is the number of holidays in the data statistics period of the i-th site, represents the period of the i-th station The total site distribution volume between .
[0093] (4) Average daily transfer volume on weekdays
[0094] x4 reflects the frequency of use of a station as a transfer hub and is a key indicator for evaluating the efficiency and importance of station transfers. The transfer volume is obtained based on the AFC card swiping data of each station, the full-day travel OD data of the entire network, and the full-day OD data is imported into the macro traffic planning software (such as TransCAD and VISUM), and the transfer passenger flow of each transfer station is obtained using the software's built-in traffic allocation algorithm. The calculation formula for this indicator is as follows:
[0095]
[0096] In the formula is the calculated value of the average daily transfer volume index of the ith station on weekdays, is the number of working days in the data statistical period of the i-th site, represents the period of the i-th station The total transfer volume.
[0097] 2) Residents’ needs
[0098] The indicators of this dimension are intended to assess the travel needs and travel habits of residents around the station, including:
[0099] (1) Population job density
[0100] This indicator represents the density of residential population and jobs within a certain influence range around the station. It is used to evaluate the potential passenger flow scale of the station service and can reflect the radiation capacity of the station to the residential or work area. Population and job data can be obtained from mobile phone signaling data. By analyzing the active locations of users at night (11 pm to 6 am the next day), the number of residents within the influence range of each station can be identified. The number of jobs can be estimated by analyzing the activity of mobile phone users in the office area during the day (usually 9 am to 6 pm). The calculation formula of this indicator is as follows:
[0101]
[0102] in Represents the calculated value of the population job density index around the i-th station, is the number of permanent residents within the influence area of the i-th station, Pi job is the number of jobs within the influence area of the i-th station, S i is the influence range of the i-th station.
[0103] (2) Average daily travel volume
[0104] This indicator represents the average daily travel volume generated by residents around the station, reflecting the actual total travel demand in the area, and can help understand the role of rail transit in the travel of residents in the area. This characteristic value can be obtained through mobile phone signaling data analysis. The calculation formula of this indicator is as follows:
[0105]
[0106] in represents the calculated value of the average daily travel scale index around the i-th station, is the number of natural days in the statistical period of mobile phone signaling data of the i-th site, is the area around the i-th station The total number of trips made by residents during the cycle.
[0107] (3) Commuting travel ratio (x7)
[0108] This indicator is used to analyze the travel purpose of residents around the station and clarify the main functions of each rail transit station. It can be obtained through mobile phone signaling data analysis. The calculation formula of this indicator is as follows:
[0109]
[0110] in represents the calculated value of the commuting travel ratio index around the i-th station, is the total number of residents’ trips in the statistical period around the i-th station, is the total number of trips by residents during the morning and evening peak hours in the statistical period around the i-th station.
[0111] 3) Area development dimension
[0112] The indicators in this dimension reflect the land use status, development vitality and economic development level of the areas surrounding rail transit stations, and are an important basis for the classification of station functions and their service objects.
[0113] (1) Land development mix (x8)
[0114] The land development mix is an indicator to evaluate the balanced distribution of different land uses around the station, which is measured by information entropy. It indicates the diversity and uniformity of different land uses such as residential, commercial, office, and industrial. A high mix means that there are multiple land uses around the station, which may bring about diverse travel needs and help improve the all-weather utilization rate of the station; while a low mix indicates a single land use, which may lead to a large difference in passenger flow between peak and non-peak hours.
[0115] This characteristic value is obtained through urban land use planning and related GIS data, and the proportion of different land uses within the influence range of the site is counted (mainly including six categories: residential, commercial, office, science and education, medical and life services), and the information entropy of land development is calculated. The formula for calculating the mixed degree of land development is as follows:
[0116]
[0117] Where x8 is the land development mix, N is the number of land use types, and p i is the proportion of the i-th land type. The value range of x8 is [0, 1]. The larger the land development mixed degree value, the more diversified the land use.
[0118] (2) Key development projects (x9)
[0119] This indicator is used to evaluate the future development potential and attractiveness of the site's surrounding area. The specific calculation formula is as follows:
[0120] x9=A1+β1A2+β2A3
[0121] Where x9 is the calculated value of the evaluation index of key development projects around the station, A1 is the area of key development projects that have been built around the station, A2 is the area of key development projects under construction around the station, A3 is the area of key development projects planned around the station, β1 and β2 are the reduction coefficients of the key development project areas, and their value range is [0, 1].
[0122] (3) Average house price (x) 10 )
[0123] This indicator reflects the economic level and real estate market activity of the area where the station is located. It is an important economic indicator to measure the development level and attractiveness of the area where the station is located. The specific calculation formula is as follows:
[0124]
[0125] where x 10 is the calculated value of the average housing price index around the station, N is the total number of real estate transactions around the station in the past year, HP iis the unit price of the i-th property transaction around the station in the past year. Real estate transaction information around the station can be obtained using the API of the Internet real estate agency website.
[0126] 4) Site connection dimension
[0127] The indicators of this dimension reflect the connection level between rail transit stations and other modes of transportation, and evaluate the station's connection capacity, service level, and adaptability to passenger travel patterns. Among them:
[0128] (1) Number of bus connecting routes
[0129] This indicator reflects the degree of connection between rail stations and conventional bus systems. Stations with high bus connection levels can gather a large number of passengers through multiple bus lines, achieve seamless transfers, and enhance the service coverage of the station. The calculation formula for this indicator is as follows:
[0130]
[0131] in represents the calculated value of the number of bus connections around the i-th station, is the number of regular bus routes around the i-th station.
[0132] (2) Accessibility of slow traffic
[0133] This indicator reflects the friendliness of the area around a rail station to slow-moving transportation modes such as walking and cycling. The convenience of slow-moving transportation can greatly increase the passenger flow source of rail transit. The calculation formula for this indicator is as follows:
[0134]
[0135] in represents the calculated value of the slow traffic accessibility index around the i-th station, is the area of the slow traffic reachable within 10 minutes around the i-th station, S i is the influence range of the i-th station.
[0136] (3) Road network density
[0137] This indicator is used to evaluate the motor vehicle accessibility of rail stations in the city. The calculation formula is as follows:
[0138]
[0139] in represents the calculated value of the road network density index around the i-th station, and are the lengths of the main roads, secondary roads and branches around the station, respectively. W1, W2 and W3 are the weights of the main roads, secondary roads and branches, respectively. The weight value range is [0, 1]. S i is the influence range of the i-th station.
[0140] (4) Parking facility density
[0141] This indicator reflects the level of motor vehicle access at rail stations, and the calculation formula is as follows:
[0142]
[0143] in represents the calculated value of the parking facility density index around the i-th station, is the number of parking spaces around the i-th station, S i is the influence range of the i-th station
[0144] 5) Dimensions of surrounding facilities
[0145] The indicators in this dimension are mainly used to evaluate the service functions and supporting facilities level around rail transit stations. The types and richness of surrounding facilities have an important impact on the service scope, passenger experience, and travel needs of the station. Among them:
[0146] (1) Commercial facilities coverage
[0147] This indicator reflects the density of commercial facilities around the site. The calculation formula is as follows:
[0148]
[0149] in represents the calculated value of the commercial facilities coverage index around the i-th station, is the commercial facility area around the i-th station, S i is the influence range of the i-th station.
[0150] (2) Public service facilities coverage
[0151] This indicator reflects the density of public service facilities such as hospitals and parks around the station. The calculation formula is as follows:
[0152]
[0153] in represents the calculated value of the public service facilities coverage index around the i-th station, Respectively represent the number of hospitals around the i-th station, is the number of park facilities around the i-th station, S i is the influence range of the i-th station.
[0154] (3) Coverage of educational facilities
[0155] This indicator reflects the density of educational facilities around the site. The calculation formula is as follows:
[0156]
[0157] in represents the calculated value of the educational facilities coverage index around the i-th station, represents the number of primary and secondary schools around the i-th station, represents the number of higher education institutions around the i-th station, W1 and W2 are the weights of the two types of schools, and the weight value range is [0, 1].
[0158] Based on the above characteristic indicators, the original characteristic data set A can be obtained for n stations in the urban rail transit network as follows:
[0159]
[0160] This example collects characteristic indicator data of urban rail transit stations in a provincial capital city. Some of the data are shown in Table 1:
[0161] Table 1 Data related to characteristic indicators of urban rail transit stations
[0162]
[0163]
[0164] After extracting the features, the feature data of all rail transit stations were constructed into a rail transit original feature data matrix, and standardized to eliminate the dimensional differences between different features, thus obtaining the rail transit standard feature data matrix. Then, using dimensionality reduction techniques such as principal component analysis (PCA), the standard feature data matrix was reduced in dimension to extract the most representative feature information, thus obtaining the rail transit standard reduced dimension feature data matrix.
[0165] After obtaining the feature data after dimensionality reduction, the K-means clustering algorithm was used to perform cluster analysis on rail transit stations. By continuously adjusting the value of the number of clusters K, and combining the dominant feature analysis with the difference feature comparison, the urban rail transit stations were finally accurately classified into seven categories: core area comprehensive stations, residential function-dominated stations, business office-dominated stations, transfer hub stations, suburban comprehensive stations, leisure and tourism function stations, and general stations. These seven categories of stations represent different functional positioning and development characteristics, thus achieving a detailed division and scientific evaluation of the functions of urban rail stations.
[0166] In one embodiment of the present invention, determining the influence range of each rail transit station includes:
[0167] According to the geographical location of the site, the city is divided into multiple hierarchical areas, including the central urban area, the sub-city center, the key development area, and the peripheral area;
[0168] According to the multiple level areas, corresponding influence radius is set for the rail transit stations in each level area; among which, the influence radius of the station in the central urban area is smaller than the influence radius of the station in the urban sub-center or the key development area, and the influence radius of the station in the urban sub-center or the key development area is smaller than the influence radius of the station in the peripheral area.
[0169] In the present invention, accurately defining the influence range of the site is one of the key steps. The influence range of the site is not only related to internal factors such as passenger flow and facility configuration, but also closely related to its geographical location. Sites in the central area of the city often serve a larger population and commercial facilities, and have a wider influence range, while sites in the peripheral areas have a relatively small influence range due to sparse passenger flow. Therefore, the traditional fixed radius method cannot effectively reflect the differences in influence of sites in different regions.
[0170] Therefore, the present invention proposes a hierarchical demarcation method based on the location of the site. This method aims to flexibly adjust the influence radius of the site according to the location conditions of the site, and overcome the limitations of the traditional fixed radius method. For example, if the site is located in the central urban area, the site influence radius is 1km; if the site is located in the sub-center of the city or the key development area, the site influence radius is 2km; if the site is located in other peripheral areas, the site influence radius is 3km.
[0171] In one embodiment of the present invention, the characteristic data of all rail transit stations are constructed into a rail transit original characteristic data matrix, and the rail transit original characteristic data matrix is standardized to obtain a rail transit standard characteristic data matrix, including:
[0172] The characteristic data of all rail transit stations are constructed into a rail transit original characteristic data matrix;
[0173] Calculate the mean value of the feature data of each feature in the rail transit original feature data matrix and the standard deviation of each feature data;
[0174] According to the feature data mean of each feature and the standard deviation of each feature data, each feature data is converted into a standard normal distribution to obtain the rail transit standard feature data matrix; the conversion formula is as follows:
[0175]
[0176] Among them, z ij represents the standardized value, x ij Represents the original feature data, μ i represents the mean of the feature data, σ i Represents the standard deviation of the feature data.
[0177] After feature extraction, data of different dimensions and indicators often have dimensional differences. The difference in dimensions will cause the impact of certain features on the results to be amplified or weakened during cluster analysis. Therefore, before cluster analysis, the data must be standardized to eliminate the interference of different dimensions on the clustering results.
[0178] This study uses the Z-score standardization method to convert each feature data into a standard normal distribution with a mean of 0 and a standard deviation of 1 by calculating the mean and standard deviation of the data. The calculation formula is as follows:
[0179]
[0180] Among them, z ij represents the standardized value, x ij Represents the original feature data, μ i represents the mean of the feature data, σ i Represents the standard deviation of the feature data.
[0181] In one embodiment of the present invention, the rail transit standard feature data matrix is subjected to dimensionality reduction processing to obtain a rail transit standard dimensionality reduction feature data matrix, including:
[0182] Calculate the covariance matrix of the rail transit standard characteristic data matrix;
[0183] Calculating the eigenvalues and eigenvectors of the covariance matrix;
[0184] Calculate the variance share of each principal component according to the size of the eigenvalue, sum the variance shares of all principal components to obtain the cumulative explained variance and draw a trend graph, and select the top k principal components with the highest contribution rate according to the trend graph;
[0185] The rail transit standard feature data matrix is projected onto the feature vectors corresponding to the first k principal components with the highest contribution rates to obtain the rail transit standard reduced-dimensional feature data matrix.
[0186] As the data dimension increases, the so-called "curse of dimensionality" problem may occur, that is, too many features will lead to data sparsity, increased computational complexity, and increased risk of model overfitting. The purpose of feature dimensionality reduction is to reduce the number of features in the data without significantly losing information, and to optimize the performance and computational efficiency of the model.
[0187] This study uses the principal component analysis (PCA) method to generate a set of new, uncorrelated principal components by projecting high-dimensional data onto a low-dimensional space. These principal components are linear combinations of the original features, and each principal component retains as much variance of the data as possible. The specific method is as follows:
[0188] 1) Calculate the covariance matrix
[0189]
[0190] Among them, S is the covariance matrix, Z is the standardized feature data matrix, and n is the number of samples.
[0191] 2) Calculate the eigenvalues and eigenvectors of the covariance matrix
[0192] 3) Calculate the variance share of each principal component, sum up to get the cumulative explained variance and draw a trend graph, and select the first k principal components with the highest contribution rate as the new eigenvalues. The variance share calculation method of the i-th principal component is as follows:
[0193]
[0194] Among them, Δ i is the variance proportion of the i-th principal component, λ i is the eigenvalue of the i-th principal component, is the sum of all principal component eigenvalues.
[0195] 4) Project the original data onto the selected k principal components to obtain the reduced-dimensional data set.
[0196] F=Z·W k
[0197] Among them, F is the dataset after dimensionality reduction, Z is the feature data matrix after standardization, and W is k is the eigenvector matrix of the selected k principal components.
[0198] In this example, after data standardization and feature dimension reduction, a data set containing 7 principal components is finally obtained. Some of the data are shown in Table 2:
[0199] Table 2 Example dataset after feature dimensionality reduction
[0200] Site Number Principal component 1 Principal component 2 Principal component 3 Principal component 4 Principal component 5 Principal component 6 Principal component 7 1 0.5 0.2 0.8 0.1 0.3 0.4 0.1 2 0.3 0.6 0.1 0.7 0.2 0.5 0.4 3 0.2 0.1 0.5 0.3 0.6 0.1 0.7 4 0.8 0.5 0.3 0.2 0.4 0.2 0.1
[0201] In one embodiment of the present invention, all rail transit stations are clustered according to the rail transit standard dimension reduction feature data matrix to obtain multiple cluster sets, including:
[0202] S1, according to the rail transit standard dimension reduction feature data matrix, calculate the intra-cluster error sum of squares WCSS under different cluster numbers, draw a WCSS curve and select the position where the error reduction speed is significantly slowed down as the optimal value to obtain the optimal cluster number;
[0203] S2, randomly selecting the optimal number of clustering feature data in the rail transit standard dimensionality reduction feature data matrix as the initial cluster center;
[0204] S3, calculating the Euclidean distance between each feature data in the rail transit standard dimensionality reduction feature data matrix and all cluster centers, and assigning each feature data to the cluster center with the closest distance;
[0205] S4, for each cluster, calculating the average value of all feature data in the cluster, and taking the average value as the new cluster center;
[0206] S5, repeatedly executing S3 and S4 until the change of the cluster center reaches a preset convergence condition or reaches a preset maximum number of iterations, and obtaining the optimal clustering number cluster sets.
[0207] K-means clustering is a commonly used unsupervised learning algorithm that is used to divide a data set into k independent clusters, so that the samples in each cluster are more similar to other samples in the same cluster, but more different from samples in other clusters. Its basic idea is to find the optimal cluster division by minimizing the square error of samples in the cluster. The specific steps of the K-means clustering algorithm used in the present invention are as follows: Figure 3 shown.
[0208] 1) Determine the number of clusters k
[0209] Using the Elbow Method, calculate the intra-cluster error sum of squares (WCSS) under different k values, draw a WCSS curve and select the position where the error reduction rate slows down significantly as the optimal k value. The intra-cluster error sum of squares is calculated as follows:
[0210]
[0211] In the formula, c j is the centroid of cluster j, Cj is the set of all data points in the jth cluster, x i For set C j The data points in , k is the number of clusters.
[0212] 2) Initialize cluster center
[0213] Randomly select k points as initial cluster centers.
[0214] 3) Sample allocation
[0215] Calculate the Euclidean distance between each sample point and all cluster centers, and then assign each sample point to the cluster center with the closest distance.
[0216] 4) Update cluster center
[0217] For each cluster, calculate the mean of all sample points in the cluster and use the mean as the new cluster center.
[0218] 5) Repeat steps 3 and 4
[0219] The sample allocation and cluster center update are repeated until the cluster center no longer changes significantly or the maximum number of iterations is reached.
[0220] This example uses the K-means clustering method to classify the sample data sites into 7 categories. The classification results are visualized using ARCGIS software. Figure 4 shown.
[0221] In one embodiment of the present invention, a plurality of cluster sets are subjected to dominant feature analysis and differential feature comparison to determine the category and functional positioning of each cluster set, including:
[0222] Performing feature analysis on the plurality of cluster sets based on the passenger flow scale dimension, the resident demand dimension, the area development dimension, the station connection dimension, and the surrounding facilities dimension to determine the dominant features of each cluster set;
[0223] Compare the characteristics of rail transit stations in multiple cluster sets to determine the difference characteristics between each cluster set and other cluster sets;
[0224] According to the difference characteristics between each cluster set and other cluster sets, and the dominant characteristics of each cluster set, the category and functional positioning of each cluster set are determined.
[0225] In the specific analysis process, by observing the characteristic distribution of samples (i.e. stations) in each cluster, the commonalities and differences of different station categories in terms of passenger flow level, residents' travel needs, and the completeness of docking facilities are clarified. At the same time, by comparing the differential characteristics between clusters, it is possible to identify which dimensions have a decisive role in station classification, such as passenger flow, docking conditions, or surrounding land use. This difference analysis not only helps to verify the rationality of the clustering results, but also provides a basis for further station optimization. After interpreting the characteristics of each cluster, each type of station is named and its functional positioning and classification are clarified.
[0226] Based on the characteristic dimensions and indicator selection of the present invention, urban rail transit stations can be mainly divided into 7 types. The specific classification and functional positioning characteristics are shown in Table 3 below.
[0227] Table 3 Classification, functional positioning and index characteristics of rail transit stations
[0228]
[0229]
[0230] After clustering according to the K-means algorithm, the feature mean, standard deviation and distribution of each category can be viewed and matched with the category indicator features defined above. By calculating the Euclidean distance between the cluster center and the given feature mean of the preset 7 types of stations, the category with the smallest distance is the classification of the clustering result. The Euclidean distance calculation method is as follows:
[0231]
[0232] Among them, x i and i are the i-th eigenvalues of the clustering category and the preset analogy respectively.
[0233] The clustering result analysis and site classification results of this embodiment are visualized and analyzed using ARCGIS software. Figure 5 shown.
[0234] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0235] Reference Figure 6, shows a structural block diagram of a rail transit station classification device based on multi-source data provided in an embodiment of the present invention, which may specifically include the following modules:
[0236] The station influence range determination module 601 is used to obtain multiple rail transit stations to be classified, determine the influence range of each rail transit station, and obtain station attribute information corresponding to each rail transit station based on the influence range of each rail transit station;
[0237] The rail transit station feature extraction module 602 is used to construct a feature index system including multiple dimensions, and based on the multiple feature indexes included in each dimension, extract the features of each rail transit station from the station attribute information corresponding to each rail transit station;
[0238] The rail transit station feature standardization module 603 is used to construct the feature data of all rail transit stations into a rail transit original feature data matrix, and perform standardization processing on the rail transit original feature data matrix to obtain a rail transit standard feature data matrix;
[0239] A rail transit station feature dimension reduction module 604 is used to perform dimension reduction processing on the rail transit standard feature data matrix to obtain a rail transit standard dimension reduction feature data matrix;
[0240] A rail transit station clustering module 605 is used to cluster all rail transit stations according to the rail transit standard dimension reduction feature data matrix to obtain multiple cluster sets;
[0241] The rail transit station category determination module 606 is used to perform dominant feature analysis and differential feature comparison on multiple cluster sets, determine the category and functional positioning of each cluster set, and determine the category of the rail transit stations in the cluster set according to the category of each cluster set.
[0242] Optionally, the site attribute information includes urban land use and planning data, rail transit operation data, geographic information data and mobile phone signaling data.
[0243] Optionally, the multiple dimensions include passenger flow scale dimension, resident demand dimension, district development dimension, station connection dimension and surrounding facilities dimension; the multiple characteristic indicators of the passenger flow scale dimension include average daily gathering and distribution volume on weekdays, gathering and distribution volume during peak hours on weekdays, average daily gathering and distribution volume on holidays, and average daily transfer volume on weekdays; the multiple characteristic indicators of the resident demand dimension include population job density, average daily travel scale, and commuting travel ratio; the multiple characteristic indicators of the district development dimension include land development mix, key development projects, and average housing prices; the multiple characteristic indicators of the station connection dimension include the number of bus connection lines, slow traffic accessibility, road network density, and parking facility density; the multiple characteristic indicators of the station connection dimension include commercial facility coverage, public service facility coverage, and educational facility coverage.
[0244] Optionally, the site influence range determination module includes:
[0245] The urban area division submodule is used to divide the city into multiple hierarchical areas according to the geographical location conditions of the site, and the hierarchical areas include the central urban area, the urban sub-center, the key development area, and the peripheral area;
[0246] The regional station influence radius determination submodule is used to set corresponding influence radii for rail transit stations in each level area according to the multiple level areas; wherein the station influence radius in the central urban area is smaller than the station influence radius in the urban sub-center or key development area, and the station influence radius in the urban sub-center or key development area is smaller than the station influence radius in the peripheral area.
[0247] Optionally, the rail transit station feature standardization module includes:
[0248] The rail transit original feature data matrix construction submodule is used to construct the feature data of all rail transit stations into a rail transit original feature data matrix;
[0249] A feature mean and standard deviation calculation submodule, used to calculate the feature data mean of each feature in the rail transit original feature data matrix and the standard deviation of each feature data;
[0250] The original feature data matrix standardization submodule is used to convert each feature data into a standard normal distribution according to the feature data mean of each feature and the standard deviation of each feature data to obtain the rail transit standard feature data matrix; the conversion formula is as follows:
[0251]
[0252] Among them, z ij represents the standardized value, x ij Represents the original feature data, μ irepresents the mean of the feature data, σ i Represents the standard deviation of the feature data.
[0253] Optionally, the rail transit station feature dimension reduction module includes:
[0254] The covariance matrix calculation submodule is used to calculate the covariance matrix of the rail transit standard characteristic data matrix;
[0255] An eigenvalue and eigenvector calculation submodule, used to calculate the eigenvalue and eigenvector of the covariance matrix;
[0256] The principal component selection submodule is used to calculate the variance share of each principal component according to the size of the eigenvalue, sum up the variance shares of all principal components to obtain the cumulative explained variance and draw a trend graph, and select the first k principal components with the highest contribution rate according to the trend graph;
[0257] The data projection submodule is used to project the rail transit standard characteristic data matrix onto the eigenvectors corresponding to the first k principal components with the highest contribution rates to obtain the rail transit standard dimensionality reduction characteristic data matrix.
[0258] Optionally, the rail transit station clustering module includes:
[0259] The optimal cluster number determination submodule is used to calculate the intra-cluster error sum of squares WCSS under different cluster numbers according to the rail transit standard dimension reduction feature data matrix, draw a WCSS curve and select the position where the error reduction speed is significantly slowed down as the optimal value to obtain the optimal cluster number;
[0260] An initial cluster center selection submodule, used for randomly selecting the optimal number of clustering feature data in the rail transit standard dimension reduction feature data matrix as the initial cluster center;
[0261] A feature data allocation submodule, used to calculate the Euclidean distance between each feature data and all cluster centers in the rail transit standard dimension reduction feature data matrix, and to allocate each feature data to the cluster center with the closest distance;
[0262] A new cluster center determination submodule is used to calculate the average value of all feature data in each cluster and use the average value as the new cluster center;
[0263] The clustering submodule is used to repeatedly execute the feature data allocation submodule and the new cluster center determination submodule until the change of the cluster center reaches a preset convergence condition or reaches a preset maximum number of iterations, thereby obtaining the optimal cluster number cluster set.
[0264] Optionally, the rail transit station category determination module includes:
[0265] A dominant feature analysis submodule is used to perform feature analysis on the plurality of cluster sets based on the passenger flow scale dimension, the resident demand dimension, the area development dimension, the station connection dimension and the surrounding facilities dimension, and determine the dominant feature of each cluster set;
[0266] The difference feature determination submodule is used to compare the features of rail transit stations in multiple cluster sets and determine the difference features between each cluster set and other cluster sets;
[0267] The category generation submodule is used to determine the category and functional positioning of each cluster set based on the difference characteristics between each cluster set and other cluster sets and the dominant characteristics of each cluster set.
[0268] Optionally, the categories of the rail transit stations include core area comprehensive stations, residential function-dominated stations, business and office-dominated stations, transfer hub stations, suburban comprehensive stations, leisure and tourism function stations, and general stations.
[0269] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0270] In addition, an embodiment of the present invention further provides an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus.
[0271] Memory, used to store computer programs;
[0272] The processor is used to implement the rail transit station classification method based on multi-source data as described in the above embodiment when executing the program stored in the memory.
[0273] In another embodiment provided by the present invention, a computer-readable storage medium is also provided, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes the rail transit station classification method based on multi-source data described in the above embodiment.
[0274] In another embodiment provided by the present invention, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute the rail transit station classification method based on multi-source data described in the above embodiment.
[0275] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0276] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0277] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A rail transit station classification method based on multi-source data, characterized in that: The method comprises: Acquire multiple rail transit stations to be classified, determine the influence range of each rail transit station, and acquire station attribute information corresponding to each rail transit station based on the influence range of each rail transit station; Construct a characteristic indicator system containing multiple dimensions, and based on the multiple characteristic indicators contained in each dimension, extract the characteristics of each rail transit station from the station attribute information corresponding to each rail transit station; The characteristic data of all rail transit stations are constructed into a rail transit original characteristic data matrix, and the original rail transit characteristic data matrix is standardized to obtain a rail transit standard characteristic data matrix; Performing dimensionality reduction processing on the rail transit standard characteristic data matrix to obtain a rail transit standard reduced dimensionality characteristic data matrix; Clustering all rail transit stations according to the rail transit standard dimension reduction feature data matrix to obtain multiple cluster sets; The dominant feature analysis and differential feature comparison of multiple cluster sets are performed to determine the category and functional positioning of each cluster set, and the category of the rail transit stations in the cluster set is determined according to the category of each cluster set.
2. The method according to claim 1, characterized in that The site attribute information includes urban land use and planning data, rail transit operation data, geographic information data and mobile phone signaling data.
3. The method according to claim 2, characterized in that The multiple dimensions include passenger flow scale dimension, resident demand dimension, district development dimension, station connection dimension and surrounding facilities dimension; the multiple characteristic indicators of the passenger flow scale dimension include the average daily gathering and distribution volume on weekdays, the gathering and distribution volume during peak hours on weekdays, the average daily gathering and distribution volume on holidays, and the average daily transfer volume on weekdays; the multiple characteristic indicators of the resident demand dimension include population job density, average daily travel scale, and commuting travel ratio; the multiple characteristic indicators of the district development dimension include land development mix, key development projects, and average housing prices; the multiple characteristic indicators of the station connection dimension include the number of bus connection lines, slow traffic accessibility, road network density, and parking facility density; the multiple characteristic indicators of the station connection dimension include commercial facility coverage, public service facility coverage, and educational facility coverage.
4. The method according to claim 3, characterized in that Determine the impact area of each rail transit station, including: According to the geographical location of the site, the city is divided into multiple hierarchical areas, including the central urban area, the sub-city center, the key development area, and the peripheral area; According to the multiple level areas, corresponding influence radius is set for the rail transit stations in each level area; among which, the influence radius of the station in the central urban area is smaller than the influence radius of the station in the urban sub-center or the key development area, and the influence radius of the station in the urban sub-center or the key development area is smaller than the influence radius of the station in the peripheral area.
5. The method according to claim 3, characterized in that: The characteristic data of all rail transit stations are constructed into a rail transit original characteristic data matrix, and the rail transit original characteristic data matrix is standardized to obtain a rail transit standard characteristic data matrix, including: The characteristic data of all rail transit stations are constructed into a rail transit original characteristic data matrix; Calculate the mean value of the feature data of each feature in the rail transit original feature data matrix and the standard deviation of each feature data; According to the feature data mean of each feature and the standard deviation of each feature data, each feature data is converted into a standard normal distribution to obtain the rail transit standard feature data matrix; the conversion formula is as follows: Among them, z ij represents the standardized value, x ij Represents the original feature data, μ i represents the mean of the feature data, σ i Represents the standard deviation of the feature data.
6. The method according to claim 3, characterized in that: The rail transit standard feature data matrix is subjected to dimensionality reduction processing to obtain a rail transit standard dimensionality reduction feature data matrix, including: Calculate the covariance matrix of the rail transit standard characteristic data matrix; Calculating the eigenvalues and eigenvectors of the covariance matrix; Calculate the variance share of each principal component according to the size of the eigenvalue, sum the variance shares of all principal components to obtain the cumulative explained variance and draw a trend graph, and select the first k principal components with the highest contribution rate according to the trend graph; The rail transit standard feature data matrix is projected onto the feature vectors corresponding to the first k principal components with the highest contribution rates to obtain the rail transit standard reduced-dimensional feature data matrix.
7. The method according to claim 3, characterized in that According to the rail transit standard dimension reduction feature data matrix, all rail transit stations are clustered to obtain multiple cluster sets, including: S1, according to the rail transit standard dimension reduction feature data matrix, calculate the intra-cluster error sum of squares WCSS under different cluster numbers, draw a WCSS curve and select the position where the error reduction speed is significantly slowed down as the optimal value to obtain the optimal cluster number; S2, randomly selecting the optimal number of clustering feature data in the rail transit standard dimensionality reduction feature data matrix as the initial cluster center; S3, calculating the Euclidean distance between each feature data in the rail transit standard dimensionality reduction feature data matrix and all cluster centers, and assigning each feature data to the cluster center with the closest distance; S4, for each cluster, calculating the average value of all feature data in the cluster, and taking the average value as the new cluster center; S5, repeatedly executing S3 and S4 until the change of the cluster center reaches a preset convergence condition or reaches a preset maximum number of iterations, and obtaining the optimal clustering number cluster sets.
8. The method according to claim 3, characterized in that Conduct dominant feature analysis and differential feature comparison on multiple cluster sets to determine the category and functional positioning of each cluster set, including: Performing feature analysis on the plurality of cluster sets based on the passenger flow scale dimension, the resident demand dimension, the area development dimension, the station connection dimension and the surrounding facilities dimension, and determining the dominant feature of each cluster set; Compare the characteristics of rail transit stations in multiple cluster sets to determine the difference characteristics between each cluster set and other cluster sets; According to the difference characteristics between each cluster set and other cluster sets, and the dominant characteristics of each cluster set, the category and functional positioning of each cluster set are determined.
9. The method according to claim 1, characterized in that: The categories of rail transit stations include core area comprehensive stations, residential function-dominated stations, business and office-dominated stations, transfer hub stations, suburban comprehensive stations, leisure and tourism function stations, and general stations.
10. A rail transit station classification device based on multi-source data, characterized in that: The device comprises: A station influence range determination module is used to obtain multiple rail transit stations to be classified, determine the influence range of each rail transit station, and obtain station attribute information corresponding to each rail transit station based on the influence range of each rail transit station; The rail transit station feature extraction module is used to construct a feature index system containing multiple dimensions, and based on the multiple feature indicators contained in each dimension, extract the features of each rail transit station from the station attribute information corresponding to each rail transit station; A rail transit station feature standardization module is used to construct the feature data of all rail transit stations into a rail transit original feature data matrix, and perform standardization processing on the rail transit original feature data matrix to obtain a rail transit standard feature data matrix; A rail transit station feature dimensionality reduction module, used to perform dimensionality reduction processing on the rail transit standard feature data matrix to obtain a rail transit standard dimensionality reduction feature data matrix; A rail transit station clustering module, used to cluster all rail transit stations according to the rail transit standard dimension reduction feature data matrix to obtain multiple cluster sets; The rail transit station category determination module is used to perform dominant feature analysis and differential feature comparison on multiple cluster sets, determine the category and functional positioning of each cluster set, and determine the category of the rail transit stations in the cluster set according to the category of each cluster set.
Citation Information
Cited By
Autonomous positioning-oriented train turn-back operation identification method and system
CN121502235A
Rail transit connection service quality evaluation and optimization method based on intelligent traffic technology
CN121724493A