Bus stop classification method based on payment data
By using a bus stop classification method based on payment data, combined with geographical location and route information, passenger travel relationships are analyzed. By utilizing spatiotemporal feature analysis and deep learning models, the problem of station matching error in traditional methods is solved, and high-precision classification of bus stops is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
- Filing Date
- 2025-12-29
- Publication Date
- 2026-05-12
AI Technical Summary
Traditional bus stop classification methods lack GPS information and cannot accurately match bus card swiping data, leading to station matching errors and affecting classification accuracy.
Based on the geographical location and route information of bus stops, payment data is used to analyze the relationship between the origin and destination of passenger trips. Combined with spatiotemporal feature analysis and deep learning models, passenger flow stops are matched and classified.
This solves the problem of station matching errors caused by missing data and dynamic route adjustments, and improves the accuracy of bus station classification.
Smart Images

Figure CN122020341A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of passenger flow analysis technology, specifically to a method for classifying bus stops based on payment data. Background Technology
[0002] Passenger flow data is a crucial basis for classifying bus stops and analyzing their passenger flow characteristics. In recent years, domestic researchers have increasingly focused on using payment data for bus passenger flow analysis. Especially in the era of big data, researchers can obtain more refined passenger flow data by utilizing payment methods such as smart cards, QR code payments, and mobile payments. By analyzing this data, researchers can more accurately identify the passenger flow characteristics, distribution, and trends of bus stops, thereby providing support for bus scheduling and route optimization. However, traditional stop classification and passenger flow analysis are limited by survey methods and costs, mainly relying on manual counting and questionnaires. The sample size of the survey data is small and has a high degree of randomness, making it difficult to accurately analyze and classify bus stops and their passenger flow characteristics.
[0003] Currently, traditional methods for classifying bus stops and their passenger flow characteristics based on passenger flow data mainly include cluster analysis, weighted regression, and a combination of qualitative and quantitative analysis. For example, Qiao Rui et al. proposed using cluster analysis and hierarchical clustering to conduct a hierarchical study of different types of bus stops. They started by examining the core functions of the bus stops, the urban location factors, and the various influencing factors of road grade, arriving at a relatively reasonable classification result. Pang Lei et al. proposed using three regression models—Ordinary Least Squares (OLS), Geographically Weighted Regression (GWR), and Multi-Scale Geographically Weighted Regression (MGWR)—to explore the influencing factors and their degree of influence on passenger flow for different types of stops. Han Xu et al. proposed using K-Means cluster analysis based on the daily passenger flow trend characteristics of different stops, classifying all rail transit stops into five types. They then used a Geographically Weighted Regression (GWR) model to study the impact of land use, built environment, and stop type on peak passenger flow during morning and evening arrival and departure times. Xia Xue et al. used the K-Means clustering algorithm to classify urban rail transit stations and analyzed passenger flow characteristics based on station classification. Deng Pingxin et al., based on AFC data, determined the number of categories by combining multiple validity indicators, and used methods such as principal component analysis, k-means clustering, and multiple linear regression to classify stations using a combination of qualitative and quantitative analysis. He Xin et al. proposed establishing a station function positioning database and used a combination of principal component analysis and cluster analysis to obtain station classification results from a quantitative perspective. Li Xiangnan proposed selecting 11 factors related to station characteristics and environmental features as initial variables for cluster analysis, using the K-means method to cluster based on extracted common factors, and finally dividing operating stations into five categories. Fu Bofeng et al. comprehensively considered the traffic function and location characteristics of stations and used a combination of qualitative and quantitative analysis to classify rail transit stations. However, the card-swiping data of buses is usually statistically analyzed based on time. Due to the lack of GPS information, traditional methods cannot determine the corresponding boarding station based on the card-swiping data of buses. Therefore, it cannot solve the problem of station matching error caused by data loss and dynamic route adjustments, which in turn affects the accuracy of bus station classification. Summary of the Invention
[0004] To address the aforementioned shortcomings in existing technologies, this invention provides a bus stop classification method based on payment data. This method analyzes the origin-destination (OD) relationship of passengers' journeys using payment data (including card swiping and QR code scanning records) based on the geographical location and route information of bus stops. Then, by combining spatiotemporal feature analysis and a deep learning model, it solves the problem of station matching errors caused by data gaps and dynamic route adjustments in traditional methods, thereby improving the accuracy of bus stop classification.
[0005] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:
[0006] A bus stop classification method based on payment data includes the following steps:
[0007] Obtain payment data, route data, and departure timetable data for public transportation vehicles;
[0008] The transaction time series of each vehicle on each bus route was extracted based on the payment data.
[0009] Calculate the estimated arrival time for each stop on each bus route based on route data and departure timetable data;
[0010] Passenger flow stations are matched based on the transaction time sequence of each vehicle and the estimated arrival time of each station to obtain the matching results of passenger flow stations and transaction times for each bus route.
[0011] Based on the matching results of passenger flow stops and transaction times in each bus route, a cluster analysis based on time series features is performed to obtain the bus stop classification results.
[0012] Optionally, the estimated arrival time for each stop on each bus route is calculated based on route data and departure timetable data, including:
[0013] Extract the travel time between stops for each bus route at different times based on the route data;
[0014] Extract the departure time of each vehicle on each bus route from the departure timetable data;
[0015] Add the departure time of each vehicle to the travel time of the corresponding stations to obtain the estimated arrival time at each station.
[0016] Optionally, the matching method for travel times between stations includes:
[0017] The matching interval for departure times is determined based on the operating time range of buses on bus routes;
[0018] Match vehicles whose departure times fall within the matching range based on the departure time;
[0019] For each vehicle, add the departure time to the travel time between stations at the midpoint of the matching interval to obtain the travel time between stations for each vehicle.
[0020] Optionally, the travel time between stops for each bus route at different times can be extracted based on the route data, including:
[0021] The site name is converted into the site's geographic location information using a geocoding method;
[0022] Based on the station's geographical location information, bus routes are obtained through bus route planning.
[0023] Extract the running time, distance between each bus route, and number of stops between each bus route at different times based on the bus routes;
[0024] The travel time between stops for each bus route at different times is extracted based on the travel time, distance between stops, and number of stops between stops for each bus route at different times.
[0025] Optionally, the travel time between stops for each bus route at different times can be extracted based on the route's travel time, distance between stops, and number of stops at different times.
[0026] The departure station is selected as the starting station. The time interval between each station and the starting station is calculated by subtracting the travel time between stations. The travel time between stations is then obtained.
[0027] Optionally, passenger flow station matching is performed based on the transaction time sequence of each vehicle and the expected arrival time of each station to obtain the matching results of passenger flow stations and transaction times for each bus route, including:
[0028] With the goal of minimizing the difference between transaction time and expected arrival time, we filter transaction time series with the same vehicle ID and the expected arrival time of each station, and match each transaction time to the corresponding passenger flow station that is closest to the expected arrival time, thus obtaining the matching result between passenger flow station and transaction time.
[0029] Optionally, based on the matching results of passenger flow stops and transaction times in each bus route, cluster analysis based on time series features is performed to obtain bus stop classification results, including:
[0030] Based on the matching results of passenger flow stations and transaction times in each bus route, a station-time point matrix is constructed and converted into long-format time series data including stations, timestamps and observations;
[0031] Extracting multidimensional time series features from long-format time series data;
[0032] Principal component analysis is used to dynamically reduce the dimensionality of multidimensional time series features.
[0033] The K-means clustering method was used to cluster the dimensionality-reduced multidimensional time series features to obtain the bus stop classification results.
[0034] Optionally, the bus stop classification results include peak-hour dense, off-peak balanced, and remote sparse bus stops.
[0035] Optionally, the payment data obtained from public transport vehicles includes public transport card data and mobile payment data.
[0036] Optionally, obtaining bus route data may include the geographical location data of bus stops.
[0037] The present invention has the following beneficial effects:
[0038] Based on the geographical location and route information of bus stops, and by analyzing the relationship between the origin and destination of passenger trips using payment data, and then combining spatiotemporal feature analysis and deep learning models, the problem of station matching error caused by data loss and dynamic route adjustments in traditional methods has been solved, thus improving the accuracy of bus stop classification. Attached Figure Description
[0039] Figure 1 This is a flowchart illustrating a bus stop classification method based on payment data. Detailed Implementation
[0040] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0041] like Figure 1 As shown, an embodiment of the present invention provides a bus stop classification method based on payment data, comprising the following steps S1 to S5:
[0042] S1. Obtain payment data, route data, and departure timetable data for public transportation vehicles;
[0043] In an optional embodiment of the present invention, step S1 uses the public transport card as the main payment data for passenger flow characteristic analysis and station classification. The public transport card is the data collected by the detector after the passenger boards the bus. Each detector is responsible for monitoring one vehicle number. The data covers the public transport payment data collected within a set time period. The dataset contains information such as payment time, date, and vehicle number from different routes, as well as the vehicle number corresponding to each payment data.
[0044] This embodiment uses a dataset extracted from 101 days of bus payment data collected from public transport smart cards, containing matching payment data for two routes. The data includes bus payment data monitored by a detector on a fixed route each time a bus departs. The data dimension is (94435, 55, 1), where 1 represents one dimension feature (traffic flow), and 55 represents the number of stops. Data is matched every other stop, and statistics are compiled hourly. Over 101 days (24 hours), 94435 = 55 * 24 * 101. Passenger flow data is further divided into two categories based on purpose: estimated arrival times at each stop and user payment data. The resulting dataset is quite redundant, consisting of public transport smart card data and mobile payment (e.g., Alipay) data, as shown in Tables 1 and 2.
[0045] Table 1. Public Transportation Smart Card Data Table
[0046]
[0047] Table 2 Mobile Payment Data Table
[0048]
[0049] S2. Extract the transaction time series of each vehicle in each bus route based on the payment data;
[0050] In an optional embodiment of the present invention, as shown in Tables 1 and 2, the dataset contains columns such as route name, route number, vehicle number, fare machine number, card type, fare, transaction date, and transaction time, which contain a lot of useless information. In this embodiment, the initial data is classified by route and the data is processed in a shallow manner. The data is divided according to the route name column, and data with the same route name are grouped into the same sheet. Furthermore, the payment data of Alipay and public transport card are merged, and the four columns of route name, vehicle number, transaction date, and transaction time are retained, as shown in Table 3.
[0051] Table 3. Data after preliminary processing
[0052]
[0053] Then, the preliminarily processed data needs to be categorized into folders by route. Since the stations on each route are different, each route station sheet needs to be extracted into an Excel file containing the route and date. To facilitate subsequent statistics, the vehicle number sheets in the route and date files are listed and distributed into the respective route sheets according to the vehicle number, as shown in Table 4.
[0054] Table 4 Results by Vehicle Number
[0055]
[0056] The first column of the dataset, representing the specific route name, is selected as the standard. The raw data is extracted and summarized, then further processed by vehicle numbering to reduce the number of rows per sheet. This example uses payment data for route 901; all subsequent estimated arrival times are based on the intervals between stations on this route.
[0057] S3. Calculate the estimated arrival time of each stop on each bus route based on the route data and departure timetable data;
[0058] In an optional embodiment of the present invention, step S3, calculating the estimated arrival time of each stop on each bus route based on route data and departure timetable data, includes:
[0059] Extract the travel time between stops for each bus route at different times based on the route data;
[0060] Extract the departure time of each vehicle on each bus route from the departure timetable data;
[0061] Add the departure time of each vehicle to the travel time of the corresponding stations to obtain the estimated arrival time at each station.
[0062] The matching methods for travel times between stations include:
[0063] The matching interval for departure times is determined based on the operating time range of buses on bus routes;
[0064] Match vehicles whose departure times fall within the matching range based on the departure time;
[0065] For each vehicle, add the departure time to the travel time between stations at the midpoint of the matching interval to obtain the travel time between stations for each vehicle.
[0066] Among them, the travel time between stops for each bus route at different times, extracted from the route data, includes:
[0067] The site name is converted into the site's geographic location information using a geocoding method;
[0068] Based on the station's geographical location information, bus routes are obtained through bus route planning.
[0069] Extract the running time, distance between each bus route, and number of stops between each bus route at different times based on the bus routes;
[0070] The travel time between stops for each bus route at different times is extracted based on the travel time, distance between stops, and number of stops between stops for each bus route at different times.
[0071] The process of extracting the travel time between stops for each bus route at different times, based on the route's travel time, distance between stops, and number of stops at different times, includes:
[0072] The departure station is selected as the starting station. The time interval between each station and the starting station is calculated by subtracting the travel time between stations. The travel time between stations is then obtained.
[0073] This embodiment uses web scraping technology to obtain travel times between stations. First, register on the Gaode Open Platform, create an application, and obtain an API Key. Then, install the necessary libraries: requests, pandas, and openpyxl (for processing Excel files). Write a function to call Gaode's geocoding API to convert station names to latitude and longitude coordinates. Write a function to call the public transport route planning API, passing in the coordinates of the origin and destination to obtain the bus route. Parse the returned JSON data to extract travel time, distance, and number of stations. Finally, store the results in a DataFrame and export them to an Excel file, as shown in Table 5.
[0074] Table 5. Results obtained from the web crawler
[0075]
[0076] As shown in Table 5, the inter-station travel interval data is somewhat messy and cannot be used directly. Therefore, it is simplified to remove duplicates. First, code is written to select the departure station as the starting station. Then, the destination station (2) and the travel time (2) are output and renamed as the station name column and the time interval column, respectively. For data that cannot be obtained directly, the time between stations is subtracted. It is also necessary to limit the difference between the data to not be too large, as shown in Table 6.
[0077] Table 6. Simplified travel timetable for stations at intervals (7 points)
[0078]
[0079] This embodiment calculates the estimated arrival time of each vehicle at each station based on the departure timetable. Like a transaction schedule, the departure timetable initially contains a lot of redundant information, so it also needs processing until only two columns remain: vehicle number and departure time, as shown in Table 7.
[0080] Table 7. Processed Departure Timetable
[0081]
[0082] The estimated arrival time at each station is obtained by adding the departure time of each vehicle to the travel time between stations. However, the travel time between stations varies in different time periods. Therefore, different travel times between stations are assigned according to the time range of the departure time. Since the running time of each bus is in the range of 60 to 90 minutes, for example, for vehicles with departure times in the range of [5:30, 6:30), the travel time to the station at 06:00 is added. The travel time between stations is roughly the same every day, so it is processed separately for each day to obtain the estimated arrival time schedule of stations, as shown in Table 8.
[0083] Table 8 Estimated arrival times for each station
[0084]
[0085] This embodiment integrates redundant data from public transport smart cards and mobile payments, retaining core fields (route name, vehicle number, transaction time), dividing the data by route and storing it in a standardized manner. It utilizes the Gaode API to crawl real-time travel times between stations, and then combines this with departure timetables to generate a dynamic estimated arrival timetable, successfully resolving the issues of messy original data formats and separation of spatiotemporal information.
[0086] S4. Based on the transaction time sequence of each vehicle and the estimated arrival time of each station, perform passenger flow station matching to obtain the matching results of passenger flow stations and transaction times for each bus route.
[0087] In an optional embodiment of the present invention, step S4 performs passenger flow station matching based on the transaction time sequence of each vehicle and the expected arrival time of each station to obtain the matching results of passenger flow stations and transaction times for each bus route, including:
[0088] With the goal of minimizing the difference between transaction time and expected arrival time, we filter transaction time series with the same vehicle ID and the expected arrival time of each station, and match each transaction time to the corresponding passenger flow station that is closest to the expected arrival time, thus obtaining the matching result between passenger flow station and transaction time.
[0089] This embodiment uses a time matching method to determine the passenger's boarding location. The core idea is to match the time information in the bus IC card swipe data with the estimated arrival time calculated based on the departure time to determine the passenger's boarding location. In this method, the swipe time record in the bus IC card data and the time node in the estimated arrival time data are compared to find the closest time point, and then the bus position corresponding to this time point is taken as the passenger's boarding location.
[0090] Specifically, the swipe time of any card swipe record r is t. r In the corresponding vehicle's estimated arrival time data, the time and t need to be found. r The closest data point, assuming the set of data points with expected arrival times is G={g1g2…,gn} and each data point g i The timestamp is t i Then the station g corresponding to the passenger's boarding location p Must meet:
[0091] ;
[0092] Use list comprehension to filter records with the same vehicle ID, then iterate through and compare time differences, and keep the bus stop corresponding to the smallest time difference, as shown in Table 9.
[0093] Table 9 Passenger Flow Station Matching Results
[0094]
[0095] If no corresponding vehicle record is found, mark it as an unknown station. Finally, save the results to a new worksheet in the original Excel file using pd.ExcelWriter's append mode, and replace the existing "Category Results" worksheet. After processing, increment b by 1 to cycle through the data for the next day.
[0096] This invention employs a time-threshold-based passenger flow matching method, comparing the passenger's card-swiping time (tr) with the vehicle's estimated arrival time, and determining the boarding station by minimizing the time difference. To address data anomalies during peak hours, a neighboring station rule is used to optimize matching accuracy.
[0097] S5. Based on the matching results of passenger flow stations and transaction times in each bus route, perform cluster analysis based on time series features to obtain the bus station classification results.
[0098] In an optional embodiment of the present invention, step S5 performs clustering analysis based on time series features according to the matching results of passenger flow stops and transaction times in each bus route to obtain bus stop classification results, including:
[0099] Based on the matching results of passenger flow stations and transaction times in each bus route, a station-time point matrix is constructed and converted into long-format time series data including stations, timestamps and observations;
[0100] Extracting multidimensional time series features from long-format time series data;
[0101] Principal component analysis is used to dynamically reduce the dimensionality of multidimensional time series features.
[0102] The K-means clustering method was used to cluster the dimensionality-reduced multidimensional time series features to obtain the bus stop classification results.
[0103] The bus stop classification results include peak-hour dense, off-peak balanced, and remote sparse bus stops.
[0104] This embodiment first performs data cleaning and structure transformation, converting the original wide-format data (such as a site-time matrix) into long-format time-series data (each row contains a site, timestamp, and observation). Based on this, the tsfresh library is used to automatically extract multidimensional time-series features, covering statistics (mean, variance), frequency domain characteristics (Fourier coefficients), and complex patterns (autocorrelation, entropy). By removing features with high missing rates and combining standardization (StandardScaler), dimensional differences and noise interference are eliminated, ultimately generating a high-quality, low-redundancy feature matrix, laying the foundation for subsequent analysis.
[0105] Tsfresh uses a systematically integrated multidisciplinary algorithm to reconstruct the original time series into a high-dimensional mathematical expression. It runs more than 70 feature calculators in parallel, combining the statistical distribution characteristics (mean, variance, skewness), frequency domain structure (spectral amplitude, phase), time series evolution pattern (linear trend), and nonlinear dynamic characteristics of the original data. Finally, the single-point counting sequence (i.e. the original low-dimensional data) is decomposed into a high-order feature matrix of more than 700 dimensions.
[0106] This embodiment then addresses the potential "curse of dimensionality" caused by high-dimensional features by employing Principal Component Analysis (PCA) for dynamic dimensionality reduction. This retains 95% of the original data variance (n_components=0.95), compressing hundreds of dimensions into core principal components (e.g., 10-20 dimensions). By plotting the cumulative variance contribution rate, the degree of information retention after dimensionality reduction is visually verified. This step not only improves computational efficiency but also suppresses noise, highlights the inherent structure of the data, and allows subsequent clustering to focus more on key differences.
[0107] The K-means algorithm is then applied to partition the data, and the robustness of the model is enhanced through multi-process acceleration and multiple centroid initializations. Finally, cluster labels are output to achieve automatic classification of similar time-series patterns. This embodiment determines the optimal number of clusters based on the dimensionality-reduced features using the elbow rule: the in-cluster sum of squares (WCSS) corresponding to different numbers of clusters is calculated, and an inflection point (such as when the WCSS decreases gradually) is selected as the classification criterion.
[0108] This embodiment uses tsfresh to extract the mean, variance and other temporal features of passenger flow. Combined with PCA dimensionality reduction and K-means clustering algorithm, the stations are divided into three categories: "peak-dense" (such as Zhongshan Road Station), "off-peak balanced" (such as Sanziqiang Station), and "remote and sparse" (such as Xiangbin Community Station), revealing the spatiotemporal heterogeneity of passenger flow.
[0109] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0110] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0111] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0112] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.
[0113] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A method for classifying bus stops based on payment data, characterized in that, Includes the following steps: Obtain payment data, route data, and departure timetable data for public transportation vehicles; The transaction time series of each vehicle on each bus route was extracted based on the payment data. Calculate the estimated arrival time for each stop on each bus route based on route data and departure timetable data; Passenger flow stations are matched based on the transaction time sequence of each vehicle and the estimated arrival time of each station to obtain the matching results of passenger flow stations and transaction times for each bus route. Based on the matching results of passenger flow stops and transaction times in each bus route, a cluster analysis based on time series features is performed to obtain the bus stop classification results.
2. The method for classifying bus stops based on payment data according to claim 1, characterized in that, The estimated arrival time for each stop on each bus route is calculated based on route data and departure timetable data, including: Extract the travel time between stops for each bus route at different times based on the route data; Extract the departure time of each vehicle on each bus route from the departure timetable data; Add the departure time of each vehicle to the travel time of the corresponding stations to obtain the estimated arrival time at each station.
3. The method for classifying bus stops based on payment data according to claim 2, characterized in that, Methods for matching travel times between stations include: The matching interval for departure times is determined based on the operating time range of buses on bus routes; Match vehicles whose departure times fall within the matching range based on the departure time; For each vehicle, add the departure time to the travel time between stations at the midpoint of the matching interval to obtain the travel time between stations for each vehicle.
4. The bus stop classification method based on payment data according to claim 2, characterized in that, Based on the route data, the travel time between stops on each bus route at different times is extracted, including: The site name is converted into the site's geographic location information using a geocoding method; Based on the station's geographical location information, bus routes are obtained through bus route planning. Extract the running time, distance between each bus route, and number of stops between each bus route at different times based on the bus routes; The travel time between stops for each bus route at different times is extracted based on the travel time, distance between stops, and number of stops between stops for each bus route at different times.
5. The bus stop classification method based on payment data according to claim 4, characterized in that, Based on the travel time, distance between stops, and number of stops for each bus route at different times, the travel time between stops for each bus route at different times is extracted, including: The departure station is selected as the starting station. The time interval between each station and the starting station is calculated by subtracting the travel time between stations. The travel time between stations is then obtained.
6. The method for classifying bus stops based on payment data according to claim 1, characterized in that, Passenger flow stations are matched based on the transaction time sequence of each vehicle and the estimated arrival time of each station to obtain the matching results of passenger flow stations and transaction times for each bus route, including: With the goal of minimizing the difference between transaction time and expected arrival time, we filter transaction time series with the same vehicle ID and the expected arrival time of each station, and match each transaction time to the corresponding passenger flow station that is closest to the expected arrival time, thus obtaining the matching result between passenger flow station and transaction time.
7. The bus stop classification method based on payment data according to claim 1, characterized in that, Based on the matching results of passenger flow stops and transaction times in each bus route, cluster analysis based on time series features is performed to obtain bus stop classification results, including: Based on the matching results of passenger flow stations and transaction times in each bus route, a station-time point matrix is constructed and converted into long-format time series data including stations, timestamps and observations; Extracting multidimensional time series features from long-format time series data; Principal component analysis is used to dynamically reduce the dimensionality of multidimensional time series features. The K-means clustering method was used to cluster the dimensionality-reduced multidimensional time series features to obtain the bus stop classification results.
8. The bus stop classification method based on payment data according to claim 7, characterized in that, The bus stop classification results include peak-hour dense, off-peak balanced, and remote and sparse bus stops.
9. A method for classifying bus stops based on payment data according to claim 1, characterized in that, The payment data obtained from public transportation vehicles includes public transportation card data and mobile payment data.
10. A method for classifying bus stops based on payment data according to claim 1, characterized in that, Obtaining bus route data includes the geographical location data of bus stops.