Subway station non-commuting passenger proportion calculation method based on random forest model
Through the random forest model and SHAP analysis, the proportion of non-commuting passenger flow at subway stations is analyzed in a refined manner, which solves the problem of difficult to accurately reflect the relationship between non-commuting and commuting passenger flow in the existing technology, and achieves more accurate resource allocation and travel experience improvement.
Patent Information
- Application Number
- CN202510113174.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology is difficult to accurately reflect the relationship between non-commuting and commuting passenger flow at subway stations, resulting in uneven allocation of transportation resources and affecting travel quality.
The random forest model was used in combination with SHapley Additive exPlanations (SHAP) analysis to refine the proportion of non-commuting passenger flow in the total passenger flow, and identify urban characteristics and physical environmental attributes that affect non-commuting travel.
It has achieved a more accurate calculation of the proportion of non-commuting passenger flow, helping traffic managers optimize resource allocation, adjust train shifts and service facilities, meet different travel needs, and improve travel experience and transportation system efficiency.
Smart Images

Figure CN119991386A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of urban planning and traffic engineering, and in particular relates to a method for calculating the proportion of non-commuting passengers at subway stations based on a random forest model. Background Art
[0002] With the rapid urbanization and population growth in urban areas, the demand for non-commuting travel (such as shopping, leisure, and visiting friends) has increased significantly. The existing subway system is mainly planned and designed to meet commuting needs, which may lead to ignoring the travel needs of non-commuting passengers, affecting travel quality and causing unbalanced allocation of transportation resources.
[0003] In order to improve this situation, it is necessary to study and understand the service capacity of subway stations for different travel purposes. Traditionally, the research on subway passenger flow has focused on commuting trips, while non-commuting trips are often ignored. Evaluating the performance of subway stations based on total passenger flow is not enough to accurately reflect the relationship between non-commuting and commuting passenger flow, because each subway station has different uses, surrounding population structures, etc.
[0004] It is necessary to conduct a detailed analysis of the proportion of non-commuting trips, which is of great significance to traffic management. It can help traffic management departments to predict the changes in passenger flow in a specific time period in advance, so as to effectively guide traffic and allocate resources. In addition, it can also assist decision makers to consider introducing intelligent transportation systems, such as real-time data analysis and automated scheduling, to enhance the responsiveness and resilience of the transportation network.
[0005] Ultimately, these measures will help improve the travel experience of all urban residents, allowing them to enjoy high-quality public transportation services regardless of their travel purposes. By paying attention to the characteristics of non-commuting travel, a more equitable, accessible and sustainable urban transportation environment can be created, ensuring that every citizen can equally enjoy convenient public transportation services, while also promoting social interaction and economic vitality.
[0006] In summary, in-depth research on the proportion of non-commuting trips and their impact on the subway system is not only the key to improving service quality, but also an important step in realizing the intelligent and humanized development of urban transportation. Through continuous technological innovation and management improvement, we can build an urban transportation system that is both efficient and inclusive. Summary of the invention
[0007] The present invention aims to conduct a refined analysis of non-commuting travel demand in the subway system to overcome the limitation that the existing evaluation method is only based on total passenger flow and cannot accurately reflect the relationship between non-commuting and commuting passenger flow. By introducing the random forest model combined with SHapley Additive exPlanations (SHAP) analysis, the present invention strives to more accurately calculate the proportion of non-commuting passenger flow in the total passenger flow, so as to truly reflect the characteristics of different types of travel behaviors. This method can not only identify which urban characteristics (such as commercial land coverage, scenic spot distribution, land use diversity, etc.) and physical environment attributes (such as road density and building density) have a significant impact on non-commuting travel, but also help traffic managers optimize resource allocation, adjust train schedules and service facility configurations according to prediction results, and ensure that subway services meet both commuting needs during peak hours and the needs of non-commuting passengers during non-peak hours. To this end, the present invention provides a method for calculating the proportion of non-commuting passengers at subway stations based on a random forest model.
[0008] A method for calculating the proportion of non-commuting passengers at a subway station based on a random forest model of the present invention comprises the following steps:
[0009] Step 1: Data collection and preprocessing.
[0010] Select a research area in the target city and collect relevant built environment data within an 800-meter buffer zone around the subway station; obtain smart card transaction data to obtain full-day travel records, covering randomly selected weekdays in each month of the year, and exclude data on holidays, weekends, and extreme weather to ensure the representativeness of the sample. Then conduct a time period-based screening to exclude passengers within the standard commuting time. And based on each passenger ID code, identify the travel frequency and time period to and from the same starting point and destination point, exclude fixed commuting outside the standard commuting time, and conduct more accurate screening to ensure that the data can truly reflect non-commuting travel modes.
[0011] Step 2: Determination of built environment variables.
[0012] Based on the literature review and preliminary analysis, important built environment variables were identified, and variables that may affect non-commuters were added in combination with local special conditions, such as green space ratio, commercial land coverage, bus line density, number of bus stops, road density, building density, and points of interest (POI) such as the number of businesses, retail stores, restaurants, entertainment, attractions, hospitals, clinics, and stadiums. These variables not only cover the traditional land use mix, but also include newly studied variables such as green space ratio and number of POIs, which capture the impact of public space on leisure travel and the diversity and availability of urban facilities, and are expected to have a unique impact on non-commuters' choice of destination. This step provides a solid foundation for subsequent model construction and ensures the comprehensiveness and pertinence of model inputs.
[0013] Step 3: Model building and training.
[0014] The prediction model was established using the random forest algorithm. The Scikit-learn module in the Python programming language was used to implement the random forest technology. The data set selected by travelers was divided into a training set and a test set. The model parameters were set as the number of decision trees n, the number of variables considered for each split m, and the maximum tree depth was not limited. Then, the model was trained and validated. The random forest model was trained using the training set data, and the performance of different model variants was evaluated using the cross-validation method. After determining the optimal model, the test set was used for validation to ensure that the model had good generalization ability, thereby effectively predicting the proportion of non-commuting travel and revealing the relationship between built environment variables and non-commuting travel behavior.
[0015] Step 4: Model interpretation and analysis.
[0016] The SHAP technique is used to analyze the output of the random forest model to clarify the importance ranking and mechanism of each built environment factor on the proportion of non-commuting passengers at each subway station. This step can reveal the complex nonlinear relationship between variables and provide global and local explanations. Through the SHAP value, it is clear which built environment characteristics have a significant impact on the proportion of non-commuting trips and how these characteristics specifically affect individual predictions. This analysis helps to deeply understand the driving factors behind non-commuting behavior and provide a scientific basis for urban planning and policy making.
[0017] Step 5: Interpretation and prediction of results.
[0018] Determine which built environment elements are most or least attractive to non-commuters based on the SHAP value; determine the numerical matrix of each built environment factor around the subway station in the scenario based on the actual scenario; substitute the above model coefficient results into the built environment factor matrix of the scenario to calculate the non-commuting travel ratio of this subway station.
[0019] Furthermore, step 1 is specifically as follows:
[0020] Step 1.1: Determine the study area and time frame.
[0021] Select case cities: Select research objects that meet the requirements, which should have mature urban rail transit systems, large traffic volumes, and significant non-commuting travel needs.
[0022] Time range: Data from recent years are specially selected, covering data from multiple working days in each month throughout the year, and excluding holidays, weekends and extreme weather to avoid bias.
[0023] Step 1.2: Get smart card transaction data.
[0024] Data source: Cooperate with rail groups or transportation departments to obtain all types of smart card transaction records generated by physical cards used by subway passengers and mobile applications.
[0025] Data fields: Make sure the dataset contains the following key fields, including card ID (to identify individual users), origin and destination station names (reflecting the passenger's travel path), start time and end time (to help determine the travel time period).
[0026] Step 1.3: Define non-commuting travel criteria.
[0027] Time period screening: Typical commuting peak hours are determined based on the working habits of the study city and the management rules and signs within the subway system; trips during all other time periods are considered potential non-commuting trips.
[0028] Removing outlier data points: To reduce the impact of noisy data, exclude passengers who always get on and off the bus at the same station during the same time period during non-commuting hours, as these may be commuters during non-regular working hours.
[0029] Step 1.4: Calculate the proportion of non-commuting trips.
[0030] Calculation method: For each subway station, the non-commuting trip ratio is calculated by dividing the number of daily non-commuting trips at the station by the total trip volume.
[0031] Averaging: Based on weekday data, the average proportion of non-commuting trips at each subway station was calculated as the basis for the final analysis.
[0032] Step 1.5: Data cleaning and verification.
[0033] Clean up duplicate or erroneous entries: Check for and remove possible duplicate records or obviously erroneous data points, such as invalid timestamps or geolocation information.
[0034] Consistency check: Ensure that the logical relationship between the starting station and the destination station is correct. For example, the starting station and destination station of the same trip are not allowed to be the same.
[0035] Step 1 can effectively extract key information about non-commuting travel from a large amount of smart card transaction data, providing a solid data foundation for the subsequent built environment analysis. This step not only ensures the quality of the data, but also ensures the reliability and accuracy of the subsequent analysis results.
[0036] Furthermore, step 2 is specifically as follows:
[0037] Step 2.1: Determine the built environment variables.
[0038] Based on previous literature and understanding of non-commuting travel behavior, the following variables were selected:
[0039] Green space ratio: Calculate the ratio of public green space area to total land area within the 800-meter buffer zone of a subway station.
[0040] Commercial permanent land coverage rate: measures the proportion of commercial land in the total land area within the 800-meter buffer zone of a subway station.
[0041] Bus route density: Count the total length of bus routes per square kilometer within the 800-meter buffer zone of the subway station;.
[0042] Number of bus stops: Record the number of bus stops within the 800-meter buffer zone of each subway station.
[0043] Road density: The ratio of the total length of all roads (including main roads and branch roads) within the 800-meter buffer zone of a subway station to the total area of the buffer zone is measured.
[0044] Building density: Assess the ratio of the total ground floor area of all buildings within the 800-meter buffer zone of the subway station to the total area of the buffer zone.
[0045] Number of POIs: Within the 800-meter buffer zone of the subway station, the number of various points of interest (such as businesses, retail stores, restaurants, entertainment facilities, attractions, hospitals, clinics, gymnasiums, etc.) is classified and counted.
[0046] Land use mix: Quantifies the degree to which different land use types are integrated.
[0047] Step 2.2: Data collection. For each variable in step 2.1, the following specific data collection and processing methods are used:
[0048] Green space ratio: Data source: topographic maps released by the government or geographic information system (GIS) data provided by urban planning departments; processing method: only include the green area of public spaces and roads, and exclude private green spaces; area measurement and ratio calculation are performed through GIS software.
[0049] Commercial land coverage rate: data source: urban land use database or public urban planning documents; processing method: according to the land use classification, extract the boundary information of commercial land, calculate its coverage area and convert it into a proportion value.
[0050] Bus route density and number of bus stops: Data source: open source map data and station distribution information provided by local transportation management departments; processing method: use GIS tools to overlay bus routes and station locations on an 800-meter buffer zone of subway stations to calculate the total length of the routes and the number of stops.
[0051] Road density: Data sources include open source map data, high-resolution remote sensing images, and urban road network databases provided by government departments; processing methods include drawing a road network map, measuring the total length of the roads, and then dividing it by the buffer area to obtain a density index.
[0052] Building density: Data source: city building permit records or identification of building outlines through satellite images; processing method: use GIS technology to outline the ground floor boundaries of buildings, summarize the area and calculate the density ratio.
[0053] Number of POIs: Data source: online map services (such as Baidu Maps, Amap) or other commercial data providers; Processing method: search and count POIs of various categories within 800 meters around the subway station according to predefined POI categories.
[0054] Land use mixture: data source: urban land use database or related research reports; processing method: refer to the calculation formula in existing literature (such as the method proposed by Lau et al.), adjust the land use type classification and category according to local actual conditions, and finally obtain the mixture score.
[0055] Step 2.3: Data cleaning and standardization.
[0056] Outlier handling: Check and remove obviously erroneous data points, such as values outside a reasonable range.
[0057] Missing value filling: For a small amount of missing data, use the average value of the neighboring area or other reasonable interpolation methods to fill it.
[0058] Step 2.4: Variable correlation analysis.
[0059] Preliminary screening: Based on theoretical assumptions and expert opinions, the variables most likely to affect the proportion of non-commuting travel are selected in advance.
[0060] Correlation test: Use statistical methods (such as Pearson correlation coefficient) to test the relationship between variables to avoid introducing highly correlated multicollinearity problems.
[0061] Feature selection: The importance scoring function of the random forest model is used to further streamline the input feature set and retain the most predictive variables for subsequent modeling.
[0062] Step 2 ensures that the selected built environment variables can accurately reflect the spatial characteristics around the subway station, and the quality and consistency of the data are guaranteed, thus providing a solid foundation for the subsequent model construction.
[0063] Furthermore, step 3 is specifically as follows:
[0064] Step 3.1: Model selection and parameter setting.
[0065] Number of decision trees (n_estimators): Set to the number of decision trees to ensure the stability and predictive power of the model.
[0066] Maximum depth (max_depth): Unlimited, allowing each tree to grow freely based on data characteristics to capture complex patterns.
[0067] Number of features per split (max_features): That is, m randomly selected features are considered at each node split, which helps improve model generalization ability and computational efficiency.
[0068] Minimum number of samples to split (min_samples_split): The default value ensures that there is enough data for each split decision.
[0069] Minimum number of leaf node samples (min_samples_leaf): The default value prevents over-segmentation from causing overfitting.
[0070] Step 3.2: Data preparation and division.
[0071] Dataset division: The collected data is divided into training set and test set, with a ratio of 70% training set and 30% test set. In addition, you can consider using the cross-validation method to further evaluate the model performance.
[0072] Standardization / normalization: All input features are standardized or normalized so that variables of different scales can be compared at the same level to prevent certain features from dominating the model output due to their large numerical range.
[0073] Encode categorical variables: If there are categorical variables (such as land use type), they need to be converted to numerical form through dummy coding or other appropriate methods so that the model can handle them.
[0074] Step 3.3: Model training.
[0075] Initialize the Random Forest model: Create a Random Forest Regressor instance based on the above parameters.
[0076] Training model: Use the data in the training set (including built environment variables and the corresponding non-commuting travel ratio) to train the model; in this process, random forest will automatically extract multiple subsamples (bootstrap samples) from the training set and generate a decision tree for each subsample.
[0077] Evaluate model performance: Use the test set to evaluate the performance of the model, with indicators including mean square error (MSE), mean absolute error (MAE), and coefficient of determination (R2); this helps understand the accuracy of the model’s predictions and how well they match the actual data.
[0078] Step 3.4: Model optimization and verification.
[0079] Hyperparameter tuning: Use grid search or randomized search combined with cross-validation to find the optimal hyperparameter combination to improve the prediction accuracy of the model.
[0080] Feature Importance Analysis: Based on the trained random forest model, the importance scores of each built environment variable are extracted to understand which factors have the most significant impact on the proportion of non-commuting trips.
[0081] Enhanced model interpretability: SHAP analysis is introduced to quantify the specific contribution of each built environment feature to a single prediction result, providing a deeper understanding.
[0082] Step 3.5: Model stability test.
[0083] Multiple training and validation: Since the random forest model contains random components, there may be differences in different training rounds. Therefore, it is recommended to conduct multiple training and record the results of each training to ensure the consistency and reliability of the model performance.
[0084] External validation: If there are additional data sources or time series data, the validity of the model can be re-tested through an external validation set to confirm whether its performance on the new data is still robust.
[0085] Through step 3, not only an efficient random forest regression model is constructed, but also the transparency and interpretability of the model are ensured, so that the research results can provide valuable insights for urban planners and support the creation of a more inclusive and sustainable urban transportation system.
[0086] Furthermore, step 4 is specifically as follows:
[0087] Step 4.1: SHAP analysis.
[0088] Shapley value: A concept derived from cooperative game theory that is used to fairly distribute the contribution of each feature to the model prediction.
[0089] Global explanation: Calculate the average SHAP value of each feature over the entire dataset to show its overall importance.
[0090] Local explanation: For a single prediction instance, it provides the specific contribution of each feature in the instance to help understand the behavior of the model in a specific situation.
[0091] Step 4.2: Global feature importance evaluation.
[0092] Mean SHAP Values: Plots a bar or column chart showing the average absolute contribution of each built environment variable to the prediction of the share of non-commuting trips; this helps identify which factors are most important to the model predictions overall.
[0093] Bee Swarm Plot: Visualizes the SHAP value distribution of each data point, reflecting how changes in feature values affect individual prediction results. For example, you can observe the trend of the impact on the proportion of non-commuting trips when the green space ratio increases.
[0094] Step 4.3: Explore nonlinear relationships.
[0095] Dependency Plot: For variables that show significant nonlinear relationships (such as commercial land coverage, green space ratio, etc.), plot the relationship curve between their SHAP values and their own values. This can reveal the positive and negative effects of variables in different intervals and their turning points.
[0096] Step 4.4: Presentation and discussion of results.
[0097] Graphical presentation: Combined with text descriptions, intuitive graphical tools (such as bar charts, bee swarm plots, dependency plots, etc.) are used to present the main findings of the model, including but not limited to which built environment variables have the greatest impact on the proportion of non-commuting travel; the nonlinear relationship between variables and possible reasons; and the unique contribution of certain variables under specific conditions.
[0098] In-depth analysis of key factors: For factors with significant positive or negative impacts, detailed exploration of the underlying mechanisms, such as why commercial land coverage above a certain threshold promotes non-commuting travel, or why green space no longer significantly increases attractiveness after reaching a certain level.
[0099] Furthermore, step 5 is specifically as follows:
[0100] Step 5.1: Result presentation:
[0101] (1) Global feature importance analysis.
[0102] Plot a bar chart: Use the mean SHAP value (mean Shapley Additive exPlanations value) to show the global importance of each built environment variable on the prediction of the proportion of non-commuting trips; the bar chart is sorted by the importance of the features, starting with the most important feature.
[0103] Explain key features: For the top features, provide detailed textual descriptions of why they have a significant impact on non-commuting trips.
[0104] (2) Analysis of the impact of local features.
[0105] Beeswarm plot display: By drawing a beeswarm plot, the specific contribution of each feature in a single sample is intuitively displayed; each point represents a data point, and its position on the x-axis reflects the impact of the feature on the prediction of that specific sample.
[0106] Case study: Select several representative subway stations and explore in depth how their local characteristics affect the proportion of non-commuting trips, using specific examples to assist in understanding the model output.
[0107] (3) Exploration of nonlinear relationships.
[0108] Dependency diagram: For features that show obvious nonlinear relationships (such as commercial land coverage, green space ratio, etc.), a segmented regression diagram is drawn to clearly show the impact trend of changes in these characteristic values on the proportion of non-commuting travel.
[0109] Threshold Effect Analysis: Identify and discuss key characteristics that have threshold effects, i.e., when a characteristic exceeds or falls below a certain critical value, its impact on non-commuting travel will change significantly.
[0110] Step 5.2: In-depth discussion of the influencing mechanisms of key factors.
[0111] Commercial land coverage: Analyze how different levels of commercial land coverage promote or restrict non-commuting activities. For example, moderate commercial land coverage may not be enough to attract a large number of non-commuters, while a higher commercial land coverage can significantly increase the proportion of non-commuting trips.
[0112] Green space ratio: Explores how the presence of green space generally promotes non-commuting travel, but why its positive effect levels off after the green space ratio increases to a certain limit.
[0113] Number of POIs: Especially for dining and scenic spot POIs, discuss how these points of interest serve as centers of attraction for non-commuters and analyze the impact of their scale effect on non-commuting travel.
[0114] Land use mix: Emphasizes how a good mix of land uses can improve accessibility and convenience for residents, thereby encouraging more non-motorized travel.
[0115] Step 5.3: Make policy recommendations.
[0116] (1) Optimize the public transportation network to meet non-commuting travel needs.
[0117] Considering the characteristics of non-commuting travel, it is recommended to adjust the subway service frequency and operating hours, especially at those stations where non-commuting travel has long been observed to be the mainstream. By analyzing the proportion of non-commuting passengers in detail, transportation managers can obtain practical tools to predict passenger flow changes and optimize train arrangements and service facility configurations. This will not only help meet commuting needs during peak hours, but also better serve non-commuting passengers during non-peak hours, improving the efficiency and service level of the overall system.
[0118] (2) Promote urban traffic management based on data analysis.
[0119] In order to achieve a more flexible and responsive operation strategy, it is recommended to introduce intelligent transportation systems, such as real-time data analysis and automated scheduling. These measures can assist decision makers in predicting passenger flow changes in a specific time period in advance, effectively divert traffic and allocate resources, and ensure that urban residents can enjoy high-quality public transportation services regardless of their travel purposes. At the same time, this also promotes social interaction and economic vitality, and promotes the sustainable development of the urban transportation environment.
[0120] The beneficial technical effects of the present invention are:
[0121] 1. Improve forecast accuracy and reliability
[0122] Introducing innovative models: This paper integrates multiple machine learning algorithms and feature engineering techniques to build a more accurate prediction model, which can effectively estimate the proportion of non-commuting passengers in subway stations with higher accuracy and reliability than traditional methods.
[0123] Enhanced data processing capabilities: Utilize big data analysis technology and advanced data preprocessing methods to ensure the quality of data input to the model, thereby further improving the accuracy of prediction results.
[0124] 2. Optimize subway operation strategy and service configuration
[0125] Flexible and responsive scheduling: Based on the refined analysis of the proportion of non-commuting trips, traffic managers can predict the changes in passenger flow during a specific time period in advance and formulate more flexible and responsive operation strategies. This includes adjusting train frequency, service hours and station service facilities to meet the needs of non-peak hours.
[0126] Improve system efficiency: By optimizing train arrangements and service facility configurations, it can not only meet commuter demand during peak hours, but also better serve non-commuter passengers during off-peak hours, thereby improving the efficiency and service level of the overall system.
[0127] 3. Promote the development of intelligent traffic management and traffic equity
[0128] Introduce intelligent transportation systems: It is recommended to adopt intelligent transportation systems such as real-time data analysis and automated scheduling to enhance the responsiveness and resilience of the transportation network and help decision makers to effectively guide traffic and allocate resources.
[0129] Scientific decision-making support: It provides urban planners with a scientific basis based on empirical research, which helps to formulate more reasonable and effective urban planning and transportation policies and promote the sustainable development of cities.
[0130] Inclusive growth: By paying attention to the characteristics of non-commuting travel, it ensures that every citizen can equally enjoy convenient public transportation services and enjoy high-quality services regardless of their travel purpose, thus promoting social equity.
[0131] In summary, this invention significantly improves the accuracy of non-commuting passenger volume prediction at subway stations by providing a novel methodology and technical means, deepens the understanding of non-commuting travel behavior, and provides strong support for urban transportation management and planning. These technical effects will work together to make an important contribution to building a more intelligent, humane, efficient and inclusive urban transportation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0132] Figure 1 It is a statistical diagram of the proportion of non-commuting travel at each subway station of the present invention.
[0133] Figure 2 This is the SHAP value analysis result diagram (global features) of the present invention.
[0134] Figure 3 This is the SHAP value analysis result diagram (local features) of the present invention.
[0135] Figure 4 It is a schematic diagram showing the impact of the key variable (commercial land coverage) of the present invention.
[0136] Figure 5 It is a schematic diagram showing the influence of the key variable (green space ratio) of the present invention.
[0137] Figure 6 It is a schematic diagram of the impact of the key variables (scenic spot POI) of the present invention.
[0138] Figure 7 It is a schematic diagram showing the influence of the key variable (land use mixture) of the present invention. DETAILED DESCRIPTION
[0139] The present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0140] Selection of research subjects
[0141] Background and Characteristics of Chengdu
[0142] Chengdu, the capital of Sichuan Province, is an important economic, cultural and transportation hub in the southwest region. The city has more than 21 million residents and is experiencing rapid urbanization and population growth. By 2024, Chengdu's subway network will expand to 13 lines with a total length of about 650 kilometers, connecting the city center with outlying suburban counties, and the average daily passenger flow is expected to reach more than 6 million. The subway system is not only the artery of urban transportation, but also an indispensable part of citizens' daily life, especially playing an important role in relieving the pressure of ground transportation.
[0143] Uniqueness of Chengdu Metro System
[0144] Chengdu is a particularly suitable city case for studying non-commuting travel modes for the following reasons:
[0145] Extensive subway coverage: The extensive coverage of Chengdu’s subway network makes it an ideal choice for studying the behavior of non-commuters. The subway lines not only run through various areas of the city, but also penetrate into many emerging development areas, providing a convenient travel option for non-commuters.
[0146] Strong attraction: Chengdu Metro system has strong appeal to non-commuting passengers, thanks to its convenience and comfort, as well as the rich commercial and entertainment facilities in the surrounding area. The built environment around subway stations is crucial to attracting non-commuting passengers, so it is important to study how these factors affect people's travel choices.
[0147] Diverse population structure: There are many types of residential areas, work areas, and leisure and entertainment areas in Chengdu, which makes the destinations of subway passengers very diverse. This diversity provides a rich sample for studying different types of non-commuting activities (such as shopping, leisure, medical treatment, etc.).
[0148] The research value of non-commuting travel modes in Chengdu
[0149] In metropolitan environments, non-commuting trips, including shopping, leisure, and social visits, constitute an important part of residents' daily travel, especially during non-working hours. Despite this, current academic research and technological applications still focus on commuting behavior, and relatively little attention has been paid to non-commuting travel modes, which to some extent limits our comprehensive understanding of urban transportation systems and their optimization potential.
[0150] This paper aims to fill the knowledge gap in this field, with a particular focus on Chengdu, a rapidly developing city. By deeply analyzing the relationship between subway stations and the surrounding built environment, we explore how these factors influence and shape non-commuting travel patterns. Specifically, we identify which built environment characteristics (such as commercial facility distribution, public space quality, land use diversity, etc.) can promote or restrict non-commuting travel choices. This analysis not only helps understand passenger behavior, but also provides empirical evidence for improving urban transportation efficiency and service quality.
[0151] Based on a detailed investigation of Chengdu subway stations and their surrounding areas, this paper aims to provide valuable insights for urban planners and policy makers, supporting them to make more informed decisions to optimize public transportation resource allocation, improve service levels during off-peak hours, and ultimately promote the development of a more efficient, fair and sustainable urban transportation ecosystem. In addition, this research will also help promote the introduction of intelligent transportation systems, improve the city's ability to cope with dynamic passenger flow changes, and enhance the overall travel experience of residents.
[0152] Data availability and representativeness
[0153] The smart card data provided by Chengdu Rail Transit Group provides a solid data foundation for the present invention. These data cover the entry and exit records of 282 subway stations, and the time span includes multiple random working days in each month in 2023, ensuring the representativeness and reliability of the data.
[0154] In summary, Chengdu was chosen as the research object not only because it has typical traffic characteristics of large cities, but also because it can provide rich and diverse examples for studying non-commuting travel modes. Through an in-depth analysis of Chengdu subway stations and the surrounding built environment, this study hopes to provide valuable references for urban planners and policymakers, and promote the development of a more inclusive and sustainable urban transportation system.
[0155] The present invention provides a method for calculating the proportion of non-commuting passengers at subway stations based on a random forest model, specifically:
[0156] Step 1: Data Collection
[0157] Data sources and selection criteria
[0158] The data of this embodiment is derived from the smart card data provided by the Chengdu Rail Transit Group, covering the entry and exit records of 282 subway stations. These data contain data from two randomly selected working days each month in 2023, for a total of 24 working days, to ensure that the data can reflect non-commuting behavior patterns outside the standard commuting time. In order to ensure the validity and representativeness of the data, data from holidays and weekends are excluded, because the vast majority of passengers during these time periods are non-commuting people, and no prediction calculation is required, which is inconsistent with the objectives of the present invention. In addition, subway stations at airports and high-speed rail stations are also excluded because the passenger flow at these stations is rarely affected by the surrounding built environment.
[0159] Data preprocessing
[0160] Data cleaning: Remove duplicate records and obviously erroneous data points, such as unusual timestamps or impossible travel paths.
[0161] Screening for non-commuting trips: Based on general practice, passengers outside the morning and evening peak hours of 7:30AM to 9:00AM and 5:30PM to 7:00PM on weekdays in Chengdu are initially classified as non-commuting people. However, in order to reduce noise data, passengers who get on and off the bus at the same time and station for a long time during non-commuting hours are excluded, because they may be commuting passengers whose working hours are not within the standard commuting time.
[0162] Calculate the proportion of non-commuting trips: Calculate by dividing the number of non-commuting trips at each subway station by the total number of trips at that station in a day. Include the number of trips entering and exiting the station, and determine the number of trips per day at each station based on the 24-day average of data. The result is as follows: Figure 1 shown.
[0163] Built environment variable data collection
[0164] In addition to smartcard data, data on multiple built environment variables were collected to fully assess their impact on non-commuting trips:
[0165] Green space ratio: the ratio of green space area (excluding green space in residential areas / institutions) to the total land area of the buffer zone. The green space areas of public spaces and along roads are extracted through the government's topographic map.
[0166] Commercial land cover: The ratio of commercial land to the total land area of the buffer zone.
[0167] Bus route density: the ratio of the total length of bus routes to the total area of the buffer zone, calculated based on the number of times bus routes overlap within the buffer zone.
[0168] Number of bus stops: The number of stops within the buffer zone for each bus line. Stops with the same name on both sides are counted as one stop.
[0169] Road density: the ratio of total road length to total buffer area.
[0170] Building density: the ratio of the total ground floor area of all buildings to the total area of the buffer zone.
[0171] Number of POIs (points of interest): includes the number of POI types such as businesses, retail stores, restaurants, entertainment, attractions, hospitals, clinics, and stadiums.
[0172] Land use mix: Calculates the comprehensiveness of different land use types within the buffer zone of each subway station, covering residential, commercial, industrial, government, green space and parks, education, health, sports and culture, and other categories.
[0173] Data Validation and Quality Control
[0174] In order to ensure the quality and reliability of the data, a series of measures have been taken:
[0175] Cross-validation: The cross-validation method is used to enhance the robustness of the model results and prevent overfitting.
[0176] Sample size check: Although the study was based on only 282 samples, the validity of the analysis was ensured by strictly controlling the number of independent variables and applying advanced statistical techniques.
[0177] External validation: Compare with similar research results in published literature to confirm the rationality of the conclusions of the present invention.
[0178] The above detailed data collection and preprocessing steps laid a solid foundation for subsequent model construction and analysis, ensuring that the data used are both representative and accurate, and can effectively reveal the impact of built environment factors on the proportion of non-commuting travel.
[0179] Step 2: Model calculation
[0180] Selection and optimization of random forest regression models
[0181] The present invention selects Random Forest (RF) as the main analysis tool because it can efficiently handle nonlinear relationships and reduce the impact of outliers, and has the ability to handle spatial autocorrelation and correlation between variables. For the continuous dependent variable of non-commuting travel proportion, the random forest regression algorithm is a suitable choice. Considering that the data set contains 282 samples, in order to deal with the calculation errors and overfitting problems that may be caused by small sample sizes, the present invention strictly controls the number of independent variables and ensures the robustness of the model results through cross-validation methods.
[0182] When building the random forest model, the following key parameters were optimized:
[0183] Number of decision trees (n_estimators = 100): Set to 100 trees to ensure that the model has enough complexity to capture the patterns in the data while avoiding excessive computational burden.
[0184] Unlimited tree depth (max_depth = None): Allow each decision tree to grow fully to better fit the data, but limit the risk of overfitting through other mechanisms such as the maximum number of features.
[0185] Number of features per split (max_features=2): Two randomly selected features are considered at each node split, which helps improve the generalization and stability of the model.
[0186] In addition, Bootstrap Sampling is used to resample multiple training samples from the original data set to construct a set of decision trees Ti(x), and the prediction f(x) of each input x is the average of the prediction values of these decision trees. This method not only enhances the robustness of the model, but also enables the model to capture the diversity within the data.
[0187] Applications of SHAP analysis
[0188] Although the random forest model has strong predictive performance, its interpretability is poor due to its ensemble nature. To this end, this paper introduces SHapley Additive exPlanations (SHAP) analysis, a game theory-based approach to quantify and visualize the specific contribution of each built environment feature to the prediction of the proportion of non-commuting trips. The SHAP value represents the average marginal contribution of each feature to the final prediction, calculated over all possible feature subsets.
[0189] For a particular model prediction f(x), the SHAP value of feature j is defined as:
[0190]
[0191] in:
[0192] N is the set of all built environment features, S is a subset of feature set N excluding feature j, f(S∈{j}) is the prediction made using feature subset S and adding feature j, f(S) is the prediction made using only feature subset S, |S| is the number of features in subset S, Φ j (f,x) is the SHAP value of feature j.
[0193] Through SHAP analysis, complex nonlinear relationships can be decomposed into easy-to-understand parts, namely the additive contribution of each feature. The model prediction f(x) is expressed as the sum of the SHAP values of all features plus a baseline value (usually the expected value of the model), that is:
[0194]
[0195] This decomposition greatly improves the transparency and interpretability of the model, allowing us to not only understand which attributes are generally most important, but also to drill down into the specific reasons behind each prediction.
[0196] Feature Importance Assessment
[0197] The random forest model combined with SHAP analysis can assess the impact of each feature of the built environment on the proportion of non-commuting trips. The mean SHAP value shows the average impact of each feature on the model prediction, revealing the global importance ( Figure 2 ); whereas the bee swarm plot shows how changes in feature values affect individual predictions, providing local explanations ( Figure 3 ). For example, features such as commercial land coverage, green space ratio, and number of POIs show significant importance, and some features such as commercial land coverage have complex nonlinear relationships.
[0198] Results interpretation and verification
[0199] In order to ensure the validity and reliability of the model, we not only rely on the prediction results output by the model, but also use graphical displays (such as Figures 2 to 7 ) intuitively presents the relationship between characteristics and the proportion of non-commuting trips. In particular, for factors that show nonlinear effects, a detailed discussion of their impact at the global and individual levels further deepens our understanding of the complex interactions between built environment factors and non-commuting travel patterns.
[0200] In summary, by combining random forest regression with SHAP analysis, the present invention not only achieves high-precision predictions, but also provides practical insights for urban planners and policymakers to support the creation of more inclusive and sustainable transportation systems.
[0201] Step 3: Result prediction and analysis
[0202] Model Performance Evaluation
[0203] By applying the random forest model and SHAP (SHapley Additive exPlanations) analysis, the present invention not only reveals the specific impact of various factors of the built environment on the proportion of non-commuting trips, but also evaluates the performance of the model. The accuracy and robustness of the model are supported by cross-validation, ensuring the reliability of the results. Specifically, the model is able to capture nonlinear relationships and effectively handle outliers, reducing the problems of spatial autocorrelation and correlation between variables. The performance of the model is evaluated by multiple indicators, including but not limited to mean square error (MSE), coefficient of determination (R2), etc.
[0204] Global Feature Importance
[0205] Figure 2 The global feature importance analysis based on the random forest model is presented, where the average absolute SHAP value of each built environment variable is used to measure its influence on the proportion of non-commuting trips. It can be clearly seen from the figure that factors such as POI-enterprises, commercial land coverage, POI-retail and stores, green space ratio, POI-attractions, and land use mix have a significant impact on the proportion of non-commuting trips. In particular, the two variables POI-enterprises and POI-retail and stores have the strongest influence in the environment around all subway stations, indicating that the economic activity centers in these areas play a key role in attracting non-commuting passengers.
[0206] Local feature influence
[0207] To gain a deeper understanding of the specific contribution of each built environment variable to individual predictions, Figure 3 A bee swarm plot is presented, showing the local impact of different variables on individual data points. Each point in the plot represents a specific sample, and its position reflects the impact of the feature on the prediction of that sample. For example, when the commercial land coverage reaches a certain threshold, its impact on the proportion of non-commuting travel will change significantly. This visualization of the impact of local features helps us better understand how each variable affects non-commuting travel patterns under different conditions.
[0208] Nonlinear effect analysis
[0209] Further research revealed nonlinear relationships between some built environment variables and the proportion of non-commuting trips, such as Figures 3 to 7 These nonlinear effects provide valuable insights into how the built environment complexly influences non-commuting behavior:
[0210] Commercial land coverage: When the commercial land coverage is less than 0.2, its increase will reduce non-commuting trips; however, when the coverage exceeds 0.25, the increase in commercial land significantly increases the non-commuting ratio. This may be because moderate commercial land is not enough to attract a large number of non-commuters, while a higher commercial land ratio creates more attraction.
[0211] Green space ratio: The positive effect of the green space ratio does not continue to increase as it increases. When the green space ratio exceeds 0.2, its positive effect tends to stabilize, indicating that a moderate level of green space is sufficient to promote non-commuting travel, and additional green space may not bring significant additional benefits.
[0212] POI-attractions: When the number of attractions in the subway station area exceeds 35, the positive impact of attractions on non-commuting travel increases significantly, showing a clear scale effect. This means that the concentrated distribution of attractions can enhance their attractiveness to non-commuters.
[0213] Land use mix: When the land use mix exceeds 0.5, its beneficial effects become stronger. A good land use mix setting improves residents' accessibility and makes it more convenient for people to engage in non-commuting activities such as dining, shopping, and leisure.
[0214] Step 4: Apply recommendations and implement strategies
[0215] Based on the in-depth analysis of non-commuting travel patterns and the application of random forest model combined with SHAP analysis in predicting and explaining the impact of built environment characteristics on the proportion of non-commuting travel, a series of application suggestions and strategy implementation frameworks for urban planners, traffic managers and policy makers are proposed. These suggestions aim to optimize the resource allocation and service quality of the subway system to better adapt to the growing non-commuting demand and promote the sustainable development of the city.
[0216] Optimize subway service configuration for different travel purposes
[0217] Based on the model prediction results, identify subway stations and their surrounding areas with a high proportion of non-commuting trips. For these stations, consider:
[0218] Adjust train frequency: Increase train frequency or shorten waiting time, especially during off-peak hours, to meet the needs of non-commuting activities such as shopping and leisure.
[0219] Extended hours of operation: Ensure that participants of nighttime entertainment and social events can use subway services safely and conveniently.
[0220] Improve station facilities: Add infrastructure such as ticket vending machines, barrier-free passages, public rest areas, etc. to enhance passenger experience.
[0221] Strengthening planning for non-commuting related built environments
[0222] The key factors revealed by this study (such as commercial land coverage, green space ratio, number of POIs, etc.) can be used to guide future urban construction and renovation projects.
[0223] Promote mixed land use: Encourage the integration of residential, commercial, cultural and other functions in the same area to reduce the transportation distance of residents' daily activities.
[0224] Increase green areas and open spaces: Reasonable layout of public green areas such as parks and squares will not only beautify the environment but also attract more citizens to engage in leisure activities.
[0225] Enrich POI types: Introduce more types of points of interest, such as catering, retail, and entertainment facilities, to enhance regional attractiveness and stimulate more non-commuting travel.
[0226] Policy support and public participation
[0227] To ensure the effective implementation of the above measures, support at the government level and extensive participation from all sectors of society are needed.
[0228] Formulate relevant policies and regulations: Introduce policies and measures that are conducive to the development of non-commuting travel, such as giving developers certain preferential conditions to build multi-functional complexes.
[0229] Carry out publicity and education activities: popularize the concept of green travel to the public, advocate a low-carbon and environmentally friendly lifestyle, and jointly create a good urban transportation atmosphere.
[0230] Establish a feedback mechanism: regularly collect citizen opinions, evaluate the effectiveness of various improvement measures, and continuously optimize the quality of urban transportation services.
[0231] In summary, the methodology and technical tools provided by this paper not only help to accurately understand non-commuting travel modes and the influencing factors behind them, but also provide a scientific decision-making basis for urban traffic management. By actively adopting the suggestions put forward by this study, we can expect a more efficient, fair and sustainable urban transportation ecosystem to gradually take shape, thereby comprehensively improving the quality of life of urban residents and the level of social and economic development.
Claims
1. A method for calculating the proportion of non-commuting passengers at subway stations based on a random forest model, characterized in that: The following steps are involved: Step 1: Data collection and preprocessing: Select a research area in the target city and collect relevant built environment data within an 800-meter buffer zone around the subway station; obtain smart card transaction data to obtain full-day travel records, covering randomly selected weekdays in each month of the year, and excluding data on holidays, weekends, and extreme weather; then perform time-based screening to exclude passengers within the standard commuting time; and identify the travel frequency and time period to and from the same starting point and destination point based on each passenger ID code, excluding fixed commuting trips outside the standard commuting time; Step 2: Determination of built environment variables: Identify important built environment variables based on literature review and preliminary analysis, and add variables that may have an impact on non-commuters based on local specific circumstances; Step 3: Model building and training: The prediction model was established using the random forest algorithm. The Scikit-learn module in the Python programming language was used to implement the random forest technology. The data set selected by the travelers was divided into a training set and a test set. The model parameters were set as the number of decision trees n, the number of variables considered for each split m, and no limit on the maximum tree depth. Then, the model was trained and validated. The random forest model was trained using the training set data, and the performance of different model variants was evaluated using the cross-validation method. After the optimal model was determined, it was validated using the test set to ensure that the model had good generalization ability. Step 4: Model interpretation and analysis: The SHAP technique was used to analyze the output of the random forest model to clarify the importance ranking and mechanism of each built environment factor on the proportion of non-commuting passengers; Step 5: Result interpretation and prediction: Determine which built environment elements are most or least attractive to non-commuters based on the SHAP value; determine the numerical matrix of each built environment factor around the subway station in the scenario based on the actual scenario; substitute the above model coefficient results into the built environment factor matrix of the scenario to calculate the non-commuting travel ratio of this subway station.
2. The method for calculating the proportion of non-commuting passengers at subway stations based on a random forest model according to claim 1, characterized in that: The step 1 is specifically as follows: Step 1.1: Determine the research area and time frame; Select case cities: Select research objects that meet the requirements, which should have mature urban rail transit systems, large traffic volumes, and significant non-commuting travel needs; Time range: data from recent years are specially selected, covering data from multiple working days in each month throughout the year, and excluding holidays, weekends and extreme weather to avoid bias; Step 1.2: Obtain smart card transaction data; Data sources: Cooperate with rail groups or transportation departments to obtain all types of smart card transaction records generated by physical cards used by subway passengers and mobile phone applications; Data fields: Ensure that the dataset contains the following key fields, including card ID, origin and destination station names, start time and end time; Step 1.3: Define non-commuting travel criteria; Time period screening: According to the working habits of the research city and the management rules and signs in the subway system, the typical commuting peak hours are determined; all other travel during the time period is regarded as potential non-commuting travel; Removing abnormal data points: To reduce the impact of noisy data, exclude passengers who always get on and off the bus at the same time and station during non-commuting hours; Step 1.4: Calculate the proportion of non-commuting trips; Calculation method: For each subway station, the non-commuting trip ratio is calculated by dividing the number of daily non-commuting trips at the station by the total trip volume; Averaging: Based on weekday data, the average non-commuting trip ratio for each subway station is calculated as the basis for the final analysis; Step 1.5: Data cleaning and verification; Clean up duplicate or erroneous entries: Check and remove possible duplicate records or obviously erroneous data points; Consistency check: Ensure that the logical relationship between the origin and destination stations is correct.
3. The method for calculating the proportion of non-commuting passengers at subway stations based on a random forest model according to claim 1, characterized in that: The step 2 is specifically as follows: Step 2.1: Determine the built environment variables; Based on previous literature and understanding of non-commuting travel behavior, the following variables were selected: Green space ratio: Calculate the ratio of public green space area to total land area within the 800-meter buffer zone of the subway station; Commercial permanent land coverage rate: measures the proportion of commercial land in the total land area within the 800-meter buffer zone of the subway station; Bus route density: Count the total length of bus routes per square kilometer within the 800-meter buffer zone of the subway station; Number of bus stops: record the number of bus stops within the 800-meter buffer zone of each subway station; Road density: the ratio of the total length of all roads within the 800-meter buffer zone of a subway station to the total area of the buffer zone is measured; Building density: evaluate the ratio of the total ground floor area of all buildings within the 800-meter buffer zone of the subway station to the total area of the buffer zone; Number of POIs: Count the number of various points of interest within the 800-meter buffer zone of the subway station; Land use mix: quantifies the degree of integration of different land use types; Step 2.2: Data collection. For each variable in step 2.1, the following specific data collection and processing methods are used: Green space ratio: Data source: topographic maps released by the government or geographic information system (GIS) data provided by the urban planning department; processing method: only include the green areas of public spaces and roads, and exclude private green areas; area measurement and ratio calculation are performed using GIS software; Commercial land coverage rate: Data source: urban land use database or public urban planning documents; Processing method: According to the land use classification, extract the boundary information of commercial land, calculate its coverage area and convert it into a proportion value; Bus route density and number of bus stops: Data source: open source map data and station distribution information provided by local transportation management departments; Processing method: Use GIS tools to overlay bus routes and station locations on an 800-meter buffer zone of subway stations to calculate the total length of the routes and the number of stops; Road density: Data sources: open source map data, high-resolution remote sensing images, and urban road network databases provided by government departments; processing methods: draw a road network map, measure the total length of the road, and then divide it by the buffer area to obtain the density index; Building density: Data source: city building permit records or identification of building outlines through satellite images; processing method: use GIS technology to outline the ground floor boundaries of buildings, summarize the area and calculate the density ratio; Number of POIs: data source, online map services or other commercial data providers; Processing method, according to the predefined POI categories, search and count POIs of each category within 800 meters around the subway station; Land use mixture: data source, urban land use database or related research reports; The processing method refers to the calculation formula in the existing literature, adjusts the land use type classification and category according to the local actual situation, and finally obtains the mixed degree score; Step 2.3: Data cleaning and standardization; Outlier handling: Check and remove obviously erroneous data points; Missing value filling: For a small amount of missing data, use the average value of the neighboring area or other reasonable interpolation methods to fill it; Step 2.4: Variable correlation analysis; Preliminary screening: Based on theoretical assumptions and expert opinions, the variables most likely to affect the proportion of non-commuting trips are selected in advance; Correlation test: Use statistical methods to test the relationship between variables to avoid introducing highly correlated multicollinearity problems; Feature selection: The importance scoring function of the random forest model is used to further streamline the input feature set and retain the most predictive variables for subsequent modeling.
4. The method for calculating the proportion of non-commuting passengers at subway stations based on a random forest model according to claim 1, characterized in that: The step 3 is specifically as follows: Step 3.1: Model selection and parameter setting; The random forest regression algorithm is used here, and the specific parameters are set as follows: Number of decision trees n_estimators: set to the number of decision trees to ensure the stability and predictive power of the model; Maximum depth max_depth: no limit, allowing each tree to grow freely according to data characteristics to capture complex patterns; The number of features per split is max_features: that is, m randomly selected features are considered at each node split; Minimum number of sample splits min_samples_split: The default value ensures that there is enough data for each split decision; Minimum number of leaf node samples min_samples_leaf: default value, to prevent over-segmentation from causing overfitting; Step 3.2: Data preparation and division; Dataset division: The collected data is divided into training set and test set, with a ratio of 70% training set and 30% test set; Standardization / normalization: Standardize or normalize all input features so that variables of different scales can be compared at the same level to prevent certain features from dominating the model output due to their large numerical range; Encoding categorical variables: If there are categorical variables, they need to be converted to numerical form through one-hot encoding so that the model can handle them; Step 3.3: Model training; Initialize the random forest model: Create a random forest regressor instance based on the above parameters; Training model: Use the data in the training set to train the model; in this process, random forest will automatically extract multiple bootstrap samples from the training set and generate a decision tree for each subsample; Evaluate model performance: Use the test set to evaluate the performance of the model, including mean square error (MSE), mean absolute error (MAE), and coefficient of determination (R2); help understand the accuracy of the model prediction and its consistency with the actual data; Step 3.4: Model optimization and verification; Hyperparameter tuning: Use grid search or random search combined with cross-validation to find the optimal hyperparameter combination to improve the prediction accuracy of the model; Feature importance analysis: Based on the trained random forest model, the importance scores of each built environment variable are extracted to understand which factors have the most significant impact on the proportion of non-commuting trips; Enhanced model interpretability: SHAP analysis is introduced to quantify the specific contribution of each built environment feature to a single prediction result, providing a deeper understanding; Step 3.5: Model stability test; Multiple training and validation: Conduct multiple trainings and record the results of each training to ensure the consistency and reliability of the model performance; External validation: If there are additional data sources or time series data, the validity of the model can be re-tested through an external validation set to confirm whether its performance on the new data is still robust.
5. The method for calculating the proportion of non-commuting passengers at subway stations based on a random forest model according to claim 1, characterized in that: The step 4 is specifically as follows: Step 4.1: SHAP analysis; Shapley value: a concept derived from cooperative game theory, used to fairly distribute the contribution of each feature to model predictions; Global explanation: Calculate the average SHAP value of each feature over the entire dataset to show its overall importance; Local explanation: For a single prediction instance, it provides the specific contribution of each feature in the instance to help understand the behavior of the model in a specific situation. Step 4.2: Global feature importance evaluation; Mean SHAP value: draw a bar chart or column chart showing the average absolute contribution of each built environment variable to the prediction of the proportion of non-commuting trips; Bee swarm diagram: visualizes the SHAP value distribution of each data point, reflecting how changes in feature values affect a single prediction result; Step 4.3: Explore nonlinear relationships; Dependency graph: For variables that show significant nonlinear relationships, plot the relationship between their SHAP values and their own values; Step 4.4: Presentation and discussion of results; Graphical presentation: Combined with textual descriptions, intuitive graphical tools are used to present the main findings of the model, including but not limited to which built environment variables have the greatest impact on the proportion of non-commuting trips; the nonlinear relationship between variables and possible reasons; the unique contribution of certain variables under specific conditions; In-depth analysis of key factors: For factors with significant positive or negative impacts, the mechanisms behind them are explored in detail.
6. The method for calculating the proportion of non-commuting passengers at subway stations based on a random forest model according to claim 1, characterized in that: The step 5 is specifically as follows: Step 5.1: Result presentation: (1) Global feature importance analysis; Plot a bar chart: Use the mean SHAP value to show the global importance of each built environment variable on the prediction of non-commuting travel share; the bar chart is sorted by the importance of the features, starting with the most important feature; Explain key features: For the top features, provide detailed textual descriptions to explain why they have a significant impact on non-commuting trips; (2) Analysis of the impact of local features; Bee swarm diagram display: By drawing a bee swarm diagram, the specific contribution of each feature in a single sample is intuitively displayed; each point represents a data point, and its position on the x-axis reflects the impact of the feature on the prediction of that specific sample; Case study: select several representative subway stations and explore in depth how their local characteristics affect the proportion of non-commuting trips, using specific examples to help understand the model output; (3) Exploration of nonlinear relationships; Dependency graph: For features that show obvious nonlinear relationships, a segmented regression graph is drawn to clearly show the impact trend of changes in these feature values on the proportion of non-commuting trips; Threshold Effect Analysis: Identify and discuss key features that have threshold effects, i.e., when a feature exceeds or falls below a certain critical value, its impact on non-commuting travel will change significantly; Step 5.2: In-depth discussion of the influencing mechanisms of key factors; Commercial land coverage: Analyze how different levels of commercial land coverage promote or restrict non-commuting activities; Green space ratio: explores how the presence of green space generally promotes non-commuting travel, but why its positive effect tends to level off after the green space ratio increases to a certain limit; Number of POIs: In particular, we discuss how these POIs serve as centers of attraction for non-commuters and analyze the impact of their scale effect on non-commuter travel; Land use mix: Emphasizes how a good mix of land uses can improve accessibility and convenience for residents, thereby encouraging more non-motorized travel; Step 5.3: Propose policy recommendations; (1) Optimize the public transportation network to accommodate non-commuting travel needs; Taking into account the characteristics of non-commuting travel, it is recommended to adjust the subway service frequency and operating hours; by fine-tuning the proportion of non-commuting passengers, transportation managers have practical tools to predict passenger flow changes and optimize train arrangements and service facility configurations; (2) Promote urban traffic management based on data analysis; In order to achieve a more flexible and responsive operation strategy, an intelligent transportation system is introduced to assist decision makers in predicting passenger flow changes within a specific time period in advance, conducting effective traffic diversion and resource allocation, and ensuring that urban residents can enjoy high-quality public transportation services regardless of their travel purposes.
Citation Information
Patent Citations
Urban rail transit short-time passenger flow intelligent prediction method
CN117455038A
Subway peripheral non-commuter destination selection prediction method based on built environment
CN119338058A
Cited By
Intelligent senile syndrome scientific research data management method and system
CN120260940A
Construction method and application of prediction model for influence of data volume threshold effect on catalytic performance
CN120611621A
Urban rail transit passenger door selection information induction method and system based on evolutionary game
CN121526138A
Carbon emission prediction and threshold identification method based on space blocking random forest
CN121563017A
Urban crowd activity and space structure dynamic deduction method and system based on single traffic flow, terminal and storage medium
CN121997283A