A rail transit connection shared bicycle use influence factor analysis method based on built environment and interpretable machine learning
By using a built-environment-based and interpretable machine learning approach, this study identifies shared bicycle connection travel behavior, acquires multi-source data, and trains XGBoost and SHAP models. This addresses the issues of ambiguous research scope and insufficient data in the connection process between shared bicycles and rail transit, optimizes the analysis of the impact of the built environment on cycling, and improves connection efficiency.
Patent Information
- Application Number
- CN202510771377.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing technologies for studying the connection between shared bicycles and rail transit suffer from problems such as vague research scope, single data dimension, and insufficient model interpretation, resulting in unscientific research on the built environment and affecting the continuity of riding and connection efficiency.
This study employs a built-environment-based and interpretable machine learning approach. By identifying shared bicycle shuttle travel behavior, multi-source data is acquired, a database of influencing factors is constructed, and the XGBoost model is used for training. The SHAP model is then used to analyze the importance and nonlinear effects of built-environment factors, thereby deeply exploring their intrinsic relationship with shared bicycle use.
It has enabled a more scientific definition of the scope of built environment research, integrated multi-source data, improved the interpretability of the model and the operability of the strategy, and optimized the use effect of shared bicycles connecting rail transit.
Smart Images

Figure CN120688890B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban and rural transportation planning technology, specifically to a method for analyzing the impact of shared bicycle usage on rail transit connections based on the built environment and interpretable machine learning. Background Technology
[0002] With the acceleration of urbanization, the growth of the urban population has placed higher demands on urban public transportation systems. Urban rail transit, as a key component of the urban transportation network, plays a vital role in improving urban transportation efficiency. However, the limitations of its station coverage and network density have led to the widespread "last mile" problem. This challenge not only affects residents' travel experience but also limits the attractiveness and efficiency of rail transit.
[0003] To address this challenge, the bike-and-ride mode, which combines rail transit and bicycles, is considered an effective strategy to increase the share of rail transit and a key model for solving the "last mile" problem in urban rail transit. Bike sharing (BS), with its flexibility and convenience, has become an indispensable part of the urban public transportation system. Bike sharing not only effectively solves the "last mile" problem but also promotes the coordinated development of rail transit and green travel.
[0004] However, shared bicycles also face numerous challenges in connecting with rail transit, such as uneven spatial and temporal distribution, a lack of dedicated non-motorized vehicle lanes and parking facilities, and unreasonable built environment elements like inappropriate land use and public transportation infrastructure layout. These factors can all affect the continuity, comfort, and efficiency of cycling. This not only hinders the healthy development of shared bicycles but also impedes the full realization of the benefits of rail transit. Therefore, a deep understanding of the influencing factors of the built environment around rail transit stations on cycling, and the subsequent planning and design of a good cycling environment around stations to increase residents' willingness to cycle and the share of urban public transportation in travel, has become a hot topic in transportation research in recent years. Exploring the impact mechanism of the built environment around rail transit stations on shared bicycle travel through massive shared bicycle cycling data is of great significance for optimizing and upgrading current urban rail transit connection facilities.
[0005] According to the search, the Chinese patent database discloses patents such as application number 202311146139.2, invention title: "Method for assessing the impact of shared bicycle use by integrating spatial heterogeneity and nonlinearity" and application number 202310051935.1, invention title: "Analysis method for the influencing factors of subway station and dockless shared bicycle connection travel".
[0006] However, the above invention still has the following limitations:
[0007] 1. Vague research scope: Traditional methods rely on fixed radius buffer zones (such as 800 meters) to define the analysis scope, without considering urban density differences and the cumulative distribution characteristics of cycling behavior, resulting in unscientific spatial definition of built environment research;
[0008] 2. Limited data dimensions: Existing studies rely heavily on POI data and lack street view image data, making it difficult to quantify the impact of spatial quality (such as green space visibility and street interface continuity) on cycling behavior;
[0009] 3. Insufficient model explanation: Existing studies mostly use linear models such as linear regression and stepwise regression, which do not adequately consider the effects of nonlinearity. When quantifying the synergistic effect between variables, the influence of a certain indicator may be underestimated or overestimated. The assumption of a linear relationship may lead to incorrect estimates of the influence of explanatory variables in certain ranges, resulting in a lack of operability in the strategy recommendations.
[0010] In view of this, a method for analyzing the influencing factors of shared bicycle use in rail transit connections based on built environment and interpretable machine learning is provided to overcome the above-mentioned shortcomings; Summary of the Invention
[0011] The purpose of this invention is to address the shortcomings of existing technologies by proposing a method for analyzing the influencing factors of shared bicycle use in rail transit connections based on built environment and interpretable machine learning.
[0012] To achieve the above objectives, the present invention adopts the following technical solution: a method for analyzing the influencing factors of shared bicycle use in rail transit connections based on built environment and interpretable machine learning, comprising:
[0013] S1: Identify traffic behaviors that use shared bicycles as a connection to rail transit, obtain shared bicycle order data within the target area, and build a shared bicycle connection usage database;
[0014] S2: Determine the research area of the built environment around each subway station, and delineate the annular buffer zone around the subway station as the research scope;
[0015] S3: Integrates multi-source built environment data, including objective built environment indicators and perceived built environment indicators;
[0016] S4: Construct a dataset of factors influencing the built environment on the use of shared bicycles based on the data from S1 and S3;
[0017] S5: Train the XGBoost model and determine its performance;
[0018] S6: The SHAP interpretation model is used to analyze the importance ranking, nonlinear effects and interaction mechanisms of built environment factors;
[0019] The dataset of influencing factors mentioned in step S4 specifically includes:
[0020] S4.1: Count the number of shared bicycles used. Use the latitude and longitude information of the shared bicycles when they are first used as the location information of the shared bicycles. Calculate the average number of shared bicycles used at each subway station during the study period as the number of shared bicycles used at each subway station. The calculation is shown in Equation (1):
[0021] Equation (1)
[0022] in, Let $i$ be the average number of shared bicycles used within subway station $i$. is the number of shared bicycles used in subway station i on day n, and N is the total number of days in the research period;
[0023] S4.2: Calculate the population density within the buffer zone of each subway station and the population density index of the built environment characteristics, as shown in Equation (2):
[0024] Equation (2)
[0025] in, Population density within subway station i It represents the population within subway station i. Let i be the area of the buffer zone of subway station i;
[0026] The employment population density within the buffer zone of each subway station was statistically analyzed, and the employment density index of the built environment characteristics was calculated as shown in equation (3):
[0027] Equation (3)
[0028] in, The density of the employed population within subway station i. It represents the number of employed people within subway station i. Let i be the area of the buffer zone of subway station i;
[0029] S4.3: The land use mixing degree within the buffer zone of each subway station is statistically analyzed, and the land use mixing degree index of built environment walkability is calculated using the entropy index model, as shown in Equation (4):
[0030] Equation (4)
[0031] in, Let i be the land use entropy within subway station i. It is the proportion of the area of land use type j within subway station i to the total area. The total number of land types within subway station i;
[0032] S4.4: Statistically determine the job-housing balance within the buffer zone of each subway station and calculate the job-housing balance index of the built environment characteristics, as shown in Equation (5):
[0033] Equation (5)
[0034] in, To improve the work-life balance within subway station i, It represents the number of POIs for different employment types within subway station i. The POI value represents the number of residents within subway station i.
[0035] S4.5: Collect POI information within the buffer zone of each subway station, including the number of restaurants, entertainment facilities, transportation facilities, and commercial facilities;
[0036] S4.6: Compile road network data within the buffer zone of each subway station, including the length of main roads, secondary roads, branch roads, sidewalks, and bicycle lanes;
[0037] S4.7: Collect statistics on bus stop and subway station data within the buffer zone of each subway station, including the number of bus stops / subway stations and the distance to the nearest bus stop / subway station;
[0038] S4.8: Statistically determine the greening rate of streets within the buffer zone of each subway station and calculate the greening rate index of the perceived built environment, as shown in Equation (6):
[0039] Equation (6)
[0040] in, The greening rate of the residential and working streets within subway station i. This represents the pixel percentage of trees in the street scene surrounding subway station i. The pixel ratio of the shrubs in the street scene around subway station i;
[0041] The street enclosure degree within the buffer zone of each subway station is statistically analyzed, and the enclosure degree index of the perceived built environment is calculated, as shown in Equation (7):
[0042] Equation (7)
[0043] in, The degree of enclosed residential and work street area within subway station i. This represents the pixel percentage of the street scene buildings surrounding subway station i. This represents the pixel percentage of the street view walls surrounding subway station i. This represents the pixel percentage of trees in the street scene surrounding subway station i. This represents the pixel percentage of the street scene roads surrounding subway station i. The pixel percentage of the sidewalks in the street view surrounding subway station i. The pixel percentage of the street scene shrubs surrounding subway station i;
[0044] S4.9: Calculate the street complexity within the buffer zone of each subway station and the complexity index of the perceived built environment, as shown in Equation (8):
[0045] Equation (8)
[0046] in, The complexity of the residential and work streets within subway station i. The pixel ratio of people in the street scene around subway station i This represents the pixel percentage of street scene road signs surrounding subway station i. The pixel percentage of the street scene lighting around subway station i. This represents the pixel percentage of trees in the street scene surrounding subway station i. This represents the pixel percentage of the street scene roads surrounding subway station i. The pixel percentage of the street scene buildings surrounding subway station i;
[0047] S4.10: Calculate the street safety level within the buffer zone of each subway station and the perceived safety index of the built environment, as shown in Equation (9):
[0048] Equation (9)
[0049] in, To assess the safety of the residential and work-life neighborhood within subway station i, The pixel ratio of people in the street scene around subway station i This represents the pixel percentage of street scene road signs surrounding subway station i. The pixel percentage of the street scene lighting around subway station i. This represents the pixel percentage of the street scene fence surrounding subway station i.
[0050] Furthermore, in this invention, the specific method for identifying connecting travel behavior in step S1 includes: creating a buffer zone at each entrance and exit of the subway station, filtering and analyzing the parking data of shared bicycles, reconstructing the travel chain, and using the KD-Tree spatial search algorithm to identify connecting travel behavior.
[0051] Furthermore, in this invention, the method for defining the research scope in step S2 includes: performing kernel density estimation on the riding distance of shared bicycles connected to rail transit, and defining the research scope of riding connection using the walking and riding connection thresholds and the 85th percentile value.
[0052] Furthermore, in step S3, the objective built environment data collected from multiple sources includes population density / employment density, land use combination / job-housing balance, POIs information, road network data, bus station data, and subway station data; the perceived built environment data includes street greening rate, boundary enclosure degree, complexity, and safety degree.
[0053] Furthermore, in step S5 of this invention, when training the XGBoost model, the multi-source dataset is randomly divided into a training set and a test set, a gradient boosting decision tree is used to train the model, and the performance of the prediction model is evaluated by mean square error, root mean square error, mean absolute error, and coefficient of determination.
[0054] Furthermore, in step S6, when using the SHAP interpretation model to analyze the importance ranking, nonlinear effects, and interaction mechanisms of built environment factors, the invention includes calculating the importance metric value of each feature and outputting a feature dependency graph.
[0055] Furthermore, in this invention, the specific prediction step of step S5 is as follows:
[0056] S5.1: Randomly divide the multi-source dataset into two parts for model testing: 70% of the training set and 30% of the test set;
[0057] S5.2: Using gradient-enhanced decision trees, this method achieves ensemble learning based on decision trees. The gradient enhancement strategy returns a more robust prediction model by iterating through multiple weak learners, steadily improving the model's accuracy.
[0058] S5.3: Model comparison experiments were conducted using the XGBoost framework to implement gradient boosting decision trees using an efficient parallel training algorithm. The objective function of the XGBoost algorithm is shown in Equation (10):
[0059] Equation (10)
[0060] in, It is a loss function. These are regularization terms, including L1 regularization terms and L2 regularization terms; It is a constant term;
[0061] S5.4: The performance of the prediction model is evaluated using four indicators: mean square error, root mean square error, mean absolute error, and coefficient of determination R².
[0062] 8. The method for analyzing the influencing factors of shared bicycle use in rail transit connections based on built environment and interpretable machine learning according to claim 1, characterized in that: the specific prediction steps of step S6 are as follows:
[0063] S6.1: The SHAP method is used to attribute the changes in the gradient boosting decision tree model to each independent variable. In a set of independent variables, the difference in prediction of this independent variable is the marginal contribution. The Shapley value can be used to estimate the marginal contribution effect of each feature. The importance quantification value of the independent variable is obtained by statistically analyzing the Shapley value of each independent variable. The calculation method is shown in the following formula (11):
[0064] Equation (11)
[0065] In the formula, Representation of features Contributions; Indicates the number of features in the set; This is the output of the original model; Let p be any subset of features that does not contain the feature p to be computed. ; In All in the non-zero item subset vector;
[0066] S6.2: Output the summary distribution of independent variables with high importance quantification values, identify the relationship between the value of each independent variable and its impact on the prediction based on the Shapley value; at the same time, output the feature dependency graph, consider the interaction of multiple features, and deeply explore the intrinsic relationship between the built environment and the use of shared bicycles.
[0067] The present invention has the following beneficial effects:
[0068] In this invention, by identifying traffic behaviors that use shared bicycles as a connection to rail transit, shared bicycle order data within the target area is obtained, and riding origin-destination (OD) points, distances, and time information are extracted to construct a shared bicycle connection usage database;
[0069] In this invention, the research area of the built environment around each subway station is determined. In order to more scientifically define the research scope of the built environment, a heuristic idea is proposed to delineate the research scope of the annular buffer zone around the subway station based on the 85th percentile of the cumulative distribution of cycling distance and the walking connection range.
[0070] In this invention, multi-source built environment data is integrated, including objective built environment indicators (density, diversity, road network design, public transport facilities and destination accessibility) and perceived built environment indicators (street greening rate, enclosure degree, complexity and safety indicators) to ensure that the research indicators are comprehensive and effective.
[0071] In this invention, a dataset of factors influencing the use of shared bicycles by the built environment is constructed based on the data from steps 1 and 3;
[0072] In this invention, the XGboost model is used for training, optimized on the validation set, and tested on the test set to determine the model performance.
[0073] In this invention, the SHAP interpretation model is used to analyze the importance ranking, nonlinear effects and interaction mechanisms of built environment factors, and to propose more specific urban planning and connectivity optimization strategies. Attached Figure Description
[0074] Figure 1 This is a schematic diagram of the overall structure of the system in this invention;
[0075] Figure 2 This is a diagram illustrating the connection of the system in this invention;
[0076] Figure 3 This is a distribution map of cycling distances in this invention;
[0077] Figure 4 This is a flowchart illustrating the steps of street view recognition in this invention;
[0078] Figure 5 This is a schematic diagram illustrating the importance ranking of built environment factors in machine learning in this invention;
[0079] Figure 6 This is a nonlinear relationship diagram of the connection between main roads and traffic facility POIs and shared bicycles used in the machine learning of this invention.
[0080] Figure 7 This is a schematic diagram of the built-in environment interaction results of the machine learning part of this invention. Detailed Implementation
[0081] With the acceleration of urbanization, urban rail transit, as a key component of urban transportation networks, plays a crucial role in improving urban transportation efficiency. However, limitations in the coverage and network density of rail transit stations have led to the widespread "last mile" problem. To address this issue, the bike-and-ride model has emerged, combining rail transit and bicycles. Among these, shared bicycles, with their flexibility and convenience, have become key to solving the "last mile" problem. However, shared bicycles also face numerous challenges in connecting with rail transit, such as uneven spatial and temporal distribution, a lack of dedicated non-motorized vehicle lanes and parking facilities, all of which can affect the continuity, comfort, and efficiency of the ride. Therefore, this invention proposes a method for analyzing the influencing factors of shared bicycle use in rail transit connections based on the built environment and interpretable machine learning. This aims to deeply understand the influencing factors of the current built environment around rail transit stations on cycling, and then plan and design a favorable cycling environment around stations to increase residents' willingness to cycle and the share of urban public transportation in the overall travel experience.
[0082] A method for analyzing the influencing factors of shared bicycle use in rail transit connections based on built environment and interpretable machine learning, the analysis method including:
[0083] S1: Identify traffic behavior that uses shared bicycles as a connection to rail transit, obtain shared bicycle order data within the target area, extract riding origin-destination (OD) points, distance, and time information, and construct a shared bicycle connection usage database;
[0084] The specific method for identification in step S1 is as follows:
[0085] S1.1: Create a 100-meter buffer zone at each entrance and exit of the subway station to filter and analyze the parking data of shared bicycles;
[0086] S1.2: Reconstruct the travel chain of shared bicycle parking data and construct a travel OD data table;
[0087] S1.3: The KD-Tree (k-dimensional Tree) spatial search algorithm is used to perform nearest neighbor search on each origin and destination in the travel OD data table to find the nearest subway station, identify connecting travel behavior, and calculate its spatial distance to the subway station.
[0088] S1.4: Append the matched subway station IDs and distance information to the OD data table to form a complete connecting travel dataset;
[0089] S2: Determine the study area of the built environment around each subway station. Based on the 85th percentile of the cumulative distribution of cycling distance and the walking connection range, delineate a 500-1800 meter ring buffer zone around the subway station as the study area.
[0090] The method for defining the scope of study in step S2 is as follows:
[0091] S2.1: Perform kernel density estimation on the riding distance of shared bicycles connecting to rail transit and generate a cumulative distribution function graph of riding distance;
[0092] S2.2: Use walking and cycling connection thresholds to define the lower limit of the research scope of cycling connection;
[0093] S2.3: The 85th percentile value is a key indicator for defining the upper limit of the research scope of cycling shuttle services;
[0094] S3: Integrates multi-source built environment data, including objective built environment indicators (POI mix, road network density) and perceived built environment indicators (street greening rate, street openness);
[0095] The objective built environment data in the multi-source dataset mentioned in step S3 includes population density / employment density, land use combination / job-housing balance, POI information, road network data, bus station data, and subway station data; the perceived built environment data includes street greening rate, boundary enclosure degree, complexity, and safety degree.
[0096] The multi-source built environment data acquisition method described in step S3 specifically includes:
[0097] S3.1: Select objective built environment based on the "5Ds" built environment, namely Density, Diversity, Road Network Design, Distance to Transit, and Destination Accessibility; also including subway station-specific attribute indicators;
[0098] S3.2: The objective built environment adopts multi-source big data, including Word Pop data, OpenStreetMap data, and Baidu Map data;
[0099] S3.3: Extract perceived built environment data based on Baidu Street View data; specifically, based on road network data, a sample point is set at every 50 m interval on each road, and street view images are acquired in four directions (0°, 90°, 180°, 270°) at each sample point;
[0100] S3.4: Perform semantic image segmentation using deep learning technology, calculate the pixel percentage in each street view image; then take the average of the pixel percentages in the street view images of a sample point in four directions to obtain the final pixel percentage; and calculate the street greening rate, enclosure degree, complexity, and safety index.
[0101] S4: Construct a dataset of factors influencing the use of shared bicycles based on the data from steps S1 and S3; the dataset of influencing factors mentioned in step S4 specifically includes:
[0102] S4.1: Count the number of shared bicycles used. Use the latitude and longitude information of the shared bicycles when they are first used as the location information of the shared bicycles. Calculate the average number of shared bicycles used at each subway station during the study period as the number of shared bicycles used at each subway station. The calculation is shown in Equation (1):
[0103] Equation (1)
[0104] in, Let $i$ be the average number of shared bicycles used within subway station $i$. is the number of shared bicycles used in subway station i on day n, and N is the total number of days in the research period;
[0105] S4.2: Calculate the population density within the buffer zone of each subway station and the population density index of the built environment characteristics, as shown in Equation (2):
[0106] Equation (2)
[0107] in, Population density within subway station i It represents the population within subway station i. Let i be the area of the buffer zone of subway station i;
[0108] The employment population density within the buffer zone of each subway station was statistically analyzed, and the employment density index of the built environment characteristics was calculated as shown in equation (3):
[0109] Equation (3)
[0110] in, The density of the employed population within subway station i. It represents the number of employed people within subway station i. Let i be the area of the buffer zone of subway station i;
[0111] S4.3: The land use mixing degree within the buffer zone of each subway station is statistically analyzed, and the land use mixing degree index of built environment walkability is calculated using the entropy index model, as shown in Equation (4):
[0112] Equation (4)
[0113] in, Let i be the land use entropy within subway station i. It is the proportion of the area of land use type j within subway station i to the total area. The total number of land types within subway station i;
[0114] S4.4: Statistically determine the job-housing balance within the buffer zone of each subway station and calculate the job-housing balance index of the built environment characteristics, as shown in Equation (5):
[0115] Equation (5)
[0116] in, To improve the work-life balance within subway station i, It represents the number of POIs for different employment types within subway station i. The POI value represents the number of residents within subway station i.
[0117] S4.5: Collect POI information within the buffer zone of each subway station, including the number of restaurants, entertainment facilities, transportation facilities, and commercial facilities;
[0118] S4.6: Compile road network data within the buffer zone of each subway station, including the length of main roads, secondary roads, branch roads, sidewalks, and bicycle lanes;
[0119] S4.7: Collect statistics on bus stop and subway station data within the buffer zone of each subway station, including the number of bus stops / subway stations and the distance to the nearest bus stop / subway station;
[0120] S4.8: Statistically determine the greening rate of streets within the buffer zone of each subway station and calculate the greening rate index of the perceived built environment, as shown in Equation (6):
[0121] Equation (6)
[0122] in, The greening rate of the residential and working streets within subway station i. This represents the pixel percentage of trees in the street scene surrounding subway station i. The pixel ratio of the shrubs in the street scene around subway station i;
[0123] The street enclosure degree within the buffer zone of each subway station is statistically analyzed, and the enclosure degree index of the perceived built environment is calculated, as shown in Equation (7):
[0124] Equation (7)
[0125] in, The degree of enclosed residential and work street area within subway station i. This represents the pixel percentage of the street scene buildings surrounding subway station i. This represents the pixel percentage of the street view walls surrounding subway station i. This represents the pixel percentage of trees in the street scene surrounding subway station i. This represents the pixel percentage of the street scene roads surrounding subway station i. The pixel percentage of the sidewalks in the street view surrounding subway station i. The pixel percentage of the street scene shrubs surrounding subway station i;
[0126] S4.9: Calculate the street complexity within the buffer zone of each subway station and the complexity index of the perceived built environment, as shown in Equation (8):
[0127] Equation (8)
[0128] in, The complexity of the residential and work streets within subway station i. The pixel ratio of people in the street scene around subway station i This represents the pixel percentage of street scene road signs surrounding subway station i. The pixel percentage of the street scene lighting around subway station i. This represents the pixel percentage of trees in the street scene surrounding subway station i. This represents the pixel percentage of the street scene roads surrounding subway station i. The pixel percentage of the street scene buildings surrounding subway station i;
[0129] S4.10: Calculate the street safety level within the buffer zone of each subway station and the perceived safety index of the built environment, as shown in Equation (9):
[0130] Equation (9)
[0131] in, To assess the safety of the residential and work-life neighborhood within subway station i, The pixel ratio of people in the street scene around subway station i This represents the pixel percentage of street scene road signs surrounding subway station i. The pixel percentage of the street scene lighting around subway station i. This represents the pixel percentage of the street scene fence surrounding subway station i.
[0132] S5: Train the XGBoost model, fine-tune it on the validation set, and test it on the test set to determine the model performance;
[0133] The specific prediction steps in step S5 are as follows:
[0134] S5.1: Randomly divide the multi-source dataset into two parts for model testing: 70% of the training set and 30% of the test set;
[0135] S5.2: Using gradient-enhanced decision trees, this method achieves ensemble learning based on decision trees. The gradient enhancement strategy returns a more robust prediction model by iterating through multiple weak learners, steadily improving the model's accuracy.
[0136] S5.3: Model comparison experiments were conducted using the XGBoost framework to implement gradient boosting decision trees using an efficient parallel training algorithm. The objective function of the XGBoost algorithm is shown in Equation (10):
[0137] Equation (10)
[0138] in, It is a loss function. These are regularization terms, including L1 regularization terms and L2 regularization terms; It is a constant term;
[0139] S5.4: The performance of the prediction model is evaluated using four indicators: mean square error, root mean square error, mean absolute error, and coefficient of determination R².
[0140] S6: The SHAP interpretation model is used to analyze the importance ranking, nonlinear effects and interaction mechanisms of built environment factors, and to generate connection optimization strategies.
[0141] The specific prediction steps for step S6 are as follows:
[0142] S6.1: The SHAP method is used to attribute the changes in the gradient boosting decision tree model to each independent variable. In a set of independent variables, the difference in prediction of this independent variable is the marginal contribution. The Shapley value can be used to estimate the marginal contribution effect of each feature. The importance quantification value of the independent variable is obtained by statistically analyzing the Shapley value of each independent variable. The calculation method is shown in the following formula (11):
[0143] Equation (11)
[0144] In the formula, Representation of features Contributions; Indicates the number of features; This is the output of the original model; for The set does not contain any feature subset of the feature p to be computed, i.e. It is a subset, not a quantity; In All in the non-zero item subset vector;
[0145] S6.2: Output the summary distribution of independent variables with high importance quantification values, identify the relationship between the value of each independent variable and its impact on the prediction based on the Shapley value; at the same time, output the feature dependency graph, consider the interaction of multiple features, and deeply explore the intrinsic relationship between the built environment and the use of shared bicycles.
[0146] This invention proposes a method for analyzing the influencing factors of shared bicycle use in rail transit connections based on built environment and interpretable machine learning. This method achieves a comprehensive analysis of the impact of the built environment surrounding rail transit stations on shared bicycle use by identifying shared bicycle connection behaviors, defining the study area, fusing multi-source built environment data, constructing an influencing factor dataset, using an XGBoost model for training and testing, and employing a SHAP interpretive model to analyze the importance ranking, nonlinear effects, and interaction mechanisms of built environment factors.
[0147] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for analyzing the influencing factors of shared bicycle use in rail transit connections based on built environment and interpretable machine learning, characterized in that, include: S1: Identify traffic behaviors that use shared bicycles as a connection to rail transit, obtain shared bicycle order data within the target area, and build a shared bicycle connection usage database; S2: Determine the research area of the built environment around each subway station, and delineate the annular buffer zone around the subway station as the research scope; S3: Integrates multi-source built environment data, including objective built environment indicators and perceived built environment indicators; S4: Construct a dataset of factors influencing the built environment on the use of shared bicycles based on the data from S1 and S3; S5: Train the XGBoost model and determine its performance; S6: The SHAP interpretation model is used to analyze the importance ranking, nonlinear effects and interaction mechanisms of built environment factors; The dataset of influencing factors mentioned in step S4 specifically includes: S4.1: Count the number of shared bicycles used. Use the latitude and longitude information of the shared bicycles when they are first used as the location information of the shared bicycles. Calculate the average number of shared bicycles used at each subway station during the study period as the number of shared bicycles used at each subway station. The calculation is shown in Equation (1): Equation (1); in, Let $i$ be the average number of shared bicycles used within subway station $i$. is the number of shared bicycles used in subway station i on day n, and N is the total number of days in the research period; S4.2: Calculate the population density within the buffer zone of each subway station and the population density index of the built environment characteristics, as shown in Equation (2): Equation (2); in, Population density within subway station i It represents the population within subway station i. Let i be the area of the buffer zone of subway station i; The employment population density within the buffer zone of each subway station was statistically analyzed, and the employment density index of the built environment characteristics was calculated as shown in equation (3): Equation (3); in, The density of the employed population within subway station i. It represents the number of employed people within subway station i. Let i be the area of the buffer zone of subway station i; S4.3: The land use mixing degree within the buffer zone of each subway station is statistically analyzed, and the land use mixing degree index of built environment walkability is calculated using the entropy index model, as shown in Equation (4): Equation (4); in, Let i be the land use entropy within subway station i. It is the proportion of the area of land use type j within subway station i to the total area. The total number of land types within subway station i; S4.4: Statistically analyze the job-housing balance within the buffer zone of each subway station and calculate the job-housing balance index of the built environment characteristics, as shown in Equation (5): Equation (5); in, To improve the work-life balance within subway station i, It represents the number of POIs for different employment types within subway station i. The POI value represents the number of residents within subway station i. S4.5: Collect POI information within the buffer zone of each subway station, including the number of restaurants, entertainment facilities, transportation facilities, and commercial facilities; S4.6: Compile road network data within the buffer zone of each subway station, including the length of main roads, secondary roads, branch roads, sidewalks, and bicycle lanes; S4.7: Collect statistics on bus stop and subway station data within the buffer zone of each subway station, including the number of bus stops / subway stations and the distance to the nearest bus stop / subway station; S4.8: Statistically determine the greening rate of streets within the buffer zone of each subway station and calculate the greening rate index of the perceived built environment, as shown in Equation (6): Equation (6); in, The greening rate of the residential and working streets within subway station i. This represents the pixel percentage of trees in the street scene surrounding subway station i. The pixel ratio of the shrubs in the street scene around subway station i; The street enclosure degree within the buffer zone of each subway station is statistically analyzed, and the enclosure degree index of the perceived built environment is calculated, as shown in Equation (7): Equation (7); in, The degree of enclosed residential and work street area within subway station i. This represents the pixel percentage of the street scene buildings surrounding subway station i. This represents the pixel percentage of the street view walls surrounding subway station i. This represents the pixel percentage of trees in the street scene surrounding subway station i. This represents the pixel percentage of the street scene roads surrounding subway station i. The pixel percentage of the sidewalks in the street view surrounding subway station i. The pixel percentage of the street scene shrubs surrounding subway station i; S4.9: Calculate the street complexity within the buffer zone of each subway station and the complexity index of the perceived built environment, as shown in Equation (8): Equation (8); in, The complexity of the residential and work streets within subway station i. The pixel ratio of people in the street scene around subway station i This represents the pixel percentage of street scene road signs surrounding subway station i. The pixel percentage of the street scene lighting around subway station i. This represents the pixel percentage of trees in the street scene surrounding subway station i. This represents the pixel percentage of the street scene roads surrounding subway station i. The pixel percentage of the street scene buildings surrounding subway station i; S4.10: Calculate the street safety level within the buffer zone of each subway station and the perceived safety index of the built environment, as shown in Equation (9): Equation (9); in, To assess the safety of the residential and work-life neighborhood within subway station i, The pixel ratio of people in the street scene around subway station i This represents the pixel percentage of street scene road signs surrounding subway station i. The pixel percentage of the street scene lighting around subway station i. This represents the pixel percentage of the street scene fence surrounding subway station i.
2. The method for analyzing the influencing factors of shared bicycle use in rail transit connections based on built environment and interpretable machine learning as described in claim 1, characterized in that, The specific methods for identifying connecting travel behavior in step S1 include: creating buffer zones at each entrance and exit of the subway station, filtering and analyzing parking data of shared bicycles, reconstructing the travel chain, and using the KD-Tree spatial search algorithm to identify connecting travel behavior.
3. The method for analyzing the influencing factors of shared bicycle use in rail transit connections based on built environment and interpretable machine learning as described in claim 1, characterized in that, The method for defining the research scope in step S2 includes: performing kernel density estimation on the riding distance of shared bicycles connected to rail transit, and using the walking and riding connection thresholds and the 85th percentile value to define the research scope of riding connection.
4. The method for analyzing the influencing factors of shared bicycle use in rail transit connections based on built environment and interpretable machine learning as described in claim 1, characterized in that, In step S3, the objective built environment data collected from multiple sources includes population density / employment density, land use combination / job-housing balance, POI information, road network data, bus station data, and subway station data; the perceived built environment data includes street greening rate, boundary enclosure degree, complexity, and safety degree.
5. The method for analyzing the influencing factors of shared bicycle use in rail transit connections based on built environment and interpretable machine learning according to claim 1, characterized in that, In step S5, when training the XGBoost model, the process includes randomly dividing the multi-source dataset into training and test sets, using gradient boosting decision trees to train the model, and evaluating the performance of the prediction model using mean squared error, root mean square error, mean absolute error, and coefficient of determination.
6. The method for analyzing the influencing factors of shared bicycle use in rail transit connections based on built environment and interpretable machine learning according to claim 1, characterized in that, In step S6, when using the SHAP interpretation model to analyze the importance ranking, nonlinear effects and interaction mechanisms of built environment factors, the important metric value of each feature is calculated and the feature dependency graph is output.
7. The method for analyzing the influencing factors of shared bicycle use in rail transit connections based on built environment and interpretable machine learning as described in claim 1, characterized in that: The specific prediction steps in step S5 are as follows: S5.1: Randomly divide the multi-source dataset into two parts for model testing: 70% of the training set and 30% of the test set; S5.2: Using gradient-enhanced decision trees, this method achieves ensemble learning based on decision trees. The gradient enhancement strategy returns a more robust prediction model by iterating through multiple weak learners, steadily improving the model's accuracy. S5.3: Model comparison experiments were conducted using the XGBoost framework to implement gradient boosting decision trees using an efficient parallel training algorithm. The objective function of the XGBoost algorithm is shown in Equation (10): Equation (10); in, It is a loss function. These are regularization terms, including L1 regularization terms and L2 regularization terms; It is a constant term; S5.4: The performance of the prediction model is evaluated using four indicators: mean square error, root mean square error, mean absolute error, and coefficient of determination R².
8. The method for analyzing the influencing factors of shared bicycle use in rail transit connections based on built environment and interpretable machine learning according to claim 1, characterized in that: The specific prediction steps for step S6 are as follows: S6.1: The SHAP method is used to attribute the changes in the gradient boosting decision tree model to each independent variable. In a set of independent variables, the difference in prediction of this independent variable is the marginal contribution. The Shapley value can be used to estimate the marginal contribution effect of each feature. The importance quantification of the independent variable is obtained by statistically analyzing the Shapley value of each independent variable. S6.2: Output the summary distribution of independent variables with high importance quantification values, identify the relationship between the value of each independent variable and its impact on the prediction based on the Shapley value; at the same time, output the feature dependency graph, consider the interaction of multiple features, and deeply explore the intrinsic relationship between the built environment and the use of shared bicycles.
Citation Information
Patent Citations
Shared bicycle use influence assessment method fusing spatial heterogeneity and nonlinearity
CN117172574A
Influence factor analysis method for connection trip of subway station and pile-free shared bicycle
CN116542419A