A Trip Generation Prediction Method and System Based on Geographically Weighted Random Forest
Patent Information
- Application Number
- CN202610732635.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-09-11
AI Technical Summary
[0008]综上所述,现有出行生成预测方法在以下方面存在明显不足:一是未能实现非线性关系与空间异质性的统一建模,导致预测精度受限;二是距离衰减效应的处理不充分或不恰当,地理权重未能与非线性模型深度融合;三是缺乏对非线性影响关系的可解释性分析,难以提供阈值、边际效应等可操作的规划依据;四是系统性全方式出行生成预测研究匮乏,难以满足城市交通规划从源头优化需求的实际需要
Smart Images

Figure CN122736004A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of urban traffic planning and artificial intelligence technology, specifically to a method and system for travel generation and prediction based on geographically weighted random forest. Background Technology
[0002] Trip generation forecasting is a core component of the four-stage urban transportation planning approach, aiming to estimate the number of trips generated and attracted within a specific spatial unit (such as a traffic analysis zone, TAZ) per unit of time. Accurate trip generation forecasting is fundamental for subsequent trip distribution, mode classification, and traffic allocation, and plays a crucial supporting role in traffic demand forecasting, land use planning, and transportation policy formulation. Traditional trip generation forecasting methods generally employ conventional statistical models such as linear regression and Poisson regression, predicting trip rates by establishing functional relationships between trip rates and built environment elements quantified based on a five-dimensional ("5Ds") or seven-dimensional ("7Ds") framework (such as density, mixed-use development, destination accessibility, transit proximity, neighborhood design, demand management, and demographic characteristics).
[0003] In recent years, with the development of machine learning technology, researchers have begun to apply nonlinear models such as Random Forest (RF) and Gradient Boosting Decision Tree (GBDT) to travel generation prediction. These methods can capture the nonlinear relationship between the built environment and travel demand (such as threshold effects, marginal effects, and non-monotonicity). Meanwhile, spatial models such as Geographically Weighted Regression (GWR) have been introduced into travel generation prediction to address the spatial heterogeneity of the built environment's influence, i.e., the impact of the same built environment characteristic variable on travel demand varies across different regions. Furthermore, some studies have attempted to combine interpretable machine learning techniques (such as SHAP and partially dependent graphs) to interpret travel generation prediction models, revealing the direction and strength of the influence of various built environment variables.
[0004] Despite the progress made by the above methods in their respective fields, the existing technologies still have the following key problems: (1) Modeling of nonlinearity and spatial heterogeneity Most existing studies explore nonlinearity or spatial heterogeneity in isolation, failing to simultaneously consider the coexistence of the two. On the one hand, nonlinear models (such as RF and GBDT) assume spatially stationary influence relationships, failing to reveal the differences in the impact of built environment factors across different regions. On the other hand, while traditional spatial models (such as GWR) can supplement the spatial heterogeneity of built environment influences, they are based on linear assumptions and struggle to effectively capture complex nonlinear characteristics. In reality, the impact of the built environment on travel demand often exhibits both nonlinearity and spatial heterogeneity; their coexistence is an inherent and complex characteristic of traffic and travel generation.
[0005] (2) The distance attenuation effect is not adequately or properly handled. Existing spatial nonlinear hybrid models have shortcomings in handling distance decay effects: some techniques (such as CN117874709A) do not introduce distance decay weights at all, only using radius to filter samples with equal weights; another type of technique (such as CN117035066A) introduces weights, but only at a shallow weighting level at the sample input level, without embedding geographical weights into the splitting process of nodes within the decision tree, and its kernel function and correction factor are designed for remote sensing specific domains, making it unsuitable for urban traffic generation analysis. These shortcomings cause the models to fail to fully comply with the first law of geography, affecting the representational ability of local models and the modeling accuracy of spatial heterogeneity.
[0006] (3) Lack of interpretability analysis for nonlinear relationships While existing nonlinear models can improve prediction accuracy, they generally suffer from the "black box" problem, making it difficult to intuitively explain the degree, direction, threshold range, and effective interval of influence of built environment characteristic variables. Some explanatory methods (such as partial dependency graphs) suffer from biased estimation problems and cannot handle the interference caused by the correlation between variables. SHAP models can only output discrete scatter plots. None of these methods are well-suited for interpreting continuous relationships that better reflect reality.
[0007] (4) Insufficient systematic modeling of all modes of travel generation Existing research largely focuses on single modes of transportation (such as shared bicycles, subways, and ride-hailing services), lacking systematic modeling of overall travel generation encompassing all modes of transportation, and its application scenarios are limited. From the perspective of urban and transportation planning practice, the total traffic generation including all modes of transportation deserves more attention. The analysis results can inspire a more strategic approach to reducing travel demand or distance at the source, but this direction has not yet been fully explored.
[0008] In summary, existing travel generation prediction methods have significant shortcomings in the following aspects: First, they fail to achieve unified modeling of nonlinear relationships and spatial heterogeneity, resulting in limited prediction accuracy; second, the distance decay effect is not adequately or appropriately handled, and geographical weights are not deeply integrated with the nonlinear model; third, they lack interpretable analysis of nonlinear influence relationships, making it difficult to provide actionable planning basis such as thresholds and marginal effects; and fourth, systematic research on all-modal travel generation prediction is scarce, making it difficult to meet the actual needs of urban transportation planning for source optimization. Therefore, there is an urgent need for a travel generation prediction method and system that can integrate geographical weighting mechanisms and the advantages of random forests, possessing both high accuracy and interpretability. Summary of the Invention
[0009] To address the shortcomings of existing technologies, this invention aims to provide a travel generation prediction method and system that simultaneously considers nonlinearity and spatial heterogeneity, incorporates distance attenuation effects, and combines high accuracy with interpretability.
[0010] To achieve the above objectives, the present invention provides the following technical solution: A trip generation and prediction method based on geographically weighted random forest includes the following steps: The study area is divided into multiple spatial analysis units. The travel rate of each spatial analysis unit is obtained as the dependent variable. Multiple built environment characteristic variables of each spatial analysis unit are obtained as independent variables. The dependent variable and the independent variables constitute the training sample, and the independent variables of the spatial analysis unit to be predicted constitute the prediction sample. A local random forest model is constructed for each target spatial analysis unit, and the geographic weight vector of the target unit is calculated. The geographic weight vector is used as the sample weight of the corresponding training sample to obtain a training sample set with geographic weights. The geographic weighted variance impurity is calculated based on the geographic weight vector, and node splitting is performed with the criterion of maximizing its reduction. The local random forest models of all spatial analysis units are integrated to obtain a geographic weighted random forest model. Using cross-validation and grid search methods, the hyperparameters of the geographic weighted random forest model are optimized based on the error evaluation criterion to obtain the optimal geographic weighted random forest model. For the predicted sample, the local random forest model corresponding to its spatial analysis unit is called to generate travel predictions and output the prediction results.
[0011] Furthermore, the travel rate is the total daily travel volume divided by the area of the corresponding spatial analysis unit.
[0012] Furthermore, the built environment characteristic variables are quantified based on a multi-dimensional indicator system, which includes seven dimensions: development density, functional mix, destination accessibility, public transport proximity, neighborhood design, demand management, and population characteristics.
[0013] Furthermore, the built environment characteristic variables include: Development density: residential land density, business office land density, commercial service land density, industrial land density, administrative land density, educational land density, medical and health land density, sports and cultural land density, and park green space density; Functional Mixing: Land Use Mixing Index; Achievable destination: Distance to the central business district; Public transport proximity: density of areas accessible by bus stops, density of areas accessible by subway stations, distance to the nearest integrated transportation hub, distance to the nearest train station, distance to the nearest long-distance bus station, and distance to the nearest airport; Street layout design: density of highways and expressways, density of arterial roads, density of secondary arterial roads, density of local roads, density of intersections, proportion of signalized intersections, proportion of crossroads, proximity to centrality; Demand management: density of roadside parking spaces, density of public parking lots; Population characteristics: population density, male population ratio, youth population ratio, middle-aged population ratio.
[0014] Furthermore, for the target space analysis unit The geographic weight vector is represented as:
[0015] in, Target space analysis unit The geographic weight vector, Representing all other spatial analysis units respectively Target space analysis unit The geographic weights are calculated using an adaptive bisquare kernel function based on geographic spatial distance:
[0016] in, It is the target space analysis unit Other spatial analysis units j Spatial distance, It is adaptive bandwidth, expressed as the bandwidth from the target space analysis unit. To its first The distance to the nearest neighbor, i.e. .
[0017] Furthermore, for the target space analysis unit The geographic weight vector is introduced when training its local random forest model. Specifically: In the target space analysis unit In the local random forest model, the impurity of the leaf nodes of its internal decision trees is calculated by the geographically weighted variance impurity, and the formula for calculating the geographically weighted variance impurity is as follows:
[0018] in, Indicates the impurity of geographically weighted variance. Represents leaf nodes The sample set, where the sample elements are represented as , , For set The cardinality; Representing sets Samples in Built environment characteristic variables, travel generation volume, and sample Target space analysis unit Geographic weight; It is a set The mean of trip generation for all samples; for each feature of a built environment And a spatial bisection threshold for this feature Candidate splits , will set Divided into two subspaces, namely the left space and right space ,in Indicates sample In terms of built environment characteristics The value; The node splitting criterion is: maximizing the reduction in geographically weighted variance impurity. ;
[0019] in, Indicated based on the splitting criterion Partitioned Sets The reduction in geographically weighted variance impurity was obtained; and They represent criteria respectively The divided left space and right space Impurity; and These represent the geometric weights based on the impurities of the left and right spaces of the samples they contain, respectively.
[0020] Furthermore, the hyperparameters are optimized using leave-one-out cross-validation and grid search, employing a two-stage strategy: In the first stage, a global random forest model is constructed based on the entire study area without considering geographical weights, and the optimal hyperparameters of the global random forest model are determined based on the root mean square error, including the number of sub-models, the proportion of randomly selected features, and the minimum number of samples for tree node splitting; In the second stage, the optimal hyperparameters from the first stage are applied to all local random forest models, and the optimal number of nearest neighbors is searched.
[0021] Furthermore, this method also includes: The performance of the geographically weighted random forest model was evaluated using error evaluation metrics and spatial autocorrelation metrics. Calculate the global and / or local relative importance of each built environment characteristic variable, and / or calculate the global and / or local cumulative effects, and output interpretable analysis results. Furthermore, the relative feature importance is calculated by accumulating the reduction in impurity of each built environment feature variable at the decision tree split node; the global relative feature importance is obtained by averaging and normalizing the relative feature importance of all local random forest models; the local relative feature importance is calculated separately for each local random forest model. The cumulative local effect map is calculated through the following steps: dividing the quantile distribution range of the selected feature into multiple intervals; calculating the local effect in each interval; accumulating and centering all local effects; and calculating the local cumulative local effect map for each local random forest model separately.
[0022] On the other hand, the present invention also provides a travel generation prediction system based on geographically weighted random forest, comprising: The data acquisition and preprocessing module is used to divide the study area into multiple spatial analysis units, obtain the travel rate of each spatial analysis unit as the dependent variable, obtain multiple built environment characteristic variables of each spatial analysis unit as independent variables, and use the dependent variable and independent variables to form training samples, and use the independent variables of the spatial analysis unit to be predicted to form prediction samples. The geographic weighted random forest model construction module is used to construct a local random forest model for each target spatial analysis unit, calculate the geographic weight vector of the target unit, use the geographic weight vector as the sample weight of the corresponding training sample to obtain a set of training samples with geographic weights, calculate the geographic weighted variance impurity based on the geographic weight vector, and perform node splitting based on maximizing its reduction, and integrate the local random forest models of all spatial analysis units to obtain the geographic weighted random forest model. Model hyperparameter optimization module: Used to optimize the hyperparameters of the geographic weighted random forest model using cross-validation and grid search methods, based on error evaluation criteria, to obtain the optimal geographic weighted random forest model; Prediction module: For the predicted sample, it calls the local random forest model corresponding to its spatial analysis unit to generate travel predictions and outputs the prediction results.
[0023] Compared with the prior art, the present invention has the following beneficial effects: 1. Significantly Improved Model Prediction Accuracy: This invention constructs a Geographically Weighted Random Forest (GWRF) model, training an independent local random forest model for each spatial analysis unit, and embedding geographical distance decay weights into the decision tree splitting process, achieving unified modeling of nonlinear relationships and spatial heterogeneity. Experimental results show that the GWRF model significantly outperforms traditional linear models, global nonlinear models, and existing spatial models in terms of RMSE, MAE, and R² evaluation metrics.
[0024] 2. Effective Elimination of Spatial Autocorrelation in Residuals: This invention introduces a geographic weight vector and a local modeling mechanism, enabling the model to actively absorb spatial structure information. Experimental results show that the global Moran's exponent of the GWRF model residuals is close to zero and statistically insignificant, proving that the model successfully eliminates the spatial autocorrelation of the residuals and avoids the coefficient estimation bias caused by spatial dependence in traditional models.
[0025] 3. Refined Characterization of Nonlinear Relationships in Built Environment Impacts: This invention employs the Accumulated Local Effects (ALE) diagram to visualize and analyze the impact patterns of built environment characteristic variables, accurately identifying linear patterns, saturation effects, and abrupt change characteristics. Specifically, some land use densities exhibit linear impacts; saturation effects: accessibility in the central business district and parking facility supply show diminishing marginal returns, with key threshold points identifiable; abrupt change characteristics: some traffic accessibility indicators and population density exhibit abrupt changes in impact intensity at specific thresholds. The identification of these nonlinear characteristics provides an operational quantitative reference for planning practice.
[0026] 4. Quantitative Analysis of the Impact of Spatial Heterogeneity: This invention achieves a quantitative analysis of the spatial heterogeneity of the impact of built environment characteristic variables, including revealing the differences in the contribution of the same characteristic variable in different regions through Local Relative Feature Importance (RFI); and revealing the differences in threshold effect and the changes in the intensity of influence in different regions through local ALE curves. This analysis provides a quantitative basis for the formulation of differentiated transportation policies, avoiding the drawbacks of "one-size-fits-all" planning.
[0027] 5. Significantly Enhanced Model Interpretability: This invention combines RFI and ALE techniques to achieve complete interpretability analysis, from "feature importance" to "influence direction and nonlinear relationship." Compared to traditional black-box models, this invention can intuitively display the contribution ranking, influence direction, threshold range, and spatial differences of various built environment feature variables, significantly improving the model's usability in practical planning decisions. Attached Figure Description
[0028] Figure 1 This is a schematic diagram of the overall process in Embodiment 2 of the present invention; Figure 2 This is an ALE curve diagram showing the relationship between development density and target attainability variables in Embodiment 2 of the present invention. Figure 2 (a) is the ALE curve of residential land density. Figure 2 (b) ALE curve diagram for commercial service land. Figure 2 (c) is an ALE curve graph showing the distance from the CBD; Figure 3 This is an ALE curve diagram of the bus proximity correlation variables in Embodiment 2 of the present invention; wherein, Figure 3 (a) is an ALE curve diagram of the density of accessible areas of bus stops. Figure 3 (b) is an ALE curve diagram of the density of accessible areas of subway stations. Figure 3 (c) is an ALE curve diagram showing the distance to the nearest integrated transportation hub. Figure 3 (d) is an ALE curve showing the distance to the nearest railway station; Figure 4 This is an ALE curve diagram of the street block design-related variables in Embodiment 2 of the present invention; wherein, Figure 4 (a) ALE curve diagram of main road density, Figure 4 (b) is the ALE curve of the density of secondary arterial roads. Figure 4 (c) is the ALE curve of branch density. Figure 4 (d) is the ALE curve of intersection density; Figure 5 This is an ALE curve diagram of population characteristics and demand management-related variables in Embodiment 2 of the present invention; wherein, Figure 5 (a) is the ALE curve of roadside parking space density. Figure 5 (b) is the ALE curve of public parking density. Figure 5 (c) is the ALE curve of population density. Detailed Implementation
[0029] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Example 1 See attached document Figure 1 This embodiment provides a travel generation and prediction method based on geographically weighted random forest, including the following steps: Step 1: Divide the study area into multiple spatial analysis units, obtain the travel rate of each spatial analysis unit as the dependent variable; obtain multiple built environment characteristic variables of each spatial analysis unit as independent variables. The dependent and independent variables of each spatial analysis unit constitute the training sample, and the independent variables of the spatial analysis unit to be predicted constitute the prediction sample. Step 2: Construct a local random forest model for each target spatial analysis unit and calculate the geographic weight vector of the target unit; use the geographic weight vector as the sample weight of the corresponding training sample to obtain a training sample set with geographic weights; calculate the geographic weighted variance impurity based on the geographic weight vector and split the node to maximize its reduction; integrate the local random forest models of all spatial analysis units to obtain the geographic weighted random forest model. Step 3: Using cross-validation and grid search methods, the hyperparameters of the geographically weighted random forest model are optimized based on the error evaluation criterion to obtain the optimal geographically weighted random forest model; Step 4: For the predicted sample, call the local random forest model corresponding to its spatial analysis unit to generate travel predictions and output the prediction results.
[0031] In step 1 of this embodiment, the research area is divided into sections based on the real-world scenario and research objectives. Each Traffic Analysis Zone (TAZ) serves as the basic unit for spatial analysis. The centroid of each TAZ is used as the spatial location point of the sample for model training and prediction.
[0032] Extract the daily travel volume of each TAZ from the smartphone location big data of a certain company, and calculate the TAZ. Trip rate as dependent variable .
[0033] Dependent variable definition: Considering the differences in area among the traffic analysis zones, the daily trip volume is divided by the area of each zone to convert it into a trip rate, which is then used as the dependent variable (unit: thousands of trips per square kilometer).
[0034] Based on the "7Ds" framework, 31 built environment characteristic variables were constructed from seven dimensions as independent variables. This includes development density (residential land density, business office land density, commercial service land density, industrial land density, administrative land density, educational land density, medical and health land density, sports and cultural land density, park and green space density), functional mix (land use mix index), destination accessibility (distance to CBD), public transport proximity (density of areas accessible by bus stops, density of areas accessible by subway stations, distance to the nearest integrated transportation hub / railway station / intercity bus station / airport), street design (density of highways and expressways, density of main roads, density of secondary roads, density of side roads, density of intersections, proportion of signalized intersections, proportion of cross intersections, proximity to centrality), demand management (density of roadside parking spaces, density of public parking lots), and demographic characteristics (population density, proportion of male population, proportion of youth population, proportion of middle-aged population).
[0035] Step 2 in this embodiment specifically includes: For each target space analysis unit Construct a local random forest model And a geographic weight vector is introduced during the training process. This achieves a deep integration of distance attenuation effect and nonlinear modeling.
[0036] (1) Geographic weight calculation For the first Each TAZ (Transcript of Local Random Forests) is used to train its local random forest model. The sample weights are vectors over the entire dataset, denoted as...
[0037] in, Target space analysis unit The geographic weight vector, Representing all other spatial analysis units respectively Target space analysis unit The geographic weights are calculated using an adaptive bisquare kernel function based on geographic spatial distance:
[0038] in, It is the target space analysis unit Other spatial analysis units j Spatial distance It is adaptive bandwidth, expressed as the bandwidth from the target space analysis unit. To its first The distance to the nearest neighbor, i.e. . The optimal number of nearest neighbors is determined through cross-validation.
[0039] (2) Training of geographically weighted random forest model For each target space analysis unit Local random forest model When constructing the sub-decision tree, geographical weights are included. Introduced into the leaf node impurity metric, this is represented as geographically weighted variance impurity, used to reflect the impact of distance decay effects on training samples in space, thus enhancing local representation capabilities. The geographically weighted variance impurity function is shown below:
[0040]
[0041] in, Indicates the impurity of geographically weighted variance. Represents leaf nodes The sample set, where the sample elements are represented as , , For set The cardinality; Represent the built environment characteristic variables, trip generation volume, and sample, respectively. Target space analysis unit Geographic weight; It is a set The mean of trip generation across all samples. For each built environment feature... And a spatial bisection threshold for this feature Candidate splits , will set Divided into two subspaces, namely the left space and right space ,in Indicates sample In terms of built environment characteristics The value of .
[0042] Split Criterion: Maximize the reduction in geographically weighted variance impurity ;
[0043] in, and They represent criteria respectively The divided left space and right space The impurity. and These represent the geometric weights based on the impurities of the left and right spaces of the samples they contain, respectively.
[0044] Decision tree pairs of subsets and The process is recursively repeated until the maximum allowed depth or the minimum number of samples required for a split node is reached. A local random forest is an ensemble of the aforementioned sub-decision trees.
[0045] (3) Local model integration After the local random forests of each TAZ are constructed, each spatial analysis unit (TAZ) corresponds to a unique local random forest model, which together form the overall geographic weighted random forest (GWRF) model.
[0046] In step 3 of this embodiment, leave-one-out cross-validation and grid search are used to optimize the hyperparameters based on the root mean square error (RMSE) criterion. A two-stage strategy is adopted: first, the hyperparameters of the global RF model are determined, and then these parameters are used to determine the spatial bandwidth.
[0047] In the first stage, three hyperparameters of the global RF model are determined: the number of sub-models, the proportion of randomly selected features, and the minimum number of samples required for tree node splitting. The optimal parameter combination is then obtained.
[0048] In the second stage, the optimal hyperparameters of these global RF models are applied to each local RF model, and then the optimal adaptive bandwidth (i.e., the number of nearest neighbors) is searched. ).
[0049] Step 4 of this embodiment specifically includes: Training process: For the first Each TAZ is used to calculate a geographic weight vector based on its spatial distance from the other TAZs. Then, a local random forest model is trained using training samples with geographical weights. And save the parameters of the local random forest model; Prediction process: For the sample to be predicted (i.e., the first) (each TAZ), and call its corresponding Local Random Forest model. Make a prediction:
[0050] in, Samples to be predicted The built environment feature vector, Indicates the sample to be predicted The geographic weight vector.
[0051] Example 2 As a further improvement to Example 1, this example also performs the following additional steps: Step 5: Evaluate the performance of the geographically weighted random forest model using error evaluation metrics and spatial autocorrelation metrics; Step 6: Calculate the relative importance of each built environment characteristic variable globally and / or locally, and / or calculate the cumulative local effects globally and / or locally, to analyze the marginal effects and spatial heterogeneity of each built environment characteristic variable on travel generation, and output interpretability analysis results.
[0052] In step 5 of this embodiment, the root mean square error (RMSE), mean absolute error (MAE), and R are used. 2 To evaluate model performance, global and local Moran indices were used to test residual space autocorrelation.
[0053] Step 6 of this embodiment includes: (1) Calculation of Relative Feature Importance (RFI) The RFI of the input variable is calculated by summing the reduction in impurity associated with nodes in the subtrees split by the variable.
[0054] Global RFI: The average RFI of all decision trees is taken and normalized to obtain the global importance score of each feature, which is used to compare the overall contribution of each built environment feature variable to trip generation.
[0055] Local RFI: For each local model Calculate the RFI separately to obtain the spatial distribution of feature importance, which can be used to analyze the differences and clustering characteristics of influencing factors in different regions.
[0056] (2) Drawing the Accumulated Local Effects (ALE) diagram ALE diagrams are used to visualize the marginal effect of various built environment characteristic variables on travel generation.
[0057] The calculation steps are as follows: a) Interval partitioning: Selecting features The distribution range is divided based on quantiles. each interval ,in Indicates the selected feature The ( ) quantiles, in particular, Each represents a selected feature. The minimum and maximum values; b) Calculation of local effects: For the first Calculate the local effects across intervals:
[0058] in, Representation of features In the Local effects in each interval Indicates falling into the first A set of data points in each interval Indicates the sample size. and These represent the prediction results after replacing the selected variables with the right and left boundaries of the interval, respectively. For the sample Other eigenvalues.
[0059] c) Accumulation and Centralization: Define this variable Local effects The cumulative effect.
[0060]
[0061] in, The constant is used to zero-mean ALE plots, making it easier to interpret deviations from the average effect; Representation of features In the Local effects within a given interval.
[0062] Global ALE curve: The ALE values of all samples are averaged to obtain the global marginal effect curve of the feature, which is used to identify nonlinear features, threshold points and effective range of action.
[0063] Local ALE curves: for each local model The ALE curves were calculated separately to obtain the spatial variation of the characteristic influence, which was used to analyze the differences in threshold effect and influence intensity in different regions.
[0064] Taking the main urban area of a certain city as an example, the specific implementation process of this embodiment is explained in detail as follows: The study areas and data sources include: (1) Study area The main urban area of a certain city is divided into 408 traffic analysis zones (TAZs) based on their actual characteristics.
[0065] (2) Travel data Using big data on smartphone location from a certain company, and selecting weekday data, the total daily trip volume for each tourist area (TAZ) was extracted. The dependent variable was defined as the trip rate. = Daily average total number of trips / Number Area of one TAZ (unit: thousand times / square kilometer).
[0066] (3) Built environment data Based on a seven-dimensional ("7D") built environment measurement framework, 31 built environment characteristic variables were constructed as independent variables. This measure is used to assess built environment elements within each Traffic Analysis Zone (TAZ). The seven dimensions are: Density, Diversity, Destination Accessibility, Distance to Transit, Design, Demand Management, and Demographics. The specific characteristics of each dimension are described below: Development Density: Measured using land use data. There are nine density variables, representing the proportion of residential land, commercial office land, commercial service land, industrial land, administrative land, educational land, medical and health land, sports and cultural land, and park green space within each TAZ.
[0067] Diversity: This dimension typically refers to land diversity, with one variable: the Land Use Mix (LUM) index. It is calculated as the entropy of the proportions of each land use type. LUM values range from 0 to 1, with higher values indicating greater land use diversity. A value of 0 indicates a completely homogeneous area dominated by a single land use type, while a value of 1 indicates a uniform distribution of all land use types.
[0068] Destination accessibility: This includes one variable, representing the distance to the key destination—the Central Business District (CBD). For each TAZ, the Euclidean distance from its centroid to the nearest CBD is calculated.
[0069] Distance to transit: Distance to transit measures the accessibility of a public transport system and includes six variables: density of bus stop access areas, density of subway station access areas, distance to the nearest integrated transportation hub, distance to the nearest train station, distance to the nearest intercity bus station, and distance to the airport. Among these, the density of bus / subway station access areas is... The calculation method is as follows:
[0070] in Indicates that it is located at the th A collection of bus / subway stations within a TAZ. This is a website Number of bus / subway lines served. This refers to the service area associated with bus / subway stations. The reachable areas are defined as a 400-meter radius buffer zone for bus stops and an 800-meter radius buffer zone for subway stations. It is the first The area of one TAZ.
[0071] Street design: Geographic information data of the road network is acquired to measure design variables. There are a total of 8 design variables, including the density of four road levels: highways and expressways, arterial roads, secondary arterial roads, and local roads; three intersection-related variables: intersection density, the proportion of signalized intersections, and the proportion of cross intersections; and the proximity centrality variable of the road network.
[0072] Demand management: Acquire parking facility location data to measure demand management variables. There are two demand management variables: roadside parking space density and public parking lot density.
[0073] Demographics: Population density data is used to measure demographic variables. There are four demographic variables: population density, male population ratio, youth population ratio, and middle-aged population ratio.
[0074] The hyperparameters were optimized using leave-one-out cross-validation and grid search. The results are as follows: the number of sub-models is 500, the proportion of randomly selected features is 0.6, the minimum number of samples for tree node splitting is 8, and the adaptive bandwidth is 390.
[0075] Comparison models: Linear Regression (LR), Poisson Regression (PR), Random Forest (RF), Geographically Weighted Regression (GWR), and Geographically Weighted Poisson Regression (GWPR).
[0076] Model comparison results: GWRF's RMSE is 5.11% lower than the best comparison model (RF); GWRF's MAE is 44.92% lower than the best comparison model (RF); GWRF's R² is 0.2% higher than the best comparison model (RF).
[0077] The global Moran's index was used to test the spatial autocorrelation of the residuals. The global Moran's index of the GWRF residuals was 0.03, indicating no significant spatial autocorrelation, which proves that the model successfully absorbed spatial structure information.
[0078] (1) Relative Feature Importance (RFI) Analysis Table 1 shows the relative characteristic importance (RFI) of influencing factors. The global RF model and the GWRF model show a high degree of consistency in measuring the contribution of built environment variables to travel generation prediction.
[0079] The top 15 built environment characteristic variables collectively contribute approximately 95% of the predictive power; The top 5 characteristic variables, sorted by RFI, are: population density, density of areas accessible by public transport stops, density of areas accessible by subway stations, density of commercial service land, and density of secondary arterial roads.
[0080] Local RFI data showed that population density was more important in the northern region and the area accessible by public transport was more important in the southern region.
[0081] Table 1. Relative feature importance of variables under Global RF and GWRF
[0082] (2) Cumulative Local Effects (ALE) Analysis Figure 2 The data shows that residential and commercial land use has a positive impact on travel generation forecasts, and its impact on overall travel demand follows an approximately linear pattern. The distance to the nearest CBD has a negative impact on travel generation and exhibits a significant marginal effect. The saturation threshold for the distance to the CBD is approximately 7 kilometers, and the impact decreases significantly within the range of 1.5 to 7 kilometers.
[0083] Figure 3 The results indicate that urban public transport accessibility (i.e., the area accessible to bus stops and subway stations) has a positive impact on trip generation and exhibits typical nonlinear characteristics, with a saturation threshold of 1.0 for the density of subway station accessibility areas. Intercity public transport accessibility (i.e., the distance to the nearest integrated transportation hub and train station) has a negative impact on trip generation, with a threshold of 8 kilometers for the distance to the integrated transportation hub.
[0084] Figure 4 This indicates that the nonlinear impact of road density on overall travel generation is more complex, especially since the patterns of arterial and secondary roads are non-monotonic. The spatial heterogeneity sensitivity of road density is higher than that of intersection density.
[0085] Figure 5 The ALE curves represent the variables related to population characteristics and demand management. Both roadside parking spaces and public parking lots have a positive impact on travel demand and exhibit saturation effects, with saturation points identified as 18 and 40, respectively. Travel rate is positively correlated with population density, with a sudden change threshold of 28,000 people / km², after which travel rate increases sharply.
[0086] Example 3 This embodiment provides a travel generation and prediction system based on geographically weighted random forest, including: The data acquisition and preprocessing module is used to divide the study area into multiple spatial analysis units, obtain the travel rate of each spatial analysis unit as the dependent variable, obtain multiple built environment characteristic variables of each spatial analysis unit as independent variables, and use the dependent variable and independent variables to form training samples, and use the independent variables of the spatial analysis unit to be predicted to form prediction samples. The geographic weighted random forest model construction module is used to construct a local random forest model for each target spatial analysis unit, calculate the geographic weight vector of the target unit, use the geographic weight vector as the sample weight of the corresponding training sample to obtain a set of training samples with geographic weights, calculate the geographic weighted variance impurity based on the geographic weight vector, and perform node splitting based on maximizing its reduction, and integrate the local random forest models of all spatial analysis units to obtain the geographic weighted random forest model. Model hyperparameter optimization module: Used to optimize the hyperparameters of the geographic weighted random forest model using cross-validation and grid search methods, based on error evaluation criteria, to obtain the optimal geographic weighted random forest model; Prediction module: For the predicted sample, it calls the local random forest model corresponding to its spatial analysis unit to generate travel predictions and outputs the prediction results.
[0087] Example 4 As a further improvement to Embodiment 3, this embodiment also includes the following additional modules: Model performance evaluation module: used to evaluate the performance of the geographically weighted random forest model using error evaluation metrics and spatial autocorrelation metrics; Interpretable Analysis Module: Used to calculate the relative importance of global and / or local characteristics of each built environment feature variable, and / or calculate the global and / or local cumulative local effects, used to analyze the marginal effects and spatial heterogeneity of each built environment feature variable on travel generation, and output interpretability analysis results; Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A trip generation and prediction method based on geographically weighted random forest, characterized in that, Includes the following steps: The study area is divided into multiple spatial analysis units. The travel rate of each spatial analysis unit is obtained as the dependent variable. Multiple built environment characteristic variables of each spatial analysis unit are obtained as independent variables. The dependent variable and the independent variables constitute the training sample, and the independent variables of the spatial analysis unit to be predicted constitute the prediction sample. A local random forest model is constructed for each target spatial analysis unit, and the geographic weight vector of the target unit is calculated. The geographic weight vector is used as the sample weight of the corresponding training sample to obtain a training sample set with geographic weights. The geographic weighted variance impurity is calculated based on the geographic weight vector, and node splitting is performed with the criterion of maximizing its reduction. The local random forest models of all spatial analysis units are integrated to obtain a geographic weighted random forest model. Using cross-validation and grid search methods, the hyperparameters of the geographic weighted random forest model are optimized based on the error evaluation criterion to obtain the optimal geographic weighted random forest model. For the predicted sample, the local random forest model corresponding to its spatial analysis unit is called to generate travel predictions and output the prediction results.
2. The travel generation and prediction method based on geographically weighted random forest according to claim 1, characterized in that, The travel rate is the total daily travel volume divided by the area of the corresponding spatial analysis unit.
3. The trip generation and prediction method based on geographically weighted random forest according to claim 1, characterized in that, The built environment characteristic variables are quantified based on a multi-dimensional indicator system, which includes seven dimensions: development density, functional mix, destination accessibility, public transport proximity, neighborhood design, demand management, and population characteristics.
4. The travel generation and prediction method based on geographically weighted random forest according to claim 3, characterized in that, The built environment characteristic variables include: Development density: residential land density, business office land density, commercial service land density, industrial land density, administrative land density, educational land density, medical and health land density, sports and cultural land density, and park green space density; Functional Mixing: Land Use Mixing Index; Achievable destination: Distance to the central business district; Public transport proximity: density of areas accessible by bus stops, density of areas accessible by subway stations, distance to the nearest integrated transportation hub, distance to the nearest train station, distance to the nearest long-distance bus station, and distance to the nearest airport; Street layout design: density of highways and expressways, density of arterial roads, density of secondary arterial roads, density of local roads, density of intersections, proportion of signalized intersections, proportion of crossroads, proximity to centrality; Demand management: density of roadside parking spaces, density of public parking lots; Population characteristics: population density, male population ratio, youth population ratio, middle-aged population ratio.
5. The trip generation and prediction method based on geographically weighted random forest according to claim 1, characterized in that, For target space analysis unit The geographic weight vector is represented as: in, Target space analysis unit The geographic weight vector, Representing all other spatial analysis units respectively Target space analysis unit The geographic weights are calculated using an adaptive bisquare kernel function based on geographic spatial distance: in, It is the target space analysis unit Other spatial analysis units j Spatial distance, It is adaptive bandwidth, expressed as the bandwidth from the target space analysis unit. To its first The distance to the nearest neighbor, i.e. .
6. The trip generation and prediction method based on geographically weighted random forest according to claim 5, characterized in that, For target space analysis unit The geographic weight vector is introduced when training its local random forest model. Specifically: In the target space analysis unit In the local random forest model, the impurity of the leaf nodes of its internal decision trees is calculated by the geographically weighted variance impurity, and the formula for calculating the geographically weighted variance impurity is as follows: in, Indicates the impurity of geographically weighted variance. Represents leaf nodes The sample set, where the sample elements are represented as , , For set The cardinality; Representing sets Samples in Built environment characteristic variables, travel generation volume, and sample Target space analysis unit Geographic weight; It is a set The mean of trip generation for all samples; for each feature of a built environment And a spatial bisection threshold for this feature Candidate splits , will set Divided into two subspaces, namely the left space and right space ,in Indicates sample In terms of built environment characteristics The value; The node splitting criterion is: maximizing the reduction in geographically weighted variance impurity. ; in, Indicated based on the splitting criterion Partitioned Sets The reduction in geographically weighted variance impurity was obtained; and They represent criteria respectively The divided left space and right space Impurity; and These represent the geometric weights based on the impurities of the left and right spaces of the samples they contain, respectively.
7. The trip generation and prediction method based on geographically weighted random forest according to claim 1, characterized in that, The hyperparameters are optimized using leave-one-out cross-validation and grid search, employing a two-stage strategy: In the first stage, a global random forest model is constructed based on the entire study area without considering geographical weights, and the optimal hyperparameters of the global random forest model are determined based on the root mean square error, including the number of sub-models, the proportion of randomly selected features, and the minimum number of samples for tree node splitting; In the second stage, the optimal hyperparameters from the first stage are applied to all local random forest models, and the optimal number of nearest neighbors is searched.
8. The trip generation and prediction method based on geographically weighted random forest according to claim 1, characterized in that, Also includes: The performance of the geographically weighted random forest model was evaluated using error evaluation metrics and spatial autocorrelation metrics. Calculate the global and / or local relative importance of each built environment characteristic variable, and / or calculate the global and / or local cumulative effects, and output interpretable analysis results.
9. The travel generation and prediction method based on geographically weighted random forest according to claim 8, characterized in that, The relative feature importance is calculated by summing the reduction in impurity of each built environment feature variable at the decision tree split node; the global relative feature importance is obtained by averaging and normalizing the relative feature importance of all local random forest models; the local relative feature importance is calculated separately for each local random forest model. The cumulative local effect map is calculated through the following steps: dividing the quantile distribution range of the selected feature into multiple intervals; calculating the local effect in each interval; accumulating and centering all local effects; and calculating the local cumulative local effect map for each local random forest model separately.
10. A trip generation and prediction system based on geographically weighted random forest, characterized in that, include: The data acquisition and preprocessing module is used to divide the study area into multiple spatial analysis units and obtain the travel rate of each spatial analysis unit as the dependent variable. Multiple built environment characteristic variables of each spatial analysis unit are obtained as independent variables. The dependent variable and the independent variables constitute the training sample, and the independent variables of the spatial analysis unit to be predicted constitute the prediction sample. The geographic weighted random forest model construction module is used to construct a local random forest model for each target spatial analysis unit, calculate the geographic weight vector of the target unit, use the geographic weight vector as the sample weight of the corresponding training sample to obtain a set of training samples with geographic weights, calculate the geographic weighted variance impurity based on the geographic weight vector, and perform node splitting based on maximizing its reduction, and integrate the local random forest models of all spatial analysis units to obtain the geographic weighted random forest model. Model hyperparameter optimization module: Used to optimize the hyperparameters of the geographic weighted random forest model using cross-validation and grid search methods, based on error evaluation criteria, to obtain the optimal geographic weighted random forest model; Prediction module: For the predicted sample, it calls the local random forest model corresponding to its spatial analysis unit to generate travel predictions and outputs the prediction results; The trip generation and prediction system based on geographically weighted random forest is used to perform the steps in the trip generation and prediction method based on geographically weighted random forest as described in any one of claims 1-9.
Citation Information
Patent Citations
Geographic weighting and random forest coupled surface temperature downscaling method
CN117035066A
Method for analyzing multi-scale influence of built environment on resident travel based on interpretable machine learning
CN117874709A