Shared bicycle passenger flow prediction method based on geographically weighted regression model and related device

By screening effective independent variables and adjusting the regression coefficients through the geographically weighted regression model, the problem of spatial heterogeneity not being taken into account in the shared bicycle passenger flow prediction was solved, and higher-precision predictions and more accurate bicycle deployment strategies were achieved.

CN120745933APending Publication Date: 2025-10-03GUANGDONG URBAN TECHNICIAN COLLEGE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510907231.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing shared bicycle passenger flow prediction methods fail to effectively consider spatial heterogeneity, resulting in low prediction accuracy, especially in the case of complex urban internal structures, which affects the formulation of bicycle deployment and scheduling strategies.

Method used

The geographically weighted regression model (GWR) is used to collect spatial data and socioeconomic data of shared bicycle deployment points, conduct spatial autocorrelation tests and multicollinearity tests, screen effective independent variables, and adjust the regression coefficients according to geographical location to capture spatial heterogeneity and improve prediction accuracy.

Benefits of technology

The accuracy of shared bicycle passenger flow forecasting has been significantly improved, which can more accurately reflect the impact characteristics of local areas and provide more targeted operational suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120745933A_ABST
    Figure CN120745933A_ABST
Patent Text Reader

Abstract

The invention discloses a shared bicycle passenger flow prediction method based on a geographically weighted regression model and a related device. The method comprises the following steps: collecting spatial data and social economic data of a shared bicycle putting point of a target city; inputting the effective independent variable data set, the spatial autocorrelation of the effective independent variables, the multi-collinearity test result and the housing price factor data into a least square regression model and a geographical weighted regression model, and outputting a regression coefficient, a significance level and a goodness of fit; and according to the regression coefficient, the significance level, the goodness of fit, the delivery point category and the actual shared bicycle passenger flow data, calculating the prediction precision and the lifting amplitude at different delivery point categories. According to the method, the spatial heterogeneity relationship between the shared bicycle passenger flow volume and the influence factors is considered through the geographically weighted regression model, the regression coefficient is adjusted according to different geographic positions, the influence characteristics of a local area are reflected more accurately, and the prediction precision is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of shared bicycle passenger flow prediction, and in particular to a shared bicycle passenger flow prediction method and device based on a geographically weighted regression model, and a computing device. Background Art

[0002] As an emerging mode of transportation, shared bikes not only provide a convenient short-distance travel option but also alleviate some of the pressure on urban traffic. However, the rapid development of shared bikes has also brought about a series of problems, such as uneven bike deployment and inaccurate passenger flow forecasts, which directly affect user experience and operational efficiency.

[0003] Traditional methods for predicting shared bike ridership rely primarily on statistical and machine learning models, such as linear regression, time series analysis, and random forests. These often overlook spatial heterogeneity, or the differences in characteristics between different regions. For example, the least squares regression model assumes that the regression coefficients are constant across the entire study area, failing to capture spatial variations between regions. This results in low prediction accuracy, particularly in complex urban structures. Spatial heterogeneity refers to significant differences in characteristics and behaviors between regions. Considering spatial heterogeneity is particularly important for predicting shared bike ridership. For example, different types of areas, such as commercial districts, residential areas, school neighborhoods, and public transportation hubs, exhibit significant differences in bike usage frequency and patterns. Ignoring these differences will lead to inaccurate predictions, which in turn will affect the formulation of bike deployment and scheduling strategies.

[0004] To solve the above problems, the present invention proposes a method for predicting shared bicycle passenger flow based on a geographically weighted regression model. Through the geographically weighted regression model, the spatial heterogeneity relationship between shared bicycle passenger flow and influencing factors is considered, and the regression coefficient is adjusted according to different geographical locations to accurately reflect the influencing characteristics of local areas and improve prediction accuracy. Summary of the Invention

[0005] In view of the above problems, the present invention provides a method and device for predicting shared bicycle passenger flow based on a geographically weighted regression model, as well as a computing device.

[0006] According to one aspect of the present invention, a method for predicting shared bicycle passenger flow using a geographically weighted regression model is provided, comprising:

[0007] Collect spatial data and socioeconomic data of shared bicycle deployment points in the target city to obtain an original data set, where the spatial data includes latitude and longitude and the location of the deployment point, and the socioeconomic data includes population density, POI data, road density, and housing prices;

[0008] Using the original data set to perform spatial autocorrelation test and multicollinearity test on the independent variables, to obtain a valid independent variable data set, spatial autocorrelation and multicollinearity test results of the valid independent variables;

[0009] Input the effective independent variable data set, the spatial autocorrelation of the effective independent variables, the multicollinearity test results, and the housing price factor data into the least squares regression model OLS and the geographically weighted regression model GWR, and output the regression coefficients, significance levels, and goodness of fit of the least squares regression model OLS and the geographically weighted regression model GWR;

[0010] Classifying the shared bicycle deployment points in the target city according to the spatial data to obtain deployment point categories, wherein the deployment point categories include commercial services, scenic spots, public services, government and corporate offices, commercial residences, and transportation facilities;

[0011] Based on the regression coefficient, significance level, goodness of fit, deployment point category and actual shared bicycle passenger flow data, the prediction accuracy and improvement of the least squares regression model OLS and the geographically weighted regression model GWR in different deployment point categories are calculated to verify the advantage of the geographically weighted regression model GWR in capturing spatial heterogeneity.

[0012] In an optional manner, the socioeconomic data includes the population size and housing prices within the buffer zone, wherein the population size is estimated by the number of households around the shared bicycle deployment point, and the housing prices are calculated by the unit price of housing in the residential area within the buffer zone.

[0013] In an optional manner, the expression of the geographically weighted regression model GWR is:

[0014] ;

[0015] in, is the shared bicycle passenger flow of the i-th sample point; is the spatial coordinate of the i-th sample point; In space coordinates The intercept term at ; is the spatial coordinate of the t-th independent variable The regression coefficient at ; is the value of the t-th independent variable at the i-th sample point; is the error term of the i-th sample point; is the number of independent variables.

[0016] In an optional manner, the expression of the spatial weight function of the geographically weighted regression model GWR is:

[0017] ;

[0018] in, is the spatial weight of point j estimated when fitting the model for shared bicycle deployment point i; is the distance between delivery points i and j; is the bandwidth, calculated using the AIC criterion.

[0019] In an optional manner, the calculation formula of the AIC criterion is:

[0020] ;

[0021] in, is the maximum likelihood estimate of the variance of the random error term, ; is a function of bandwidth b; n is the number of samples; The sum of the squares of the differences between the actual value and the model prediction for each observation point.

[0022] In an optional manner, performing a spatial autocorrelation test and a multicollinearity test on the independent variables using the original data set to obtain a valid independent variable data set and the spatial autocorrelation and multicollinearity test results of the valid independent variables further includes:

[0023] The Moran index is used to test the spatial autocorrelation of each independent variable in the original data set, and the Moran index value and significance level of each independent variable are calculated. If the Moran index is significant, the independent variable has significant spatial autocorrelation.

[0024] Calculate the variance inflation factor of all independent variables in the original data set. If the variance inflation factor of an independent variable exceeds the threshold, there is multicollinearity between the independent variable and other independent variables.

[0025] According to the results of the spatial autocorrelation test and the multicollinearity test, the effective independent variables are screened out, and the screened independent variable data set, the Moran's index value and the variance inflation factor of each effective independent variable are obtained.

[0026] In an alternative approach, when the sample size n is small, the modified AIC criterion AIC is used. C Determine the bandwidth b, the AIC C The calculation formula is:

[0027] ;

[0028] in, is the number of independent variables.

[0029] In an optional manner, the threshold of the variance inflation factor is 10. When the variance inflation factor of an independent variable exceeds 10, there is serious multicollinearity between the independent variable and other independent variables. The independent variable with the largest variance inflation factor is eliminated until the variance inflation factors of all independent variables are less than 10.

[0030] According to another aspect of the present invention, a device for predicting shared bicycle passenger flow based on a geographically weighted regression model is provided, comprising:

[0031] A data collection module is used to collect spatial data and socioeconomic data of shared bicycle deployment points in the target city to obtain an original data set, wherein the spatial data includes latitude and longitude and the location of the deployment point, and the socioeconomic data includes population density, POI data, road density, and housing prices;

[0032] A variable testing module is used to perform a spatial autocorrelation test and a multicollinearity test on the independent variables using the original data set to obtain a valid independent variable data set and the spatial autocorrelation and multicollinearity test results of the valid independent variables;

[0033] A model regression module is used to input the effective independent variable data set, the spatial autocorrelation of the effective independent variables, the multicollinearity test results, and the housing price factor data into the least squares regression model OLS and the geographically weighted regression model GWR, and output the regression coefficients, significance levels, and goodness of fit of the least squares regression model OLS and the geographically weighted regression model GWR;

[0034] a classification module for classifying the shared bicycle deployment points in the target city according to the spatial data to obtain deployment point categories, wherein the deployment point categories include commercial services, scenic spots, public services, government and corporate offices, commercial residences, and transportation facilities;

[0035] The accuracy evaluation module is used to calculate the prediction accuracy and improvement of the least squares regression model OLS and the geographically weighted regression model GWR in different deployment point categories based on the regression coefficient, significance level, goodness of fit, deployment point category and actual shared bicycle passenger flow data, so as to verify the advantage of the geographically weighted regression model GWR in capturing spatial heterogeneity.

[0036] According to another aspect of the present invention, there is provided a computing device, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus;

[0037] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the shared bicycle passenger flow prediction method based on the above-mentioned geographically weighted regression model.

[0038] According to the solution provided by the present invention, the spatial data and socioeconomic data of the shared bicycle deployment points in the target city are collected to obtain an original data set, wherein the spatial data include latitude and longitude and the location of the deployment point, and the socioeconomic data include population density, POI data, road density and housing prices; the original data set is used to perform spatial autocorrelation test and multicollinearity test on the independent variables to obtain an effective independent variable data set, spatial autocorrelation and multicollinearity test results of the effective independent variables; the effective independent variable data set, spatial autocorrelation of the effective independent variables, multicollinearity test results and housing price factor data are input into the least squares regression model OLS and the geographically weighted regression model GWR, and the least squares regression model is output. The regression coefficient, significance level and goodness of fit of the regression model OLS and the geographically weighted regression model GWR are multiplied; the target city's shared bicycle delivery points are divided into delivery point categories according to the spatial data to obtain delivery point categories, and the delivery point categories include commercial services, scenic spots, public services, government and corporate offices, commercial residences and transportation facilities; according to the regression coefficient, significance level, goodness of fit, delivery point categories and actual shared bicycle passenger flow data, the prediction accuracy and improvement of the least squares regression model OLS and the geographically weighted regression model GWR in different delivery point categories are calculated to verify the advantage of the geographically weighted regression model GWR in capturing spatial heterogeneity. The present invention takes into account the spatial heterogeneity relationship between shared bicycle passenger flow and influencing factors through the geographically weighted regression model, and can adjust the regression coefficient according to different geographical locations, thereby more accurately reflecting the influence characteristics of the local area and significantly improving the prediction accuracy.

[0039] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0041] Figure 1 A schematic diagram showing a flow chart of a method for predicting shared bicycle passenger flow using a geographically weighted regression model according to an embodiment of the present invention;

[0042] Figure 2A schematic diagram showing the framework of a device for predicting shared bicycle passenger flow using a geographically weighted regression model according to an embodiment of the present invention is shown;

[0043] Figure 3 A schematic structural diagram of a computing device according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0044] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0045] Figure 1 The flowchart of the method for predicting the passenger flow of shared bicycles based on the geographically weighted regression model according to an embodiment of the present invention is shown. Figure 1 As shown, the following steps are included:

[0046] Step S101: Collect spatial data and socioeconomic data of shared bicycle deployment points in the target city to obtain an original data set, wherein the spatial data includes latitude and longitude and the location of the deployment point, and the socioeconomic data includes population density, POI data, road density, and housing prices.

[0047] In this embodiment, the socioeconomic data include the population and housing prices within the buffer zone, wherein the population is estimated by the number of households around the shared bicycle deployment point, and the housing prices are calculated by the unit price of housing in the residential area within the buffer zone. The population within the buffer zone is closely related to the passenger flow of shared bicycles and is one of the important factors affecting the passenger flow of shared bicycle deployment points. Reasonable housing prices help optimize urban resource allocation, improve land use efficiency, optimize the built environment, and thus affect the passenger flow of the site. The population is estimated by the number of households around the shared bicycle deployment point. For example, .in, is the average housing price of the i-th community among the t communities within the buffer zone, is the number of households in cell i out of a total of t cells within the buffer zone. Land use type POI is a major component of the built environment. The present invention uses Python to obtain POI data for Shenzhen, China through the AMAP open platform. AMAP divides POI data into 23 major categories. According to the POI type and combined with the urban land use standard 54, land use is divided into six categories: commercial services, scenic spots, public services, government and corporate offices, commercial residences, and transportation facilities, corresponding to variables X3 to X8 respectively. The building volume ratio reflects the intensity of land use development and, to a certain extent, can reflect the economic level and residents' activities in the area, recorded as X9. , where F is the total area of ​​all above-ground buildings within the rail station buffer zone; S is the area of ​​the shared bicycle deployment point buffer zone. The road network density reflects the external accessibility of the shared bicycle deployment point, reflects the convenience of residents' travel, and also affects the passenger flow of the shared bicycle deployment point. The road network density is recorded as X 11 , , S is the buffer area of ​​the shared bicycle deployment point; L is the length of each road section within the buffer area.

[0048] Step S102 : performing a spatial autocorrelation test and a multicollinearity test on the independent variables using the original data set to obtain a valid independent variable data set and spatial autocorrelation and multicollinearity test results of the valid independent variables.

[0049] In this example, we consider the potential correlation in ridership between geographically adjacent locations, avoiding the direct inclusion of spatially dependent variables in the model, thereby reducing model bias. Multicollinearity testing avoids the problem of high correlation between independent variables leading to unstable model parameter estimates and reduced model interpretability. Eliminating collinear independent variables results in more robust model parameters. By selecting effective independent variables, the model makes it easier to understand the key factors influencing shared bike ridership.

[0050] In an optional manner, performing a spatial autocorrelation test and a multicollinearity test on the independent variables using the original data set to obtain a valid independent variable data set and the spatial autocorrelation and multicollinearity test results of the valid independent variables further includes:

[0051] The Moran index is used to test the spatial autocorrelation of each independent variable in the original data set, and the Moran index value and significance level of each independent variable are calculated. If the Moran index is significant, the independent variable has significant spatial autocorrelation.

[0052] Calculate the variance inflation factor of all independent variables in the original data set. If the variance inflation factor of an independent variable exceeds the threshold, there is multicollinearity between the independent variable and other independent variables.

[0053] According to the results of the spatial autocorrelation test and the multicollinearity test, the effective independent variables are screened out, and the screened independent variable data set, the Moran's index value and the variance inflation factor of each effective independent variable are obtained.

[0054] In this embodiment, for example, the Moran index is used to test whether each independent variable has spatial autocorrelation, and the results shown in Table 1 are obtained.

[0055] Table 1

[0056] Independent variable Moran Index P-value Significance population density 0.35 0.01 significant Number of catering POIs 0.28 0.03 significant Shopping POI quantity 0.15 0.10 Not significant Number of leisure and entertainment POIs 0.42 0.001 significant Road density 0.08 0.25 Not significant Housing prices 0.51 0.0001 significant

[0057] As can be seen from Table 1, there is significant spatial autocorrelation among population density, number of catering POIs, number of leisure and entertainment POIs, and housing prices.

[0058] The VIF (variance inflation factor) of all independent variables was calculated, and the results were shown in Table 2. All VIFs were less than 10, so there was no serious multicollinearity.

[0059] Table 2

[0060] Independent variable VIF population density 4.5 Number of catering POIs 7.2 Shopping POI quantity 8.9 Number of leisure and entertainment POIs 6.1 Road density 3.2 Housing prices 2.8

[0061] Because population density, the number of dining POIs, the number of leisure and entertainment POIs, and housing prices all exhibit significant spatial autocorrelation, these variables were removed to avoid bias in the OLS regression model. Ultimately, the number of shopping POIs and road density were selected as valid independent variables. The final dataset of valid independent variables includes the coordinates (latitude and longitude) of shared bike deployment points, the number of shopping POIs, and road density.

[0062] In an optional manner, the threshold of the variance inflation factor is 10. When the variance inflation factor of an independent variable exceeds 10, there is serious multicollinearity between the independent variable and other independent variables. The independent variable with the largest variance inflation factor is eliminated until the variance inflation factors of all independent variables are less than 10.

[0063] In this example, multicollinearity refers to high correlations between independent variables, which can lead to unstable coefficients in the OLS regression model, excessive variance, or even opposite variances, affecting the model's predictive accuracy and interpretability. By setting a VIF threshold and iteratively removing independent variables with high VIFs, the impact of multicollinearity on the model can be effectively reduced.

[0064] For example, a preliminary calculation revealed that the VIF value for population density was 12, exceeding the threshold of 10. In the first iteration, the population density variable was removed, and the VIF was recalculated for the remaining variables (POI data, road density, housing prices, and green area). In subsequent iterations, assuming the VIF value for "housing prices" is now 11, the "housing prices" variable was removed again. Repeat these steps until the VIF values ​​of all remaining variables are less than 10. Ultimately, the only variables remaining are "POI data," "road density," and "green area." Through iterative elimination, "population density" and "housing prices" were removed. While these variables themselves may be correlated with shared bike ridership, they are highly correlated with other variables, making the model unstable. The remaining variables, "POI data," "road density," and "green area," provide more stable and reliable forecasts.

[0065] Step S103: Input the effective independent variable data set, the spatial autocorrelation of the effective independent variables, the multicollinearity test results, and the housing price factor data into the least squares regression model OLS and the geographically weighted regression model GWR, and output the regression coefficients, significance levels, and goodness of fit of the least squares regression model OLS and the geographically weighted regression model GWR.

[0066] In this example, by using both the OLS and GWR models, we can comprehensively compare their ability to capture shared bike ridership. The OLS model assumes that the regression coefficients are constant across the entire spatial range, while the GWR model allows for spatial variation in the regression coefficients, thereby better capturing spatial heterogeneity. By comparing the regression coefficients, significance levels, and goodness of fit of the two models, we can assess the GWR model's superiority in capturing spatial heterogeneity.

[0067] In an optional manner, the expression of the geographically weighted regression model GWR is:

[0068] ;

[0069] in, is the shared bicycle passenger flow of the i-th sample point; is the spatial coordinate of the i-th sample point; In space coordinates The intercept term at ; is the spatial coordinate of the t-th independent variable The regression coefficient at ; is the value of the t-th independent variable at the i-th sample point; is the error term of the i-th sample point; is the number of independent variables.

[0070] In this example, the traditional least squares regression model assumes that the regression coefficient is constant across the entire study area, while the GWR model allows the regression coefficient to vary across different spatial locations, thereby better capturing spatial heterogeneity. By introducing a spatial weight function, the GWR model can more accurately reflect the spatial relationships between different deployment points, thereby improving the model's goodness of fit and prediction accuracy. Dynamically adjusting the regression coefficient based on the characteristics of actual data to accommodate the unique circumstances of different regions provides greater flexibility and adaptability.

[0071] In an optional manner, the expression of the spatial weight function of the geographically weighted regression model GWR is:

[0072] ;

[0073] in, is the spatial weight of point j estimated when fitting the model for shared bicycle deployment point i; is the distance between delivery points i and j; is the bandwidth, calculated using the AIC criterion.

[0074] In this embodiment, the spatial weight function considers the impact of spatial proximity on the model. Nearby observation points have a greater influence on the estimated regression coefficient of the target point, while more distant observation points have a smaller influence. Adaptively selecting the bandwidth (b) using the AIC criterion automatically adjusts the influence of the spatial weight based on the spatial distribution characteristics of the data, improving the adaptability of the model.

[0075] In an optional manner, the calculation formula of the AIC criterion is:

[0076] ;

[0077] in, is the maximum likelihood estimate of the variance of the random error term, ; is a function of bandwidth b; n is the number of samples; The sum of the squares of the differences between the actual value and the model prediction for each observation point.

[0078] In this example, the Akaike Information Criterion (AIC) avoids bias caused by subjective model selection and favors relatively simple models that fit the data well, thereby reducing the risk of overfitting. The formula includes a penalty for model complexity. While adding more independent variables to the model improves the model's goodness of fit (reducing RSS), it also increases model complexity, leading to an increase in the AIC value. The AIC criterion balances model goodness of fit and model complexity. In the GWR model, a bandwidth that is too large smooths out local variations, preventing the model from capturing spatial heterogeneity. A bandwidth that is too small can lead to model overfitting and reduce the generalization of predictions. The AIC criterion is used to select the bandwidth value that achieves the optimal balance between goodness of fit and model complexity.

[0079] In an alternative approach, when the sample size n is small, the modified AIC criterion AIC is used. C Determine the bandwidth b, the AIC C The calculation formula is:

[0080] ;

[0081] in, is the number of independent variables.

[0082] In this example, standard AIC tends to underestimate model complexity when the sample size is small and the model parameters are large, thus tending to select overfitted models. AICC introduces a correction term to correct for the bias in small sample sizes, making model selection more accurate. AICC also imposes a greater penalty on model complexity, thus tending to select simpler models in small sample sizes.

[0083] Step S104: Classify the shared bicycle deployment points in the target city according to the spatial data to obtain deployment point categories, where the deployment point categories include commercial services, scenic spots, public services, government and corporate offices, commercial residences, and transportation facilities.

[0084] In this embodiment, considering the categories of delivery points can better distinguish the usage patterns of shared bicycles in different areas. For example, there will be significant differences in the passenger flow patterns in commercial areas and residential areas. By classifying the delivery points, it is possible to model the characteristics of different types of areas and improve the prediction accuracy of the model in different areas. The goal of the GWR model is to capture spatial heterogeneity, but the classification of delivery points can further reflect this heterogeneity. The passenger flow of different types of delivery points is affected by different factors to different degrees. By classifying the delivery point types, these differences can be identified more accurately. By analyzing the passenger flow patterns of different delivery point categories, more targeted suggestions can be provided for shared bicycle operations. For example, more vehicles can be deployed in commercial areas, or night services can be provided at transportation hubs.

[0085] Step S105, based on the regression coefficient, significance level, goodness of fit, deployment point category and actual shared bicycle passenger flow data, calculate the prediction accuracy and improvement of the least squares regression model OLS and the geographically weighted regression model GWR in different deployment point categories to verify the advantage of the geographically weighted regression model GWR in capturing spatial heterogeneity.

[0086] In this embodiment, by comparing the prediction accuracy of OLS and GWR in different categories, the impact of spatial heterogeneity on the prediction results is inferred. If GWR performs significantly better than OLS in categories with strong spatial heterogeneity (such as a mixture of commercial and residential areas), the importance of spatial heterogeneity is verified. For example, GWR shows a clear advantage in the two categories of commercial services and residential areas, with an improvement of 20%. This shows that in these two areas, spatial heterogeneity is relatively strong, and the GWR model can better capture these differences. In transportation hubs, scenic spots, public services, and government and corporate office areas, the improvement of GWR is relatively small, indicating that the spatial heterogeneity in these areas is relatively weak, and the OLS model can already make good predictions. By analyzing the regression coefficients, we can further understand the spatial differences captured by the GWR model in different areas. For example, it may be found that in commercial areas, POI density has a greater impact on passenger flow, while in residential areas, population density has a greater impact on passenger flow.

[0087] According to the solution provided by the present invention, the spatial data and socioeconomic data of the shared bicycle deployment points in the target city are collected to obtain an original data set, wherein the spatial data include latitude and longitude and the location of the deployment point, and the socioeconomic data include population density, POI data, road density and housing prices; the original data set is used to perform spatial autocorrelation test and multicollinearity test on the independent variables to obtain an effective independent variable data set, spatial autocorrelation and multicollinearity test results of the effective independent variables; the effective independent variable data set, spatial autocorrelation of the effective independent variables, multicollinearity test results and housing price factor data are input into the least squares regression model OLS and the geographically weighted regression model GWR, and the least squares regression model is output. The regression coefficient, significance level and goodness of fit of the regression model OLS and the geographically weighted regression model GWR are multiplied; the target city's shared bicycle delivery points are divided into delivery point categories according to the spatial data to obtain delivery point categories, and the delivery point categories include commercial services, scenic spots, public services, government and corporate offices, commercial residences and transportation facilities; according to the regression coefficient, significance level, goodness of fit, delivery point categories and actual shared bicycle passenger flow data, the prediction accuracy and improvement of the least squares regression model OLS and the geographically weighted regression model GWR in different delivery point categories are calculated to verify the advantage of the geographically weighted regression model GWR in capturing spatial heterogeneity. The present invention takes into account the spatial heterogeneity relationship between shared bicycle passenger flow and influencing factors through the geographically weighted regression model, and can adjust the regression coefficient according to different geographical locations, thereby more accurately reflecting the influence characteristics of the local area and significantly improving the prediction accuracy.

[0088] Figure 2 The following is a schematic diagram showing the framework of a device for predicting the passenger flow of shared bicycles using a geographically weighted regression model according to an embodiment of the present invention. The device for predicting the passenger flow of shared bicycles using a geographically weighted regression model includes:

[0089] Data collection module 210 is used to collect spatial data and socioeconomic data of shared bicycle deployment points in the target city to obtain an original data set, wherein the spatial data includes latitude and longitude and the location of the deployment point, and the socioeconomic data includes population density, POI data, road density, and housing prices;

[0090] A variable testing module 220 is configured to perform a spatial autocorrelation test and a multicollinearity test on the independent variables using the original data set, and obtain a valid independent variable data set and spatial autocorrelation and multicollinearity test results of the valid independent variables;

[0091] The model regression module 230 is used to input the effective independent variable dataset, the spatial autocorrelation of the effective independent variables, the multicollinearity test results, and the housing price factor data into the least squares regression model OLS and the geographically weighted regression model GWR, and output the regression coefficients, significance levels, and goodness of fit of the least squares regression model OLS and the geographically weighted regression model GWR;

[0092] A classification module 240 is configured to classify the shared bicycle deployment points in the target city according to the spatial data to obtain deployment point categories, wherein the deployment point categories include commercial services, scenic spots, public services, government and corporate offices, commercial residences, and transportation facilities;

[0093] The accuracy evaluation module 250 is used to calculate the prediction accuracy and improvement of the least squares regression model OLS and the geographically weighted regression model GWR in different deployment point categories based on the regression coefficient, significance level, goodness of fit, deployment point category and actual shared bicycle passenger flow data, so as to verify the advantage of the geographically weighted regression model GWR in capturing spatial heterogeneity.

[0094] Figure 3 The schematic diagram of the structure of the computing device embodiment of the present invention is shown. The specific embodiment of the present invention does not limit the specific implementation of the computing device.

[0095] like Figure 3 As shown, the computing device may include: a processor (processor) 302, a communication interface (Communications Interface) 304, a memory (memory) 306, and a communication bus 308.

[0096] Processor 302, communication interface 304, and memory 306 communicate with each other via communication bus 308. Communication interface 304 is used to communicate with other devices, such as clients or other server network elements. Processor 302 is used to execute program 310, specifically the steps described in the embodiment of the method for predicting shared bicycle passenger flow using a geographically weighted regression model.

[0097] Specifically, the program 310 may include program codes, which include computer operation instructions.

[0098] Processor 302 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The one or more processors included in a computing device may be of the same type, such as one or more CPUs, or different types, such as one or more CPUs and one or more ASICs.

[0099] The memory 306 is used to store the program 310. The memory 306 may include a high-speed RAM memory, or may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0100] According to the solution provided by the present invention, the spatial data and socioeconomic data of the shared bicycle deployment points in the target city are collected to obtain an original data set, wherein the spatial data include latitude and longitude and the location of the deployment point, and the socioeconomic data include population density, POI data, road density and housing prices; the original data set is used to perform spatial autocorrelation test and multicollinearity test on the independent variables to obtain an effective independent variable data set, spatial autocorrelation and multicollinearity test results of the effective independent variables; the effective independent variable data set, spatial autocorrelation of the effective independent variables, multicollinearity test results and housing price factor data are input into the least squares regression model OLS and the geographically weighted regression model GWR, and the least squares regression model is output. The regression coefficient, significance level and goodness of fit of the regression model OLS and the geographically weighted regression model GWR are multiplied; the target city's shared bicycle delivery points are divided into delivery point categories according to the spatial data to obtain delivery point categories, and the delivery point categories include commercial services, scenic spots, public services, government and corporate offices, commercial residences and transportation facilities; according to the regression coefficient, significance level, goodness of fit, delivery point categories and actual shared bicycle passenger flow data, the prediction accuracy and improvement of the least squares regression model OLS and the geographically weighted regression model GWR in different delivery point categories are calculated to verify the advantage of the geographically weighted regression model GWR in capturing spatial heterogeneity. The present invention takes into account the spatial heterogeneity relationship between shared bicycle passenger flow and influencing factors through the geographically weighted regression model, and can adjust the regression coefficient according to different geographical locations, thereby more accurately reflecting the influence characteristics of the local area and significantly improving the prediction accuracy.

[0101] Those skilled in the art will appreciate that the modules in the devices of the embodiments can be adaptively modified and deployed in one or more devices different from the embodiments. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and furthermore, they can be divided into multiple sub-modules, sub-units, or sub-components. All features disclosed in this specification (including the accompanying claims, abstract, and drawings), as well as all processes or units of any method or device disclosed therein, can be combined in any combination, unless at least some of such features and / or processes or units are mutually exclusive. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature serving the same, equivalent, or similar purpose. Furthermore, those skilled in the art will appreciate that although some embodiments herein include certain features and not others included in other embodiments, combinations of features from different embodiments are intended to fall within the scope of the present invention and form different embodiments. For example, in the claims below, any of the claimed embodiments may be used in any combination. The present invention may be implemented by means of hardware comprising several distinct elements and by means of a suitably programmed computer. In a unit claim that lists several means, several of these means may be embodied by the same hardware item. Unless otherwise specified, the steps in the above embodiments should not be understood as limiting the order of execution.

Claims

1. Collect spatial data and socioeconomic data of shared bicycle deployment points in the target city to obtain the original data set, where: The spatial data includes latitude and longitude and the location of the delivery point, and the socioeconomic data includes population density, POI data, road density and housing prices; Using the original data set to perform spatial autocorrelation test and multicollinearity test on the independent variables, to obtain a valid independent variable data set, spatial autocorrelation and multicollinearity test results of the valid independent variables; Input the effective independent variable data set, the spatial autocorrelation of the effective independent variables, the multicollinearity test results, and the housing price factor data into the least squares regression model OLS and the geographically weighted regression model GWR, and output the regression coefficients, significance levels, and goodness of fit of the least squares regression model OLS and the geographically weighted regression model GWR; Classifying the shared bicycle deployment points in the target city according to the spatial data to obtain deployment point categories, wherein the deployment point categories include commercial services, scenic spots, public services, government and corporate offices, commercial residences, and transportation facilities; Based on the regression coefficient, significance level, goodness of fit, deployment point category and actual shared bicycle passenger flow data, the prediction accuracy and improvement of the least squares regression model OLS and the geographically weighted regression model GWR in different deployment point categories are calculated to verify the advantage of the geographically weighted regression model GWR in capturing spatial heterogeneity.

2. The method for predicting shared bicycle passenger flow based on a geographically weighted regression model according to claim 1 is characterized in that: The socioeconomic data includes the population and housing prices within the buffer zone, wherein the population is estimated by the number of households around the shared bicycle deployment point, and the housing prices are calculated by the unit price of housing in the residential areas within the buffer zone.

3. The method for predicting shared bicycle passenger flow based on a geographically weighted regression model according to claim 1, characterized in that: The expression of the geographically weighted regression model GWR is: ; in, is the shared bicycle passenger flow of the i-th sample point; is the spatial coordinate of the i-th sample point; In space coordinates The intercept term at ; is the spatial coordinate of the t-th independent variable The regression coefficient at ; is the value of the t-th independent variable at the i-th sample point; is the error term of the i-th sample point; is the number of independent variables.

4. The method for predicting shared bicycle passenger flow based on a geographically weighted regression model according to claim 1, characterized in that: The expression of the spatial weight function of the geographically weighted regression model GWR is: ; in, is the spatial weight of point j estimated when fitting the model for shared bicycle deployment point i; is the distance between delivery points i and j; is the bandwidth, calculated using the AIC criterion.

5. The method for predicting shared bicycle passenger flow based on a geographically weighted regression model according to claim 4 is characterized in that: The calculation formula of the AIC criterion is: ; in, is the maximum likelihood estimate of the variance of the random error term, ; is a function of bandwidth b; n is the number of samples; The sum of the squares of the differences between the actual value and the model prediction for each observation point.

6. The method for predicting shared bicycle passenger flow based on a geographically weighted regression model according to claim 1, characterized in that: The performing of a spatial autocorrelation test and a multicollinearity test on the independent variables using the original data set to obtain a valid independent variable data set and the spatial autocorrelation and multicollinearity test results of the valid independent variables further comprises: The Moran index is used to test the spatial autocorrelation of each independent variable in the original data set, and the Moran index value and significance level of each independent variable are calculated. If the Moran index is significant, the independent variable has significant spatial autocorrelation. Calculate the variance inflation factor of all independent variables in the original data set. If the variance inflation factor of an independent variable exceeds the threshold, there is multicollinearity between the independent variable and other independent variables. According to the results of the spatial autocorrelation test and the multicollinearity test, the effective independent variables are screened out, and the screened independent variable data set, the Moran's index value and the variance inflation factor of each effective independent variable are obtained.

7. The method for predicting shared bicycle passenger flow based on a geographically weighted regression model according to claim 5, characterized in that: When the sample size n is small, the modified AIC criterion AIC is used. C Determine the bandwidth b, the AIC C The calculation formula is: ; in, is the number of independent variables.

8. The method for predicting shared bicycle passenger flow based on a geographically weighted regression model according to claim 6, characterized in that: The threshold of the variance inflation factor is 10. When the variance inflation factor of an independent variable exceeds 10, there is serious multicollinearity between the independent variable and other independent variables. The independent variable with the largest variance inflation factor is eliminated until the variance inflation factors of all independent variables are less than 10.

9. A shared bicycle passenger flow prediction device based on a geographically weighted regression model, characterized in that: include: A data collection module is used to collect spatial data and socioeconomic data of shared bicycle deployment points in the target city to obtain an original data set, wherein the spatial data includes latitude and longitude and the location of the deployment point, and the socioeconomic data includes population density, POI data, road density, and housing prices; A variable testing module is used to perform a spatial autocorrelation test and a multicollinearity test on the independent variables using the original data set to obtain a valid independent variable data set and the spatial autocorrelation and multicollinearity test results of the valid independent variables; A model regression module is used to input the effective independent variable data set, the spatial autocorrelation of the effective independent variables, the multicollinearity test results, and the housing price factor data into the least squares regression model OLS and the geographically weighted regression model GWR, and output the regression coefficients, significance levels, and goodness of fit of the least squares regression model OLS and the geographically weighted regression model GWR; a classification module for classifying the shared bicycle deployment points in the target city according to the spatial data to obtain deployment point categories, wherein the deployment point categories include commercial services, scenic spots, public services, government and corporate offices, commercial residences, and transportation facilities; The accuracy evaluation module is used to calculate the prediction accuracy and improvement of the least squares regression model OLS and the geographically weighted regression model GWR in different deployment point categories based on the regression coefficient, significance level, goodness of fit, deployment point category and actual shared bicycle passenger flow data, so as to verify the advantage of the geographically weighted regression model GWR in capturing spatial heterogeneity.

10. A computing device comprising: A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the shared bicycle passenger flow prediction method based on the above-mentioned geographically weighted regression model.

Citation Information

Cited By

  • Method, device and system for monitoring suitable growth area of harmful insects in grassland, and storage medium

    CN121959256A