Urban street space analysis method based on extreme gradient lifting algorithm
The spatial characteristics of urban streets are quantified through extreme gradient enhancement algorithms, and the problem of low efficiency in automatic identification of urban streets in the existing technology is solved, and efficient and reliable street space recognition is achieved.
Patent Information
- Application Number
- CN202510714076.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The prior art is difficult to efficiently and automatically identify the spatial range of urban streets, especially in large-scale data and high-dimensional feature processing, and traditional methods are costly and inefficient.
The extreme gradient enhancement algorithm is used to transform the road network and POI data into a graph structure containing equidistant road network nodes, and quantify the road center, business structure complexity and business spatial distribution characteristics, and use the extreme gradient enhancement algorithm to train the urban street spatial recognition model to automatically identify the street spatial scope.
It significantly improves the efficiency and reliability of automatic identification of urban street spaces, and can identify the location and scope of street spaces on a large scale, which greatly improves efficiency compared with traditional methods.
Smart Images

Figure CN120256751A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of urban design and artificial intelligence, and particularly relates to a method for analyzing urban street space based on the extreme gradient boosting algorithm. Background Art
[0002] Urban public space refers to open places accessible to all the public, mainly used for citizens' daily life and social activities, including outdoor spaces such as urban squares, streets, parks, etc. The core of determining urban public space lies in its public nature. The criteria for determining the public nature of public space generally include institutional and institutional factors (i.e., whether the ownership and management methods are public), accessibility (whether it is open to everyone, whether it can accommodate diverse activities, and whether it has good physical and symbolic accessibility), social value, etc.
[0003] Urban public space is complex, multi - dimensional and dynamic. Effective definition of its scope is particularly important and is the basis for further research. However, for a long time, there have been great difficulties in defining the spatial scope of urban public space. Existing methods mostly rely on planning texts combined with a large number of on - site surveys and rely on manual subjective judgment, resulting in high costs and low efficiency.
[0004] Streets are the most important part of urban public space. On the plane map, the morphological characteristics of streets and roads are highly similar. Therefore, they are also the type of public space with the most difficult automatic recognition of spatial boundaries.
[0005] With the continuous development of machine learning algorithms, many classifiers have been widely applied to various practical problems, such as support vector machines (SVM), random forests (Random Forest), traditional gradient - boosting decision trees (GBDT), etc. Although these methods perform well in different scenarios, there are still some limitations, especially in dealing with large - scale data, high - dimensional features and training efficiency. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for analyzing urban street space based on the extreme gradient boosting algorithm, which can analyze the spatial scope of urban streets based on the extreme gradient boosting algorithm and can significantly improve the analysis efficiency and reliability. The technical solution adopted by the present invention is as follows.
[0007] On the one hand, the present invention provides a method for analyzing urban street space based on the extreme gradient boosting algorithm, including:
[0008] Obtaining road network vector data and POI (Point of Interest) vector data of the target analysis area;
[0009] Based on the road network vector data, converting the road network into a graph structure containing equidistant road network nodes according to the road level;
[0010] Based on the POI vector data and the graph structure containing equidistant road network nodes, the POI is assigned to the nearest road node to obtain a graph structure containing POI information and equidistant road network nodes;
[0011] Based on the graph structure containing POI information and equidistant road network nodes, obtain the centrality index of each node for measuring the road centrality, the business structure complexity index for measuring the complexity of the business structure, and the number of target POI paths reachable for measuring the spatial distribution characteristics of the business;
[0012] The longitude and latitude coordinates, proximity centrality and betweenness centrality, weighted Hill diversity index, and the number of path reachable along the road distance of each node are input into the urban street space recognition model pre-trained by the extreme gradient boosting algorithm to obtain the urban street point set of the target parsing area;
[0013] Based on the urban street point set, according to the road grade and building outline data in the road network vector data of the target parsing area, the urban street space with boundaries in the target parsing area is obtained.
[0014] Optionally, based on the road network vector data, converting the road network into a graph structure containing equidistant road network nodes according to the road level includes:
[0015] Generate road network Shapefile based on road network vector data;
[0016] The road network vector data is converted into the df_streets variable of the GeoDataFrame object through the read_file function in the Python geopandas library;
[0017] Convert the road network GeoDataFrame to the MultiGraph data structure of the NetworkX library and define it as the clipped_momepy variable;
[0018] Among them, in the MultiGraph data structure of the road network, road intersections and road segmentation points are used as network nodes, and road entities are used as connecting edges to generate a graph structure containing equidistant road network nodes;
[0019] The road segmentation points are obtained in the following manner: for a road between any two adjacent road intersections, if its length exceeds a set segment length threshold, the road is equidistantly segmented to obtain corresponding road segmentation points.
[0020] The above technical solution can completely retain the geometric connectivity characteristics and multi-edge features of the original road network, providing a data basis for the reliability of subsequent analysis.
[0021] Optionally, the source of the road network vector data includes the road network data of the open-source map website OpenStreetMap and the commercial road network information of the Baidu Map Developer Platform;
[0022] The step of converting the road network into a graph structure containing equidistant road network nodes further includes: after converting the road network vector data into the df_streets variable of the GeoDataFrame object, calling the to_crs method to unify the coordinate systems of data from different sources to UTM Zone 50N. This implementation can ensure the projection consistency of multi-source road network vector data during spatial analysis, significantly improving the calculation accuracy of subsequent quantitative analysis while maintaining the topological integrity of the road network.
[0023] Optionally, the assignment of POIs to the nearest road nodes includes:
[0024] Converting the Shapefile file data of the POI vector data into the data_gdf variable of the GeoDataFrame object through the read_file function of the Python geopandas library, and performing preprocessing operations including projection normalization and outlier filtering;
[0025] Based on the road network MultiGraph data structure, for each POI, according to the nearest road edge, assign it to the two nearest road network nodes of the nearest road, obtaining a graph structure containing POI information and equidistant road network nodes, and a POI data_gdf variable containing adjacent street information;
[0026] In this graph structure, each node includes road network node information and brief POI information assigned to this road network node, and this brief POI information can be mapped to the POI vector data in the data_gdf variable;
[0027] In the POI data_gdf variable containing adjacent street information, each POI also records the brief information of the road network node to which this POI is assigned.
[0028] Optionally, based on the graph structure containing POI information and equidistant road network nodes, obtaining various node centrality indicators for measuring road centrality includes:
[0029] Convert the data of the clipped_momepy variable in the MultiGraph data structure into the data structures required for network analysis: the nodes_gdf variable in the GeoDataFrame format of node data, the edges_gdf variable in the GeoDataFrame format of connected edge data, and the network_structure variable that describes the topological relationship, where the network_structure variable is used to record the topological relationship information between the nodes of the road network and the POIs.
[0030] Conduct road centrality analysis: Traverse each node in the graph structure. Taking each node as the center, delimit the boundary of the local window network within the preset walking distance threshold range, and calculate the closeness centrality and betweenness centrality of the central node of the local window network as the centrality indicators.
[0031] The formula for closeness centrality is:
[0032]
[0033] In the formula, represents the closeness centrality of node , n is the total number of nodes in the local window network, is the distance between node and node in the local window network.
[0034] The formula for betweenness centrality is:
[0035]
[0036] In the formula, is the betweenness centrality of node , is the number of the shortest paths between point and point in the local window network, is the number of the shortest paths between point and point passing through point in the local window network.
[0037] In the above technical solutions, the closeness centrality is based on the distance of the shortest path to measure the closeness between points. The betweenness centrality is used to measure the extent to which a point controls the connections between other points. By delimiting the local window network to calculate the centrality of each node, the network range associated with the current node is dynamically defined, which improves the calculation efficiency while avoiding the urban boundary effect.
[0038] Optionally, the step of defining the local window network boundary within a preset walking distance threshold range includes: for each node, taking the node as the center, defining the local window network boundary within multiple preset walking distance threshold ranges;
[0039] The method of obtaining the centrality index of each node for measuring the road centrality also includes: for each node, respectively calculating its proximity centrality and betweenness centrality in multiple local window networks corresponding to different walking distance threshold ranges, to obtain multiple groups of centrality index data.
[0040] The above scheme sets multiple distance thresholds to ensure that window division can cover different spatial scales while maintaining data integrity, and can balance calculation accuracy and performance.
[0041] Optionally, based on the graph structure containing POI information and equidistant road network nodes, a business structure complexity index for measuring the complexity of the business structure is obtained, including:
[0042] Based on the POI data_gdf variable containing the adjacent street information and the network_structure variable containing the street adjacent POI information, for each road network node, the weighted Hill diversity index is calculated as follows:
[0043]
[0044] in:
[0045]
[0046]
[0047] Where: Represents a road network node The weighted Hill diversity index; and Represents the nodes assigned to the road network No. POI type and POI types; Represents the nodes assigned to the road network The total number of POI types; Represents the nodes assigned to the road network No. The number of POIs of each POI type; Indicates that at the node Pedestrians walk to the The decay weight of each POI, Indicates The mean of the decay weights corresponding to all POIs under the POI type; The parameter is used to control the degree of emphasis on the richness of POI types and the balance of POI types; e is the natural constant, is the road network node to the POI node distance, is the preset walking distance threshold.
[0048] Optionally, the target POI path reachable quantity for measuring the spatial distribution characteristics of business forms is, for each road network node, the number of path instances that can reach the specified POI type along the road within the set distance threshold.
[0049] Optionally, based on the set of urban street points, according to the road grades and building contour data in the road network vector data of the target analysis area, the urban street space with boundaries in the target analysis area is obtained, including:
[0050] According to the node ID field of the point elements in the set of urban street points, establish a topological connection with the start and end point IDs of the road segments in the road network, and reconstruct the street network of the target analysis area;
[0051] According to the preset buffer radius mapped by the road grade field in the target analysis area vector data, perform a spatial difference operation on the buffer and the building polygon data to remove the road-side building encroachment area; output the street space polygon set and store it in the Shapefile format to obtain the final street space position and range.
[0052] In a second aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. The feature is that when the computer program is executed by a processor, the steps of the method for analyzing urban street space based on the extreme gradient boosting algorithm as described in the first aspect are implemented.
[0053] Beneficial effects
[0054] The present invention realizes the correlation modeling of the urban road network characteristics and the spatial distribution characteristics of business forms by converting the road network and POI data into a graph structure containing POI information and equidistant road network nodes, and achieves the research idea of quantifying the business forms by calculating the distance along the road extension direction. Based on the constructed graph structure, quantitatively analyze the key influencing factors of the urban street space, including road centrality, spatial distribution characteristics of business forms, and complexity of business form structure, and use the extreme gradient boosting algorithm to train an automatic recognition model to locate the urban street space based on the quantified key influencing factors, and further judge the scope and boundary of the urban street space, realizing the intelligent determination of the position and scope of the urban street space. It is of great significance for large-scale automatic recognition of the position and scope of urban street space, and the efficiency is greatly improved compared with traditional methods. Description of the drawings
[0055] Figure 1 It is a schematic flowchart of an embodiment of the present invention;
[0056] Figure 2 It is a visualization schematic diagram of the graph structure after converting the road network within a certain area in Nanjing City as the target analysis area at a refined distance of 10 m in an embodiment of the present invention;
[0057] Figure 3 It is a schematic diagram of the position and range results of the urban street space obtained by analyzing the target analysis area in an embodiment of the present invention. Detailed implementation manners
[0058] The following is further described in conjunction with the accompanying drawings and specific embodiments.
[0059] Extreme Gradient Boosting (XGBoost) is an efficient and powerful machine learning algorithm, which belongs to the extension of gradient boosting decision trees. It constructs a strong learner by combining multiple weak decision trees (usually regression trees) and is widely used in classification, regression, and ranking tasks, featuring high efficiency, flexibility, and convenience.
[0060] Considering the complex, multi-dimensional, and dynamic characteristics of urban public spaces, the present invention selects the extreme gradient boosting algorithm as the training algorithm for the urban street space recognition model to achieve the automatic recognition of urban street spaces; compared with other machine learning classifier algorithms such as the random forest algorithm, the extreme gradient boosting algorithm has higher accuracy, precision, recall rate, and F1 score, and can overcome the limitations of existing classification algorithms in dealing with large-scale data, high-dimensional features, and training efficiency.
[0061] On a plan view, the morphological features of streets and roads are highly similar, and the main difference between them is that compared with roads, streets are more prominent in terms of enclosure and slow traffic characteristics, tend to carry composite functions, and should have a series of interconnected places for people to stay and communicate. The important basis for distinguishing streets and roads is the characteristics of road traffic flow, the land use nature on both sides of the road, and the main service objects. Therefore, the present invention applies heuristic algorithms, shortest path algorithms, combined with machine learning classifier algorithms, to achieve the automatic discrimination of urban street spaces.
[0062] Embodiment 1
[0063] This embodiment introduces a method for analyzing urban street spaces based on the extreme gradient boosting algorithm. The key to its technical solution lies in how to establish an associated model between urban road network features and business space distribution features, quantify and construct them into a data structure that can be processed by machine learning, and reverse-transform the recognition results output by the algorithm into street space vector data. The method includes:
[0064] Obtain the road network vector data and POI (Point of Interest) vector data of the target parsing area;
[0065] Based on the road network vector data, converting the road network into a graph structure containing equidistant road network nodes according to the road level;
[0066] Based on the POI vector data and the graph structure containing equidistant road network nodes, the POI is assigned to the nearest road node to obtain a graph structure containing POI information and equidistant road network nodes;
[0067] Based on the graph structure containing POI information and equidistant road network nodes, obtain the centrality index of each node for measuring the road centrality, the business structure complexity index for measuring the complexity of the business structure, and the number of target POI paths reachable for measuring the spatial distribution characteristics of the business;
[0068] The closeness centrality and betweenness centrality of each node, the weighted Hill diversity index, and the number of path reachable along the road distance are input into the urban street space recognition model pre-trained by the extreme gradient boosting algorithm to obtain the urban street point set of the target parsing area;
[0069] Based on the urban street point set, according to the road grade and building outline data in the road network vector data of the target parsing area, the urban street space with boundaries in the target parsing area is obtained.
[0070] The specific implementation of the above technical solution involves the following aspects.
[0071] It should be noted here that the extreme gradient boosting algorithm is used to pre-train the urban street space recognition model, which is based on the marked urban street coordinate points, the proximity centrality and betweenness centrality of each node in the training area, the weighted Hill diversity index, and the path reachability data along the road distance. The following is a unified introduction to the data processing content involved in the training process and the actual detection process.
[0072] 1. Road network data and POI data processing
[0073] In this embodiment, the road network vector data can be downloaded from an open source map website, and the road network can be converted into a graph structure containing equidistant road network nodes according to the road level. Specifically:
[0074] For the road network sample data required for model training, we obtained the training sample area vector road network data from the open source map website OpenStreetMap, and combined it with the commercial road network information of the Baidu Map Developer Platform to generate the training sample area road network Shapefile file. We converted the road network vector data into the df_streets variable of the GeoDataFrame object through the read_file function of the geopandas library.
[0075] To ensure the projection consistency of spatial analysis, the to_crs method can be called to convert the coordinate system of the road network vector data to UTM Zone 50N (EPSG:32650). Then the road network GeoDataFrame is converted to the MultiGraph data structure of the NetworkX library, defined as the variable clipped_momepy, to support the expression of undirected multi-edge networks, that is, using road intersections as network nodes and road entities as connecting edges to generate a graph structure containing equidistant road network nodes, so as to fully preserve the geometric connectivity and multi-edge characteristics of the original road network.
[0076] In order to eliminate the scale effect of uneven density of original road network nodes on the subsequent calculation of centrality index, this embodiment uses a road network segmentation algorithm for refined processing when generating a graph structure containing equidistant road network nodes, that is: for any road between two adjacent road intersections, if its length exceeds the set segment length threshold, the road is equidistantly segmented to obtain the corresponding road segmentation points. The segment length threshold can be set as needed, such as setting 10 meters as the maximum segmentation threshold, and equidistantly segmenting the road edges that exceed the length to generate a graph structure containing equidistant road network nodes. This operation significantly improves the calculation accuracy of subsequent quantitative analysis while maintaining the topological integrity of the road network.
[0077] For the POI data required for model training, first download the POI vector dataset of the training sample area from the open source map website, and perform data preprocessing operations including filtering outliers and filling missing values. POI is called Point of Interest, which covers multiple urban functional elements such as residential areas, commercial facilities, and transportation stations. Based on the preprocessed data, this embodiment loads the Shapefile data of POI into the data_gdf variable of the GeoDataFrame object through the read_file function of the geopandas library, and performs projection normalization operations to unify the data to the EPSG:32650 coordinate system.
[0078] For the area of urban street space to be analyzed, refer to the above process to obtain the road network vector data and POI vector data of the corresponding area, and perform corresponding processing.
[0079] II. Association Processing of Road Network Vector Data and POI Data
[0080] In this part, in this embodiment, the shortest path algorithm is used to allocate POI data to the two nearest road network nodes according to the nearest adjacent street edge, obtaining an equidistant node graph structure of POI data calculated along the road. This allocation mechanism simulates the directional characteristics of POI accessibility from the perspective of pedestrians, facilitating subsequent precise quantitative analysis.
[0081] The specific operations for allocating POIs to the nearest road nodes in this embodiment include the following:
[0082] Based on the road network MultiGraph data structure, for each POI, according to the nearest road edge, it is allocated to the two nearest road network nodes of this nearest road, obtaining a graph structure containing POI information and equidistant road network nodes. Here, it is still the clipped_momepy variable, as well as the POI data_gdf variable containing adjacent street information.
[0083] In the graph structure containing POI information and equidistant road network nodes, each node includes road network node information and brief POI information allocated to this road network node. This brief POI information can be mapped to the POI vector data in the data_gdf variable, and based on this, the specific detailed information stored in the data_gdf can be found. In the POI data_gdf variable containing adjacent street information, at each POI, the brief information of the road network node to which this POI is allocated is also recorded.
[0084] So far, the clipped_momepy data can be converted into three structures required for network analysis: the GeoDataFrame format nodes_gdf of road network node data, the GeoDataFrame format edges_gdf of connection edge data, and the network_structure describing the topological relationship, for subsequent calculations. Among them, the network_structure variable is used to record the topological relationship information between road network nodes and POIs.
[0085] The principle of the association processing method for training sample data and the data of the area to be parsed is the same.
[0086] III. Feature Parameter Extraction
[0087] In order to effectively improve the efficiency of urban street space analysis, the present invention adopts neural network technology and the extreme gradient boosting algorithm. Based on the feature parameter extraction results of the training sample area and the marked labels of known urban street space points, the neural network model is trained. The feature parameters include the node centrality indicators for measuring road centrality of each road network node, the business format structure complexity indicators for measuring the complexity of the business format structure, and the target POI path accessibility quantity for measuring the business format spatial distribution characteristics. The extraction methods of each feature parameter are as follows.
[0088] 3.1) Extraction of centrality indicators
[0089] In this embodiment, a heuristic algorithm is used to calculate the closeness centrality and betweenness centrality of each road network node on the graph structure containing equidistant road network nodes, for measuring road centrality.
[0090] In the road centrality analysis, this embodiment adopts a spatial analysis method based on a "moving window": traverse each node in the network, with the current node as the center, and delimit the network boundary within a preset range of multiple walking distance thresholds, such as within ranges of 50m, 100m, 200m, 400m, etc., and calculate the centrality of the network within the local window.
[0091] Based on the GeoDataFrame format nodes_gdf of the aforementioned road network node data, the GeoDataFrame format edges_gdf of the connection edge data, and the network_structure describing the topological relationship, road centrality analysis is carried out. Specifically:
[0092] Traverse each node in the graph structure, respectively take each node as the center, delimit the local window network boundary within a preset range of multiple walking distance thresholds, and calculate the closeness centrality and betweenness centrality of the central node of the local window network as the centrality indicators;
[0093] The formula for calculating the closeness centrality is:
[0094]
[0095] In the formula, represents the closeness centrality of node , n is the total number of nodes in the local window network, is the distance between node and node in the local window network;
[0096] The formula for calculating the betweenness centrality is:
[0097]
[0098] In the formula, Is a node The betweenness centrality, Yes and Point The number of shortest paths between Yes and Point Passing point The number of shortest paths.
[0099] After the above processing, for each node in the graph structure, a set of multiple closeness centrality and betweenness centrality characteristic parameters can be obtained as part of the neural network model input. This implementation method realizes the dynamic definition of the network range associated with the current node by setting multiple distance thresholds, ensuring that the window division can cover different spatial scales while maintaining data integrity, improving the calculation efficiency while avoiding the city boundary effect, and balancing the calculation accuracy and reliability performance.
[0100] 3.2 Extraction of business structure complexity indicators
[0101] This embodiment calculates the weighted Hill diversity index for the road network node graph structure containing POI information to measure the complexity of the business structure, that is, the data_gdf containing adjacent street information and the network structure network_structure containing adjacent POI information are used to calculate the complexity of the business structure and the spatial distribution characteristics of the business. In the process of calculating the complexity of the business structure and the spatial distribution characteristics of the business, the aforementioned "moving window" spatial analysis method is also used.
[0102] Considering that pedestrians’ willingness to walk to POI gradually decreases with increasing distance, a negative exponential decay function is used to simulate spatial impedance. Pedestrians walk to the The decay weight of a POI is:
[0103]
[0104] When calculating the complexity of the business structure, the weighted Hill diversity index is used, and the road network nodes The weighted Hill diversity index is expressed as:
[0105]
[0106] For road network nodes The first The corresponding average attenuation weight for each POI type is expressed as:
[0107]
[0108] In the formula: represents the weighted Hill diversity index of the road network node ; represents the th POI type assigned to the road network node represents the total number of all POI types assigned to the road network node ; represents the ratio of the number of POIs of the th POI type assigned to the road network node to the total number of POIs assigned to the road network node The parameter is used to control the degree of emphasis on the richness and balance of POI types, and can be set according to experience and needs; e is the natural constant, is the distance from the road network node to the POI node ; is the preset walking distance threshold. To optimize the operation speed and ensure data integrity, the distance threshold can be selected as 50m, 100m, 200m, and 400m.
[0109] 3.3 Extraction of the spatial distribution characteristics of business formats
[0110] In this embodiment, when calculating the spatial distribution characteristics of business formats, the reachable quantity index is used, which represents the number of path instances of a given POI type corresponding to a road network node that can reach the specified POI type along the road within the specified distance threshold.
[0111] IV. Training of the urban street space recognition model
[0112] Based on the node coordinates and feature parameters of the training samples extracted above, that is, reading the information of the network nodes nodes_gdf, the connecting edges edges_gdf, the network structure network_structure, etc., a reference set graph structure is jointly constructed and divided into two parts: a training set and a validation set. The urban street space boundary is manually marked as the supervision label. The longitude and latitude coordinates, proximity centrality and betweenness centrality, negative exponential weighted Hill coefficient, and reachable distance number of the road network nodes in the former graph structure are used as input data, and an extreme gradient boosting algorithm is used to train an urban street space recognition model.
[0113] The validation set is used to verify the prediction accuracy, precision, recall rate, and F1 score of the model to determine the quality of the model. Finally, the parameters of the extreme gradient boosting algorithm that can achieve better results are adjusted, and an urban street space recognition model with a higher score is obtained with these parameters.
[0114] V. Identification of Urban Street Spaces in the Area to be Parsed
[0115] When identifying the area of the urban street space to be parsed, input the longitude and latitude coordinates of the nodes and the feature parameter data of the target area map structure after quantitative preprocessing into the urban street space identification model, and the point set goal_nodes_gdf of the urban street space in the target area output by the model can be obtained.
[0116] VI. Data Post-Processing
[0117] Based on the urban street point set obtained in the fifth step, and according to the road grades and building contour data in the road network vector data of the target parsing area, the urban street space with boundaries in the target parsing area is obtained. Specifically:
[0118] According to the node ID field of the goal_nodes_gdf point feature in the urban street point set data, establish a topological connection with the start and end point IDs of the road segments in the road network, and reconstruct the street network of the target parsing area;
[0119] Map the preset buffer radius according to the road grade field in the target parsing area vector data, perform a spatial difference operation on the buffer and the building polygon data, and remove the area occupied by roadside buildings; output the street space polygon set and store it in the Shapefile format to obtain the final street space position and range.
[0120] Example 2
[0121] In this example, Gulou District of Nanjing is used as the reference set, and the area of the main urban area of Nanjing except Gulou District is used as the target set to illustrate the method of automatically identifying urban street spaces using the extreme gradient boosting algorithm:
[0122] First, obtain the vector road network data of the target area in Nanjing from the open-source map website OpenStreetMap, and supplement it with the commercial road network information from the Baidu Map Developer Platform. Finally, generate the Shapefile file of the road network in the training area. Convert the road network vector data into a GeoDataFrame object (variable df_streets) using the read_file function of the geopandas library. Subsequently, call the to_crs method to unify its coordinate system to UTM Zone 50N (EPSG:32650) to ensure the projection consistency of spatial analysis. Convert the road network GeoDataFrame into a MultiGraph data structure of the NetworkX library (supporting undirected multi-edge network representation): use road intersections as network nodes and road entities as connecting edges, fully retaining the geometric connectivity characteristics and multi-edge features of the original road network. To eliminate the scale effect caused by uneven node density in the original road network on the calculation of centrality indicators, a road network segmentation algorithm is used for fine processing. Set 10 meters as the maximum segmentation threshold, and equally divide the road edges that exceed the length to generate a graph structure (variable clipped_momepy) with equally spaced road network nodes, as Figure 2 shown. This operation significantly improves the calculation accuracy of subsequent quantitative analysis while maintaining the topological integrity of the road network.
[0123] Obtain the POI vector dataset of the target area in Nanjing from the open-source map website, covering various urban functional elements such as residential areas, commercial facilities, and transportation stations. Load the Shapefile data of the POIs into a GeoDataFrame object (variable data_gdf) using the read_file function of the geopandas library, and perform preprocessing operations such as projection normalization (EPSG:32650) and outlier filtering. Subsequently, based on the refined road network MultiGraph model (variable clipped_momepy) constructed by the NetworkX library, use the bidirectional adjacent node allocation algorithm to allocate the POI information to the two nearest road network nodes based on the nearest adjacent street edge, obtaining an equally spaced node graph structure of the POI data calculated along the road. This allocation mechanism simulates the directional characteristics of POI accessibility from the perspective of pedestrians, facilitating subsequent precise quantitative analysis.
[0124] Convert clipped_momepy into a GeoDataFrame format nodes_gdf of network nodes, a GeoDataFrame format edges_gdf of connected edges, and a network structure network_structure for subsequent calculations. During the calculation of road centrality, a "moving window" - style spatial analysis method is adopted, that is, each node in the network is visited in turn, and the network is isolated within the specified walking distance threshold range of the currently selected node, and then the road centrality can be calculated for the locally windowed environment. Its advantage is that it can clearly and consistently define the network boundary related to the current analysis point, thereby optimizing the calculation speed and avoiding a series of problems such as the precise definition of the urban boundary. To accurately divide the "moving window" and optimize the operation speed and ensure data integrity as much as possible, the distance thresholds are selected as 50m, 100m, 200m, and 400m. Calculate the road centrality indicators, namely: closeness centrality, which is used to measure the closeness between points based on the distance of the shortest path; and betweenness centrality, which is used to measure the extent to which a point controls the connections between other points. The specific formulas refer to Embodiment 1
[0125] Use data_gdf containing adjacent street information and network structure network_structure containing adjacent POI information to calculate the business format structure complexity and the business format spatial distribution characteristics. The above - mentioned "moving window" - style spatial analysis method is also used during the calculation of the business format structure complexity and the business format spatial distribution characteristics. To optimize the operation speed and ensure data integrity, the distance thresholds are selected as 50m, 100m, 200m, and 400m. Considering that as the distance increases, the willingness of pedestrians to walk to POIs gradually decreases, a negative exponential decay function is adopted to simulate the spatial impedance, and then when calculating the business format structure complexity, a weighted Hill diversity index is used
[0126] When calculating the business format spatial distribution characteristics, use the reachable quantity index , representing the number of reachable instances of a given POI type (within the specified distance threshold) for the corresponding road network nodes
[0127] Read the information of network nodes nodes_gdf, connection edges edges_gdf, and network structure network_structure, and jointly construct the reference set graph structure. Randomly divide the reference set graph structure into two parts: a training set and a validation set. Input the longitude and latitude coordinates, closeness centrality, betweenness centrality, negative exponential weighted Hill coefficient, and reachable distance number of the road network nodes in the former structure into the extreme gradient boosting algorithm. Use the manually marked urban street space boundaries in the training set area as the supervised labels to train an automatic recognition model. Use the validation set to verify the prediction accuracy, precision, recall rate, and F1 score of the model to determine the quality of the model. Finally, adjust the parameters of the extreme gradient boosting algorithm to achieve better results, such as: the maximum depth of the tree is 20, the learning rate is 0.01, and the number of training rounds is 2000. And train an automatic recognition model with higher scores using these parameters.
[0128] After experiments, comparing the urban street space analysis algorithm of the present invention with other existing machine learning classifier algorithms, such as the random forest algorithm, etc., the urban street space recognition model trained by the present invention using the extreme gradient boosting algorithm has higher accuracy, precision, recall rate, and F1 score, as shown in the following table:
[0129]
[0130] Quantitatively preprocess the road network and POI vector data of the target area according to the above method to obtain the target area graph structure. Input the longitude and latitude coordinates, closeness centrality, betweenness centrality, negative exponential weighted Hill coefficient, and reachable distance number of the road network nodes in the target area graph structure into the automatic recognition model to obtain the urban street point set goal_nodes_gdf of Nanjing Gulou District.
[0131] Use the urban street point set goal_nodes_gdf to navigate back to the road network in the target area to intercept the road network, generate a buffer zone with reference to the road grades in the target area vector data, and remove the building outlines on both sides of the road to generate a surface set, which is saved as a Shapefile file, that is, the position and range of the urban street space in the target area of Nanjing Gulou District obtained by parsing, as Figure 3 shown.
[0132] Embodiment 3
[0133] Based on the same inventive concept as Embodiment 1, this embodiment introduces a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the urban street space parsing method based on the extreme gradient boosting algorithm as described in Embodiment 1.
[0134] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0135] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0136] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0137] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0138] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the present invention and the claims. All of these are within the protection scope of the present invention.
Claims
1. An urban street space analysis method based on the extreme gradient boosting algorithm, characterized in that include: Obtain the road network vector data and POI vector data of the target parsing area; Based on the road network vector data, converting the road network into a graph structure containing equidistant road network nodes according to the road level; Based on the POI vector data and the graph structure containing equidistant road network nodes, the POI is assigned to the nearest road node to obtain a graph structure containing POI information and equidistant road network nodes; Based on the graph structure containing POI information and equidistant road network nodes, obtain the centrality index of each node for measuring the road centrality, the business structure complexity index for measuring the complexity of the business structure, and the number of target POI paths reachable for measuring the spatial distribution characteristics of the business; The longitude and latitude coordinates, proximity centrality and betweenness centrality, weighted Hill diversity index, and the number of path reachable along the road distance of each node are input into the urban street space recognition model pre-trained by the extreme gradient boosting algorithm to obtain the urban street point set of the target parsing area; Based on the urban street point set, according to the road grade and building outline data in the road network vector data of the target parsing area, the urban street space with boundaries in the target parsing area is obtained.
2. The method according to claim 1, characterized in that, Based on the road network vector data, the road network is converted into a graph structure containing equidistant road network nodes according to the road level, including: Generate road network Shapefile based on road network vector data; The road network vector data is converted into the df_streets variable of the GeoDataFrame object through the read_file function in the Python geopandas library; Convert the road network GeoDataFrame to the MultiGraph data structure of the NetworkX library and define it as the clipped_momepy variable; Among them, in the MultiGraph data structure of the road network, road intersections and road segmentation points are used as network nodes, and road entities are used as connecting edges to generate a graph structure containing equidistant road network nodes; The road segmentation points are obtained in the following manner: for a road between any two adjacent road intersections, if its length exceeds a set segment length threshold, the road is equidistantly segmented to obtain corresponding road segmentation points.
3. The method according to claim 2, wherein The sources of the road network vector data include road network data from the open source map website OpenStreetMap and commercial road network information from the Baidu Map Developer Platform; The step of converting the road network into a graph structure containing equidistant road network nodes also includes: after converting the road network vector data into the df_streets variable of the GeoDataFrame object, calling the to_crs method to convert the coordinate systems of data from different sources to UTM Zone 50N.
4. The method according to claim 2, characterized in that, The step of allocating the POI to the nearest road node comprises: The Shapefile data of the POI vector data is converted into the data_gdf variable of the GeoDataFrame object through the read_file function of the Python geopandas library, and preprocessing operations including projection normalization and outlier filtering are performed; Based on the road network MultiGraph data structure, for each POI, according to the nearest road edge, it is assigned to the two nearest road network nodes of the nearest road, obtaining a graph structure containing POI information and equidistant road network nodes, and a POI data_gdf variable containing adjacent street information; In the graph structure containing POI information and equidistant road network nodes, each node includes road network node information and brief POI information assigned to the road network node; In the POI data_gdf variable containing adjacent street information, brief information of the road network nodes assigned to the POI is also recorded at each POI.
5. The method according to claim 4, characterized in that, Based on the graph structure containing POI information and equidistant road network nodes, each node centrality index for measuring road centrality is obtained, including: The data of the clipped_momepy variable in the MultiGraph data structure is converted into the data structures required for network analysis: the nodes_gdf variable in the GeoDataFrame format of node data, the edges_gdf variable in the GeoDataFrame format of connection edge data, and the network_structure variable for describing the topological relationship, where the network_structure variable is used to record the topological relationship information between road network nodes and POIs; Perform road centrality analysis: Traverse each node in the graph structure, respectively take each node as the center, delimit the local window network boundary within the preset walking distance threshold range, and calculate the closeness centrality and betweenness centrality of the central node of the local window network as the centrality index; The formula for closeness centrality is: In the formula, represents the closeness centrality of the node , n is the total number of nodes in the local window network, is the distance between node and node in the local window network; The formula for betweenness centrality is: In the formula, is the betweenness centrality of node . is the number of the shortest paths between point and point . is the number of the shortest paths passing through point from point to point .
6. The method according to claim 5, characterized in that, The delimiting of the local window network boundary within the preset walking distance threshold range includes: for each node, respectively take it as the center and delimit the local window network boundary within the preset multiple walking distance threshold ranges; The obtaining of each node centrality index for measuring road centrality also includes: for each node, respectively calculate its closeness centrality and betweenness centrality in the local window networks corresponding to multiple different walking distance threshold ranges, obtaining multiple groups of centrality index data.
7. The method according to claim 5, characterized in that, Based on the graph structure containing POI information and equidistant road network nodes, the business format structure complexity index for measuring the complexity of the business format structure is obtained, including: Based on the POI data_gdf variable containing adjacent street information and the network_structure variable containing street adjacent POI information, for each road network node, calculate the weighted Hill diversity index, and the formula is: Where: Wherein: represents the weighted Hill diversity index of the road network nodes ; and respectively represent the th th POI types assigned to the road network node ; represents the total number of all POI types assigned to the road network node ; represents the number of POIs of the th POI type assigned to the road network node ; represents the attenuation weight of pedestrians walking to the th POI at the node ; represents the average value of the attenuation weights corresponding to all POIs under the th POI type; is a preset walking distance threshold.
8. The method according to any one of claims 1 to 7, characterized in that, The target POI path reachable quantity for measuring the spatial distribution characteristics of business formats is the number of path instances that can reach a specified POI type along roads within a set distance threshold for each road network node.
9. The method according to any one of claims 1 to 7, characterized in that Based on the set of urban street points, the urban street space with boundaries in the target analysis area is obtained according to the road grades and building contour data in the road network vector data of the target analysis area, including: According to the node ID field of the point elements in the set of urban street points, a topological connection is established with the start and end point IDs of the road segments in the road network to reconstruct the street network of the target analysis area; According to the preset buffer radius mapped by the road grade field in the road network vector data of the target analysis area, a spatial difference operation is performed between the buffer zone and the building polygon data to remove the area occupied by roadside buildings; the street space polygon set is output and stored in the Shapefile format to obtain the final street space position and range.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the urban street space analysis method based on the extreme gradient boosting algorithm as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Rural network node radiation domain-oriented rural residential area renovation zoning method
CN103971312A
Urban business district boundary identification method and system, computer equipment and storage medium
CN112766718A
Urban building space data-based built-up area boundary identification method and device
WO2020233152A1