An urban green land cooling benefit evaluation and optimization method, system, device and medium based on an interpretable spatial machine learning
By employing an interpretable spatial machine learning approach and utilizing a geographically weighted random forest model to construct an interpretable model of the cooling effect of urban green space characteristic indicators, this approach addresses the issues of simplistic methods and insufficient spatial adaptability in the evaluation and optimization of urban green space cooling effects. It enables the precise formulation and efficient implementation of green space optimization strategies, thereby enhancing the city's climate adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN UNIV
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies for evaluating and optimizing the cooling benefits of urban green spaces suffer from limited methods, vague control strategies, and insufficient spatial adaptability, making it difficult to accurately characterize the complex nonlinear relationships and spatial heterogeneity of the cooling effect of green spaces.
This study employs an interpretable spatial machine learning approach, using a geographically weighted random forest model to construct an interpretable model of the cooling effect of urban green space characteristic indicators. By combining the geographically weighted random forest model with a local random forest sub-model, and considering the nonlinear relationship between urban green space characteristic indicators and surface temperature, the study identifies feature importance and reliable thresholds, thereby enabling the precise formulation of green space optimization strategies.
It enables precise assessment and optimization of the cooling benefits of urban green spaces, and can be implemented efficiently in multiple dimensions such as spatial location, effect characteristics and intervention intensity, thereby enhancing the city's ability to adapt to climate change.
Smart Images

Figure CN122114712A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of urban ecological planning, and specifically relates to a method, system, equipment and medium for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning. Background Technology
[0002] Existing technologies are increasingly sophisticated in exploring the cooling mechanisms of urban green spaces. For example, numerous studies have revealed the threshold effect of green space cooling based on the nonlinear relationship between urban green space size and temperature. However, these relationships mostly focus on the cooling benefits of basic indicators such as green space ratio or patch size. Research on the complex cooling mechanisms of other characteristic indicators of urban green spaces (such as shape complexity, fragmentation, and diversity) remains unclear, and a unified and comparable research framework is still lacking regarding the impact mechanisms of key indicators of urban green spaces on temperature. On the other hand, influenced by factors such as regional climate, topography, and urban morphology, the cooling mechanisms of green spaces exhibit strong spatial heterogeneity across different regions. This makes it difficult to simply transfer local experiences from one region to other regions, limiting the ability to extract universal laws.
[0003] In summary, the current assessment and optimization of the cooling effect of urban green spaces still faces key bottlenecks such as a single methodology, vague control strategies, insufficient spatial adaptability, and limited practical application, making it difficult to accurately characterize the complex nonlinear relationships and significant spatial heterogeneity of the cooling effect of green spaces. Summary of the Invention
[0004] To overcome the key bottlenecks in the current assessment and optimization of urban green space cooling benefits, such as the simplistic approach, vague control strategies, insufficient spatial adaptability, and limited practical application, which make it difficult to accurately characterize the complex nonlinear relationships and significant spatial heterogeneity of green space cooling effects, this invention provides a method, system, equipment, and medium for assessing and optimizing urban green space cooling benefits based on interpretable spatial machine learning. By considering the numerous types of urban green space characteristic indicators, this invention comprehensively considers all indicators affecting the cooling benefits of urban green spaces. Furthermore, it proposes a geographically weighted random forest model that can accurately analyze the complex nonlinear relationships and potential spatial heterogeneity of urban green space cooling. Based on the geographically weighted random forest model, it constructs interpretable models of the cooling benefits of various urban green space characteristic indicators, realizing a technical method for the precise formulation and efficient implementation of urban green space optimization strategies.
[0005] According to one aspect of the present invention, a method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning is provided, comprising:
[0006] Based on urban boundary data, land use data and surface temperature data of the target area, an urban green space dataset and a corresponding urban surface temperature dataset are constructed.
[0007] For each spatial unit in the urban green space dataset, calculate several corresponding urban green space characteristic indicators.
[0008] A geographically weighted random forest model is constructed to fit the nonlinear relationship between the latitude and longitude of all spatial units and several urban green space characteristic indicators and the surface temperature data of the corresponding spatial units in the urban surface temperature dataset. The geographically weighted random forest model integrates several local random forest sub-models. The local random forest sub-models take the latitude and longitude of each spatial unit and several urban green space characteristic indicators as inputs and output the surface temperature data of that spatial unit.
[0009] The target area is divided into zones, and based on the geographically weighted random forest model, an interpretable model of the cooling effect of each urban green space characteristic index within the target area and within each zone of the target area is constructed, including: the characteristic importance of each urban green space characteristic index, the cumulative local effect map, and the reliable threshold.
[0010] Based on the interpretable modeling of the cooling effect of the green space characteristic indicators of each city, the recommended value range and the value range to be avoided for each green space characteristic indicator in the target area and in each sub-area of the target area are determined, so as to obtain the optimization decision of the cooling effect of green space in the target area.
[0011] As a further technical solution, the urban green space dataset internally stores an urban green space mask, and the construction process includes:
[0012] Within the urban boundaries defined by urban boundary data, the urban green space definition area masks in land use data are merged to obtain an urban green space mask, and an urban green space dataset is constructed accordingly.
[0013] As a further technical solution, urban green space characteristic indicators include, but are not limited to, green space ratio, patch size, maximum patch index, patch density, area-weighted average shape index, average nearest neighbor distance, Shannon diversity index, Moran's I, and overall connectivity index.
[0014] As a further technical solution, the construction process of the geographically weighted random forest model includes:
[0015] Select a single spatial unit as the target spatial unit and construct a random forest sub-model for it. Repeat this process to traverse the spatial units and obtain several random forest sub-models. Combine the latitude and longitude of each spatial unit, the characteristic indicators of urban green space, and the corresponding urban surface temperature data in the urban surface temperature dataset to obtain training samples. In this way, the training samples corresponding to each spatial unit are obtained.
[0016] The spatial distance between the target spatial unit and other spatial units of the random forest sub-model is calculated based on latitude and longitude. The spatial weight is calculated using the Gaussian kernel function. The training samples corresponding to each spatial unit are weighted using the spatial weight to obtain the geographically weighted sample training set of the random forest sub-model. The random forest sub-model is trained using the geographically weighted sample training set. After training, a local random forest sub-model is obtained. Several local random forest sub-models are obtained in this way.
[0017] As a further technical solution, the process of obtaining the characteristic importance of various urban green space characteristic indicators includes:
[0018] The importance scores of each urban green space feature index in the local random forest sub-model of the corresponding spatial unit in the geographic weighted random forest model are obtained. The importance score vector is constructed as the local variable importance of the spatial unit, and the local variable importance of each spatial unit is obtained accordingly.
[0019] The importance of local variables of all spatial units within the region is obtained, the proportion of the importance of local variables of each urban green space characteristic index to all urban green space characteristic indices is calculated, and the characteristic importance of each urban green space characteristic index within the region is statistically obtained.
[0020] As a further technical solution, the process of obtaining the cumulative local effect map and the reliable threshold includes:
[0021] The latitude and longitude of spatial units within the region and several urban green space characteristic indicators are input into a geographically weighted random forest model. The predicted surface temperature data is output, and a cumulative local effect map of each urban green space characteristic indicator within the region is plotted. The urban green space characteristic indicator values represented by the slope abrupt change points in the cumulative local effect map are selected as candidate thresholds for urban green space characteristic indicators within the region. Candidate thresholds that meet the triple selection principle are selected as reliable thresholds for urban green space characteristic indicators within the region.
[0022] As a further technical solution, the reliable threshold of the urban green space characteristic indicators in the region simultaneously meets the triple screening principle, including: the density of local spatial units at the reliable threshold is not less than 50% of the average point density of the entire variable interval; the preset neighborhood interval of the reliable threshold contains no less than a preset number of spatial units; and the length of the 95% confidence interval of the fitted cumulative local effect curve at the reliable threshold is not greater than 20% of the entire prediction range.
[0023] According to another aspect of this specification, 8. A system for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning is provided, characterized in that it includes:
[0024] The data preparation module is used to construct an urban green space dataset and a corresponding urban surface temperature dataset based on urban boundary data, land use data and surface temperature data of the target area.
[0025] The urban green space characteristic index calculation module is used to calculate several urban green space characteristic indices for each spatial unit in the urban green space dataset.
[0026] The geographic weighted random forest model construction module is used to construct a geographic weighted random forest model to fit the nonlinear relationship between the latitude and longitude of all spatial units and several urban green space characteristic indicators and the surface temperature data of the corresponding spatial units in the urban surface temperature dataset. The geographic weighted random forest model integrates several local random forest sub-models. The local random forest sub-models take the latitude and longitude of each spatial unit and several urban green space characteristic indicators as inputs and output the surface temperature data of that spatial unit.
[0027] The cooling effect modeling module is used to divide the target area into zones and, based on the geographically weighted random forest model, construct an interpretable model of the cooling effect of each urban green space characteristic index within the target area and within each zone of the target area. This includes: the characteristic importance of each urban green space characteristic index, the cumulative local effect map, and the reliable threshold.
[0028] The green space cooling benefit optimization decision module is used to interpretably model the cooling benefits of each city's green space characteristic indicators, determine the recommended value range and the value range to be avoided for each green space characteristic indicator in the target area and in each sub-area of the target area, and obtain the green space cooling benefit optimization decision for the target area.
[0029] According to another aspect of this specification, an electronic device is provided, including a memory and a processor, the memory storing program instructions executed by the processor, the processor invoking the program instructions to execute a method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning.
[0030] According to another aspect of this specification, a non-transitory computer-readable storage medium is provided, the non-transitory computer-readable storage medium storing computer instructions that cause the computer to execute a method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning.
[0031] Compared with the prior art, the beneficial effects of this invention are as follows: It constructs an interpretable spatial machine learning analysis framework: by considering the complex types of urban green space characteristic indicators, it comprehensively considers various indicators that affect the cooling effect of urban green spaces, and further proposes a geographically weighted random forest model that can accurately analyze the complex nonlinear relationship of urban green space cooling and its potential spatial heterogeneity. Based on the geographically weighted random forest model, it constructs an interpretable model of the cooling effect of various urban green space characteristic indicators, and scientifically and accurately explores the nonlinear relationship of urban green space cooling mechanism and its spatial heterogeneity.
[0032] Based on the reliable threshold identification and feature importance calculation based on the triple principle, the spatial heterogeneity of the whole area and the zone in multiple dimensions such as spatial location, function characteristics and intervention intensity can be compared. This can effectively realize the accurate implementation of green space optimization strategies in multiple dimensions such as spatial location, function characteristics and intervention intensity, and provide methodological and technical support for improving the city's climate adaptation capacity based on nature solutions.
[0033] This invention comprehensively explores the nonlinear cooling mechanism and potential spatial heterogeneity of urban green spaces in multiple dimensions, and then realizes the precise formulation and efficient implementation of urban green space optimization strategies in multiple dimensions such as spatial location, action characteristics and intervention intensity, providing methodological and technical support for improving the city's climate adaptation capacity based on nature solutions. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 A flowchart illustrating an embodiment of the present invention for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning;
[0036] Figure 2 This is a schematic diagram of a reliable threshold identification method based on the triple principle in an embodiment of the present invention;
[0037] Figure 3 This is an ALE diagram of various urban green space characteristic indicators affecting the urban surface thermal environment in China, as shown in this embodiment of the invention.
[0038] Figure 4 ALE diagrams of five major urban green space characteristic indicators that affect the urban surface thermal environment in different climate zones, as shown in this embodiment of the invention.
[0039] Figure 5This is a schematic diagram of the structure of an urban green space cooling benefit evaluation and optimization system based on interpretable spatial machine learning, provided in an embodiment of the present invention.
[0040] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0041] like Figure 1 As shown, a method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning includes the following steps:
[0042] Step 1: Based on the urban boundary data, land use data and surface temperature data of the target area, construct an urban green space dataset and a corresponding urban surface temperature dataset.
[0043] Step 2: For each spatial unit in the urban green space dataset shown, calculate its corresponding urban green space characteristic indicators.
[0044] Step 3: Construct a geographic weighted random forest model to fit the nonlinear relationship between the latitude and longitude of all spatial units and several urban green space characteristic indicators and the surface temperature data of the corresponding spatial units in the urban surface temperature dataset; the geographic weighted random forest model integrates several local random forest sub-models, and the local random forest sub-models take the latitude and longitude of each spatial unit and several urban green space characteristic indicators as inputs and output the surface temperature data of that spatial unit.
[0045] Step 4: Divide the target area into zones. Based on the geographically weighted random forest model, construct an interpretable model of the cooling effect of each urban green space characteristic index within the target area and within each zone of the target area. This includes: the characteristic importance of each urban green space characteristic index, the cumulative local effect map, and the reliable threshold.
[0046] Step 5: Based on the interpretable modeling of the cooling effect of the green space characteristic indicators of each city, determine the recommended value range and the value range to be avoided for each green space characteristic indicator in the target area and in each sub-area of the target area, so as to obtain the optimization decision of the cooling effect of green space in the target area.
[0047] In step 1, the urban green space dataset stores urban green space masks. The construction process includes:
[0048] Within the urban boundaries defined by urban boundary data, different urban green space definition areas in land use (cover) data are merged into urban green space, and an urban green space dataset is constructed accordingly.
[0049] Commonly defined green areas include, but are not limited to, grasslands, woodlands, shrublands, and arable land.
[0050] Preferably, in the process of constructing the urban green space dataset, it is considered that the urban cooling effect not only comes from the green space in the built-up area, but also from the green space in the suburbs that is highly accessible to residents, which plays an important role in alleviating the urban thermal environment. Therefore, when delineating the research scope or urban boundary, it should be appropriately expanded outward to include the suburban green space outside the urban core area and accessible to residents' daily activities in the analysis scope, so as to more comprehensively reflect the overall regulatory capacity of the green space system on the urban thermal environment.
[0051] Preferably, when conducting large-sample studies at the urban scale, the urban surface temperature dataset uses surface heat island intensity, rather than surface temperature, as an indicator to characterize the urban thermal environment, in order to enhance the comparability between different cities.
[0052] In step 1, the urban surface temperature dataset stores urban area surface temperature data. The construction process is as follows: obtain the surface temperature data of each urban green space mask in the corresponding urban green space dataset from the surface temperature data to obtain the urban surface temperature dataset.
[0053] In step 2, the spatial unit is the benchmark for regional spatial division. If it is at a micro scale, i.e., focusing on a single city, with a fixed grid size as the calculation unit, then each unit must participate in the analysis. If it is at a macro scale, i.e. focusing on a region / country, with a single city as the research object, the minimum threshold for each city studied is different, such as 10, 20, 30, 50, or 100 km. 2 They have them all.
[0054] Urban green space characteristic indicators include, but are not limited to, green space ratio, patch size, maximum patch index, patch density, area-weighted average shape index, average nearest neighbor distance, Shannon diversity index, Moran's I, and overall connectivity index.
[0055] Preferably, Pearson correlation analysis is performed on multi-dimensional urban green space characteristic indicators to screen out variables with significant correlations below P < 0.01. To avoid collinearity among the screened variables, SPSS software is used to introduce the variance inflation factor (VIF) to diagnose collinearity of the variables, and urban green space characteristic indicators with VIF < 7.5 are retained for subsequent modeling.
[0056] Optionally, in step 2, the urban green space characteristic indicators are calculated using software such as ArcGIS, Fragstats, Matlab, and Conefor:
[0057] Green space ratio, patch size, maximum patch index, patch density, area-weighted average shape index, average nearest neighbor distance, and Shannon diversity index are commonly used indicators in landscape ecology, which can be calculated from raster data in urban green space datasets using Fragstats software.
[0058] Moran's I is used to measure the overall structuring level of urban green spaces, and its calculation formula is as follows:
[0059]
[0060] In the formula, n is the total number of grids in the city (the grid size depends on the research scale, and is generally the spatial resolution of temperature data). Let be the percentage of green space in the i-th grid. This represents the average percentage of green space across all grid grids in the city. Indicates spatial weight (when grid i is adjacent to grid j, =1, otherwise =0). The larger the Moran's I value, the higher the level of green space structure and the more concentrated the distribution. The specific calculation steps of this indicator are: (1) Generate a grid in the study area according to the "Create Fishing Net" tool in ArcGIS; (2) Calculate the proportion of green space in each grid; (3) Obtain the Moran's I result through the "Spatial Autocorrelation (Moran's I)" tool in ArcGIS.
[0061] The Integral Index of Connectivity (IIC) is used to measure the overall structuring level of urban green spaces. Its calculation formula is as follows:
[0062]
[0063] In the formula, n represents the number of green patches in the city. and Let i and j be the areas of patches i and j, respectively. IIC represents the number of links in the shortest path (topological distance) between patches i and j, and A represents the total area of the entire urban landscape. The larger the IIC value, the higher the network connectivity of the green space. The specific calculation steps for this index are as follows: (1) First, calculate the shortest topological distance between different green space patches using the Conefor plugin in ArcGIS; (2) Then, input the obtained patch areas and shortest distances into the Conefor software, set the maximum connectivity distance, and run to obtain the final result.
[0064] In step 3, the geographically weighted random forest (GWRF) model is constructed by using the latitude and longitude of the spatial unit and the urban green space characteristic indicators obtained in step 2 as independent variables, and the surface temperature of the spatial unit calculated in step 1 as the dependent variable. The model's built-in output includes the sub-model parameters of each spatial unit, the importance of local variables, and the model's predicted values.
[0065] The GWRF model combines the local sensitivity of Geographically Weighted Regression (GWR) with the powerful nonlinear prediction capabilities of Random Forest (RF), enabling it to effectively capture spatial heterogeneity and nonlinear relationships in geospatial data, thereby improving the prediction accuracy of complex geographical processes such as ecology.
[0066] The geographically weighted random forest model integrates several local random forest sub-models. These sub-models take the latitude and longitude of spatial units and urban green space characteristic indicators as input, and output the surface temperature data of the spatial units. The training process includes:
[0067] Data preparation: Select a single spatial unit as the target spatial unit and construct a random forest sub-model for it. Repeat this process to traverse the spatial units and obtain several random forest sub-models. Combine the latitude and longitude of the spatial unit, the characteristic indicators of urban green space, and the corresponding urban surface temperature data in the urban surface temperature dataset to obtain training samples. In this way, the training samples corresponding to each spatial unit are obtained.
[0068] Training process: Calculate the spatial distance between the target spatial unit and other spatial units of the random forest sub-model based on latitude and longitude, and calculate the spatial weights using the Gaussian kernel function. Use the spatial weights to weight the training samples corresponding to each spatial unit to obtain the geographically weighted sample training set of the random forest sub-model. Use the geographically weighted sample training set to train the random forest sub-model. After training, a local random forest sub-model is obtained, and several local random forest sub-models are obtained in this way.
[0069] In the above process, multiple local Random Forest (RF) sub-models are constructed and run. These sub-models make no assumptions about the data distribution, and each sub-model is fitted to a set of neighboring spatial units. These spatial units are assigned different spatial weights using a Gaussian kernel function based on their distance from the target location. This ensures that samples that are farther away have less influence on local predictions, thus effectively representing spatial heterogeneity. Furthermore, by utilizing the ensemble learning framework of Random Forest (RF), GWRF demonstrates stronger robustness in handling multicollinearity and overfitting.
[0070] The built-in outputs of the GWRF model include, but are not limited to, sub-model parameters for each spatial unit, local variable importance, and model predictions.
[0071] Step 4 involves result interpretation and visualization: Based on the importance of local variables output by the GWRF model, the characteristic importance of urban green space indicators is quantified. The relationship between urban green space characteristic indicators and model predicted values (urban surface temperature data) is plotted using the Accumulated Local Effects (ALE) plot. This is used to visualize the impact of urban green space characteristic indicators on cooling benefits across the entire target area and in each sub-area. The slope abrupt change points in the ALE plot are also identified for subsequent threshold identification based on the triple principle, thus enabling interpretable modeling of cooling benefits.
[0072] Optionally, in step 4, the target area is spatially partitioned according to the heterogeneity to be explored (macro-background climate, economic zones, geographical zones, etc., as well as local climate zones (LCZ), urban functional zones (UFZ), etc.), and the feature importance and predicted values obtained from the GWRF model are classified and statistically analyzed to obtain the feature importance and ALE diagram of each partition, which are used for subsequent analysis of the commonalities and differences between different partitions and between the partitions and the whole area.
[0073] It is worth noting that, to ensure the reliability of the analysis results, preferably, each partition has at least 20 spatial units. For partitions that do not meet this requirement, they can be appropriately merged with regions of similar characteristics to enhance the robustness of the analysis results.
[0074] Step 4, the process of obtaining the characteristic importance of each city's green space characteristic indicators includes:
[0075] The importance scores of each urban green space feature index in the local random forest sub-model of the corresponding spatial unit in the geographic weighted random forest model are obtained. The importance score vector is constructed as the local variable importance of the spatial unit, and the local variable importance of each spatial unit is obtained accordingly.
[0076] The importance of local variables of all spatial units within the region is obtained, the proportion of the importance of local variables of each urban green space characteristic index to all urban green space characteristic indices is calculated, and the characteristic importance of each urban green space characteristic index within the region is statistically obtained.
[0077] Specifically, the importance of features within a region is determined based on the local variable importance provided by the Geographically Weighted Random Forest (GWEF) model. This is specifically expressed as the proportion of the local variable importance of each urban green space feature indicator in each spatial unit to the total importance of all indicators in that spatial unit. Optionally, the importance score is obtained as follows: within each local random forest sub-model, the variable contribution of each tree is evaluated using a standard RF method (such as permutation importance), and the importance score (LVI) of each input is calculated as the local variable importance of the target spatial unit corresponding to the sub-model. Preferably, the local variable importance is represented in LVI vector form.
[0078] Step 4, the process of obtaining the cumulative local effect map and the reliability threshold includes:
[0079] The latitude and longitude of spatial units within the region and several urban green space characteristic indicators are input into a geographically weighted random forest model. The predicted surface temperature data is output, and a cumulative local effect map of each urban green space characteristic indicator within the region is plotted. The urban green space characteristic indicator values represented by the slope abrupt change points in the cumulative local effect map are selected as candidate thresholds for urban green space characteristic indicators within the region. Candidate thresholds that meet the triple selection principle are selected as reliable thresholds for urban green space characteristic indicators within the region.
[0080] Among them, the reliable threshold of the urban green space characteristic indicators in the region simultaneously meets the triple screening principle, including: the density of local spatial units at the reliable threshold is not less than 50% of the average point density of the entire variable interval; the preset neighborhood interval of the reliable threshold contains no less than a preset number of spatial units; and the length of the 95% confidence interval of the fitted ALE curve at the reliable threshold is not greater than 20% of the entire prediction range.
[0081] Specifically, the Accumulated Local Effects (ALE) plot is used to visualize the nonlinear relationship between urban green space characteristic indicators and output urban temperature data. The visualization of the nonlinear relationship is based on the predicted values output by the model. The ALE plot averages the changes in green space characteristics and then accumulates them on a grid. The theoretical formula is shown below:
[0082]
[0083] In the formula, This represents a black-box machine learning model, specifically a local random forest sub-model. The input includes all features, such as latitude and longitude and urban green space characteristics, and the output is the predicted value of the surface temperature data. It is used to plot the features of interest in ALE (one of the urban green space feature indicators). Other characteristics (latitude and longitude, and other urban green space characteristic indicators); The starting point for integration is usually taken as the minimum or mean value of the features of interest in the feature set S; Indicates that in a given Under the condition, other characteristics Conditional distribution; Representation Model Features of interest The partial derivatives; Indicates from the starting point to The total effect is obtained by cumulatively integrating the local effects; constant is the mean of the entire ALE function under the data distribution.
[0084] It can be seen that the ALE method does not directly average the prediction results, but integrates the derivative of the target feature to obtain the estimated value of the result. The operation of differentiation and integration isolates the interference caused by feature correlation.
[0085] Based on this, the present invention develops a reliable threshold identification method based on a triple screening principle, optionally, as follows: Figure 2 As shown, the triple screening principle in this invention is as follows: (1) the local spatial cell density at the threshold is at least 50% of the average point density of the entire variable interval; (2) the threshold neighborhood interval is at least 5 spatial cells; (3) the length of the 95% confidence interval of the fitted curve at the threshold is at most 20% of the entire prediction range. Only when the threshold satisfies the above triple principle is the threshold considered reliable and robust, and will it be included in the subsequent analysis.
[0086] Step 5 essentially involves optimizing the urban green space spatial pattern according to local conditions. It involves determining the recommended and avoided value ranges for each green space characteristic indicator within the target area and its sub-areas; and by comparing and analyzing the commonalities and differences in the multi-dimensional cooling mechanisms of urban green spaces across the entire region and different sub-areas, it provides scientific decision support for optimizing the cooling efficiency of green spaces, tailored to local conditions, for subsequent targeted and cost-effective urban green space planning aimed at improving cooling benefits.
[0087] Furthermore, step 5 involves analyzing and summarizing the impact relationships and reliable thresholds of urban green space characteristic indicators on surface temperature at the overall and regional scales. It also statistically analyzes the importance of different urban green space indicator characteristics across the entire region and within each region, outputting the dominant influencing factors on the urban thermal environment within the region (top-K urban green space characteristic indicators ranked by feature importance). A thorough analysis of the commonalities and regional differences in the interpretable modeling of the cooling benefits of various urban green space characteristic indicators leads to the proposal of site-specific urban green space planning schemes aimed at improving cooling benefits. This effectively enables the precise implementation of urban green space optimization strategies across multiple dimensions, including spatial location, functional characteristics, and intervention intensity, providing methodological and technical support for enhancing urban climate adaptation capabilities based on nature-based solutions.
[0088] Under the same technical concept as described above, the present invention provides an optional embodiment, including the following steps:
[0089] S1. Data Acquisition: Acquire urban boundary data, land use / cover data, and surface temperature data of the target area, and process them to obtain urban green space dataset and surface temperature dataset.
[0090] In practice, in order to ensure that green spaces with high social and ecological benefits that are highly accessible to residents in the suburbs are included in the analysis, the Global Urban Boundary (GUB) data produced by Professor Gong Peng's team at the University of Hong Kong was selected as the initial data source. On this basis, a strict data processing procedure was adopted to control the interference of leapfrog development on boundary delineation, and finally the global / national urban boundary dataset was obtained.
[0091] When selecting land use / cover data, various factors such as the size of the study area and the level of detail in green space characterization should be considered to select data with an appropriate spatial resolution to define urban green space. Based on the dataset production process and principles, this invention adopts the most widely used definition of urban green space in current research, merging grassland, woodland, shrubland, and cultivated land into one category, and obtaining an urban green space dataset based on ArcGIS clipping tools.
[0092] The selection of surface temperature data also needs to be determined based on the study area. For exploring the cooling mechanism of urban green spaces, 30 m resolution Landsat data is sufficient. It is worth noting that the weather on the day the Landsat image was acquired and the two days prior must be sunny, cloudless, and with light winds (wind force ≤ 3). For large-scale studies such as urban agglomerations or countries, MODIS surface temperature data (spatial resolution 1 km), which is widely used for large-scale analysis, is selected to assess the urban thermal environment level.
[0093] S2. Indicator Calculation: Based on the acquired urban green space dataset, urban green space characteristic indicators are calculated using software such as ArcGIS, Fragstats, Matlab, and Conefor. These indicators include green space ratio, patch size, maximum patch index, patch density, area-weighted average shape index, average nearest neighbor distance, Shannon diversity index, Moran's I, and overall connectivity index.
[0094] S3. Model Construction: Using the latitude and longitude of all spatial units and the urban green space characteristic indicators obtained in S2 as independent variables, and the surface temperature / surface heat island intensity of the corresponding spatial units calculated in S1 as dependent variables, a geographically weighted random forest (GWRF) model is constructed.
[0095] In practice, GWRF can be implemented directly using the SpatialML function package in R Studio. Its built-in `grf.bw` function iteratively determines the optimal bandwidth for the model, and the model is then retrained using this optimal bandwidth to obtain the final trained model. The model's built-in outputs include sub-model parameters for each spatial unit, local variable importance, and model predictions. The coefficient of determination R-squared is used. 2 The root mean square error (RMSE) and mean absolute error (MAE) are used as indicators to evaluate model accuracy. The specific calculation formulas are as follows:
[0096]
[0097]
[0098]
[0099] In the formula, n is the sample size. For the observation value of the i-th sample, The average of all sample observations. Let be the predicted value for the i-th sample. R is the average of the predicted values for all samples. 2 The closer the value is to 1, the better the model fits. The lower the values of RMSE and MAE, the more accurate the regression model is.
[0100] S4. Results Interpretation and Visualization: Based on feature importance, the contribution of urban green space features to model decision-making is quantified. The response relationship between different urban green space features and the urban thermal environment is visualized through the cumulative local effects ALE diagram. Reliable thresholds are identified based on the triple principle threshold identification method, thus achieving interpretable modeling.
[0101] In practical implementation, regarding feature importance, GWRF can output the importance of local variables for different characteristic indicators of urban green space at each spatial unit. At the global scale, the sum of the importance of local variables for each spatial unit is set to 1. By calculating the proportion of each indicator's importance to the total importance of all indicators, the comparability of the results of the local sub-models for different spatial units is ensured. The mean of the importance proportions of each indicator across all spatial units is used as the feature importance of that indicator across the entire region, identifying the dominant urban green space characteristic indicators. Furthermore, it is necessary to spatially visualize the most significant indicator for each spatial unit to more intuitively understand the spatial heterogeneity of the urban green space cooling mechanism.
[0102] Regarding nonlinear relationships, ALE, as an unbiased method, can effectively handle collinearity of indicators and accurately visualize the impact curves of urban green space characteristic indicators on land surface temperature. Specifically, the ALE plot averages the changes in green space characteristics and then accumulates them on a grid. This algorithm calculates the change in the average predicted value, not the predicted value itself: instead of directly averaging the prediction results, it integrates the derivative of the target feature to obtain an estimate of the result. The differentiation and integration operations isolate interference caused by feature correlations. The short vertical lines at the bottom represent probability density; denser distributions indicate a greater number of spatial units. It is worth noting that only the ALE curve trend in relatively dense data distribution intervals is considered relatively reliable.
[0103] Furthermore, a reliable threshold identification method specifically includes: (1) the local point density at the threshold is at least 50% of the average point density of the entire variable interval; (2) the threshold neighborhood interval is at least 5 spatial units; and (3) the length of the 95% confidence interval of the fitted curve at the threshold is at most 20% of the entire prediction range. Only when the threshold simultaneously meets the above three principles is it considered reliable and robust, and will it be included in subsequent analysis. Figure 3 As shown, in the given "W"-shaped ALE curves, only thresholds ① that simultaneously satisfy all three principles are considered robust and reliable and will be included in subsequent analyses.
[0104] S5. Analysis of different spatial driving factors: Based on different zones (background climate, economic zone, geographical zone, etc.), the predicted values of surface temperature are classified and statistically analyzed to obtain the dominant influencing factors of each zone and their corresponding ALE maps, as well as their respective reliable thresholds.
[0105] In practice, based on the type of spatial heterogeneity to be explored (macro-level background climate, economic zones, geographical zones, etc., and local climate zones (LCZ), urban functional zones (UFZ), etc.), the feature importance and predicted values obtained from the GWRF model are classified and statistically analyzed to obtain the feature importance and ALE plots for each zone. It is worth noting that to ensure the reliability of the analysis results, the sample size for each zone should be at least 20. For regions that do not meet this requirement, they can be appropriately merged with regions of similar characteristics to enhance the robustness of the analysis results. Furthermore, considering that the analysis involves the intersection of multi-dimensional indicator systems and multiple spatial regions, making the content somewhat complex, it is recommended to analyze the most significant top few indicators for each region (at least 50% of the total number of indicators should be analyzed).
[0106] S6. Optimization of urban green space spatial pattern according to local conditions: Comparative analysis of the commonalities and differences in the cooling mechanism of urban green space in the whole region and different zones, so as to provide scientific decision support for subsequent targeted and cost-effective urban green space planning to improve cooling efficiency.
[0107] In practice, the system systematically sorts out and summarizes the influence relationship and threshold of various urban green space characteristic indicators on surface temperature under the whole area and different zones, deeply analyzes their common laws and regional differences, and combines the importance of the indicators to propose differentiated green space planning strategies that maximize cooling benefits and take into account the overall planning of the whole area and the adaptation of zones, so as to achieve the optimization of urban green space that is tailored to local conditions and cost-effective.
[0108] To better understand the basic scheme of this invention, the following uses 222 cities in a certain country as examples, combined with relevant experimental data, to further illustrate the method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning provided by this invention. Specifically, it includes:
[0109] S01. Data Acquisition.
[0110] To ensure the inclusion of highly accessible, socio-ecologically beneficial green spaces in the analysis, the Global Urban Boundary (GUB) data produced by Professor Gong Peng's team at the University of Hong Kong was used as the initial data source. In fact, the extraction of urban green spaces and the subsequent measurement of their diverse characteristics are highly sensitive to the method of urban boundary delineation. Therefore, this study employed a systematic preprocessing procedure to address the challenges posed by rapidly expanding suburbs and ensure the reliability of the analysis. In these suburbs, fragmented expansion often disrupts the consistency of the determined urban boundaries, thus interfering with the characterization of urban land use features. Given that urban land use features in smaller cities may be more susceptible to pixel variations in boundary delineation, we only included cities with an area exceeding 10 square kilometers in the study. This area threshold also helps to achieve robust urban boundary delineation, thereby accurately calculating the SUHII (Sustainable Urban Area Index). Furthermore, to maintain the integrity of the urban boundary morphology while minimizing suburban interference, only key suburban areas were included in the final urban scope. Specifically, a city area can only be merged into its city if its area exceeds 1 square kilometer and its distance from the largest city area within the city does not exceed twice the data resolution (60 meters in this study).
[0111] Green space was defined using land use / cover data. Data with a spatial resolution of 30m was selected to balance the measurement needs of overall green space characteristics and fine structure of small patches. The most widely used definition of urban green space in current research was adopted. The four categories of grassland, woodland, shrubland and cultivated land in the land use / cover data were merged into green space using ArcGIS reclassification tools. The merged green space data was then cropped using urban boundary data to obtain the urban green space dataset.
[0112] MODIS surface temperature data (spatial resolution of 1 km), widely used for large-scale analysis, was selected to assess the urban thermal environment level. The urban surface temperature dataset was obtained by cropping the surface temperature data using urban boundary data. Considering that the absolute value of urban surface temperature is influenced by both human activity and background climate, it cannot be directly used to characterize urban impacts (including the cooling effect of green spaces, which is the focus of this invention). In contrast, surface heat island intensity, as the difference between urban and rural average surface temperatures, can largely eliminate the influence of background climate and has been widely used as a measure of urban thermal environment in large-scale comparative studies. To further enhance the comparability of surface heat island intensity among different cities, this invention uses a strict method for delineating rural reference areas. Specifically, to avoid the influence of urban temperatures, buffer zones were created outwards from each city, extending outwards by its area. It was found that the temperature in the rural reference area tended to stabilize after the eighth buffer zone; therefore, a buffer zone of equal area, located farthest from the city, was chosen as the initial range for the rural reference area. To control interference introduced by differentiated rural reference areas, water bodies within the buffer zones, areas exceeding the city's average elevation by ±50m, and the heat island footprint of neighboring cities were further removed from the initial range. By calculating the difference between urban surface temperature and the surface temperature of the rural reference area, a dataset of urban surface heat island intensity within the study area was obtained.
[0113] S02. Indicator Calculation.
[0114] Based on the acquired urban green space dataset, and through multidisciplinary collaboration, this study calculates multidimensional characteristic indicators of urban green space spatial patterns. These include spatial composition and landscape morphology dimensions commonly used in existing local-scale green space cooling studies, as well as layout structure and network connectivity dimensions for measuring the structuring and networking levels of green spaces. A total of 32 candidate indicators were ultimately selected: Coverage, Area-weighted Mean Patch Size (Area_AM), Largest Patch Index (LPI), Area-weighted Mean Shape Index (AWMSI), Patch Density (PD), Area-weighted Mean Euclidean Nearest Neighbor Index (ENN_AM), Shannon's Diversity Index (SHDI), Moran's I, and Integral Index of Connectivity (IIC).
[0115] Through collinearity analysis, the following urban green space characteristic indicators were selected for subsequent modeling: Shannon Diversity Index (SHDI), Patch Density (PD), Overall Connectivity Index (IIC), Moran's Index (Moran I), Area-Weighted Nearest Neighbor Index (ENN_AM), Green Space Percentage (Coverage), Area-Weighted Average Patch Size (Area_AM), Largest Patch Index (LPI), and Area-Weighted Shape Index (AWMSI).
[0116] S03. Model Construction. Using the latitude and longitude of the spatial units and the urban green space characteristic indicators obtained in S02 as independent variables, and the surface heat island intensity of the spatial units calculated in S01 as the dependent variable, a geographically weighted random forest (GWRF) model is constructed.
[0117] In practice, GWRF is implemented directly using the SpatialML function package in R Studio. In this embodiment, the optimal bandwidth for the model is determined to be 38. The model is refitted using this bandwidth, and the coefficient of determination R is used. 2 The root mean square error (RMSE) and mean absolute deviation (MAE) were used as evaluation metrics for model accuracy. The model in this embodiment achieved very good fitting results, specifically: R0 2 =0.92, RMSE=0.15℃, MAE=0.10℃.
[0118] S04. Results Interpretation and Visualization. GWRF was used to explore the feature importance of urban green space characteristic indicators. The feature importance of each spatial unit was obtained by calculating the proportion of each indicator to the local feature importance of all indicators. Regional statistics were used to obtain the ranking of the feature importance of urban green space characteristic indicators in different regions. For example, the importance ranking at the national scale is as follows: Shannon Diversity Index (SHDI) > Patch Density (PD) > Overall Connectivity Index (IIC) > Moran's Index (Moran I) > Area-Weighted Nearest Neighbor Index (ENN_AM) > Green Space Percentage (Coverage) > Area-Weighted Average Patch Size (Area_AM) > Largest Patch Index (LPI) > Area-Weighted Shape Index (AWMSI) (Table 1).
[0119] Table 1. The Significance of Various Urban Green Space Characteristics in National and Different Climate Zones
[0120]
[0121] Visualize the impact curves of different characteristic indicators of urban green space on the intensity of surface heat island in China using ALE plots, such as... Figure 3As shown, the blue dashed line represents the reliable threshold identified by the threshold identification method based on the triple principle. The specific contents of the triple principle are as follows: (1) The local point density at the threshold is at least 50% of the average point density of the entire variable interval; (2) The threshold neighborhood interval is at least 5 spatial units; (3) The length of the 95% confidence interval of the fitted curve at the threshold is at most 20% of the entire prediction range. Among them, the local point density and the neighborhood interval are (1) the existing data are first objectively binned based on the Friedman-Diaconis rule, and (2) the point density and the number of spatial units in the bin where the threshold is located are then calculated. Only thresholds that meet all three principles are considered robust and reliable, which helps to guide urban green space planning more accurately. Overall, the ALE curve generally presents five types of influence curves, namely U-shaped, inverted U-shaped, N-shaped, inverted N-shaped, and continuously decreasing. Specifically, SHDI, PD, and AWMSI are inverted U-shaped, with their respective thresholds of 0.12, 13.00, and 6.00. ENN_AM and Area_AM exhibit a relatively complex N-type behavior, while Moran's I exhibits an inverted N-type behavior.
[0122] S05. Analysis of Different Spatial Driving Factors: Based on the predicted values of surface heat island intensity in different spatial zones (background climate, economic zones, geographical zones, etc.), statistical analysis was conducted to obtain key influencing indicators for each zone, their corresponding ALE diagrams, and their respective impact thresholds. This embodiment mainly focuses on the nonlinear cooling mechanism and spatial heterogeneity of urban green spaces in different climate zones. As shown in Table 1, the main influencing indicators in different climate zones exhibit significant spatial heterogeneity. For example, SHDI is the most significant dominant indicator in the warm temperate zone, while PD is the dominant indicator in the northern subtropical zone. In the temperate region, SHDI and LPI are the most prevalent dominant indicators, while in the northern subtropical zone, PD, Coverage, and IIC have the strongest impact in more cities. The ALE curves of significant influencing indicators in different climate zones are shown below. Figure 4 As shown. Overall, IIC and Area_AM show strong similarity, while other indicators show significant heterogeneity. For example, SHDI shows a continuous downward trend in temperate and northern subtropical zones, while it shows an inverted N-shaped curve in warm temperate zones and an N-shaped curve in the tropics. Moran's I shows an N-shaped curve similar to the national model in warm temperate and central subtropical zones, while it shows a continuous downward trend in temperate and tropical zones.
[0123] S06. Optimization of Urban Green Space Spatial Pattern Based on Local Conditions: This study systematically reviews and summarizes the influence relationships and thresholds of multi-dimensional characteristic indicators of urban green space on the intensity of the surface heat island at the national and regional levels, and deeply analyzes their common patterns and regional differences. The thresholds of the influencing indicators at the national and regional levels are summarized, and the results are shown in Table 2. Overall, except for coverage and IIC, all UGI indicators at the city level show obvious thresholds and significant regional differences. In most cases, the threshold for achieving optimal cooling effects in a specific climate zone (marked with *) is lower than the threshold at the national level. For example, in temperate and warm temperate regions, Area_AM reaches 1300 hectares to produce the best cooling effect, while national-level studies indicate that an additional 200 hectares are needed. Similar differences exist between the national and regional optimal thresholds for ENN_AM and Moran's I, but the magnitude of the difference is smaller.
[0124] Table 2 Summary of threshold values for influencing indicators in different regions
[0125]
[0126] In Table 2, + indicates the threshold at which the urban heat island effect changes from warming to cooling, - indicates the threshold at which the opposite change occurs (i.e., from cooling to warming), # indicates the threshold at which the cooling effect intensifies (i.e., from mild to severe), and ~ indicates the threshold at which the cooling effect weakens (i.e., from severe to mild). The thresholds for achieving the optimal cooling effect are also marked with *.
[0127] Based on the differentiated characteristics, importance, impact curves, and thresholds of different regions, this paper proposes site-specific optimization strategies for urban green spaces. At the national level, it recommends constructing diverse, highly networked, and structured urban green space networks. In tropical regions facing severe heat problems, it recommends establishing green spaces with an LPI ≥ 58% and a PD ≤ 12.4 No. / 100ha. In arid and temperate climate zones, it is more recommended to design green space networks composed of small or medium-sized green patches. Specifically, in warm temperate regions, green patches should be designed to be less than 500ha or around 1300ha to achieve more effective heat mitigation; while in temperate regions, the threshold for small patches should be adjusted to below 600ha.
[0128] As can be seen from the above, the method of this invention is simple and highly operable. The proposed method for evaluating and optimizing the cooling benefits of urban green space systems based on interpretable spatial machine learning is more in line with real-world needs. Compared with currently applied methods for evaluating and optimizing the cooling benefits of urban green spaces:
[0129] This invention constructs an interpretable spatial machine learning analytical framework, scientifically and accurately uncovering the nonlinear relationships and spatial heterogeneity of urban green space cooling mechanisms. It also develops a reliable threshold identification method based on a triple principle, enabling the precise implementation of green space optimization strategies across multiple dimensions, including spatial location, effect characteristics, and intervention intensity. This provides methodological and technical support for enhancing urban climate adaptation capabilities through nature-based solutions.
[0130] This invention constructs an interpretable spatial machine learning analysis framework combining geographically weighted random forests and cumulative local effects maps. Based on a large-sample example at the city scale, the operational process of this framework is detailed, and the nonlinear relationships and impact thresholds of multi-dimensional characteristic indicators of urban green space spatial patterns on the surface thermal environment are explored, along with their similarities and differences across different regions. Based on the scientific and precise exploration of the cooling mechanisms and spatial heterogeneity of urban green spaces, site-specific urban green space planning strategies are proposed, providing scientific support for enhancing urban climate adaptation capabilities based on nature-based solutions.
[0131] The implementation of the various embodiments of this invention is based on programmed processing through a system with processor functionality. Therefore, in practical engineering, the technical solutions and functions of the various embodiments of this invention are encapsulated into various modules. Based on this reality, and building upon the above embodiments, the embodiments of this invention provide a system for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning. This system is used to execute a method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning from the above method embodiments.
[0132] See Figure 5 The system includes:
[0133] The data preparation module is used to construct an urban green space dataset and a corresponding urban surface temperature dataset based on urban boundary data, land use data and surface temperature data of the target area.
[0134] The urban green space characteristic index calculation module is used to calculate several urban green space characteristic indices for each spatial unit in the urban green space dataset.
[0135] The geographic weighted random forest model construction module is used to construct a geographic weighted random forest model to fit the nonlinear relationship between the latitude and longitude of all spatial units and several urban green space characteristic indicators and the surface temperature data of the corresponding spatial units in the urban surface temperature dataset. The geographic weighted random forest model integrates several local random forest sub-models. The local random forest sub-models take the latitude and longitude of each spatial unit and several urban green space characteristic indicators as inputs and output the surface temperature data of that spatial unit.
[0136] The cooling effect modeling module is used to divide the target area into zones and, based on the geographically weighted random forest model, construct an interpretable model of the cooling effect of each urban green space characteristic index within the target area and within each zone of the target area. This includes: the characteristic importance of each urban green space characteristic index, the cumulative local effect map, and the reliable threshold.
[0137] The green space cooling benefit optimization decision module is used to interpretably model the cooling benefits of each city's green space characteristic indicators, determine the recommended value range and the value range to be avoided for each green space characteristic indicator in the target area and in each sub-area of the target area, and obtain the green space cooling benefit optimization decision for the target area.
[0138] It should be noted that the system embodiments provided by the present invention are used not only to implement the methods in the above method embodiments, but also to implement the methods in other method embodiments provided by the present invention. The only difference is that corresponding functional modules are set. The principle is basically the same as that of the above system embodiments provided by the present invention. As long as those skilled in the art can improve the modules in the above system embodiments by referring to the specific technical solutions in other method embodiments and combining technical features to obtain corresponding technical means and technical solutions composed of these technical means, on the basis of the above system embodiments, and on the premise of ensuring the practicality of the technical solutions, they can obtain corresponding system-like embodiments for implementing the methods in other method-like embodiments.
[0139] The method in this embodiment of the invention is implemented using an electronic device; therefore, it is necessary to introduce the relevant electronic device. For this purpose, embodiments of the present invention provide an electronic device, such as... Figure 6 As shown, the electronic device includes: at least one processor, a communication interface, at least one memory, and a communication bus, wherein the at least one processor, the communication interface, and the at least one memory communicate with each other via the communication bus. The at least one processor invokes logical instructions stored in the at least one memory to execute all or part of the steps of the methods provided in the foregoing method embodiments.
[0140] Furthermore, when the logical instructions in at least one of the aforementioned memories are implemented as software functional units and sold or used as independent products, they are stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, is embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (a personal computer, server, or network device) to execute all or part of the steps of the methods described in the various method embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks—various media for storing program code.
[0141] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, located in one place, or distributed across multiple network units. The purpose of this embodiment is achieved by selecting some or all of the modules according to actual needs. Those skilled in the art will understand and implement this without any inventive effort.
[0142] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0143] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0144] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0145] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0146] Based on the same technical concept as the foregoing embodiments, the present invention provides a non-transitory computer-readable storage medium that stores computer instructions that cause the computer to execute a method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning.
[0147] In summary, this invention discloses a method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning. The method includes the following steps: acquiring basic data for the target area; characterizing multi-dimensional features of urban green spaces; constructing a geographically weighted random forest model, where urban green space indicators and the latitude and longitude of spatial units are independent variables, and surface temperature is the dependent variable. The model can output results such as the importance of local variables and predicted values for each spatial unit; interpreting and visualizing the results: quantifying the contribution of various green space features to the thermal environment across the entire region based on the importance of local variables, revealing the nonlinear relationship between green space features and cooling effects using a cumulative local effect (ALE) plot, and determining a robust and reliable cooling threshold using a triple-principle threshold identification method; analyzing different spatial driving factors; and optimizing urban green spaces according to local conditions. This invention is highly operable and can scientifically and accurately uncover the nonlinear relationships and spatial heterogeneity of the multi-dimensional cooling mechanisms of urban green spaces. It effectively realizes the precise implementation of green space optimization strategies in multiple dimensions such as spatial location, effect characteristics, and intervention intensity, providing methodological and technical support for improving urban climate adaptation capabilities based on nature-based solutions.
[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning, characterized in that, include: Based on urban boundary data, land use data and surface temperature data of the target area, an urban green space dataset and a corresponding urban surface temperature dataset are constructed. For each spatial unit in the urban green space dataset, calculate several corresponding urban green space characteristic indicators. A geographically weighted random forest model is constructed to fit the nonlinear relationship between the latitude and longitude of all spatial units and several urban green space characteristic indicators and the surface temperature data of the corresponding spatial units in the urban surface temperature dataset. The geographically weighted random forest model integrates several local random forest sub-models. The local random forest sub-models take the latitude and longitude of each spatial unit and several urban green space characteristic indicators as inputs and output the surface temperature data of that spatial unit. The target area is divided into zones, and based on the geographically weighted random forest model, an interpretable model of the cooling effect of each urban green space characteristic index within the target area and within each zone of the target area is constructed, including: the characteristic importance of each urban green space characteristic index, the cumulative local effect map, and the reliable threshold. Based on the interpretable modeling of the cooling effect of the green space characteristic indicators of each city, the recommended value range and the value range to be avoided for each green space characteristic indicator in the target area and in each sub-area of the target area are determined, so as to obtain the optimization decision of the cooling effect of green space in the target area.
2. The method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning as described in claim 1, characterized in that, The urban green space dataset stores urban green space masks, and the construction process includes: Within the urban boundaries defined by urban boundary data, the urban green space definition area masks in land use data are merged to obtain an urban green space mask, and an urban green space dataset is constructed accordingly.
3. The method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning as described in claim 1, characterized in that, The urban green space characteristic indicators include, but are not limited to, green space ratio, patch size, maximum patch index, patch density, area-weighted average shape index, average nearest neighbor distance, Shannon diversity index, Moran's I, and overall connectivity index.
4. The method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning as described in claim 1, characterized in that, The construction process of the geographically weighted random forest model includes: Select a single spatial unit as the target spatial unit and construct a random forest sub-model for it. Repeat this process for all spatial units to obtain several random forest sub-models. Combine the latitude and longitude of each spatial unit, urban green space characteristic indicators and the corresponding urban surface temperature data in the urban surface temperature dataset to obtain training samples. In this way, the training samples corresponding to each spatial unit are obtained. The spatial distance between the target spatial unit and other spatial units of the random forest sub-model is calculated based on latitude and longitude. The spatial weight is calculated using the Gaussian kernel function. The training samples corresponding to each spatial unit are weighted using the spatial weight to obtain the geographically weighted sample training set of the random forest sub-model. The random forest sub-model is trained using the geographically weighted sample training set. After training, a local random forest sub-model is obtained. Several local random forest sub-models are obtained in this way.
5. The method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning as described in claim 1, characterized in that, The process of obtaining the characteristic importance of each urban green space characteristic indicator includes: To obtain the importance scores of each urban green space feature index in the local random forest sub-model of each spatial unit in the geographically weighted random forest model, construct an importance score vector as the local variable importance of the spatial unit, and thus obtain the local variable importance of each spatial unit. The importance of local variables of all spatial units within the region is obtained, the proportion of the importance of local variables of each urban green space characteristic index to all urban green space characteristic indices is calculated, and the characteristic importance of each urban green space characteristic index within the region is statistically obtained.
6. The method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning as described in claim 1, wherein the process of obtaining the cumulative local effect map and the reliable threshold includes: The latitude and longitude of all spatial units in the region and several urban green space characteristic indicators are input into a geographic weighted random forest model. The predicted surface temperature data results are output, and a cumulative local effect map of each urban green space characteristic indicator in the region is drawn. The urban green space characteristic indicator values represented by the slope change points in the cumulative local effect map are selected as candidate thresholds for urban green space characteristic indicators in the region. Candidate thresholds that meet the triple selection principle are selected as reliable thresholds for urban green space characteristic indicators in the region.
7. The method for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning as described in claim 1, characterized in that, The reliable thresholds for urban green space characteristic indicators within the region simultaneously meet the triple screening principle, including: the density of local spatial units at the reliable threshold is not less than 50% of the average point density of the entire variable interval; the preset neighborhood interval of the reliable threshold contains no less than a preset number of spatial units; and the length of the 95% confidence interval of the fitted cumulative local effect curve at the reliable threshold is not greater than 20% of the entire prediction range.
8. A system for evaluating and optimizing the cooling benefits of urban green spaces based on interpretable spatial machine learning, characterized in that, include: The data preparation module is used to construct an urban green space dataset and a corresponding urban surface temperature dataset based on urban boundary data, land use data and surface temperature data of the target area. The urban green space characteristic index calculation module is used to calculate several urban green space characteristic indices for each spatial unit in the urban green space dataset. The geographic weighted random forest model construction module is used to construct a geographic weighted random forest model to fit the nonlinear relationship between the latitude and longitude of all spatial units and several urban green space characteristic indicators and the surface temperature data of the corresponding spatial units in the urban surface temperature dataset. The geographic weighted random forest model integrates several local random forest sub-models. The local random forest sub-models take the latitude and longitude of each spatial unit and several urban green space characteristic indicators as inputs and output the surface temperature data of that spatial unit. The cooling effect modeling module is used to divide the target area into zones and, based on the geographically weighted random forest model, construct an interpretable model of the cooling effect of each urban green space characteristic index within the target area and within each zone of the target area. This includes: the characteristic importance of each urban green space characteristic index, the cumulative local effect map, and the reliable threshold. The green space cooling benefit optimization decision module is used to interpretably model the cooling benefits of each city's green space characteristic indicators, determine the recommended value range and the value range to be avoided for each green space characteristic indicator in the target area and in each sub-area of the target area, and obtain the green space cooling benefit optimization decision for the target area.
9. An electronic device, characterized in that, The method includes a memory and a processor, the memory storing program instructions that are executed by the processor, the processor invoking the program instructions to perform the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions that cause the computer to perform the method described in any one of claims 1 to 7.