Urban ecological risk prediction method combined with random forest algorithm

By combining the random forest algorithm and Bagging strategy, processing multi-source data and combining the ARIMA model, a high-precision urban ecological risk prediction model is built, solving the shortcomings of traditional methods in data processing and model accuracy, and achieving more accurate and dynamic risk assessment.

CN120013264APending Publication Date: 2025-05-16CHINESE RES ACAD OF ENVIRONMENTAL SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510503786.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional urban ecological risk prediction methods are difficult to maintain data integrity when processing multi-source and high-dimensional data, and the model lacks prediction accuracy and dynamic prediction capabilities when facing complex nonlinear relationships and dynamically changing urban ecosystems.

Method used

The random forest algorithm is used in combination with Bagging strategy, and the multi-source data is standardized and spatially aligned through the data preprocessing module, and the core variables are screened and the ARIMA model is combined to predict the change trend of elements. A high-precision urban ecological risk prediction model is constructed, and a decision support solution is output through the GIS visualization platform.

Benefits of technology

It significantly improves the accuracy and stability of urban ecological risk prediction, can more accurately reflect the true status and future trends of urban ecosystems, and provides real-time and dynamic risk assessment support to help decision makers more effectively identify and regulate risk areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013264A_ABST
    Figure CN120013264A_ABST
Patent Text Reader

Abstract

The invention provides an urban ecological risk prediction method combined with a random forest algorithm, and belongs to the technical field of crossing of urban ecological risk assessment and machine learning. The data preprocessing module carries out standardization and space alignment on multi-source city data, and constructs a feature set containing natural factors and human activity elements. And screening core variables significantly related to the ecological response by using multiple regression and an ARIMA model, and predicting the change trend of the core variables. A random forest model is constructed through a Bagging strategy, input variables are optimized according to feature importance, and model parameters are adjusted through grid search and cross validation. The model can convert real-time data into a risk level distribution diagram, and a visualization scheme is provided for decision support by using a GIS visualization platform. According to the method, a statistical method and a machine learning algorithm are combined, high-accuracy urban ecological risk prediction is realized, and the method has an important value for urban planning and ecological management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross-technical field of urban ecological risk assessment and machine learning, and more specifically relates to an urban ecological risk prediction method combined with a random forest algorithm. Background Art

[0002] Urban ecological risk prediction is an important topic in the fields of urban planning, ecological management, and sustainable development. Traditional prediction methods mainly rely on statistical analysis and simple machine learning models. However, these methods gradually expose many limitations when faced with complex and changing urban ecosystems.

[0003] First, in terms of data integration, traditional methods often have difficulty in efficiently processing multi-source, high-dimensional data. Urban ecosystems include natural elements such as mountains, rivers, forests, fields, lakes, and grasses, as well as human activity data such as population, GDP, and industrial development intensity. These data have different sources, formats, and scales. When processing these data, traditional methods often require a lot of preprocessing and simplification work, which may lead to the loss of key information, thus affecting the accuracy of the prediction results.

[0004] Secondly, in terms of model accuracy, traditional statistical methods such as multivariate regression analysis are weak in capturing nonlinear relationships and complex interactions. Single machine learning models, such as decision trees, can handle nonlinear relationships to a certain extent, but are easily disturbed by noisy data and have poor generalization performance. This results in traditional methods often failing to accurately reflect the true state and future trends of the ecosystem when predicting urban ecological risks.

[0005] Finally, in terms of dynamic prediction capabilities, existing methods generally lack effective fitting of time series data and accurate prediction of future trends. The urban ecosystem is a dynamically changing system, with various ecological elements and human activities constantly changing. However, traditional methods can often only analyze based on static data snapshots and cannot effectively simulate the dynamic interactions between multiple factors, thus failing to provide real-time, dynamic prediction support for urban planning and ecological management. Summary of the invention

[0006] This paper proposes an urban ecological risk prediction method combined with a random forest algorithm, aiming to integrate big data analysis and machine learning technology to achieve high-precision prediction and dynamic assessment of urban ecological risks to adapt to the complex and changeable urban ecosystem.

[0007] In order to achieve the above object, the present invention is implemented by adopting the following technical scheme: the method comprises: The multi-source urban data are standardized and spatially aligned through the data preprocessing module to construct a feature set containing natural factors and human activity elements; Multiple regression analysis was used to screen the core variables significantly related to ecological response, and the ARIMA model was used to predict the trend of factor changes; The random forest model was constructed using the bagging strategy, the input variable set was optimized through feature importance evaluation, and the number and depth parameters of decision trees were adjusted using grid search and cross-validation. Real-time data is input into the optimization model to generate a risk level distribution map, risk hotspots are identified based on spatial autocorrelation analysis and density clustering, and decision support solutions are output through the GIS visualization platform.

[0008] In one solution, the data preprocessing includes: performing radiation correction and atmospheric correction on remote sensing image data, spatially downscaling socioeconomic statistical data according to population density weights, unifying the geographic coordinate system to CGCS2000, and filling missing values ​​using the KNN algorithm.

[0009] In one embodiment, in the multiple regression analysis, variance inflation factor VIF>10 is used as the multicollinearity elimination standard, and the regression coefficient is corrected by Bonferroni to retain significant variables with p<0.01.

[0010] In one embodiment, the ARIMA model uses an ADF test to determine the difference order d, determines the autoregressive order p by the truncation of the partial autocorrelation plot, and uses the Bayesian Information Criterion BIC to select the optimal (q, p, d) parameter combination.

[0011] In one embodiment, the feature importance evaluation uses a permutation importance algorithm to perform 100 random perturbations on each feature, and calculates the weighted average of the decrease in model accuracy as the importance score.

[0012] In one embodiment, the grid search sets the number of decision trees to a range of 200,800 with a step size of 50, the maximum depth range to 8,15, and the minimum number of leaf node samples to 0.5% of the total sample size.

[0013] In one solution, the GIS visualization platform integrates ArcGIS Engine components, and the risk level distribution map uses Kriging interpolation to generate isosurfaces, overlays the OpenStreetMap base map, and supports three-dimensional terrain perspective rendering.

[0014] In one solution, the decision support solution includes a risk diffusion simulation module, which uses a cellular automaton model to set transfer rules for industrial land expansion and vegetation coverage attenuation, and combines the Monte Carlo method to generate a risk probability cloud map.

[0015] Beneficial effects of the present invention: 1. The present invention adopts a standardized and spatially aligned data preprocessing method, which effectively overcomes the difficulties of traditional methods in processing multi-source, high-dimensional data. The correlation between human activity elements and natural factors is clearer, greatly improving the accuracy of model prediction.

[0016] 2. Through multivariate regression analysis, core variables significantly correlated with ecological responses were screened, and the ARIMA model was combined to predict the changing trends of environmental factors. This not only captured the linear relationship of the data, but also took into account the impact of time series, providing more comprehensive and accurate input parameters for risk prediction.

[0017] 3. The present invention adopts the random forest model and the bagging strategy, which extends the capabilities of a single decision tree model, significantly improves the model accuracy and stability, resists the intrusion of noise data, and improves the generalization performance of the model.

[0018] 4. Identifying risk hotspots based on spatial autocorrelation analysis and density clustering can intuitively reflect the areas with higher risks in the city, and visualize the distribution of risk levels through the GIS visualization platform, which greatly facilitates decision makers to identify and regulate risk areas.

[0019] 5. The decision support scheme generated by the present invention includes a risk diffusion simulation module, which can be used to predict future risk distribution and provide a basis for scientific decision-making for decision makers, and has high practical value.

[0020] In summary, the present invention combines statistical methods with machine learning algorithms, fully utilizes the advantages of various data and models, effectively improves the accuracy and real-time performance of urban ecological risk prediction, and has broad application prospects in urban planning and ecological management. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0022] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Typical embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described in the present invention. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0023] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as those understood by a person skilled in the art of the present invention. The terms used in the present invention in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Typical embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described in the present invention. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0024] This paper proposes a method for predicting urban ecological risks by combining a random forest algorithm. This method aims to achieve high-precision prediction and dynamic assessment of urban ecological risks by integrating multi-source data with advanced machine learning technology. The following are the specific implementation steps of this method: like Figure 1 As shown, the present invention proposes a method for predicting urban ecological risks by combining a random forest algorithm. The method aims to achieve high-precision prediction and dynamic assessment of urban ecological risks by integrating multi-source data with advanced machine learning technology. The following are the specific implementation steps of the method: Step 1: Data collection and interpretation. First, a multi-source data collection system needs to be established to systematically obtain multi-dimensional information of urban ecosystems. For natural element data, it is necessary to integrate land survey data, forestry resource census data, and hydrological monitoring data from the water conservancy department through the geographic information system (GIS) platform, including spatial vector data such as terrain elevation model (DEM), water system distribution map, vegetation coverage index (NDVI), soil type database, etc. At the same time, it is necessary to connect with the biodiversity monitoring site data of the environmental protection department to build an ecological background database covering "mountains, rivers, forests, fields, lakes and grasses". For human activity data, it is necessary to coordinate with the statistical department to obtain population census grid data, regional GDP data in the economic statistics yearbook, integrate the geographic coordinates of industrial enterprises and pollution permit data in the industrial and commercial registration information, generate industrial development intensity heat maps through spatial interpolation technology, and use mobile phone signaling data or public transportation card swipe records to invert the dynamic distribution characteristics of the population. The processing of long-term satellite remote sensing data requires the use of ENVI or Google Earth Engine platform to perform radiation correction, atmospheric correction and geometric precision correction on Landsat and Sentinel series satellite images, and use object-oriented classification algorithms to extract the trajectories of land use changes such as urban built-up area expansion, cultivated land conversion, and green space fragmentation between 1990 and 2020, and quantify the spatial structure evolution through landscape pattern indices (such as patch density and aggregation index).

[0025] In the data preprocessing stage, it is necessary to build an automated data cleaning pipeline, use tools such as Pandas to fill missing values ​​(such as KNN interpolation method) and remove outliers (based on IQR interval or 3σ principle) for structured data, use FME data conversion tools to uniformly project and convert spatial data in different coordinate systems (such as WGS84 and CGCS2000), and implement field alignment and unit standardization (such as uniform conversion to hectares, ten thousand yuan, etc.) for Excel table data provided by the statistical department. For multi-source heterogeneous data, it is necessary to establish a spatiotemporal benchmark matching mechanism, align the monthly economic data, annual land change survey data with the daily updated remote sensing data in time series, and ensure that all data layers have the same spatial resolution and grid division standards through spatial overlay analysis. Data standardization processing requires selecting appropriate methods according to the data type: continuous variables (such as GDP values) are standardized using Z-score, categorical variables (such as land use types) are One-Hot encoded, and spatial raster data are subject to pixel value normalization (Min-Max Scaling). Ultimately, a multidimensional data cube with complete spatiotemporal dimensions and unified indicator scales is constructed and stored in the PostgreSQL spatial database or Hadoop distributed file system to provide high-quality input for subsequent modeling.

[0026] Step 2: Model construction and optimization: Establish a quantitative relationship model between human activity factors and ecological environment through multiple regression analysis. Let the ecological environment response variable be Y (such as biodiversity index or pollutant concentration), and the human activity explanatory variables constitute the matrix (including pre-processed standardized data such as population density and industrial output value), and construct a multivariate linear regression model , the least squares method is used to estimate the parameters The significant variables (p<0.05) were screened by t-test, and strong effect factors with the absolute value of standardized regression coefficient greater than 0.1 were retained. The ARIMA model was applied to time series data for factor trend forecasting, and the Order-difference autoregressive moving average model ,in is the lag operator, is an autoregressive polynomial, It is a moving average polynomial, and the optimal parameter combination is determined by the AIC criterion to predict the change trajectory of each factor in the next five years.

[0027] In the construction of the random forest model, the ecological risk level R is set as a categorical target variable (low, medium, high risk) or a continuous risk index, and the input feature space contains the preprocessed natural element data. Human activity factors screened by regression analysis Each decision tree uses the Gini impurity minimization criterion when splitting nodes , where \( p_k \) is the proportion of samples of the kth class in the node, and the training subset is generated by Bootstrap sampling and constructed

[0028] A decision tree forms a forest. The feature importance is calculated by calculating the The weighted average reduction in Gini impurity is evaluated: ,in Indicates the number of trees based on The node set is split, and the features with the top 20% importance scores are retained to build a streamlined model. The bagging strategy is used in the model optimization stage, and 63.2% of the original data of the training set is sampled each time, and the out-of-bag error (OOB error) is used to calculate the original data.

[0029] Evaluate model performance in real time, where is the majority vote result of using only trees that do not contain sample i for prediction. Use 5-fold cross validation to adjust hyperparameters and define the parameter search space: number of decision trees , Maximum Depth , minimum number of split samples , find the cross-validation loss function through grid search The optimal parameter combination that minimizes the mean square error of the validation set , and finally output the optimized model with the smallest OOB error and the highest Kappa coefficient.

[0030] Step 3: Risk prediction and output: Import real-time monitoring data (such as real-time readings of air quality sensors, traffic flow data) and simulation scenario data (such as planned industrial zone expansion plans) into the optimized random forest model through API interfaces or direct database connections. The input data must undergo a standardized processing process consistent with the training data, including unified conversion of the spatial coordinate system WGS84, comparable price conversion of GDP data, and gridded resampling of population density to ensure that the data format fully matches the tensor structure of the model input layer. The model performs parallel calculations on each geographic grid unit (such as 100m×100m grids), outputs ecological risk index values, and divides them into three risk levels: low (0-0.3), medium (0.3-0.6), and high (0.6-1.0) according to the natural breakpoint method. At the same time, the main contributing factors in each unit are recorded (such as industrial emission intensity exceeding the threshold is marked as a red warning factor).

[0031] When generating the risk level distribution map, ArcGIS Pro's spatial analysis module is used to spatially connect the CSV format risk index table output by the model with the original geographic vector boundary, and the discrete point data is converted into a continuous surface through the Kriging interpolation method. Basic geographic elements such as road networks and water system layers are superimposed, and the color visual variable design principle is applied - low-risk areas use cold colors (blue-green gradients), high-risk areas use warm colors (yellow-red gradients), and dynamic legends and risk value annotations are added. For complex areas such as urban built-up areas, 3D voxel visualization technology is used to construct a risk stereo model, and the risk gradient changes at different height levels are displayed through transparency mapping.

[0032] Risk hotspot identification uses a method that combines spatial autocorrelation analysis with density clustering. First, Moran's I is calculated to evaluate the spatial aggregation of risk distribution. When I>0.4, it indicates a significant hotspot effect. Then, the Getis-Ord Gi* statistic is used to scan each grid unit and its neighborhood, and the Z score is calculated to identify statistically significant (p<0.01) high-value clustering areas. At the same time, the DBSCAN algorithm is used to detect contiguous areas with risk values ​​exceeding the threshold and continuous areas greater than 5 square kilometers as core hotspots. The prediction of potential risk areas combines the trend of factor changes output by the ARIMA model, simulates the risk diffusion path in the next three years, and uses cellular automata to set driving rules such as industrial land expansion and population migration to generate a probability cloud map of risk spread.

[0033] The decision support system achieves output by building a WebGIS platform, integrating three modules: risk, hot spot list and prediction simulation. The backend of the platform uses GeoServer to publish WMS services, and the front end uses the Leaflet framework to achieve interactive browsing, allowing the planning department to automatically generate a diagnostic report after selecting a specific area, including the historical risk evolution curve of the area, the contribution pie chart of the dominant influencing factors, and the governance case library of the adjacent areas. For the identified high-risk plots, the system calls the spatial optimization algorithm to generate a multi-objective restoration plan: the optimal path of the ecological corridor is calculated based on the minimum cost path model, the genetic algorithm is used to solve the Pareto optimal solution set of green space repair, and the risk mitigation effect after the implementation of different plans is visualized through the three-dimensional digital twin engine, and finally a set of decision-making documents including priority governance zoning, ecological control red lines and project library lists are formed.

[0034] Example: A specific embodiment of the present invention takes Shenzhen as the implementation object, integrating Landsat 8 remote sensing images from 2020 to 2023, PM2.5, SO2, and NO2 monitoring data released by the Ministry of Ecology and Environment, and socio-economic statistical data such as industrial output value, population density, road length, and green area released by the Shenzhen Statistics Bureau.

[0035] First, the remote sensing images were radiometrically calibrated and FLAASH atmospheric corrected to convert the reflectance of each band into the spectral characteristics of the surface material. The industrial output value, road length and other data were spatially segmented according to administrative boundaries, and the population density was used as the interpolation weight. The socioeconomic data were spatially aligned with the 500m×500m grid. For the missing traffic flow data, the KNN algorithm was used to fill in the average value of the adjacent grids.

[0036] Then, a multivariate regression analysis was conducted to test the correlation between factors such as PM2.5, SO2, NO2 concentration, industrial land proportion, population density, road length, and green area and ecological risk. Through the test of variance inflation factor VIF, the road length factor with VIF greater than 10 was excluded, and PM2.5 concentration, industrial land proportion, and population density were finally determined as core variables. The ARIMA (1,0,2) (1,1,0) 12 model was used to predict the green area. It is predicted that by 2025, the green area in Shenzhen will decrease by 2.7%, and the proportion of industrial land will increase by 8.3%.

[0037] Then, the parameters of the random forest model were determined through grid search and cross-validation, setting the number of decision trees between 200-800, the step size of 50, the maximum depth of 8-15, and the minimum number of leaf node samples to 0.5% of the total sample size. After training, a random forest model with 400 trees and a depth of 10 was finally determined. The importance of each variable was evaluated using the permutation importance evaluation method, and it was found that PM2.5 concentration was the most important, followed by the proportion of industrial land and population density.

[0038] According to the derived model, the real-time data in 2024 was input into the model for prediction, and the ecological risk level distribution map of Shenzhen was generated. It was found that the risks in Luohu District, Futian District and Nanshan District were relatively high and needed attention.

[0039] Finally, the risk level distribution map was published as a WMS service through ArcGIS server, and a public WebGIS application was established on the ArcGIS Online platform, providing functional modules such as risk warning, risk depression analysis, and risk spatiotemporal trend analysis. For areas with higher risks, optimization measures such as adjusting industrial layout, optimizing traffic design, and increasing green space coverage were proposed, and the risk changes after implementing these measures were simulated through cellular automation and Monte Carlo methods, providing a scientific basis for decision makers.

[0040] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).

[0041] It should be understood that the detailed description of the technical solutions of the present invention by means of the preferred embodiments is illustrative rather than restrictive. A person skilled in the art may modify the technical solutions described in the embodiments, or replace some of the technical features by equivalents, based on reading the specification of the present invention; and these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting urban ecological risks combined with a random forest algorithm, characterized in that: The method includes: The multi-source urban data are standardized and spatially aligned through the data preprocessing module to construct a feature set containing natural factors and human activity elements; Multiple regression analysis was used to screen the core variables significantly related to ecological response, and the ARIMA model was used to predict the trend of factor changes; The random forest model was constructed using the bagging strategy, the input variable set was optimized through feature importance evaluation, and the number and depth parameters of decision trees were adjusted using grid search and cross-validation. Real-time data is input into the optimization model to generate a risk level distribution map, risk hotspots are identified based on spatial autocorrelation analysis and density clustering, and decision support solutions are output through the GIS visualization platform.

2. The urban ecological risk prediction method combined with the random forest algorithm according to claim 1 is characterized in that: The data preprocessing includes: performing radiation correction and atmospheric correction on remote sensing image data, spatially downscaling and distributing socioeconomic statistical data according to population density weights, unifying the geographic coordinate system to CGCS2000, and filling missing values ​​with the KNN algorithm.

3. The urban ecological risk prediction method combined with the random forest algorithm according to claim 1 is characterized in that: In the multivariate regression analysis, variance inflation factor VIF>10 was used as the multicollinearity elimination standard, and the regression coefficients were corrected by Bonferroni to retain significant variables with p<0.

01.

4. The urban ecological risk prediction method combined with the random forest algorithm according to claim 1 is characterized in that: The ARIMA model uses the ADF test to determine the difference order d, determines the autoregressive order p by the truncation of the partial autocorrelation diagram, and uses the Bayesian Information Criterion BIC to select the optimal (q, p, d) parameter combination.

5. The urban ecological risk prediction method combined with the random forest algorithm according to claim 1 is characterized in that: The feature importance evaluation adopts the permutation importance algorithm, performs 100 random perturbations on each feature, and calculates the weighted average of the decrease in model accuracy as the importance score.

6. The urban ecological risk prediction method combined with the random forest algorithm according to claim 1 is characterized by: The grid search sets the number of decision trees to a range of 200, 800 with a step size of 50, the maximum depth range to 8, 15, and the minimum number of leaf node samples to 0.5% of the total sample size.

7. The urban ecological risk prediction method combined with random forest algorithm according to claim 1 is characterized by: The GIS visualization platform integrates ArcGIS Engine components, and the risk level distribution map uses Kriging interpolation to generate isosurfaces, overlays the OpenStreetMap base map and supports three-dimensional terrain perspective rendering.

8. The urban ecological risk prediction method combined with random forest algorithm according to claim 1 is characterized by: The decision support scheme includes a risk diffusion simulation module, which uses a cellular automaton model to set the transfer rules of industrial land expansion and vegetation coverage attenuation, and combines the Monte Carlo method to generate a risk probability cloud map.

Citation Information

Patent Citations

  • Cellular automaton urban growth simulating method based on random forest

    CN104156537A

  • Wetland ecological risk index rapid estimation method based on remote sensing technology

    CN114118867A

  • Ecological system simulation and forest cultivation optimization method

    CN117892550A

  • Water quality evaluation method based on data analysis

    CN119624228A

  • Systems and methods for automatic environmental planning and decision support using artificial intelligence and data fusion techniques on distributed sensor network data

    US20230259798A1