A method for precise management of territorial space based on key driver threshold

By integrating multi-source data and machine learning models, combined with constraint line models, the problems of data integration and threshold identification in the assessment of the supply and demand balance of ecosystem services were solved, realizing a closed-loop technology chain from data input to management recommendations, and ensuring the systematicness and operability of the assessment results.

CN122114677APending Publication Date: 2026-05-29YUNNAN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YUNNAN UNIV
Filing Date
2026-01-22
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies for assessing the supply and demand balance of ecosystem services suffer from insufficient capacity for multi-factor and cross-scale integration, limitations in data heterogeneity and assimilation methods, reliance on linear or empirical judgments for threshold identification, and inadequate model interpretability and uncertainty quantification. These limitations make it difficult to translate scientific results into actionable management measures, leading to difficulties in policy implementation.

Method used

By employing multi-source data fusion, self-organizing neural network models for cluster analysis, and XGBoost regression models to identify key driving factors, combined with constraint line models to determine management thresholds, and through standardized vector grids and logical judgments to execute logical rules, a complete technical closed loop from data input to evaluation recommendations is achieved.

Benefits of technology

It enables precise management of the supply and demand balance of ecosystem services, improves the systematicness, objectivity and precision of assessment results, ensures the integrity and operability of assessment and monitoring methods, and provides actionable management strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114677A_ABST
    Figure CN122114677A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of space optimization management, and discloses a land space precise management method based on key driving factor thresholds, which adopts technical coupling and comparison of the leading ESs supply-demand types identified by SOM and the key driving factor thresholds extracted by constraint lines. The method solves the technical problem that the data collection, model analysis and policy making links are disconnected in the existing land space evaluation and monitoring, resulting in the difficulty in converting the evaluation results into specific instructions, realizes the full-chain technical closed loop from data input to evaluation suggestion output, ensures the systematicness and integrity of the evaluation and monitoring method, adopts the XGBoost model to process the complex interaction between variables, and innovatively adopts the constraint line model to analyze the constraint boundary of data distribution, so that the technical problem that the traditional evaluation method relies on a linear model and is difficult to capture nonlinear relationships and the evaluation suggestions are mostly qualitative descriptions is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spatial optimization management technology, specifically to a method for precise management of national land space based on key driving factor thresholds. Background Technology

[0002] With rapid urbanization and increased human activity in regions, the mismatch between the supply and demand of ecosystem services (ESs) has become one of the core issues restricting regional sustainable development. In recent years, policy requirements that emphasize both ecological protection and economic development (such as ecological red lines, territorial spatial planning, and carbon peaking and carbon neutrality targets) have created an urgent need for technological means that can provide operable, quantifiable, and enforceable spatial management thresholds. However, existing related technologies and research have significant shortcomings and deficiencies in meeting policy and decision-making level applications: insufficient multi-factor and cross-scale integration capabilities, limitations in data heterogeneity and assimilation methods, reliance on linear or empirical judgments for threshold identification, insufficient model interpretability and uncertainty quantification, and unclear transformation paths from scientific results to enforceable management measures. These issues collectively restrict the policy implementation and practical application of research findings. Specifically, existing methods mostly focus on estimating a single or a few ESSs, making it difficult to quantify the supply and demand of multiple types of services and their supply-demand relationships on the same raster or grid cell; the thresholds for driving factors are usually calculated based on linear regression, making it difficult to reveal the complex nonlinear relationship between driving factors and the supply-demand balance of ESSs; although machine learning methods can improve prediction accuracy, their black-box nature and lack of robust uncertainty assessment limit their interpretive support for policymakers; even if key influencing factors are identified, it is difficult to translate them into actionable and verifiable spatial control guidelines. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a method for precise management of national land space based on key driving factor thresholds, which has advantages such as ensuring the systematicness and integrity of the assessment and monitoring methods, and solves the aforementioned technical problems.

[0004] To achieve the above objectives, the present invention provides the following technical solution: a method for precise management of national land space based on key driving factor thresholds, comprising the following steps:

[0005] Step 1: Acquire and preprocess geographic data;

[0006] Step 2: Based on the preprocessed geographic data from Step 1, construct the supply and demand of ecosystem services, and calculate the corresponding supply-demand ratios. ;

[0007] Step 3: Generate a standardized vector grid based on ArcGIS software, perform statistical analysis and average calculations for each grid, assign the average to the corresponding grid cells, and output a comprehensive attribute table;

[0008] Step 4: After Z-score standardization based on the comprehensive attribute table output in Step 3, input it into the self-organizing neural network model for cluster analysis, obtain the cluster labels after clustering, associate them with the IDs of the grid cells, and input them into ArcGIS software;

[0009] Step 5: Input the Z-score-standardized comprehensive attribute table from Step 4 into the established XGBoost regression model to filter out the top-ranked data. Key driving factors;

[0010] Step Six: Use the key driving factors output in Step Five as independent variables, and their corresponding... The value is used as the dependent variable, and the beneficial influence threshold is obtained by using a constraint line fitting method.

[0011] Step 7: Obtain the measured values ​​of key driving factors and factor control thresholds in the current grid cell, calculate the deviation, and perform logical judgment based on the deviation and the preset logical rule base.

[0012] As a preferred technical solution of the present invention, the supply-demand ratio in step two The specific expression is as follows:

[0013]

[0014] in, Indicates supply, Indicate demand, Indicates maximum supply. This indicates the maximum demand.

[0015] As a preferred technical solution of the present invention, the supply and demand of ecosystem services in step two include the supply and demand of food production, the supply and demand of carbon sequestration, the supply and demand of soil conservation, the supply and demand of water production services, the supply and demand of habitat quality and the supply and demand of recreational services.

[0016] The supply of the grain production is allocated to spatial plots based on the total output in the statistical yearbook according to the vegetation index NDVI, and the demand for the grain production is obtained by multiplying the regional population by the per capita energy intake standard.

[0017] The supply of carbon sequestration is directly calculated from the net primary productivity of vegetation (NPP), and the demand for carbon sequestration is spatially calculated by correlating energy consumption with nighttime light data.

[0018] The supply of soil conservation is assessed using the RUSLE soil erosion model, and the demand for soil conservation is assessed.

[0019] The supply of the water production service is obtained by simulating rainfall, evapotranspiration and land use patterns using the InVEST model; the demand for the water production service is also obtained.

[0020] The supply of habitat quality, and the demand for habitat quality;

[0021] The supply of the leisure and recreation services, and the demand for the leisure and recreation services.

[0022] As a preferred technical solution of the present invention, the self-organizing neural network model in step three is constructed and trained using the kohonen package in the R language.

[0023] As a preferred technical solution of the present invention, the parameters optimized during the training process of the XGBoost regression model in step five include the number of trees, the maximum depth of the trees, the learning rate, and the regularization parameter.

[0024] Compared with existing technologies, this invention provides a method for precise management of national land space based on key driving factor thresholds, which has the following beneficial effects:

[0025] 1. This invention solves the technical problem in existing land space assessment and monitoring where data collection, model analysis, and policy formulation are disconnected, making it difficult to translate assessment results into specific instructions. It adopts seven interconnected technical steps, including multi-source data fusion, spatial intelligent zoning, key factor identification, and management threshold determination. This achieves a closed-loop technology chain from data input to assessment recommendation output, ensuring the systematic nature and integrity of the assessment and monitoring method.

[0026] 2. This invention uses standardized fishing net units as the basic unit for assessment and monitoring, and combines a self-organizing neural network model to perform unsupervised clustering of multi-dimensional ESDR data. This solves the problem of plastic area units caused by the reliance on administrative divisions in existing assessment methods, as well as the technical problems of strong subjectivity in traditional zoning methods, which leads to distorted analysis results. It realizes ecological zoning based on the inherent laws of data, and significantly enhances the objectivity of spatial analysis and the precision of assessment results.

[0027] 3. This invention uses the XGBoost model to handle complex interactions between variables and innovatively employs a constraint line model to analyze the constraint boundaries of data distribution. This solves the technical problems of traditional assessment methods relying on linear models and making it difficult to capture nonlinear relationships, and the assessment recommendations being mostly qualitative descriptions. It enables the efficient identification of key driving factors from massive driving factors and accurately extracts the quantitative thresholds that the key driving factors must satisfy in terms of numerical range or inflection points to achieve the balance between supply and demand of ecosystem services. This makes the decision-making basis for assessment and monitoring more accurate and operational.

[0028] 4. This invention employs a technical coupling and comparison method that combines the dominant ESS supply and demand type identified by SOM with the key driving factor threshold extracted from the constraint line. This solves the technical problem that traditional evaluation methods can only provide macroscopic and qualitative suggestions, but cannot link the suggestions with specific problems and quantitative indicators of spatial units. Attached Figure Description

[0029] Figure 1 A technical roadmap for a precise land space management method based on key driving factor thresholds;

[0030] Figure 2 This is a schematic diagram illustrating the basic principles of the XGBoost model;

[0031] Figure 3 This is a map showing the ecological zoning results of the national land space.

[0032] Figure 4 A ranking plot of the importance of driving factors;

[0033] Figure 5 The threshold results for key driving factors are shown in the figure. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] The abbreviations and translations of the terms mentioned below are as follows: ESs (Ecosystem services); ESDR (Ecosystem services supply-demand ratio); SOM (The self-organizing map); XGBoost (The Extreme Gradient Boosting algorithm); NDVI (Normalized Differential Vegetation Index); MAUP (Modifiable Area Unit Problem); FP (Food Production); CS (Carbon Sequestration); SR (Soil Retention); WY (Water Yield); HQ (Habitat Quality); LR (Leisure and Recreation) Leisure and recreation); POP (Density of population); GDP (Density of GDP); NT (Nightlight index); DFB (Distance from building land); DFT (Distance from trunk roads); DFR (Distance from railways); POC (Proportion of cultivated land); POF (Proportion of forest land); POB (Proportion of building land); TEM (Average annual temperature); PRE (Total annual precipitation); ALT (Altitude); SLP (Slope).

[0036] Please see Figures 1-5 A method for precise land space management based on key driving factor thresholds includes the following steps:

[0037] Step 1: Acquire and preprocess geographic data. This step aims to build a standardized, multi-source spatial database that can be used for all subsequent analyses, ensuring that all data are completely consistent in coordinate system, spatial resolution, and geographic extent.

[0038] Specifically, the study first systematically acquired natural environmental data from major scientific research data centers, and then collected socio-economic indicators and infrastructure data in conjunction with local statistical yearbooks. For example, digital elevation models (DEMs), land use / land cover data, and meteorological data were downloaded from the China Resources and Environment Science and Data Center; normalized difference vegetation index (NDVI), net primary productivity (NPP), and nighttime light remote sensing data were obtained from the Tibetan Plateau Science Data Center of the Chinese Academy of Sciences; simultaneously, socio-economic indicators such as population and GDP were collected from local statistical yearbooks and water resources bulletins. After data aggregation, rigorous preprocessing was performed using the ArcGIS software platform: all data were unified to the WGS_1984_UTM projection coordinate system using projection transformation tools; then, the spatial resolution of all raster data was standardized to 1000 meters using resampling tools; finally, based on the administrative boundaries of the study area, all data layers were masked and cropped to form a comprehensive "socio-ecological" system database that reflects the complex interaction between natural processes and human activities, laying a solid foundation for subsequent precise analysis.

[0039] Step 2: Based on the preprocessed geographic data from Step 1, construct the supply and demand of ecosystem services, and calculate the corresponding supply-demand ratios. ;

[0040] This invention quantifies the supply, demand, and balance of ecosystem services. Through multi-model integration and parameter localization, it aims to accurately quantify the supply and demand of six core ecosystem services (ESs)—food production, carbon sequestration, soil conservation, water production services, habitat quality, and recreation—on a unified grid cell, and assess their supply-demand balance. Compared to traditional methods that only assess supply, this invention comprehensively introduces and quantifies human demand for ecosystem services, constructing a more complete assessment system. Furthermore, it localizes all model parameters to reflect the specific conditions of the study area, significantly improving the accuracy and practical guidance of the assessment results.

[0041] Specifically, to accurately assess the health status of regional ecosystem services, this invention first systematically quantifies the supply and demand of six core services. Regarding food production, the supply is allocated to spatial plots based on the productivity levels reflected by the vegetation index (NDVI) in the statistical yearbook, while the demand is obtained by multiplying the regional population by the per capita energy intake standard. Specifically, for carbon sequestration services, the supply is quantified using the CASA model. Net primary productivity (NPP) is calculated by multiplying the photosynthetically active radiation of vegetation and light energy utilization rate using meteorological data, and then converted into carbon sequestration based on the principle of photosynthesis fixing 1.63 kg of carbon dioxide per kg of dry matter. The demand is spatially estimated by establishing a mapping relationship between nighttime light data and energy consumption, precisely allocating the total carbon emissions at the county scale according to the weight of each grid cell's light intensity value (DN value) relative to the total intensity. For soil conservation services, the Modified Universal Soil Erosion Equation (RUSLE) is used for assessment. This involves a comprehensive calculation of rainfall erosivity, soil erodibility, slope length and gradient, vegetation cover, and soil and water conservation measures to determine the amount of soil erosion that vegetation can prevent, which is then considered the supply. The calculated actual soil erosion is treated as the objective demand requiring remediation. For water production services, the supply is simulated using the InVEST model's water production module. Based on the water balance principle, the annual rainfall is subtracted from the actual evapotranspiration, determined by vegetation root depth, soil depth, and the plant's effective water use coefficient, to obtain the grid annual water production in millimeters (mm). The demand is determined by spatially weighting the total agricultural, ecological, industrial, and domestic water use within the region according to the NDVI distribution of cultivated land, the distribution of forest and grassland, the GDP share of industrial and mining land, and the population density of residential land. For habitat quality services, the supply is calculated using the InVEST model. By setting threat sources such as farmland, towns, and roads, along with their maximum impact distance and weights, and considering the sensitivity of different land use types, habitat degradation is calculated. The specific habitat quality value is then quantified using a half-saturation constant formula. Demand is calculated using the mean-difference method, defined as the difference between the average habitat quality standard of the study area and the actual supply value of each grid. For recreational services, the supply is first determined by assessing the potential recreational value per unit area of ​​green spaces such as woodlands and grasslands. Then, the effective supply is calculated by adjusting the grid's distance from main roads (accessibility factor), distance from water bodies (landscape factor), and topographic slope factor. Demand is quantified by multiplying the grid's population density by the government-planned per capita green space area standard and the average value of regional green space landscape, thus achieving a refined spatial expression of demand.

[0042] After independently quantifying the supply (S) and demand (D) of various services, this invention uses the Ecosystem Service Supply-Demand Ratio (ESDR) to ultimately measure and visualize the spatial matching relationship between the two. This indicator is calculated using the formula ESDR = 2(SD) / (Smax + Dmax), and the results intuitively reveal the supply and demand situation: when the ESDR value is greater than or equal to 0, it indicates that the supply of ecosystem services in the region can meet or even exceed human needs, indicating a state of balance or surplus; conversely, when the ESDR value is less than 0, it means insufficient supply, manifesting as an ecological deficit. This quantitative result provides a core basis for subsequent scientific zoning and precise management.

[0043] Step 3: Generate a standardized vector grid based on ArcGIS software, perform statistical analysis and average calculations for each grid, assign the average to the corresponding grid cells, and output a comprehensive attribute table;

[0044] Creating a fishing net and zoning statistics: To provide formatted and standardized input data for subsequent machine learning model analysis, this step utilizes a regular fishing net as a unified spatial analysis unit. The aim is to efficiently integrate and transform multi-source, fragmented geospatial data into a structured dataset with clearly defined causal relationships. This fundamentally overcomes the malleable area unit (MAUP) problem caused by differences in the shape and size of the analysis units, ensuring the objectivity of the model input and the scientific rigor of the analysis results. It serves as a crucial bridge connecting spatial data preprocessing and advanced model analysis.

[0045] This step begins by using the "Create Fishing Net" tool in the ArcGIS software environment to generate a standardized 1000m × 1000m vector grid with a unique ID, using the study area as a template. This grid forms the unified spatial analysis unit for all subsequent calculations. Then, using this fishing net as the basis for partitioning, the "Partition Statistics" tool is used for efficient data aggregation. On one hand, the feature values ​​of 14 driving factors within each grid cell are systematically calculated and extracted. For continuous raster data such as altitude, slope, temperature, and NDVI, their average values ​​within each grid cell are directly calculated. For distance-type factors such as roads and railways, a distance raster layer is first generated using the "Euclidean Distance" tool, and then the grid average value is calculated. For land use structure factors such as cultivated land and forest land, the land use layer is reclassified and partitioned to accurately determine the area proportion of each type of land use within each grid cell.

[0046] While aggregating the driving factors, the same partitioned statistical method was used to precisely assign the average values ​​of the six ecosystem service supply-demand ratios (ESDRs) calculated in step two—food production, carbon sequestration, soil conservation, water yield, habitat quality, and recreation—to each corresponding grid cell. Finally, all attribute data generated through partitioned statistics were spatially connected using unique grid IDs, integrating them into a comprehensive attribute table. In this table, each row represents a grid cell (i.e., an analysis sample), and each column corresponds to its unique ID, the quantified values ​​of 14 driving factors (as independent variables in the model), and the quantified values ​​of the six ESDRs (as dependent variables in the model), thus successfully transforming the complex spatial relationships into a standardized two-dimensional dataset that can be directly input into subsequent SOM and XGBoost models. The R code is as follows:

[0047] #Unsupervisedself-OrganizingMaps

[0048] #install.packages("kohonen")

[0049] library (kohonen)

[0050] #filesetting

[0051] setwd("F: / som / DZ_CLUSTER")

[0052] library(readxl)

[0053] #watershedscaleinputdata

[0054] data<-read.csv("GRID2010.csv",header=T)

[0055] #datapre-processing

[0056] str(data)

[0057] X<-scale(data[,-1])

[0058] summary(X)

[0059] #SOM

[0060] set.seed(2022)

[0061] g<-somgrid(xdim=5,ydim=1,topo="rectangular")

[0062] map<- som(X,grid=g,radius=1)

[0063] #mapresults

[0064] plot(map, type="counts")

[0065] plot(map,type="codes")

[0066] #outputresults

[0067] map$codes

[0068] map$unit.classif

[0069] output<-cbind(data,map$unit.classif)

[0070] write.table(output,file="Results_GRID2010.csv",sep=",",row.names=FALSE)

[0071] Step 4: After Z-score standardization based on the comprehensive attribute table output in Step 3, input it into the self-organizing neural network model for cluster analysis, obtain the cluster labels after clustering, associate them with the IDs of the grid cells, and input them into ArcGIS software;

[0072] Ecological functional zoning based on the SOM model: To achieve objective and scientific land spatial zoning based on multi-dimensional ecosystem service supply and demand characteristics, this step introduces the Self-Organizing Neural Network (SOM) model for unsupervised clustering analysis. The core advantage of this method is that it can automatically identify the inherent structure and similarity patterns in complex data composed of six ecosystem service supply-demand ratios (ESDRs) without pre-setting any prior knowledge, and group spatial units (fish grids) with similar supply and demand characteristics into one category, thereby forming a data-driven ecological functional zoning that can truly reflect the functional characteristics of the regional "socio-ecological" system.

[0073] Specifically, this step first exports the comprehensive attribute table generated in the previous step to CSV format and processes it in the RStudio data analysis environment. To eliminate model bias that may be caused by differences in units and numerical ranges among different ecosystem services (ESs), the six columns of ESDR data used as the basis for clustering are first Z-score standardized, so that each ESDR variable is transformed into a standard normal distribution with a mean of 0 and a standard deviation of 1. Next, the powerful kohonen package in R is used to build and train the SOM model. During model training, a two-dimensional neural network iteratively learns and gradually adjusts its weight vector to optimally map the topology of the input high-dimensional ESDR data.

[0074] To gain a deeper understanding of the specific ecological function characteristics of each partition, the system generates a fan-shaped ES supply and demand clustering map. The map displays the composition and relative size of ESS supply and demand within each ES cluster. Longer segments represent higher ESDR values, while shorter segments represent lower ESDR values. Based on the "weakest link" principle, this invention identifies the shorter segment type in the ESDRB as the dominant ESS supply and demand type for that partition, such as "high pressure on food production - ecological function deficit area" and "core conservation of habitat quality - comprehensive service surplus area." Finally, the clustering results carrying partition labels are linked to GIS spatial data using the unique ID of the fishing net unit, and then symbolized and colored in software such as ArcGIS to generate a clear and intuitive spatial distribution map of ecosystem service supply and demand clusters. The main ES supply and demand types for each region are shown in Table 1.

[0075] Table 1. Main ES Supply and Demand Types in Each Region

[0076] Step 5: Input the Z-score-standardized comprehensive attribute table from Step 4 into the established XGBoost regression model to filter out the top-ranked data. Key driving factors;

[0077] Identifying Key Driving Factors Based on the XGBoost Model: To accurately identify and quantify the key driving factors affecting the supply and demand balance (ESDR) of various ecosystem services, this step employs the Extreme Gradient Boosting (XGBoost) regression model, renowned for its high accuracy and efficiency in machine learning. The advantage of this method lies in its ability to effectively capture and handle the complex nonlinear relationships and interactions between driving factors. Compared to traditional linear models, it can reveal the true driving mechanisms within the "socio-ecological" system more profoundly and accurately, providing a robust and reliable scientific basis for the formulation of subsequent management strategies.

[0078] Specifically, this step first uses the comprehensive attribute table generated in step three, containing 14 driving factors and 6 ESDR values, as the input dataset. For each ecosystem service (e.g., food production ESDR), an independent prediction model will be built. In the RStudio environment, using packages such as caret, the dataset is divided into training and test sets in a 7:3 ratio through stratified random sampling to ensure consistency in the distribution of the dependent variable between the two subsets.

[0079] Next, an XGBoost regression model is constructed, and the key hyperparameters are systematically optimized using a rigorous grid search combined with 7-fold cross-validation. The optimized parameters mainly include the number of trees (nrounds), the maximum tree depth (max_depth), the learning rate (eta), and the regularization parameter. Specifically, the parameter optimization space is set as follows: the search set for the number of iterations (nrounds) is set to {500, 1000, 2000} to ensure sufficient model convergence; the maximum tree depth (max_depth) is set to {3, 5, 6} to achieve a balance between model complexity and fitting ability; the learning rate (eta) is set to {0.01, 0.05, 0.1} to finely control the weight update step size; and a regularization parameter (gamma) is set to {0, 0.1} and a minimum leaf node weight (min_child_weight) is set to {1, 3} to effectively suppress overfitting. The goal of optimization is to find a set of parameters that minimizes the model's prediction error (e.g., root mean square error RMSE) on the validation set. Then, using this optimal set of parameters, the training sample set is input into the XGBoost model for supervised learning training, thereby establishing a complex predictive relationship between 14 driving factors and a specific ESDR.

[0080] After training, the model's performance must undergo rigorous evaluation. This invention validates the trained model using reserved test set data that the model has never seen before, and evaluates the model's predictive accuracy by calculating key performance indicators such as the coefficient of determination (R²) and root mean square error (RMSE). Only when the model's accuracy reaches a preset standard (e.g., R² > 0.65) is the model considered reliable. Finally, based on the validated model, XGBoost's built-in feature importance calculation method (centered on the Gain metric, which measures the average gain of each feature to all decision tree splits in the model) is applied to quantify the contribution of each driving factor to the target ESDR change. The driving factors are sorted in descending order of contribution scores, and the highest-ranking and most stable driving factors (e.g., consistently ranking in the top five in multiple repeated experiments) are selected as key driving factors affecting the supply and demand balance of this ecosystem service. A visual importance ranking bubble chart is output, providing precise variable input for the next step of threshold analysis. The importance ranking of driving factors is shown in the table below:

[0081] Table 2 Ranking of Driving Factors by Importance

[0082]

[0083] The R code is as follows:

[0084] library(xgboost)

[0085] library(caret)

[0086] library(pdp)

[0087] library(cowplot)

[0088] #Read dataset

[0089] pre<-read.csv("F: / driving / driving / driving1.csv")

[0090] #Split the dataset using cross-validation

[0091] set.seed(42)

[0092] train_control<-trainControl(method="cv",number=5) #5-fold cross-validation

[0093] `trains <- createDataPartition(y=pre$D,p=0.7,list=FALSE)` # Replaces the dependent variable "A" with "D".

[0094] data_train <- pre[trains, ]

[0095] data_test <- pre[-trains, ]

[0096] # Data preparation

[0097] data_trainx <- model.matrix(~cultiva + forest + grass + water + const + unused + river + rail + road + slope + dem + ndvi + pop + nt + gdp + sol + tmp + pre, data = data_train)

[0098] data_trainy <- data_train$D

[0099] data_testx <- model.matrix(~cultiva + forest + grass + water + const + unused + river + rail + road + slope + dem + ndvi + pop + nt + gdp + sol + tmp + pre, data = data_test)

[0100] data_testy <- data_test$D

[0101] # Create DMatrix objects using the xgb.DMatrix function

[0102] dtrain <- xgb.DMatrix(data = data_trainx, label = data_trainy)

[0103] dtest <- xgb.DMatrix(data = data_testx, label = data_testy)

[0104] # Train the model

[0105] fit_xgb_reg <- train([

[0106] x = data_trainx,

[0107] y = data_trainy,

[0108] method = "xgbTree",

[0109] trControl=train_control,

[0110] verbose=FALSE,

[0111] tuneGrid = expand.grid(

[0112] nrounds=1000,

[0113] eta=0.3,

[0114] max_depth=2,

[0115] gamma=0.001,

[0116] min_child_weight=1, # Add the min_child_weight parameter

[0117] subsample=0.7,

[0118] colsample_bytree=0.4 ) )

[0121] #Output Model

[0122] print(fit_xgb_reg)

[0123] #Calculate and visualize the importance of features in XGBoost

[0124] importance_matrix<-xgb.importance(model=fit_xgb_reg$finalModel)

[0125] print(importance_matrix)

[0126] xgb.plot.importance(importance_matrix=importance_matrix,measure="Cover")

[0127] #Prediction Results

[0128] predictions<-predict(fit_xgb_reg,data_testx)

[0129] #Calculate MAE

[0130] mae_value<-mean(abs(predictions-data_testy))

[0131] cat("MeanAbsoluteError(MAE):",mae_value,"\n")

[0132] #Calculate MSE

[0133] mse_value<-mean((predictions-data_testy)^2)

[0134] cat("MeanSquaredError(MSE):",mse_value,"\n")

[0135] #Calculate RMSE

[0136] rmse_value<-sqrt(mean((predictions-data_testy)^2))

[0137] cat("RootMeanSquaredError(RMSE):",rmse_value,"\n")

[0138] #Calculate R^2

[0139] r_squared_value<-1-mean((data_testy-predictions)^2) / var(data_testy)

[0140] cat("VarianceExplained(R^2):",r_squared_value,"\n")

[0141] Step Six: Use the key driving factors output in Step Five as independent variables, and their corresponding... The value is used as the dependent variable, and the beneficial influence threshold is obtained by using a constraint line fitting method.

[0142] This step innovatively employs a constraint line model to reveal the impact thresholds of key driving factors on the ecosystem service supply-demand balance (ESDR) and transform the analysis results of the preceding model from "relationship descriptions" into "quantitative management indicators." The innovative perspective of this method lies in its departure from focusing on the "average trend" response relationship between driving factors and ecosystem services. Instead, it explores the "maximum potential" or "maximum constraint" that ecosystem services can achieve under specific driving factor conditions by fitting the "upper boundary" of the data scatter distribution. This allows for the precise extraction of the favorable numerical range or critical inflection point required for key driving factors to maintain ecosystem services in a healthy and sustainable state (i.e., supply-demand balance or surplus).

[0143] Specifically, this step conducts an independent threshold analysis for each key driver identified in step five for each ecosystem service. First, a two-dimensional scatter plot is constructed using all raster values ​​of a key driver (e.g., "slope") as the independent variable (X) and its corresponding ecosystem service supply-demand ratio (e.g., "soil retention ESDR") as the dependent variable (Y). Next, algorithms such as binning-maximum or upper quantile regression are used to fit the constraint curve. For example, the range of values ​​for the independent variable X can be divided into several continuous intervals. Within each interval, the maximum value (or the 95th or 99th percentile) of the dependent variable Y is selected. Then, nonlinear curve fitting (such as piecewise regression or polynomial fitting) is performed on these selected boundary points to obtain a constraint curve equation that represents the upper boundary of the data distribution.

[0144] After obtaining a constraint line equation with good fit (evaluable by metrics such as R² or cross-validation), the threshold can be precisely extracted. By analyzing this fitted constraint curve, the system automatically calculates the value of the independent variable (key driving factor) when the dependent variable ESDR equals 0 (i.e., the critical point where supply and demand shift from a "deficit" to a "balance"). This value is defined as a key management inflection point. Furthermore, the system identifies the range of independent variable values ​​corresponding to all portions of the constraint curve where ESDR ≥ 0; this range is defined as the threshold range of beneficial influences that promote the healthy maintenance of the ecosystem service. Figure 5 As shown in the figure, this embodiment identifies key threshold characteristics of supply and demand balance for different ecosystem services using the above method: for food supply and demand balance ( Figure 5 a) identified a limiting threshold of 31.27 people / km² for population density, indicating that food services in the region can maintain a surplus if and only if population density is kept below this value; for carbon sequestration services ( Figure 5 b) and water conservation services Figure 5 c) identified a nighttime light index of 18.74 and a population density of 95.47 people / km² as critical thresholds. Management must control the intensity of disturbances within these thresholds to avoid a service deficit. Regarding soil conservation services ( Figure 5 d) The critical threshold for the Normalized Difference Vegetation Index (NDVI) was identified as 0.20, and the constraint line showed a positive driving relationship, meaning that the NDVI needs to be maintained above 0.20 to ensure that the soil maintains a supply-demand balance; regarding habitat quality ( Figure 5 e), the distance threshold from the building site is 1.00 km, meaning that ecological space must maintain a buffer distance of at least 1 km from the building site to ensure a habitat quality surplus; while for recreational services ( Figure 5 f), with a threshold of 55.05 km from built-up areas, reveals the effective spatial radius of recreational service provision. Ultimately, this step outputs a series of clear, actionable, and quantitative management guidelines (e.g., "To ensure no deficit in soil conservation services, the threshold of the key driver factor 'NDVI' must be controlled below 0.2"), providing direct and robust data support and decision-making basis for the final "one policy per area" precise management plan. The R language code is as follows:

[0145] library(ggplot2)

[0146] library(ggpmisc)

[0147] library(gridExtra)

[0148] library(grid) # Used for layout

[0149] #Set working directory

[0150] setwd("F: / DZ_driving_factors / driving")

[0151] #Read in data

[0152] data1<-read.csv("F: / DZ_driving_factors / driving / driving2020county.csv")

[0153] head(data1)

[0154] # Sort by column PC

[0155] f2<-data1[order(data1$PC),]

[0156] # Define grouping parameters

[0157] n<-6

[0158] lv<-rep(1:ceiling(nrow(f2) / n),each=n,length.out=nrow(f2))

[0159] n100<-split(f2$F,lv)

[0160] n100_2<-split(f2,lv)

[0161] #Calculate the maximum value and boundary values ​​for each data set.

[0162] fh<-sapply(n100_2,function(x){

[0163] y<-quantile(f2$F,0.960)

[0164] hghgh<-subset(x,F<=y)

[0165] fh2<-hghgh[which.max(hghgh$F),]

[0166] ESC2<-fh2$PC

[0167] ESN2<-fh2$F

[0168] return(c(ESC2,ESN2))

[0169] })

[0170] # Convert to a data frame

[0171] kk<-as.data.frame(t(as.data.frame(fh)))

[0172] names(kk) <- c("ESC2","ESN2")

[0173] # Define style

[0174] theme2<-theme(

[0175] axis.text=element_text(size=12,family="sans",colour="black"),

[0176] axis.text.x=element_text(angle=0,hjust=0.5),

[0177] axis.ticks=element_line(size=1.2),

[0178] axis.line=element_line(colour="black",size=1.2),

[0179] axis.ticks.length=unit(0.2,"cm"),

[0180] panel.grid.major=element_line(colour=NA),

[0181] panel.border=element_blank(),

[0182] panel.background=element_blank(),

[0183] plot.background=element_blank(),

[0184] panel.grid.minor=element_blank(),

[0185] axis.title=element_text(size=12,family="sans",color="black") )

[0187] #Define Formula

[0188] my.formula<-ESN2~poly(ESC2,2,raw=TRUE)

[0189] # Fit the model and calculate the x value when y=0

[0190] fit<-lm(my.formula,data=kk)

[0191] coeff <- coef(fit)

[0192] #Calculate the value of x when y=0

[0193] a<-coeff[3]#coefficient of the quadratic term

[0194] b<-coeff[2]#coefficient of the linear term

[0195] c<-coeff[1]#constant term

[0196] #Use the quadratic formula to find the value of x.

[0197] # Calculate the discriminant <- b^2 - 4 * a * c # Calculate the discriminant

[0198] #Calculate the value of x when y=0

[0199] x_values<-NULL

[0200] if(discriminant>=0){

[0201] x_values<-c((-b+sqrt(discriminant)) / (2*a),(-b-sqrt(discriminant)) / (2*a))

[0202] }

[0203] # Only keep x values ​​greater than 0

[0204] x_values<-x_values[x_values>0]

[0205] # Get a summary of the fitting results

[0206] summary_fit <- summary(fit)

[0207] #Creating a graphic

[0208] a<-ggplot(data=kk,aes(x=ESC2,y=ESN2))+

[0209] geom_point(data=f2,aes(x=PC,y=F),color="#00AFBB",size=3,alpha=1)+geom_point(color="#FC4E07",size=3,alpha=1)+

[0210] geom_smooth(method="lm",se=FALSE,fullrange=TRUE,color="#FF1C1C",

[0211] formula=y~poly(x,2,raw=TRUE),linewidth=1)+# Replace size with linewidth

[0212] geom_hline(yintercept=0,linetype="dashed",color="#000000",linewidth=1)+#Adds an auxiliary line with y=0.

[0213] {if(length(x_values)>0)geom_vline(xintercept=x_values[1],linetype="dashed",color="#FF1C1C",linewidth=1)}+#Add auxiliary line for the first x value

[0214] {if(length(x_values)>1)geom_vline(xintercept=x_values[2],linetype="dashed",color="#FF1C1C",linewidth=1)}+#Add auxiliary line for the second x value

[0215] {if(length(x_values)>0)annotate("text",x=x_values[1],y=max(kk$ESN2),label=paste(round(x_values[1],2)),

[0216] color="#FF1C1C", size=5, hjust=-0.1, vjust=1.5, family="sans")}+# marks the first x value.

[0217] {if(length(x_values)>1)annotate("text",x=x_values[2],y=max(kk$ESN2)*0.9,label=paste(round(x_values[2],2)),

[0218] color="#FF1C1C", size=5, hjust=-0.1, vjust=1.5, family="sans")}+# marks the second x value.

[0219] stat_poly_eq(method="lm",formula=y~poly(x,2,raw=TRUE),

[0220] aes(label=paste(after_stat(rr.label),"P<0.01",sep="~~~~")),

[0221] parse=TRUE,size=4,family="sans",color="black")+

[0222] theme2+

[0223] ylab(expression("ESDRFP(2020)"))+

[0224] xlab("PC(2020)")

[0225] #Print graphics and model summaries

[0226] print(a)

[0227] print(summary_fit);

[0228] Step 7: Obtain the measured values ​​of key driving factors and the "beneficial influence threshold" in the current grid cell, calculate the deviation, and perform logical judgment based on the deviation and the preset logical rule base;

[0229] First, a "Zoned Management Measures Rule Database" is constructed. A rule database is pre-established containing mappings between "Zoned Type," "Driving Factor Type," "Status Determination," and "Management Strategy Text." This database covers rules for determining the status of driving factors: if the driving factor is positive (e.g., NDVI) and the measured value is less than the threshold, the system calls the "Supplement / Repair" template, outputting a strategy of "a factor in this area is insufficient, requiring ecological restoration to raise it above the threshold"; if the driving factor is negative (e.g., population density) and the measured value is greater than the threshold, the system calls the "Reduction / Control" template, outputting a strategy of "a factor in this area is overloaded, requiring reduction to below the threshold." Simultaneously, the database includes keynote rules for zoning characteristics, such as "Strict protection, prohibition of destructive development" for the B1 ecological enhancement zone and "Optimization of existing resources, increase of green infrastructure" for the B5 urban core zone.

[0230] Secondly, the system performs automatic diagnosis, threshold deviation calculation, and logical judgment of spatial units. For each grid unit or administrative unit within the study area, the system automatically extracts the zoning and dominant ecological defects determined in step four, the key driving factors determined in step five, and the current measured factor values ​​(V) for that unit. real ) and the factor control threshold (V) calculated by the constraint line model in step six. threshold The system automatically calculates the deviation ∆=V. real- V threshold And perform logical judgment: for negative factors (such as population density), if V real >V thresholdIf the condition is determined to be "overloaded," the system automatically retrieves the "reduction / control" measure template from the database and embeds the specific threshold value into the text; for positive factors (such as NDVI), if V real Less than V threshold If the condition is determined to be in a "deficient state", the "supplement / repair" action template will be automatically invoked.

[0231] Finally, the system automatically synthesizes and outputs structured policy texts. It synthesizes the above judgment results into text and outputs a complete report for each partition, containing "problem diagnosis + quantitative targets + action instructions." The output format is "[Partition Name] + [Main Ecological Problems] + [Key Issues] + [Specific Control Instructions (including thresholds)]." Based on the above technical solution, specific output examples for five typical partitions within the study area are shown in Table 3 below:

[0232] Table 3 Spatial Control Targets

[0233]

[0234] (1) For the B1-ecological enhancement zone: The system identifies that this area is mainly distributed in forest land and key water conservation areas. Its dominant ecosystem service supply and demand characteristics are manifested at different scales as potential demand for habitat quality (HQ) and recreation (LR), or demand for soil conservation (SR). Based on this diagnosis, the system automatically generates control measures: For the grid scale, it outputs "Strictly restrict the expansion of construction land, ensure that the ecological core area is far away from the interference of construction land, and require that the distance between ecological sensitive points and construction land be kept above the threshold of 1km in order to maintain the balance between habitat quality supply and demand"; at the same time, combined with natural resources, it outputs "Develop ecotourism and under-forest economy in an orderly manner within the environmental carrying capacity, and establish an ecological agriculture demonstration zone".

[0235] (2) For the B2-Comprehensive Ecological Conservation Zone: The system identifies the characteristics of this area as follows: all ecosystem services except recreation (LR) are in a state of supply and demand balance, and it has high-quality vegetation cover. Based on the "protection first" rule, the system automatically generates control measures: output "delineate strict ecological protection red lines, take the maintenance of the integrity of the ecosystem structure as the primary goal, and prohibit any form of destructive development"; for vegetation management, output "establish ecological buffer zones and green isolation belts to ensure that the vegetation cover in the area is stable above the threshold of 0.2 for a long time to prevent soil erosion and maintain water conservation function"; for the shortcomings of recreation, output "under the premise of not destroying authenticity, build ecological trails and science popularization facilities appropriately based on existing landscape resources".

[0236] (3) For B3-multifunctional zones: The system identifies that these zones are typically distributed along main roads, combining ecological functions with human transportation needs, with soil conservation (SR) as the dominant ecosystem service. In response to the dual pressures of road construction and ecological protection, the system automatically generates control measures: it outputs the implementation of a green infrastructure construction strategy, focusing on the construction of protective forest belts and ecological corridors along roads and around towns; based on the threshold of key driving factors, it outputs strict monitoring of the impact of road construction on surrounding vegetation, adopting a combination of engineering and biological measures to ensure that the NDVI value of key areas is not lower than 0.20, so as to enhance the soil and water conservation capacity of the road system; for urban internal spaces, it outputs the promotion of vertical greening, rooftop greening and pocket park construction, and optimization of the proportion of impermeable surfaces.

[0237] (4) For the B4-ecological restoration zone: The system identifies that this area is located in the urban-rural fringe and is subjected to the dual pressure of urban expansion and agricultural development. The main ecological defects are concentrated in soil conservation (SR) and habitat quality (HQ). Based on the rule of "ecological restoration and urban-rural integration", the system automatically generates control measures: output "establish an urban-rural ecological compensation mechanism and support agricultural non-point source pollution control and damaged habitat restoration through fiscal transfer payments"; for key driving factors, output "strictly control the intensity of agricultural development, implement the conversion of sloping farmland to forest and grassland, and make mandatory requirements to keep the distance from building land above the threshold of 1km in order to reverse the soil conservation deficit"; at the same time, output "use artificial wetlands or natural wetland systems to improve water and soil quality and construct a ring-city ecological recreation belt".

[0238] (5) For B5-Urban Core Area: The system identifies this area as a population and economic activity cluster, with severe deficits in food production (FP), carbon sequestration (CS), and water production (WY), and the measured values ​​of key driving factors (such as population density) are generally higher than the threshold. Based on the "reduction and quality improvement" rule, the system automatically generates control measures: output "The main problem in this area is the ecological deficit caused by high-intensity development, and the key constraint is population density. Management measures: strictly control the urban development boundary and construction intensity, implement population relocation and industrial upgrading strategies, and the mandatory control target is to control the population density within the ecological carrying capacity threshold of 31.27 people / km² to alleviate the pressure on food production and other services"; for the carbon sequestration deficit, output "Implement high-standard urban greening projects, increase vertical greening to make up for the carbon sink gap, and promote clean energy and low-carbon lifestyles".

[0239] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for precise land space management based on key driving factor thresholds, characterized in that: Includes the following steps: Step 1: Acquire and preprocess geographic data; Step 2: Based on the preprocessed geographic data from Step 1, construct the supply and demand of ecosystem services, and calculate the corresponding supply-demand ratios. ; Step 3: Generate a standardized vector grid based on ArcGIS software, perform statistical analysis and average calculations for each grid, assign the average to the corresponding grid cells, and output a comprehensive attribute table; Step 4: After Z-score standardization based on the comprehensive attribute table output in Step 3, input it into the self-organizing neural network model for cluster analysis, obtain the cluster labels after clustering, associate them with the IDs of the grid cells, and input them into ArcGIS software; Step 5: Input the Z-score-standardized comprehensive attribute table from Step 4 into the established XGBoost regression model to filter out the top-ranked data. Key driving factors; Step Six: Use the key driving factors output in Step Five as independent variables, and their corresponding... The value is used as the dependent variable, and the beneficial influence threshold is obtained by using a constraint line fitting method. Step 7: Obtain the measured values ​​of key driving factors and factor control thresholds in the current grid cell, calculate the deviation, and perform logical judgment based on the deviation and the preset logical rule base.

2. The method for precise management of national land space based on key driving factor thresholds according to claim 1, characterized in that: The supply-demand ratio in step two The specific expression is as follows: in, Indicates supply, Indicate demand, Indicates maximum supply. This indicates the maximum demand.

3. The method for precise management of national land space based on key driving factor thresholds according to claim 2, characterized in that: The supply and demand of ecosystem services in step two include the supply and demand of food production, carbon sequestration, soil conservation, water production services, habitat quality, and recreational services. The supply of the grain production is allocated to spatial plots based on the total output in the statistical yearbook according to the vegetation index NDVI, and the demand for the grain production is obtained by multiplying the regional population by the per capita energy intake standard. The supply of carbon sequestration is directly calculated from the net primary productivity of vegetation (NPP), and the demand for carbon sequestration is spatially calculated by correlating energy consumption with nighttime light data. The supply of soil conservation is assessed using the RUSLE soil erosion model, and the demand for soil conservation is assessed. The supply of the water production service is obtained by simulating rainfall, evapotranspiration and land use patterns using the InVEST model; the demand for the water production service is also obtained. The supply of habitat quality, and the demand for habitat quality; The supply of the leisure and recreation services, and the demand for the leisure and recreation services.

4. The method for precise management of national land space based on key driving factor thresholds according to claim 3, characterized in that: In step three, the self-organizing neural network model is constructed and trained using the kohonen package in the R language.

5. A method for precise management of national land space based on key driving factor thresholds according to claim 4, characterized in that: The parameters optimized during the training process of the XGBoost regression model in step five include the number of trees, the maximum depth of the trees, the learning rate, and the regularization parameter.