Geographic big data and residual analysis-based underground water critical burial depth calculation method
By introducing geographical big data and residual analysis methods into the critical buried depth calculation of groundwater, combined with random sampling and multi-model modeling, the problem of nonlinear influence of multiple environmental factors and spatial autocorrelation of vegetation distribution is solved, and a more efficient and accurate groundwater critical buried depth calculation is achieved.
Patent Information
- Application Number
- CN202510198101.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively deal with the nonlinear influence of multiple environmental factors and the spatial autocorrelation of vegetation distribution when calculating the critical buried depth of groundwater, resulting in insufficient reliability and accuracy of the calculation results.
The calculation method based on geographic big data and residual analysis is adopted, and the processing ability of nonlinear impact indicators and multi-model modeling strategies is enhanced by introducing new groundwater burial depth impact indicators and multi-model modeling strategies, combining random sampling and spatial statistical methods.
It significantly improves the accuracy and reliability of the calculation of critical buried depth of groundwater, and can more accurately quantify the critical buried depth of vegetation ecological response to groundwater level changes, supporting accurate solutions for ecological hydrological research.
Smart Images

Figure CN120068640A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of eco-hydrogeology and geospatial big data analysis, and specifically to a method for calculating the critical groundwater depth based on geospatial big data and residual analysis. Background Art
[0002] In the arid and semi-arid regions of northwestern China, groundwater resources are crucial for maintaining the ecosystem. However, overexploitation of groundwater may lead to a decline in the groundwater level, thereby triggering vegetation degradation and ecosystem instability. In order to achieve the sustainable utilization of groundwater resources during regional development while maintaining the health of the ecosystem, it is urgent to determine the critical groundwater depth, that is, keeping the groundwater level above the groundwater level corresponding to the critical depth can maintain the normal function of the vegetation ecosystem and prevent obvious degradation of the vegetation. The complexity of this problem lies in that, on the one hand, it is necessary to quantify the independent impact of groundwater depth on vegetation, excluding the interference of environmental factors such as climate, terrain, soil, and the spatial autocorrelation of vegetation on vegetation distribution; on the other hand, it is necessary to construct an efficient analysis framework, based on multi-source geospatial big data, to reveal the precise relationship between groundwater depth and vegetation, and finally quantify the critical depth of the vegetation ecological response to groundwater level changes.
[0003] The sum of the root length of groundwater-dependent vegetation and the capillary rise height supported by the water table can be used as a simple method for estimating the critical groundwater depth, but this method is more difficult to apply at the basin scale. The simple mean difference method estimates the causal effect of the treatment variable on the outcome variable by comparing the mean differences between the treatment group and the control group. Due to its simple and easy-to-implement characteristics, it is widely used to calculate the critical depth threshold of vegetation dependence on groundwater at the regional scale. The groundwater depth variable values of the samples are divided into multiple treatment groups and a control group according to a fixed grouping interval. The groundwater depth condition of the control group is usually set to an empirical value (such as 15 m) exceeding the maximum root length of the vegetation. By analyzing the change of the mean difference of vegetation cover (treatment group minus control group) with groundwater depth, the groundwater depth when the difference approaches 0 is used as the ecological groundwater level threshold.
[0004] However, the simple mean difference method has some deficiencies in dealing with the interference of confounding factors. First, the simple mean difference method cannot control other confounding variables that may affect vegetation cover (such as temperature, precipitation, elevation, or adjacent vegetation). Second, this method assumes that the treatment group and the control group are similar before treatment allocation (randomized grouping), but this assumption is usually difficult to meet in observational data, resulting in the mean difference being unable to accurately reflect the true impact of groundwater depth. Finally, this method uses a fixed groundwater depth threshold to divide the control group, which is only applicable to the case where all samples depend on groundwater, while in reality, only some samples depend on groundwater. Summary of the Invention
[0005] The present application provides a method for calculating the critical groundwater depth based on geographical big data and residual analysis. By introducing new influencing indicators of groundwater depth and a multi-model modeling strategy based on random sampling, the processing ability for the non-linear influence of multi-environmental factors and the spatial autocorrelation of vegetation distribution is enhanced, and the reliability and accuracy of the calculation results are improved, effectively solving the problems in the background art.
[0006] To achieve the above object, the present application provides the following technical solutions: The critical groundwater depth based on geographical big data and residual analysis includes the following steps:
[0007] Step 1: Calculation of groundwater depth data:
[0008] First, perform Kriging interpolation on the groundwater level monitoring data.
[0009] According to the aquifer type in the study area and the data density of the groundwater level observation data, screen the interpolation results; if the data density in a certain area is lower than the preset threshold, set the groundwater level in that area as a null value.
[0010] Use the raster calculation tool to calculate the difference between the surface elevation raster and the groundwater level raster to generate groundwater depth raster data.
[0011] Use the conditional screening tool to set the pixel values less than 0 in the groundwater depth raster to 0 to eliminate the influence of outliers and ensure data rationality.
[0012] Step 2: Calculation of neighboring covariate data of environmental factors:
[0013] Calculate the neighboring covariates of environmental factors by the inverse distance weighting method.
[0014] The neighboring covariate is the inverse distance weighted mean of environmental variables within a given radius r around the sample, and the calculation formula is:
[0015] ;
[0016] where, is the environmental variable value of the i-th raster pixel, is the distance between the current pixel and the -th neighboring pixel, is the distance attenuation coefficient, usually set to 2, is the number of non-missing pixels within the radius neighborhood around the pixel.
[0017] Step 3: Spatial matching of raster data:
[0018] Adjust raster data from different sources to the same resolution and coordinate system to ensure the accuracy of subsequent analysis.
[0019] The specific steps are as follows:
[0020] Projection transformation: Convert all raster data to the same coordinate system.
[0021] Resampling operation: Resample all raster data to the same resolution.
[0022] Filling missing areas: For the missing areas of other environmental factor rasters, use spatial interpolation methods to fill them. For the missing areas of the groundwater depth raster, use a fixed value nv to fill them.
[0023] Overlay raster data: Overlay all variables one by one into the same raster data, and extract all variable values corresponding to the center points of each raster cell.
[0024] Generate the basic dataset D: Generate the basic dataset D based on the complete variable information of each raster cell.
[0025] Step 4. Random sampling to construct the background dataset:
[0026] To construct the background dataset, samples need to be randomly drawn from the basic dataset D. The specific steps are as follows:
[0027] Set the random seed seed: Ensure the repeatability of each sampling.
[0028] Random sampling: Randomly draw N samples without replacement from the basic dataset D to generate the background data value D bg .
[0029] Determine the sample size N: By analyzing the impact of the sample quantity on the performance of the background model, refer to the sample size threshold at the performance convergence point to determine the value of the parameter N.
[0030] Step 5. Establish a vegetation coverage prediction model:
[0031] Select a machine learning algorithm, using the background dataset D bg as the input to establish a model M for predicting vegetation coverage.
[0032] Step 6. Calculate the impact amount of groundwater depth:
[0033] Apply the model M trained in Step 5 to predict the samples in the basic data D where the groundwater depth variable is not equal to the missing value nv, and calculate the residual between the model prediction result and the actual observation value , which is used to reflect the impact of groundwater depth on vegetation coverage. The calculation formula is as follows:
[0034] ;
[0035] Among them, is the true vegetation coverage, is the predicted vegetation coverage.
[0036] Step 7: Repeat multiple times until convergence, and calculate the mean residual:
[0037] Change the random seed seed, and repeat Steps 4 to 6 multiple times until the mean model residual at each pixel location tends to be stable and meets the convergence criterion.
[0038] Step 8: Fit the empirical model and calculate the critical groundwater depth:
[0039] Fit the empirical model.
[0040] Data grouping: Group the samples of the basic dataset with a groundwater depth interval of 1 m;
[0041] Calculate the standard deviation and mean: For each group, calculate the standard deviation within the group and the corresponding mean groundwater depth.
[0042] Exponential function model fitting: Use the mean groundwater depth as the independent variable x, and the standard deviation as the dependent variable y, and fit the relationship between the two with an exponential function model. The model expression is as follows:
[0043] ;
[0044] where, , , are model parameters estimated by the nonlinear least squares method.
[0045] The parameter determines the rate of decline of the exponential function. At a given confidence level of 1-α, the confidence interval of the parameter , is calculated by the following formula:
[0046] ;
[0047] ;
[0048] where, is the critical value of the t-distribution, which depends on the confidence level α and the groundwater depth value that defines the decline amplitude to reach the total amplitude The groundwater depth value is the ecological critical groundwater depth , and the calculation formula is:
[0049] ;
[0050] The confidence interval of the critical burial depth can be indirectly obtained by substituting the upper and lower limits of the parameter b. , Indirectly obtained.
[0051] Preferably, the data density screening conditions for the high-confidence interpolation regions corresponding to different water-bearing media in step 1 are as follows: there are no less than 3 observation points per square kilometer in the pore water aquifer; there are no less than 2 observation points per square kilometer in the karst water; and there are no less than 1 observation point per square kilometer in the fissure water.
[0052] Preferably, the environmental factors in step 2 include factors related to climate, terrain, soil, water bodies, and human activities.
[0053] Preferably, the specific steps for calculating the neighboring covariates of environmental factors by the inverse distance weighting method are as follows: determine all non-missing pixels around the current grid cell and their distances; for each non-missing pixel, calculate the power of the value of its environmental variable divided by the distance; sum all the results and divide by the sum of the powers of the distances to obtain the value of the neighboring covariate.
[0054] Preferably, in step 5, any one of the random forest and extreme gradient boosting tree is selected as the machine learning algorithm.
[0055] Preferably, the convergence criterion in step 7 is that the relative change rate of the residual mean of all 90% of the pixels is lower than 5%.
[0056] Compared with the prior art, the beneficial effects of the present application are as follows:
[0057] 1. The innovation of this solution lies in the organic combination of random sampling, multi-model calculation, and spatial statistical methods, and their application in the quantitative analysis of the relationship between groundwater depth and vegetation, forming a set of efficient and accurate calculation methods for groundwater critical burial depth;
[0058] 2. Multiple random samplings and background model construction: Multiple background data sets are generated through random sampling, combined with spatial neighboring variables, to reduce the influence of spatial autocorrelation and accurately quantify the baseline influence of environmental factors on vegetation coverage;
[0059] 3. Multi-model calculation of residual mean: Based on the machine learning model algorithm, the residual mean is calculated iteratively multiple times and used as the influence amount of groundwater depth, significantly improving the stability and accuracy of the calculation;
[0060] 4. Groundwater critical burial depth calculation framework: Innovatively combine the high-confidence groundwater level interpolation region with the residual mean to construct a new calculation method for groundwater critical burial depth, providing an accurate solution for eco-hydrological research. Description of the Drawings
[0061] Figure 1 This is the overall flowchart of the method;
[0062] Figure 2 This is a schematic diagram for calculating the critical groundwater depth;
[0063] Figure 3 This is a graph showing the relationship between the search radius (r) in the calculation of neighboring covariates and the model performance;
[0064] Figure 4 This is a graph showing the relationship between the sample size (N) of the background dataset and the model performance;
[0065] Figure 5 This is a schematic diagram for calculating the critical groundwater depth and the confidence range; Specific implementation manners
[0066] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0067] Please refer to Figures 1-5 , the present application provides the following technical solutions: A method for calculating the critical groundwater depth based on geospatial big data and residual analysis, including:
[0068] Step 1: Calculation of groundwater depth data:
[0069] First, perform Kriging interpolation on the groundwater level monitoring data. This is a geostatistical method used to estimate values at unknown locations in space. By using the observed data at known points to predict the data at unknown points, the non-rasterized data is converted into groundwater level raster data with a specified spatial resolution.
[0070] According to the aquifer type in the study area and the data density of the groundwater level observation data, screen the interpolation results; if the data density in a certain area is lower than the preset threshold, set the groundwater level in that area to a null value.
[0071] Use the raster calculation tool to calculate the difference between the surface elevation raster and the groundwater level raster to generate groundwater depth raster data.
[0072] This can help us understand the distribution of groundwater underground and the differences in groundwater burial depths in different regions.
[0073] Use the conditional screening tool to set the pixel values less than 0 in the groundwater depth raster to 0 to eliminate the influence of outliers and ensure data rationality.
[0074] This can avoid inaccurate analysis results or misleading conclusions caused by outliers.
[0075] Step 2: Calculation of neighboring covariate data for environmental factors:
[0076] Calculate the neighboring covariates of environmental factors by the inverse distance weighting method.
[0077] The neighboring covariate is the inverse distance weighted mean of environmental variables within a given radius r around the sample, and the calculation formula is:
[0078] ;
[0079] where is the environmental variable value of the i-th grid cell, is the distance between the current cell and the -th neighboring cell, is the distance decay coefficient, usually set to 2, is the number of non-missing cells within the radius neighborhood around the cell.
[0080] Step 3: Spatial matching of raster data:
[0081] Adjust raster data from different sources to the same resolution and coordinate system to ensure the accuracy of subsequent analysis.
[0082] The specific steps are as follows:
[0083] Projection transformation: Convert all raster data to the same coordinate system.
[0084] Resampling operation: Resample all raster data to the same resolution. This can be achieved by interpolation methods (such as bilinear interpolation).
[0085] Fill in missing areas: Use spatial interpolation methods to fill in the missing areas of other environmental factor rasters, and use a fixed value nv (such as -999) to fill in the missing areas of the groundwater depth raster.
[0086] Overlay raster data: Overlay all variables one by one into the same raster data, and extract all variable values corresponding to the center point positions of each raster cell.
[0087] Generate the basic dataset D: Generate the basic dataset D based on the complete variable information of each raster cell.
[0088] Step 4: Random sampling to construct the background dataset:
[0089] To construct the background dataset, samples need to be randomly drawn from the basic dataset D. The specific steps are as follows:
[0090] Set the random seed `seed`: Ensure the repeatability of each sampling.
[0091] Random sampling: Randomly draw N samples without replacement from the basic dataset D to generate the background data value D bg .
[0092] Determine the sample size N: By analyzing the impact of the sample quantity on the performance of the background model, refer to the sample size threshold at the performance convergence point to determine the value of the parameter N.
[0093] The above steps involve complex data processing and mathematical calculations, aiming to provide a high-quality data basis for subsequent groundwater level analysis and prediction.
[0094] Step 5: Establish a vegetation coverage prediction model:
[0095] Select a machine learning algorithm, using the background dataset D bg as the input to establish a model M for predicting vegetation coverage.
[0096] Step 6: Calculate the influence amount of groundwater depth:
[0097] Apply the model M trained in Step 5 to the samples in the prediction basic dataset D where the groundwater depth variable is not equal to the missing value nv, and calculate the residual between the model prediction result and the actual observed value , which is used to reflect the influence of groundwater depth on vegetation coverage. The calculation formula is as follows:
[0098] ;
[0099] where is the true vegetation coverage, is the predicted vegetation coverage, and the actual observed value is the vegetation coverage data observed by remote sensing means.
[0100] Step 7: Repeat multiple times until convergence, and calculate the mean residual:
[0101] Change the random seed `seed`, and repeat Steps 4 to 6 multiple times until the mean model residual at each pixel location tends to be stable and meets the convergence criterion.
[0102] Step 8: Fit the empirical model and calculate the ecological water depth threshold:
[0103] Using a groundwater depth of 1m as the grouping interval, group the samples in the basic dataset (such as 0 - 1m, 1 - 2m, 2 - 3m, and so on). For each group, calculate the standard deviation within the group and the corresponding mean groundwater depth. Taking the mean groundwater depth as the independent variable x, The standard deviation of is the dependent variable y, and the relationship between the two is fitted with an exponential function model. The model expression is as follows:
[0104] ;
[0105] Among them, , , are model parameters and are estimated by the nonlinear least squares method.
[0106] Parameter determines the rate of decline of the exponential function. At a given confidence level of 1-α (such as 95%), the confidence interval of parameter , is calculated by the following formula:
[0107] ;
[0108] ;
[0109] Among them, is the critical value of the t-distribution, which depends on the confidence level α and the groundwater depth value that defines the decline amplitude to reach p% (the recommended value of p is 80) of the total amplitude as the critical groundwater depth. The calculation formula is:
[0110] ;
[0111] Example 1:
[0112] This example is an analysis of a certain simulation scenario. The dataset of the simulation scenario is obtained through the following methods:
[0113] 1. Data generation process:
[0114] The vegetation coverage Y is composed of the environmental factor influence term f(X), the groundwater depth influence term g(GD), and the normal noise term σ, as shown in the formula: ;
[0115] The groundwater depth influence term g(GD) adopts an exponential decay function to simulate the dependence of vegetation on groundwater: ;
[0116] Among them, the parameter is set to 0.20, is set to 0.8, the change range of the function value is 0.2, and the groundwater depth value corresponding to 80% of the maximum change amplitude is 2.01m.
[0117] The noise term σ: follows a normal distribution with a mean of 0 and a standard deviation of 0.02, simulating the observation error and unmodeled factors.
[0118] The environmental factor influence term f(X) is set as a non - linear function of air temperature, precipitation, potential evapotranspiration, saturated soil water content, and the proportion of arable land.
[0119] By setting the random seed values from 1 to 100, 100 sub - datasets (D1 - D100) are generated, each dataset containing 100,000 samples to ensure the statistical robustness of the results.
[0120] 2. Model training and parameter setting:
[0121] The parameter settings for the data generation process of the simulation scenario are as follows: successively set the random seed values from 1 to 100, and the number of sub - datasets is 100,000, obtaining a total of 100 datasets, numbered D1, D2,..., D100.
[0122] For the randomly generated datasets above, random sampling is performed. A background dataset with a sample size of N is drawn, and the parameter N is set to 50,000. The background datasets are abbreviated as Dbg1, Dbg2, Dbg3,..., Dbg100.
[0123] Random forest is selected as the model algorithm for predicting the vegetation coverage of the background dataset. The selected features are the same as the environmental variables used in the data generation process, including environmental factors (air temperature, precipitation, potential evapotranspiration, saturated soil water content, proportion of arable land) and a total of 10 variables of the adjacent covariates corresponding to the environmental factors.
[0124] Successively using the background datasets Dbg1, Dbg2, Dbg3,..., Dbg100 as inputs, 100 models are trained. Each model is applied to the corresponding dataset to calculate the residuals, and the mean of the residuals is obtained by summarization.
[0125] Summarize the mean of the residuals of all models and the corresponding groundwater depth data pairs. Set the grouping interval of the groundwater depth to 1m, and calculate the standard deviation, the mean x of the groundwater depth in the 0 - 30m interval, and the standard deviation y of Figure 2 . The x and y together form 30 data point pairs, as shown by the black circles in the appendix of the specification.
[0126] x y x y x y 0.362 0.085 10.505 0.004 20.478 0.004 1.514 0.039 11.484 0.004 21.494 0.004 2.478 0.021 12.511 0.005 22.507 0.004 3.481 0.012 13.517 0.005 23.500 0.004 4.489 0.007 14.506 0.005 24.503 0.004 5.484 0.006 15.486 0.004 25.503 0.004 6.484 0.005 16.494 0.004 26.519 0.004 7.488 0.005 17.483 0.004 27.506 0.003 8.499 0.005 18.507 0.004 28.521 0.004 9.498 0.005 19.486 0.004 29.870 0.004
[0127] Table 1 shows the coordinate data points of the mean of the residuals x and the groundwater depth y in Example 1. In the table, x is the mean of the groundwater depth (m), and y is the standard deviation of the residuals.
[0128] 3. Exponential model fitting and critical depth calculation:
[0129] The non - linear model parameters obtained by least - squares fitting are: a = 0.119, b = 0.797, c = 0.004, and the coefficient of determination R of the model fitting 2 is 0.988. The fitting curves ( Figure 2 the black lines in it) all show that as x increases, y shows an obvious downward trend.
[0130] At the 95% confidence level, the confidence interval of parameter b is [0.775, 0.818], and the confidence intervals of parameters a and c are [0.117, 0.122] and [0.0039, 0.0044] respectively. Set the decline rate as p%, p = 80.
[0131] According to the above formula, we can calculate the critical groundwater depth :
[0132] ;
[0133] The confidence interval of and The corresponding calculation results are respectively:
[0134] (b = 0.775)= 2.08m;
[0135] (b = 0.818)= 1.97m;
[0136] Therefore, The confidence interval of
[0137] Example 2:
[0138] Calculate the critical groundwater depth of Mu Us Sandy Land in the northern part of the Ordos Plateau, China. The aquifer type in this area is a pore aquifer. Most of the surface area is covered by aeolian sand, and the maximum capillary rise height is about 1.0m. The main root systems of typical groundwater - dependent plants Salix psammophila and Salix matsudana are 4 - 5m and 5 - 6m respectively. In addition to the groundwater depth, the distribution of vegetation coverage in this area is also affected by meteorology (such as the spatial heterogeneity of precipitation), topography (topographic undulation), soil (soil water), water bodies (river water supply to vegetation), irrigation and other factors.
[0139] Based on the field survey data of groundwater levels and 1 km elevation raster data, 1 km groundwater depth data was calculated. Since the aquifer type in this area is a pore aquifer, the data density screening condition for the high-confidence groundwater level interpolation result area is greater than 2 per 100 km 2 .
[0140] The environmental factors involved in the modeling include a total of 12 factors related to climate (air temperature, precipitation, potential evapotranspiration, wind speed), topography (elevation, aspect, slope), soil (soil moisture content, clay content), water bodies (distance to lakes, distance to rivers), and human activities (proportion of cultivated land). By analyzing the influence of the parameter r in the inverse distance weighting method on the model performance, the optimal value of the parameter r was determined to be 15 km (as shown in the attached Figure 3 instruction manual). Set = 2, and use the inverse distance weighting method to generate the neighboring covariate raster data corresponding to these environmental factors.
[0141] Perform spatial matching on the groundwater depth, vegetation coverage, other environmental factors, and their covariate rasters at a 1 km spatial resolution. After excluding the low-confidence groundwater depth data, a total of 100,610 data points are obtained to form the dataset D.
[0142] Extract a subset of size N from the overall dataset D for training to obtain a model for predicting vegetation coverage. Analyzing the influence of the parameter N on the model performance shows that as the parameter N increases, the root mean square error of the ten-fold cross-validation of the model continuously decreases. When the sample size is higher than 50,000 data points, the relative change rate of the model performance is less than 5%. Based on this, the value of the parameter N is determined to be 50,000 (as shown in the attached Figure 4 instruction manual).
[0143] Set the initial seed value to 1, and randomly extract 50,000 samples from D to form the background dataset. Using the background dataset as the input, establish an extreme gradient boosting tree model with environmental factors and their corresponding neighboring covariates as independent variables and vegetation coverage as the dependent variable, and calculate the residuals of this model applied to the high-confidence groundwater depth area.
[0144] After repeating the above random extraction process 100 times, the relative change rate of the mean of the model residual calculation results is less than 5% for more than 90% of the samples, reaching the convergence criterion, stopping the iteration, and calculating the mean residual .
[0145] Set the groundwater depth grouping interval to 1 m, and statistically calculate the standard deviation of those belonging to the same groundwater depth group, the mean x of the groundwater depth in the 0 - 30 m interval, and The standard deviation of y. A total of 30 data point pairs are composed of x and y, as shown by the black circles in the attached instructions Figure 5 as shown.
[0146] x y x y x y 0.318 0.035 10.486 0.025 20.493 0.024 1.507 0.032 11.486 0.026 21.505 0.026 2.493 0.030 12.506 0.025 22.517 0.026 3.495 0.030 13.483 0.025 23.511 0.024 4.479 0.028 14.499 0.027 24.502 0.026 5.474 0.027 15.510 0.025 25.509 0.024 6.487 0.027 16.525 0.024 26.497 0.026 7.478 0.029 17.467 0.024 27.513 0.024 8.495 0.026 18.508 0.026 28.528 0.025 9.505 0.027 19.504 0.025 29.737 0.024
[0147] Table 2 shows the coordinates of the residual mean x and the groundwater depth y data points for Example 2. In the table, x is the mean groundwater depth (m), and y is the residual standard deviation.
[0148] The model parameters obtained by fitting through the nonlinear least squares method are: = 0.011, = 0.230, = 0.025. The coefficient of determination R 2 of the model fitting is 0.850. The fitting curve ( Figure 3 the black line in) all shows that as x increases, y shows an obvious downward trend.
[0149] At the 95% confidence level, the confidence interval of the parameter is [0.152, 0.308], and the confidence intervals of the parameters , are [0.009, 0.013] and [0.024, 0.025] respectively.
[0150] Set the downward amplitude p%, p = 80.
[0151] According to the above formula, we can calculate the ecological critical groundwater depth :
[0152] = 7.00m;
[0153] The confidence interval of and The corresponding calculation results are:
[0154] (b = 0.152) = 10.59m;
[0155] (b = 0.308) = 5.23m;
[0156] Therefore, The calculation result is 7.00 m, and the confidence interval is from 5.23 m to 10.59 m, indicating that the critical groundwater depth in this area is 7.00 m. When the depth exceeds this value, vegetation degradation will occur. This calculation result is relatively close to the sum (7.0 m) of the maximum root length (6 m) and the capillary rise height (1.0 m) of the vegetation in the area where the embodiment is located, further verifying the rationality and credibility of the critical groundwater depth value calculated by this method.
[0157] Compared with the prior art, the present application has at least the following advantages:
[0158] 1. Comprehensively consider the combined effects of complex multi-environmental factors: This method uses machine learning algorithms (such as random forest, extreme gradient boosting tree) to replace the traditional linear fitting method, and can more accurately capture the complex and non-linear influence mechanism of multi-environmental factors (such as climate, terrain, soil) on vegetation coverage.
[0159] 2. Effectively handle the spatial autocorrelation of vegetation: Introduce the neighboring covariates of environmental factors as features into model construction to eliminate the prediction bias caused by the spatial autocorrelation of vegetation coverage, thereby significantly improving the model prediction accuracy.
[0160] 3. Enhance the reliability of results: Calculate the groundwater depth influence amount (residual mean) by means of multiple random samplings and independent modelings until the results converge. The summary of the results of multiple iterations significantly improves the stability of the estimation and reduces the error caused by a single model.
[0161] 4. Improve statistical power and optimize resource utilization: By analyzing the influence of sample size on model performance, determine the sample size threshold at the performance convergence point, enhance the statistical power analysis, and ensure the reliability of model results. Multiple iterations provide additional stability tests, making the analysis results more robust and efficient. This method not only improves statistical power but also optimizes resource utilization, providing a strong guarantee for the scientificity and practicality of the method.
[0162] 5. Flexibility and scalability: This method has no strict requirements on the number of environmental variables and sample size, is applicable to a wider range of eco-hydrological research scenarios, and can be easily extended to the analysis of different regions and complex systems.
[0163] In summary, the core innovation of the present invention lies in establishing a calculation framework that integrates multi-factor influences, considers spatial autocorrelation and uncertainty analysis, provides a new solution for the improved eco-hydrological depth threshold calculation method, and is significantly superior to the existing methods.
[0164] Although embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application. The scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A method for calculating the critical depth of groundwater based on geographic big data and residual analysis, characterized by: The method comprises the following steps: Step 1: Calculation of groundwater depth data: Firstly, Kriging interpolation is performed on the groundwater level monitoring data; The interpolation results are screened according to the aquifer type and the data density of groundwater level observation data in the study area; if the data density of a certain area is lower than the preset threshold, the groundwater level of the area is set to a null value; Use the raster calculation tool to calculate the difference between the surface elevation raster and the groundwater level raster to generate groundwater depth raster data; Using the conditional screening tool, set the pixel values less than 0 in the groundwater depth raster to 0 to eliminate the influence of outliers and ensure the rationality of the data; Step 2: Calculation of neighboring covariate data of environmental factors: Neighborhood covariates of environmental factors were calculated using the inverse distance weighting method; The neighboring covariate is the inverse distance weighted mean of the environmental variables within a given radius r around the sample, and the calculation formula is: ; in, is the environmental variable value of the i-th grid pixel, The current pixel and the The distance between adjacent pixels, is the distance attenuation coefficient, usually set to 2. is the number of non-missing pixels within the radius neighborhood around the pixel; Step 3: Raster data spatial matching: Adjust raster data from different sources to the same resolution and coordinate system to ensure the accuracy of subsequent analysis; The specific steps are as follows: Projection transformation: transform all raster data to the same coordinate system; Resampling operation: resample all raster data to the same resolution; Filling missing areas: The missing areas of other environmental factor grids are filled using spatial interpolation methods, and the missing areas of groundwater depth grids are filled using a fixed value nv; Overlay raster data: Overlay all variables one by one onto the same raster data, and extract all variable values corresponding to the center point of each raster pixel; Generate basic dataset D: Generate basic dataset D based on the complete variable information of each raster pixel; Step 4: Randomly sample to construct a background data set: In order to construct the background data set, it is necessary to randomly extract samples from the basic data set D. The specific steps are as follows: Set the random seed: ensure the repeatability of each sampling; Random sampling: N samples are randomly selected from the basic data set D without replacement to generate background data value D bg ; Determine the sample size N: By analyzing the impact of the number of samples on the performance of the background model, the value of the parameter N is determined by referring to the sample size threshold that reaches the performance convergence point; Step 5: Establish vegetation coverage prediction model: Select a machine learning algorithm based on the background dataset D bg As input, a model M for predicting vegetation coverage is established; Step 6: Calculate the impact of groundwater depth: The model M trained in step 5 is applied to the samples in the prediction basic data D where the groundwater depth variable is not equal to the missing value nv, and the residual between the model prediction result and the actual observation value is calculated. , which is used to reflect the impact of groundwater depth on vegetation coverage. The calculation formula is as follows: ; in, is the actual vegetation coverage, is the predicted vegetation cover; Step 7. Repeat several times until convergence and calculate the residual mean: Change the random seed and repeat steps 4 to 6 multiple times until the model residual mean at each pixel location is tends to be stable and satisfies the convergence criterion; Step 8: Fit the empirical model and calculate the critical depth of groundwater: Empirical model fitting; Data grouping: The samples of the basic data set are grouped with the groundwater depth of 1m as the grouping interval; Calculate the standard deviation and mean: For each group, calculate the The standard deviation of and the corresponding mean of groundwater depth; Exponential function model fitting: the mean groundwater depth is taken as the independent variable x, The standard deviation of is the dependent variable y, and the exponential function model is used to fit the relationship between the two. The model expression is as follows: ; in, , , are model parameters, estimated by nonlinear least squares method; parameter Determines how fast the exponential function decreases. At a given confidence level of 1-α, the parameter The confidence interval of , ] is calculated by the following formula: ; ; in, is the critical value of the t distribution, depending on the confidence level α and the definition of the decline to reach the total amplitude The groundwater depth is the critical groundwater depth. , the calculation formula is: ; The confidence interval can be obtained by substituting the upper and lower limits of parameter b , Obtained indirectly.
2. The method for calculating the critical depth of groundwater based on geographic big data and residual analysis according to claim 1 is characterized by: The data density threshold corresponding to the high confidence interpolation result in step 1 is: no less than 3 observation points per square kilometer of porous water aquifer; no less than 2 observation points per square kilometer of karst water; and no less than 1 observation point per square kilometer of fissure water.
3. The method for calculating the critical depth of groundwater based on geographic big data and residual analysis according to claim 1 is characterized in that: The environmental factors in step 2 include factors related to climate, topography, soil, water bodies, and human activities.
4. The method for calculating the critical depth of groundwater based on geographic big data and residual analysis according to claim 1 is characterized in that: The specific steps of calculating the neighboring covariates of environmental factors by the inverse distance weighted method are as follows: determine all non-missing pixels and their distances around the current grid cell; for each non-missing pixel, calculate the value of its environmental variable divided by the distance power; add all the results and divide by the distance The sum of the powers gives the value of the adjacent covariate.
5. The method for calculating the critical depth of groundwater based on geographic big data and residual analysis according to claim 1 is characterized in that: In step 5, the machine learning algorithm is selected from any one of random forest and extreme gradient boosting tree.
6. The method for calculating the critical depth of groundwater based on geographic big data and residual analysis according to claim 1 is characterized by: The convergence criterion in step 7 is that the relative change rate of the residual mean of all 90% pixels is less than 5%.
Citation Information
Cited By
Oasis ecological hydrological simulation method, device and equipment and storage medium
CN120874606A
Oasis eco-hydrological simulation method, device, equipment and storage medium
CN120874606B