Underground water fluorine space prediction method based on LGBM-SHAP model
By combining multiple indicators with the LGBM-SHAP model for spatial prediction of fluoride concentration, the problems of sparse groundwater sample distribution and model disorder in traditional methods are solved, generating more accurate fluoride concentration and vulnerability maps, and supporting sustainable groundwater management.
Patent Information
- Application Number
- CN202510621181.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-02-20
- Filing Date
- 2025-05-14
- Publication Date
- 2025-11-07
AI Technical Summary
Traditional methods for spatial prediction of groundwater fluoride concentration are limited by the sparse distribution of groundwater samples and the inability to consider geological, hydrological and environmental variables, resulting in reduced accuracy near the boundary. Furthermore, machine learning models suffer from insufficient exploration of model and algorithm types, disorder in the selection of input indicators and lack of interpretability, thus failing to effectively support sustainable groundwater management.
The Light Gradient Bumper (LGBM) regression model was combined with SHAP analysis to integrate hydrological, geological, environmental, climatic, hydrochemical and human activity indicators. Twenty-six indicators were selected through the pressure-state-response framework to predict fluoride concentration spatially, and sustainable management strategies were formulated based on the inherent vulnerability of aquifers and groundwater reserves.
It enables more accurate spatial prediction of fluoride concentration, generates more refined maps of inherent aquifer vulnerability, identifies areas with different groundwater management intensities, and supports sustainable groundwater management decisions.
Smart Images

Figure CN120913682A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of groundwater chemical component prediction, more specifically, relates to a groundwater fluorine spatial prediction method based on an LGBM-SHAP model. BACKGROUND
[0002] Groundwater (GW) is an important resource for sustaining agricultural, industrial, and domestic activities, while supporting ecosystems and global economic health. However, groundwater quality is increasingly threatened by various pollutants, with fluoride (F-) in groundwater being the most notable due to its dual benefits and hazards. Sources of F- in groundwater include natural and anthropogenic factors. From a geological perspective, it originates from the dissolution of major fluoride-containing minerals. Anthropogenic sources include emissions from the aluminum and coal industries, agricultural fertilizers, and various manufacturing processes. Low concentrations of F- help prevent tooth decay and strengthen bones, but excessive exposure can lead to serious health problems such as dental fluorosis and skeletal fluorosis. Long-term exposure to F- can cause harm to communities, particularly those lacking resources to implement expensive water treatment technologies. Therefore, effective management of F- concentrations in groundwater is crucial for global health and environmental sustainability.
[0003] Spatial prediction of fluoride concentrations (SPFC) provides groundwater managers with critical information to identify priority areas and develop targeted and effective management strategies. Traditional groundwater SPFC methods mainly rely on spatial interpolation techniques based on collected groundwater samples, such as the Kriging method. However, these methods are significantly limited by the sparsity of groundwater sample distribution, the inability to consider geological, hydrological, and environmental variables, and the reduced accuracy near the boundaries of the study area due to boundary effects. Given these limitations, it is necessary to adopt advanced spatial prediction methods that can effectively integrate multiple indicators to significantly improve accuracy and address spatial variability issues in SPFC.
[0004] Currently, machine learning (ML) models are widely used for spatial prediction of groundwater pollutants due to their high accuracy, efficiency, and ability to integrate diverse indicators, thereby improving spatial prediction capabilities. They have significantly improved the prediction accuracy of groundwater pollutants such as arsenic, nitrate, ammonia, iron, and manganese, as well as organic substances. Although these studies have contributed to SPFC through ML models, they still face three significant challenges: insufficient exploration of model and algorithm types, disordered selection of input indicators, limited application of explainable machine learning techniques (EMLTs), and insufficient support for future sustainable groundwater management. SUMMARY
[0005] In order to solve the above technical problems, the present application provides a groundwater fluoride spatial prediction method based on LGBM-SHAP model, which adopts a light gradient boosting machine (LGBM) regression model, and combines hydrology, geology, environment, climate, water chemistry and human activity indicators, and an inherent aquifer vulnerability (IAV) model to predict the spatial distribution of groundwater fluoride concentration. At the same time, SHAP analysis, as one of the explainable machine learning techniques (EMLTs), is combined with the LGBM model to visualize the contribution of each indicator to the result. The selection of indicators is based on the pressure-state-response (PSR) framework to consider the impact of human activities on fluoride concentration. More importantly, the present application combines the results of groundwater reserves (GWS), IAV and SPFC to propose a matrix-based framework for developing sustainable groundwater management strategies.
[0006] The purpose and effect of the groundwater fluoride spatial prediction method based on LGBM-SHAP model of the present application are achieved by the following specific technical means:
[0007] The groundwater fluoride spatial prediction method based on LGBM-SHAP model comprises the following steps: 1. LGBM model construction, calibration and verification; 2. index importance analysis by SHAP; 3. combination of inherent aquifer vulnerability (IAV) and groundwater reserves (GSW) to develop sustainable groundwater management strategies for fluoride pollution.
[0008] The method further comprises index selection, i.e. a total of 26 indicators (pressure, state, response, water chemistry and others) in 5 categories are adopted, which consider both natural and human sources of fluoride pollution in groundwater.
[0009] In step one, data preprocessing, IAV weight determination, LGBM model, hyperparameter space and optimization are included.
[0010] In step one, data preprocessing includes encoding of categorical variables, standardization of indicators, and replacing F- concentrations below 0.1 mg / L (detection limit) with a smaller value (0.05 mg / L).
[0011] In step one, the LGBM model includes two innovations: gradient-based one-sided sampling (GOSS) and exclusive feature bundling (EFB);
[0012] GOSS improves training efficiency by prioritizing samples with larger gradients, allowing the model to focus on optimizing important data points with information gain; that is, the first a% of instances in samples with the largest gradient are initially grouped into subset A, and for the remaining sample set B containing lower gradients, a new subset B' is created by randomly sampling b% of instances from B, then the variance gain IG of data instances is calculated, which considers both subsets, and the optimal split criterion is determined using its weighted gradient;
[0013]
[0014] wherein:
[0015] N is the total number of samples;
[0016] gi is the gradient of the loss function with respect to instance iii;
[0017] A and B are subsets of the sample set;
[0018] A' and B' are subsets obtained after splitting A and B;
[0019] is the coefficient of gradient normalization;
[0020] and are the number of samples in the left and right nodes after splitting using decision variable d. on feature j;
[0021] EFB (Exclusive Feature Bundling) effectively reduces the number of features by grouping features in sparse space, it considers this problem as a graph coloring problem, combining those rarely active features into a feature group, this method optimizes computational efficiency without affecting model accuracy; in addition, the LGBM model adopts a leaf-based tree growth strategy, which splits nodes based on maximum information gain, its mathematical expression is:
[0022]
[0023] wherein:
[0024] G L and G R are the total sum of gradients in the left and right nodes, respectively;
[0025] H L and H R are the total sum of second-order gradients (Hessian values) in the left and right nodes, respectively;
[0026] λ is a regularization parameter that controls the complexity of the tree;
[0027] gamma is the regularization term for the leaf nodes to prevent overfitting.
[0028] The present application comprises at least the following advantages:
[0029] The present application adopts LGBM-SHAP regression model, combines multiple indicators to carry out spatial prediction of fluorine concentration (SPFC). Through SHAP analysis (an interpretable machine learning technique, EMLT), the visualization of the contribution of each indicator in the model is realized. In addition, the present application develops a matrix-based sustainable groundwater management framework, which integrates IAV, GWS and SPFC, and identifies seven areas with different groundwater management intensities. The main findings of the present application include the following aspects:
[0030] 1. A more accurate IVA map is generated. However, the correlation analysis shows that there is no single common groundwater quality parameter that shows strong correlation, which highlights the necessity of using multiple methods to verify the IAV model.
[0031] 2. Through the TPE optimization method, the LGBM model performs well in SPFC. The predicted F - concentration is highly consistent with the observed value, and the granular and local dispersed distribution emphasizes the difference from the interpolation method.
[0032] 3. In addition to water chemical indicators, the importance of population density and GDP is relatively high, while the correlation of DRASTIC indicators is low. This indicates that point source pollution may be the main source of F - pollution in the study area, and highlights the limitations of the DRASTIC model based on the "source-path-target" framework in predicting natural sources of F - . Additional indicators such as NDVI, elevation, TWI and distance to rivers significantly improve the model performance.
[0033] 4. SHAP analysis not only visualizes the contribution of indicators to the model, but also supports F - water chemical analysis and pollution source inference. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 is the research area of the present application;
[0035] Figure 2 is a schematic diagram of the methodological framework of the present application;
[0036] Figure 3 is a schematic diagram of the matrix-based groundwater fluorine pollution management strategy framework of the present application ((a) high fluorine concentration; (b) medium fluorine concentration; (c) low fluorine concentration);
[0037] Figure 4is a map of the IVA distribution of the study area and its correlation with groundwater quality parameters;
[0038] Figure 5 is a schematic diagram of SPFC and model performance evaluation;
[0039] Figure 6 Indicator analysis: (a) SHAP analysis (red dots indicate high feature values, blue dots indicate low feature values); (b) count-based feature importance analysis (normalized);
[0040] Figure 7 is a structural diagram of the self-attention mechanism of the present application. DETAILED DESCRIPTION
[0041] The embodiments of the present application are further described in detail below by way of examples. The following examples are used to illustrate the present application, but cannot be used to limit the scope of the present application.
[0042] In the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more; the orientation or positional relationship indicated by the terms "coaxial", "bottom", "one end", "top", "middle", "the other end", "upper", "one side", "top", "inner", "front", "central", "both ends" and the like is based on the orientation or positional relationship shown in the drawings, which is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second", "third" and the like are only for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0043] In the description of the present application, it should be noted that, unless otherwise specified and limited, the terms "mounting", "setting", "connecting", "fixing", "rotating", "inner diameter", "outer diameter" and the like should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or it can be integrated; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium; it can be the internal communication of two elements or the interaction relationship between two elements; it can be a smaller diameter, or it can be a larger diameter; unless otherwise explicitly limited, the above-mentioned terms in the present application can be understood according to the specific meaning of the above-mentioned terms in the present application according to the specific circumstances.
[0044] Embodiment:
[0045] 1. Regional overview of the embodiment:
[0046] The Ordos Basin is located in northwestern China, covering a total area of approximately 263,977 square kilometers, spanning five provinces and autonomous regions: Shaanxi, Shanxi, Gansu, Inner Mongolia Autonomous Region, and Ningxia Hui Autonomous Region. Its geographical coordinates range from 105°45' to 112°13' east longitude and 33°42' to 41°49' north latitude. Figure 1 This region has a temperate continental monsoon climate, with an average annual precipitation of 150–770 mm, an average annual temperature of 3.6–13.6℃, and an average annual evaporation of 1000–3500 mm, making it a typical semi-arid region. Except for the endorheic basin in the central part of the northern desert plateau, the Ordos Basin belongs to the Yellow River basin, with major rivers including the Molin River, the Dos River, and the Wuding River.
[0047] Topographically, the Ordos Basin slopes from northwest to southeast, with an elevation ranging from 319 meters to 3748 meters. Based on landform type, the Ordos Basin can be divided into desert plateau, loess plateau, and mountainous areas. Geologically, the Ordos Basin is mainly composed of the following lithologies: Precambrian crystalline schist and slightly metamorphosed clastic-carbonate rocks, Cambrian-Ordovician carbonate rocks interbedded with clastic rocks, Carboniferous-Jurassic coal-bearing clastic rocks, Cretaceous clastic rocks, Cenozoic Tertiary-Quaternary argillaceous clastic rocks, and Quaternary loose sediments.
[0048] As a typical semi-arid basin, groundwater is often a vital water resource for drinking water, irrigation, and various municipal and industrial activities in the study area. Although research on inherent aquifer vulnerability (IAV), groundwater systems, groundwater health risks, groundwater recharge and discharge, and hydrogeochemical processes is progressing in the Ordos Basin, the protection and sustainable use of groundwater remains an ongoing challenge in this study area.
[0049] 2. Methodology
[0050] The method framework of this invention is as follows: Figure 2 As shown, this includes indicator selection, LGBM model construction, calibration and validation, indicator importance analysis via SHAP, and the development of sustainable groundwater management strategies for fluoride pollution by combining inherent aquifer vulnerability (IAV) and groundwater storage (GSW).
[0051] 2.1 Description of Groundwater Samples and Fluoride Concentration
[0052] The groundwater samples used in this embodiment of the invention are from a study by Hongyun et al., who collected 742 groundwater samples from the Ordos Basin, of which 704 samples were located within the study area. Detailed information on the characteristics of the groundwater samples can be found in Hongyun et al. Of these 742 groundwater samples, 158 samples had a fluoride concentration exceeding 1.0 mg / L (China's Class III groundwater standard), and 102 samples exceeded 1.5 mg / L (WHO recommended limit).
[0053] 2.2 Indicator selection
[0054] The present study adopted a total of 26 indicators (stress, state, response, water chemistry, and others) in five categories in SPFC, which considered both natural and anthropogenic sources of fluoride pollution in groundwater. Due to the existence of vector and raster data with different scales and resolutions, the vector data were converted to 500-meter resolution using the vector-to-raster tool in GIS. For raster data with inconsistent resolutions, the resampling technique was used to unify the resolution to 500 meters. Although this method may cause some errors, it is widely accepted and used in this field. Specifically, for the nine water chemistry indicators, the Kriging interpolation method was used to draw their spatial distribution maps. To improve the accuracy of interpolation at the boundary and reduce the boundary effect, the groundwater samples outside the study area were included in the interpolation process. Table 1 provides detailed information on these indicators.
[0055] Table 1. Sources, data types, and scales of indicators
[0056]
[0057] Notes:
[0058] NBS: National Bureau of Statistics
[0059] WRB: Water Resources Bulletin
[0060] CGS: China Geological Survey
[0061] CMA: China Meteorological Administration
[0062] HWSD v2.0: Harmonized World Soil Database version 2.0
[0063] NESDC: National Ecological System Science Data Center
[0064] HDI: Human Development Index
[0065] GDP: Gross Domestic Product
[0066] CMSH: Ca 2+ +Mg 2+ -SO4 2- -HCO3 -
[0067] NKCL: Na + +K + -Cl-
[0068] CAI: Chlor-alkali Index
[0069] SAR: Sodium Adsorption Ratio
[0070] TDS: total dissolved solids
[0071] 2.3 Model process
[0072] 2.3.1 Data pre-processing
[0073] In addition to the resampling method mentioned earlier, the data pre-processing step of the present invention includes encoding of categorical variables, standardization of indicators, and replacing F- concentrations below 0.1 mg / L (detection limit) with a smaller value (0.05 mg / L). Encoding of categorical variables converts categorical indicators into numerical form for subsequent analysis. For indicators in the DRASTIC model, ordinal encoding is used to rank and score the indicators. For other categorical indicators, label encoding is used for numerical conversion. Standardization scales the values of indicators and predictor variables into a common range to prevent individual features from having an undue influence on the model.
[0074] In the present invention, different standardization methods are used for F- concentrations and indicators. For F- concentrations, the standardization method is based on the study of spatial prediction by Guo et al. In this method, groundwater samples with F- concentrations greater than 5 mg / L are standardized to 1, while sample values with concentrations less than 5 mg / L are divided by 5. The rationality of this method lies in that extremely high concentration values are considered as outliers, which may be influenced by unique environmental factors. These extremely high values are grouped with high concentration values, which helps the model better identify feature patterns. For indicator standardization, the Max-Min standardization method is used, which is widely used in the current study. In addition, there are 34 groundwater samples in the database with F- concentrations below the detection limit (<0.1 mg / L). To avoid potential anomalies caused by zero values in logarithmic operations or calculations sensitive to zero, these values are replaced with 0.05 mg / L. In addition, 80% of the groundwater samples are used as the training set and 20% as the validation set.
[0075] 2.3.2 IAV weight determination
[0076] The traditional DRASTIC model uses fixed weights, with the order of allocation being 5, 4, 3, 2, 1, 5, and 3. Many studies have shown that combining the AHP method with the traditional IVA model can improve model performance. The AHP model was developed by Saaty and is a decision-making technique used to solve complex problems by breaking them down into hierarchical components. It evaluates multiple criteria through pairwise comparisons, generating a judgment matrix where each element reflects the importance of one criterion relative to another. The consistency of these comparisons is checked by calculating the consistency index (CI, Equation 1) and the consistency ratio (CR, Equation 2).
[0077]
[0078] where λ max represents the maximum eigenvalue, n represents the number of indicators, and RI is the average random uniformity indicator value CI calculated from a randomly generated matrix of size n x n.
[0079] 2.3.3 LGBM model
[0080] The LGBM model developed by Ke et al. at Microsoft is a gradient boosting decision tree (GBDT) framework designed to efficiently handle large-scale, high-dimensional datasets. Since the present invention involves 26 indicators, this model is very suitable. In the present invention, as a highly efficient ensemble learning model, LGBM is widely used in the field of groundwater research, including groundwater potential prediction, mineral spring identification, water storage downscaling processing based on GRACE data, and pollution source estimation.
[0081] The LGBM model contains two key innovations: gradient-based one-sided sampling (GOSS) and exclusive feature bundling (EFB). GOSS improves training efficiency by prioritizing samples with larger gradients, allowing the model to focus on optimizing important data points with high information gain. Specifically, the top a% of instances in the sample with the largest gradient are initially grouped into subset A. For the remaining sample set B containing lower gradients, a new subset B' is created by randomly sampling b% of instances from B. Then, the variance gain IG (Formula 1) is calculated for the data instances, considering both subsets, and the weighted gradient is used to determine the optimal split criterion.
[0082]
[0083] where:
[0084] N is the total number of samples;
[0085] gi is the gradient of the loss function with respect to instance iii;
[0086] A and B are subsets of the sample set;
[0087] A' and B' are subsets obtained after splitting A and B;
[0088] is the coefficient of gradient normalization;
[0089] and and
[0090] EFB (Exclusive Feature Bundling) effectively reduces the number of features by grouping features in sparse space. It treats this problem as a graph coloring problem, combining features that are rarely active at the same time into a feature group. This method optimizes computational efficiency without affecting model accuracy. In addition, the LGBM model adopts a leaf-based tree growing strategy, which splits nodes based on maximum information gain, whose mathematical expression is:
[0091]
[0092] where:
[0093] G L and G R are the total sum of gradients in the left and right nodes, respectively;
[0094] H L and H R are the total sum of second-order gradients (Hessian values) in the left and right nodes, respectively;
[0095] λ is a regularization parameter that controls the complexity of the tree;
[0096] γ is a regularization term for leaf nodes to prevent overfitting.
[0097] 2.3.4 Hyperparameter Space and Optimization
[0098] Table 2 Hyperparameters, Hyperparameter Space, and Optimal Hyperparameters
[0099]
[0100] Hyperparameter optimization plays a key role in improving model prediction accuracy, as it addresses the inefficiency of traditional trial-and-error methods in optimization. The selected hyperparameters and hyperparameter space in this invention are detailed in Table 2. It is worth noting that "colsample_bytree", "reg_alpha" (L1 regularization) and "reg_lambda" (L2 regularization) are included in the selection of hyperparameters, as they play an important role in solving the problem of overfitting and multicollinearity. At the same time, as a tree-based non-parametric model, the LGBM model itself can partially overcome the problem of multicollinearity between indicators.
[0101] Subsequently, the Tree-structured Parzen Estimator (TPE) method is used for hyperparameter optimization. TPE is a global Bayesian optimization technique that focuses on the most promising areas based on past evaluation results to optimize search for optimal settings. In this invention, the TPE method is used to minimize the Root Mean Square Error (RMSE), and the hyperparameters are optimized through the conditional probability P(y∣x)P(y|x)P(y∣x), which is defined as follows:
[0102]
[0103] where x represents a hyperparameter, y is the observed RMSE, y□ is the RMSE threshold based on the observed data. l(x) represents the probability density of the observed values less than the threshold, while g(x) represents the probability density of the observed values greater than or equal to the threshold. Subsequently, the expected improvement is calculated by formula Eq(5).
[0104]
[0105] Let γ represent the fraction less than the threshold, γ = p(y < y□), and EI * y *(x) represents the fraction greater than or equal to the threshold, then it can be expressed as:
[0106]
[0107] To maximize the expected improvement (EI), g(x) / l(x) is continuously reduced until a preset number of iterations is reached. In the present invention, the number of iterations is set to 1000, which is generally higher than the number of iterations required to reach stability. If instability still exists, further adjustment may be needed to increase the number of iterations.
[0108] 2.4 Model performance evaluation
[0109] The present invention uses four key indicators to evaluate the accuracy and reliability of the prediction regression model: the coefficient of determination (R2), the root mean square error (RMSE), the mean absolute error (MAE), and the explained variance (EV).
[0110] R 2 : measures the degree of fitting of the model to the observed results, indicating the proportion of the total variation of the results explained by the model.
[0111] RMSE: provides a clear indication of model precision by measuring the average magnitude of prediction errors.
[0112] MAE: quantifies the average over-prediction or under-prediction without considering direction, providing an intuitive understanding of prediction errors expressed in the same unit.
[0113] EV: measures the degree of explanation of the model to the variability of the target variable.
[0114] These indicators are widely used in regression models to evaluate their performance, and the calculation formulas of these indicators are as follows:
[0115]
[0116] Where:
[0117] n is the number of groundwater samples;
[0118] y i is the actual value;
[0119] is the predicted value;
[0120] Var(y) is the variance of the actual value, whose calculation logic is: higher R 2 indicates a better model fitting, while lower RMSE, MAE and EV values indicate higher model performance.
[0121] 2.5 Indicator Analysis and SHAP Observation
[0122] Feature Importance (FI) is a useful tool for indicator analysis, used to determine the relative importance of each indicator in the model. In the LGBM model, the calculation of FI is based on the frequency of a certain feature being used as a split in all trees and the contribution of these splits to the improvement of model performance. This is one of the key advantages of the LGBM model, as it can quickly and efficiently complete the FI calculation in the Python environment.
[0123] SHAP analysis goes further than FI by providing a systematic method to explain the contribution of each feature (indicator) to the model prediction. SHAP analysis is widely used in ensemble learning models, helping to quantify the impact of each feature by assigning a SHAP value to each feature to represent the degree of influence of the feature on the model output. SHAP analysis includes two key components: SHAP value and feature value.
[0124] SHAP value: quantifies the contribution of each feature to the model prediction, indicating how much it shifts the predicted value from the baseline value. Positive or negative SHAP values indicate that the feature increases or decreases the likelihood of a given outcome, respectively.
[0125] Feature value: represents the actual observed value of the indicator used in the model.
[0126] The SHAP value formula for feature j is:
[0127]
[0128] where:
[0129] ·φ i is the SHAP value of feature i;
[0130] ·|S| is the number of features in the subset S;
[0131] ·M is the total number of features;
[0132] ·f(S) is the prediction of the model only containing feature subset S;
[0133] • f(S U {i}) is the prediction of the model when the feature subset S includes feature i.
[0134] 2.6 Fluoride concentration map and groundwater fluoride management
[0135] Groundwater management needs to consider not only pollution, but also groundwater storage (GWS) and inherent vulnerability (IVA). Based on the optimized hyperparameters, a fluoride concentration map can be drawn in combination with the selected 26 indicators. The fluoride concentration is classified with reference to the Chinese Class III groundwater standard (GBT 14848-2017): areas with fluoride concentration exceeding 1.0 mg / L are classified as high-pollution areas, while the remaining areas are divided into low-pollution areas and moderate-pollution areas with 0.5 mg / L as the threshold.
[0136] The data for the GWS map comes from the water resources bulletins of each county in China, while the IVA map is generated based on the weighted superposition of the seven “state” indicators in Table 2. Both maps are classified into low, medium, and high categories using the natural break method. Based on IVA, GWS, and SPFC, the study area can be divided into seven different groundwater fluoride management regions, namely priority repair region, general repair region, key monitoring region, general monitoring region, low fluoride protection region, low fluoride sustainable utilization region, and low GMS protection region. The specific zoning method is shown in Figure 3 .
[0137] Since fluoride is the main pollutant in the study area, this framework is developed to guide groundwater management strategies. In fact, this framework can also be adapted to the management of other important feature pollutants or overall water quality issues.
[0138] 3. Results and discussion
[0139] 3.1 IAV and the basis for using groundwater quality parameters as verification indicators
[0140] Figure 4 a shows the IVA map of the study area, which is classified into high, medium, and low categories using the natural break method. It can be observed that the high IVA regions are mainly concentrated in the north-central region, where the groundwater level is relatively high, the slope is gentle, and the net recharge is significant. The aquifer medium in this region is mainly composed of sand and gravel, which contributes to a higher IVA. Low IVA is mainly distributed in the southeast, due to significant groundwater depth and aquifer composed of sandy mudstone, with an air zone of loess. However, compared with previous studies, Figure 4 a shows significant differences in the results, which may be due to the unclear description of the hydrogeological and environmental conditions in the study area in the early studies. The results of the present invention are more accurate by combining the latest hydrogeological conditions of the Ordos Basin.
[0141] Figure 4 b to Figure 4 g shows IVA distribution and three different correlation coefficients between IVA and six groundwater quality parameters (F - , TDS, Cl - , SO4 2- , NO3 - , and HCO3 - ). It is worth noting that the DRASTIC index has very low correlation with the concentrations of these parameters. This is because IVA reflects the basic hydrogeological characteristics of the study area, thus it is called a “state” indicator. A positive correlation between the groundwater quality parameter level and the IVA index would only occur if the loads were evenly distributed. However, due to the significant differences in geochemical background values and human activities, the even distribution of pollution loads is almost impossible to achieve.
[0142] In this case, it is not reasonable to use groundwater quality parameters as the validation indicator of the IAV model, although NO3 - and TDS are widely used as representative validation methods. Therefore, it is suggested that the correlation between groundwater quality parameters and the IAV model index should not be the only validation standard for the IAV model. In addition, when using this method, it should be ensured that the selected groundwater quality parameters are as spatially uniform as possible. For example, outliers or sampling points close to point source pollution should be excluded.
[0143] In addition, other methods such as numerical simulation, expert judgment, tracer testing, and field investigation should also be incorporated into the validation process of the IAV model.
[0144] 3.2 Model performance and spatial prediction of fluoride concentration (SPFC)
[0145] The optimal hyperparameters obtained after tuning using the TPE algorithm are shown in Table 2, under which the RMSE reaches the minimum value, and the model performs best. The model performs very well in fitting the groundwater samples in the training set (R 2 = 0.9180, RMSE = 0.0582, MAE = 0.0309, EV = 0.9202) Figure 5 b), and the performance on the validation set is slightly lower but still very high (R 2 = 0.7579, RMSE = 0.0748, MAE = 0.0467, EV = 0.7581) Figure 5 c). Figure 5 d and Figure 5 e show the comparison of the standardized actual values (blue dots) and predicted values (red dots) of each groundwater sample, which further verifies the high fitting degree and performance of the model (including the training set and validation set of groundwater samples).
[0146] Appendix Figure 5To illustrate SPFC and model performance evaluation, (a) SPFC map; (b) model performance in the training GW sample (normalized values), 95% confidence interval in red shade; (c) model performance for the validation GW sample (normalized values), orange shade indicates 95% confidence interval); (d) comparison of normalized actual and predicted F-concentration for the training GW sample (blue for actual values, red for predicted values); (e) comparison of normalized actual and predicted F-concentration for the validation GW sample (blue for actual values, red for predicted values).
[0147] The model not only runs efficiently with limited time consumption, but also outperforms other machine learning regression models in groundwater quality research. In addition, the advantage of LGBM model in dealing with multi-dimensional indicators (a total of 26 indicators) is also very obvious. Compared with the research using classification algorithm for SPFC, the LGBM regression model used in this invention effectively solves the problem of unbalanced data set. Given that there are currently fewer studies using regression advanced ensemble learning models for SPFC, it is recommended to develop more advanced and hybrid models for this field. This not only helps to further improve the performance of the model, but also promotes the comparative study between regression models and classification models.
[0148] Figure 5 The red area in a represents the high fluoride concentration area (>1 mg / L). It can be observed that these areas present a patchy aggregation pattern, accounting for about 17.48% of the total area. The blue area represents the medium fluoride concentration area (0.5 to 1 mg / L), which is usually distributed around the high fluoride concentration area (accounting for about 39.05% of the total area). This indicates that the fluoride concentration is distributed in a gradient overall, which is consistent with the hydrogeological conditions of the study area, which is dominated by porous aquifer. The remaining area is the low fluoride concentration area (<0.5 mg / L), accounting for about 43.47% of the total area.
[0149] An important finding is that the SPFC generated by the LGBM model in this invention is quite different from the smooth curve and aggregated pattern observed in the interpolation results. On the contrary, due to the model running once based on the optimal parameters of each grid cell to predict fluoride concentration, certain grid cells show granularization and local dispersion distribution. This method integrates the specific index conditions of hydrogeology, climate, and human activities at each grid location. Therefore, granularization pattern is also observed in many studies that use these indicators and machine learning models for spatial prediction of groundwater quality or contaminant concentration. This method effectively solves the problem of spatial heterogeneity that cannot be achieved by interpolation techniques.
[0150] Another important finding is that the LGBM model accurately predicts the fluoride concentration in the training set and validation set (R 2 = 0.9180 and 0.7579), which is much higher than the interpolation results (R Figure 4The results of the correlation analysis based on IVA in b (almost no correlation) form a stark contrast. This indicates that groundwater quality parameters are influenced by a variety of factors, such as hydrogeology, topography, climate, and human activities. While the IAV model takes into account some of these indicators, it does not adequately consider human activities and other factors such as NDVI, TWI, and river network. This further supports the point mentioned in Section 3.1 that it is unreasonable to use groundwater quality parameters as the only validation indicator for the IAV model. Instead, multiple validation methods should be considered in combination.
[0151] In addition, it is also observed that some studies use machine learning models to use groundwater quality parameters as training and validation datasets and generate IAV maps in combination with IAV indicators, which leads to a high correlation between groundwater quality parameters and IAV indices. However, it is important to emphasize that the spatial distribution obtained in this way represents the distribution of a specific pollutant, and renaming it as "vulnerability" can lead to confusion. For example, in the present invention, the map generated using the LGBM model reflects the SPFC Figure 5 a) which is significantly different from the IAV map Figure 4 a). Therefore, it is recommended that future studies distinguish the terminology of IAV and groundwater quality concentration distribution to avoid conceptual confusion.
[0152] 3.3 Indicator analysis, SHAP observation and discussion
[0153] Figure 6 a and Figure 6 b show the SHAP analysis results of the selected indicators and the feature importance analysis based on counting, respectively. It can be observed that the water chemistry indicators, NDVI (normalized vegetation index), elevation, river, TWI (soil moisture index), population density, and GDP have higher importance, while the extraction intensity based on IAV, HDI (human development index), and state indicators have lower importance.
[0154] Among these water chemistry indicators, the high feature values of pH, Na + , NKCL, SAR, and NKCA are associated with high SHAP values, indicating that areas with high values of these indicators tend to have high concentrations of F - . High-fluoride groundwater areas are usually alkaline, as high pH increases the solubility of F - by promoting the decomposition of silicate minerals, while under acidic conditions, F - tends to be adsorbed by clay. In addition, the positive correlation between the water chemistry indicators related to Na + (Na + , NKCL, SAR, NKCA, and CAI) and SHAP values indicates that Na + is associated with F -forms high solubility NaF, thus increasing the concentration of F - in groundwater. Na + may also ion exchange with other cations (such as Ca 2+ and Mg 2+ ), further promoting the release of F - and indirectly increasing the concentration of F - By examining the water chemistry data, it was found that the Na / Cl ratio was significantly greater than 1, and most groundwater samples showed negative CAI-1 and CAI-2 values, indicating that cation exchange did indeed occur, thus increasing the concentration of F - Meanwhile, indicators related to Ca 2+ were negatively correlated with SHAP values and feature values, because Ca 2+ combined with F - to form insoluble CaF2, inhibiting the dissolution of F - and thus reducing the concentration of F - in groundwater. Therefore, in addition to commonly used Piper and Gibbs diagrams, SHAP analysis visualization can also serve as supplementary evidence for regional water chemistry analysis, especially when focusing on specific pollution ions in spatial analysis.
[0155] Among the indicators identified based on the PSR framework, the indicators of population density and GDP showed higher importance, while the seven IAV model indicators designed as state indicators and the importance of extraction intensity were lower Figure 7 b). In combination with SHAP analysis, the feature values of high population density and GDP were respectively associated with high and low F - concentration. A reasonable speculation is that human activities (especially point source pollution such as industry, landfills, and mining) are the main sources of F - pollution in the study area. Although these high-GDP areas are usually accompanied by high population density and numerous pollution sources, their high GDP levels enable them to implement effective management measures, which may have reduced the concentration of F - Conversely, areas with lower GDP may not have implemented such measures, resulting in higher levels of F - pollution. Therefore, accurately identifying pollution sources is a priority for addressing F - pollution.
[0156] The indicators of the IAV model showed low importance, partly due to the assumption of uniform distribution of contaminants, which does not hold in this study as mentioned in Section 3.1. On the other hand, DRASTIC, as a source vulnerability model based on the "source-path- target" framework, cannot effectively simulate the release of contaminants from minerals. Moreover, selecting groundwater quality indicators from the perspective of PSR can provide a comprehensive understanding of how human activities and natural conditions interact with groundwater. Although this framework has been widely used in many fields, its application in groundwater quality research is still limited. This study combined with SHAP analysis, demonstrated how the selected indicators contributed to the model and inferred the factors that affected pollution. However, current groundwater quality mapping studies usually select indicators based on groundwater quality parameters and environmental conditions, with insufficient consideration of IAV and human activity impacts. This is insufficient because human activities-induced pollution is a major factor of poor groundwater quality, while IAV reflects the inherent vulnerability of the region to pollution. Both are essential indicators in groundwater quality assessment.
[0157] Another important finding is that terrain-related indicators, including elevation, TWI, and distance to rivers, showed high importance in SPFC. These indicators mainly affect F - concentration by influencing the dissolution of fluorine-containing minerals. Low-elevation areas usually accumulate groundwater, increasing the contact time with fluorine-containing minerals, leading to higher F - concentration. In addition, high TWI values indicate areas with higher soil moisture and water accumulation, while areas close to rivers can increase groundwater recharge. These conditions promote the dissolution of fluorine-containing minerals, thereby increasing the F - concentration in groundwater. The results show that although IAV models like DRASTIC cannot accurately identify F - release from minerals, the introduction of additional indicators (such as TWI, rivers, and NDVI) can effectively compensate for this limitation. However, the importance of NDVI is still unclear, which may be related to human activities or vegetation types, so it is difficult to determine the exact reason. For F - pollution with dual sources (natural and human activities), introducing new indicators to improve modeling accuracy is a feasible option.
[0158] 3.4 Groundwater Fluoride Management Considering Groundwater Reserves (GWS) and Intrinsic Aquifer Vulnerability (IVA)
[0159] Most existing studies focus on achieving accurate spatial distribution of fluoride concentration, optimizing model performance. However, as pointed out by some recent studies, sustainable groundwater management is the ultimate goal of groundwater research. Under the framework of the United Nations Sustainable Development Goals (SDGs), there is now a need to integrate various aspects of groundwater research to collectively serve sustainable groundwater management. However, the application of groundwater quality distribution to sustainable groundwater management remains a long-standing challenge. This invention attempts to address this issue by designing a matrix framework that combines GWS and IAV, which identifies seven zones requiring different management intensities, including repair, monitoring, utilization, and protection zones Figure 7 e).
[0160] This framework not only applies to the management of groundwater fluoride, but can also be applied to other water quality parameters or overall water quality indices, such as nitrate (NO3 - ), sulfate (SO4 2- ), heavy metals, and water quality index (WQI). These different management intensities support the efficient allocation of funds and resources, providing valuable guidance for groundwater management decision-makers to support future sustainable groundwater management.
[0161] This example takes the Ordos Basin as an example, uses the LGBM-SHAP regression model, and combines multiple indicators to perform spatial prediction of fluoride concentration (SPFC). Through SHAP analysis (an explainable machine learning technique, EMLT), visualization of the contribution of each indicator in the model is achieved. In addition, a matrix-based sustainable groundwater management framework is developed, integrating IAV, GWS, and SPFC, and identifying seven zones with different groundwater management intensities. The main findings of the study include the following:
[0162] 1. A more accurate IVA map is generated for the Ordos Basin. However, correlation analysis shows that no single common groundwater quality parameter shows strong correlation, highlighting the need to validate the IAV model using multiple methods.
[0163] 2. Through the TPE optimization method, the LGBM model performs well in SPFC. The predicted F - concentration is highly consistent with the observed values, and the granular and locally dispersed distribution emphasizes the difference from interpolation methods.
[0164] 3. In addition to water chemistry indicators, the importance of population density and GDP is relatively high, while the correlation of DRASTIC indicators is low. This suggests that point source pollution may be the main source of F - contamination in the study area, and highlights the limitations of the DRASTIC model based on the "source-path-target" framework in predicting natural sources of F -Limitations of the aspects. Additional indicators such as NDVI, elevation, TWI, and distance to rivers significantly improved model performance.
[0165] 4. The SHAP analysis not only visualizes the contribution of indicators to the model, but also supports the F - Chemical analysis of contaminated water and pollution source inference.
[0166] The invention is not limited in its application to the details of construction and the arrangement of components set forth in the description or illustrated in the drawings. The invention is capable of other embodiments and of being practiced or being carried out in various ways.
[0167] Embodiments of the invention are presented for the purpose of illustration and description and not limitation. Numerous modifications and variations are apparent to those skilled in the art in light of this disclosure. The embodiments are chosen and described in order to best explain the principles of the invention and its practical application to thereby enable others skilled in the art to best utilize the invention in various embodiments and with various modifications as are suited to the particular use contemplated.
Claims
1. A groundwater fluoride spatial prediction method based on LGBM-SHAP model, characterized in that: The method steps are:
1. LGBM model construction, calibration and verification; 2. index importance analysis by SHAP; 3. combined with inherent aquifer vulnerability (IAV) and groundwater storage (GSW) to develop sustainable groundwater management strategies for fluorine pollution.
2. The groundwater fluoride spatial prediction method based on the LGBM-SHAP model according to claim 1, characterized in that: The method also includes index selection, i.e. 26 indicators in 5 categories are adopted, including pressure, state, response, water chemistry, which considers both natural and human sources of fluorine pollution in groundwater.
3. The groundwater fluoride spatial prediction method based on the LGBM-SHAP model according to claim 1, characterized in that: In step one, data preprocessing, IAV weight determination, LGBM model, hyperparameter space and optimization are included.
4. The groundwater fluoride spatial prediction method based on the LGBM-SHAP model according to claim 3, characterized in that: In step one, data preprocessing includes encoding of categorical variables, standardization of indicators, and replacing F- concentrations below 0.1 mg / L (detection limit) with a smaller value (0.05 mg / L).
5. The groundwater fluoride spatial prediction method based on the LGBM-SHAP model according to claim 3, characterized in that: In step one, the LGBM model includes two innovations: gradient-based one-sided sampling (GOSS) and exclusive feature bundling (EFB); GOSS improves training efficiency by prioritizing samples with larger gradients, allowing the model to focus on optimizing important data points with maximum information gain; that is, the first a% of instances with the largest gradient are initially grouped into subset A, and for the remaining sample set B containing lower gradients, a new subset B' is created by randomly sampling b% of instances from B, then the variance gain IG of data instances is calculated, which considers both subsets, and the optimal split criterion is determined using its weighted gradient; Where: N is the total number of samples; gi is the gradient of the loss function with respect to instance iii; A and B are subsets of the sample set; A' and B' are subsets obtained by splitting A and B; is a coefficient of gradient normalization; and are the number of samples of the left and right nodes after splitting on feature j using decision variable d. EFB (exclusive feature bundling) effectively reduces the number of features by grouping features in sparse space, it considers this problem as a graph coloring problem, combining those rarely active features into a feature group, this method optimizes computational efficiency without affecting model accuracy; In addition, the LGBM model uses a leaf node-based tree growth strategy, which splits nodes based on maximum information gain, with the mathematical expression as follows: Where: G L and G R are the total sum of gradients in the left and right nodes, respectively; H L and H R are the sum of the second order gradients (Hessian values) in the left and right nodes, respectively; λ is a regularization parameter that controls the complexity of the tree; γ is a regularization term for leaf nodes to prevent overfitting.
Citation Information
Patent Citations
Gradient boosting decision tree model construction optimization method based on dynamic sampling
CN113537497A
Method for evaluating vulnerability of organic pollutants in underground water
CN114974458A
Method for evaluating vulnerability of regional groundwater based on GA-BP
CN116662760A
Explanatable landslide surface displacement prediction method based on LightGBM and SHAP
CN118536032A