Karst region karst pipeline identification method based on machine learning and groundwater multi-isotope tracing
By combining multi-isotope data and machine learning models, the problems of large data volume and lack of analysis methods in the identification of karst pipelines in karst areas were solved, and the accurate identification of karst pipelines and the prediction of pollutant migration paths were achieved, providing a scientific basis for water resources management.
Patent Information
- Application Number
- CN202510802194.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies for identifying karst pipelines in karst areas suffer from large amounts of data and lack effective analysis methods, resulting in inaccurate pollution control measures and reduced effectiveness of environmental policy implementation.
Combining multi-isotope data and machine learning models, we obtain groundwater flow models and hydrogeological characteristics, train quantitative models, adjust isotope feature weights, and identify karst conduits.
It has achieved accurate identification of karst pipelines in karst areas, provided a scientific basis, and enabled fast and effective data processing capabilities for water resources management and pollution prevention and control.
Smart Images

Figure CN120705543A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of groundwater basin pollution analysis, and in particular to a karst pipeline identification method in karst areas based on machine learning and groundwater multi-isotope tracing. Background Art
[0002] Sustainable water supply in karst regions is one of the major challenges facing human society today. Groundwater provides drinking water for 25% of the world's population. However, pollutants can secretly migrate into groundwater resources through karst conduits, causing rapid contamination and concentrated discharge. This poses a serious threat to public health and ecological safety. The complex karst conduits and ambiguous groundwater migration pathways in karst regions significantly increase the technical and management challenges of water pollution prevention and control.
[0003] In recent years, with the rapid development of geophysical exploration, information technology, and geochemical tracing techniques, these technologies have become important means for exploring groundwater contamination pathways and tracing pollution sources. However, due to the hidden nature and complexity of karst groundwater contamination, previous reliance on single isotopes to track groundwater circulation processes and pollutant migration pathways has resulted in significant errors. Research methods based on multi-isotope monitoring combined with multi-index analysis can more accurately characterize the water cycle characteristics within the complex geological structures of karst areas, thereby providing a scientific basis for identifying groundwater migration processes in karst areas. However, this method has the following disadvantages: the large amount of data without uncertainty verification, the lack of mutual support between data, and the lack of effective analytical methods to accurately process massive amounts of data. These can lead to inaccurate control measures, resulting in reduced implementation of environmental policies and governance effectiveness. Summary of the Invention
[0004] In order to solve the problems existing in the above-mentioned prior art, the purpose of the present invention is to propose a karst pipeline identification method in karst areas based on machine learning and groundwater multi-isotope tracing. By combining the high-precision characterization capability of multi-isotope data and the efficient data processing capability of machine learning models, it is possible to have a deeper understanding of the complex behavior of karst groundwater systems and provide a basis for the scientific management of water resources.
[0005] To achieve the above object, the present invention provides the following solutions:
[0006] A karst conduit identification method in karst areas based on machine learning and groundwater multi-isotope tracing includes:
[0007] Acquire multi-isotope data in a region to be detected, input the multi-isotope data into a groundwater flow model, obtain spatial distribution characteristics of groundwater isotope composition, and use the spatial distribution characteristics to obtain karst conduits in the karst region; the groundwater flow model is obtained by training a quantitative model using a first training set, the first training set including: multi-isotope initial data;
[0008] During the training of the quantitative model, the sensitivity coefficient of the output data of the quantitative model to the multi-isotope initial data is calculated, the weights of the individual isotope characteristics are adjusted, and the model parameters are trained.
[0009] Optionally, acquiring the multi-isotope data in the area to be detected includes:
[0010] Obtaining hydrogeological characteristics of the area to be detected, inputting the hydrogeological characteristics into a random forest model to obtain the multi-isotope data; the random forest model is trained using a second training set, wherein the second training set includes: original hydrogeological characteristics;
[0011] The multiple decision trees of the random forest model are:
[0012]
[0013] Where T is the number of trees, f t (x) is the prediction result of the t-th tree.
[0014] Optionally, constructing the quantitative model includes:
[0015] The multi-isotope initial data is preprocessed to obtain isotope characteristics, and the correlation coefficient between the isotope characteristics and the target variable is calculated. Feature selection is performed based on the correlation coefficient, and the quantitative model is selected based on the isotope characteristics after feature selection and the characteristics of the isotope fractionation structure principle.
[0016] Optionally, obtaining the isotopic signature includes:
[0017] Preprocessing the multi-isotope initial data includes: data cleaning, missing value and outlier processing, data standardization and point data interpolation processing, to generate a spatial distribution map of isotope composition and a spatial distribution map of hydrogeological parameters;
[0018] The isotope characteristics are obtained based on the spatial distribution map of the isotope composition and the spatial distribution map of the hydrogeological parameters.
[0019] Optionally, adjusting the weight of each isotope feature includes:
[0020] The sensitivity coefficient of the output data of the quantitative model to the input data is calculated, and the indicators affecting the environmental process are screened out according to the sensitivity coefficient, and the weights of the various isotope characteristics are adjusted according to the indicators.
[0021] Optionally, calculating the sensitivity coefficient of the output data of the quantitative model to the input data includes:
[0022]
[0023] Among them, S is the sensitivity coefficient, which indicates the sensitivity of the model output t to the change of the input parameter i. The larger the absolute value of S, the more sensitive the model output is to the change of the input parameter. t is the output variable of the model, and i is the input parameter of the model.
[0024] Optionally, screening out the indicators affecting the environmental process includes:
[0025]
[0026] Among them, φ i is the Shapley value, which represents the contribution of each feature to the model prediction result. F is the set of all features, and S is a subset of F, which does not include feature x. i The feature combination of x, f(S) represents the predicted value of the model when the feature subset S is used, and f(S∪{i}) represents the addition of feature x i The predicted value of the posterior model, |S| is the size of the subset S, |F| is the number of all features, and ! represents the factorial.
[0027] Optionally, training the model parameters includes:
[0028] The multi-isotope initial data is input into the quantitative model to obtain a prediction result. Based on the prediction result, the isotope composition or the migration path of the surface water body is obtained. The migration path of the unknown water sample is validated by k-fold cross validation to train the model parameters.
[0029] The beneficial effects of the present invention are:
[0030] The present invention uses machine learning methods to quickly process data sets containing a large number of features and identify the dominant indicators that contribute most to the target variable.
[0031] The present invention uses machine learning models combined with multi-isotope data to quickly identify pollution sources and predict the migration paths of pollutants, providing a scientific basis for pollution prevention and control.
[0032] The present invention combines isotope data with machine learning models to more accurately depict the complex structure and water circulation process of the karst groundwater system. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0034] Figure 1 Schematic diagram of water chemical ion composition in the study area of an embodiment of the present invention;
[0035] Figure 2 This is a schematic diagram of HO isotope variation in the study area according to an embodiment of the present invention;
[0036] Figure 3 This is a schematic diagram of the S isotope variation in the study area according to an embodiment of the present invention;
[0037] Figure 4 This is a schematic diagram of Sr isotope variation in the study area according to an embodiment of the present invention;
[0038] Figure 5 This is a schematic diagram of the changes in phosphate oxygen isotopes in the study area according to an embodiment of the present invention;
[0039] Figure 6 This is a schematic diagram of groundwater migration paths ascertained based on hydrochemical and isotopic information according to an embodiment of the present invention;
[0040] Figure 7 A schematic diagram of a groundwater verification path based on numerical simulation according to an embodiment of the present invention;
[0041] Figure 8 A schematic diagram of a groundwater transport pipeline and recharge method in a karst area according to an embodiment of the present invention;
[0042] Figure 9 This invention discloses a method for identifying karst pipelines in karst areas based on machine learning and groundwater multi-isotope tracing. DETAILED DESCRIPTION
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0044] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0045] like Figure 9 As shown, this embodiment discloses a method for identifying karst conduits in karst areas based on machine learning and groundwater multi-isotope tracing, including: obtaining multi-isotope data in the area to be detected, inputting the multi-isotope data into a groundwater flow model, obtaining the spatial distribution characteristics of the groundwater isotope composition, and using the spatial distribution characteristics to obtain karst conduits in the karst area; the groundwater flow model is obtained by training a quantitative model using a first training set, the first training set including: multi-isotope initial data; during the training of the quantitative model, calculating the sensitivity coefficient of the output data of the quantitative model to the multi-isotope initial data, adjusting the weight of each isotope feature, and training the model parameters.
[0046] Specifically, this embodiment discloses a method for identifying karst conduits in karst areas based on machine learning and multi-isotope tracing of groundwater. The method comprises: Step S1: systematically collecting data and comprehensively analyzing information on hydrogeological conditions and isotope responses to natural processes. Step S2: data preprocessing. Step S3: constructing a quantitative model using characteristic values of isotope fractionation. Step S4: calculating the sensitivity of the model output to input parameters to optimize the model's prediction accuracy. Step S5: training the quantitative model using machine learning. Step S6: using the quantitative model to predict groundwater migration paths and infer karst conduits in karst areas.
[0047] Furthermore, constructing a quantitative model includes: preprocessing the multi-isotope initial data, obtaining isotope characteristics, calculating the correlation coefficient between the isotope characteristics and the target variable, performing feature selection based on the correlation coefficient, and selecting a quantitative model based on the isotope characteristics after feature selection and the characteristics of the isotope fractionation structure principle.
[0048] Obtaining isotope characteristics includes: preprocessing the initial multi-isotope data, including: data cleaning, missing value and outlier processing, data standardization and point data interpolation processing, generating the spatial distribution map of the isotope composition and the spatial distribution map of the hydrogeological parameters; based on the spatial distribution map of the isotope composition and the spatial distribution map of the hydrogeological parameters, obtaining the isotope characteristics.
[0049] Specifically, in step S1, the system collects data and comprehensively analyzes information on hydrogeological conditions and isotope responses to natural processes.
[0050] The specific method of step S1 is:
[0051] Step S11: Obtain the current hydrogeological conditions of the study area from available sampling points (outcropping springs and rivers), remote sensing, and historical data from environmental stations. Specifically, this includes (1) hydrological, topographic, geomorphological, and monitoring well information; (2) lithology and spatial distribution characteristics of aquifers and aquicludes; and (3) groundwater recharge, runoff, and discharge patterns.
[0052] Arrange the hydrological data of the study area and record it as D hydrology ={m1,m2,…,m n}, where D hydrology is the hydrological dataset, m j represents the jth hydrogeological feature.
[0053] Step S13: Collect marked samples in the study area through drilling, sampling, monitoring, etc.
[0054] Step S14: laboratory analysis to obtain isotope data, for example, using a mass spectrometer to measure the isotope ratio in the water sample.
[0055] Step S15: Collect multi-isotope data of the study area ( 2 H. 18 O. 87 Sr / 86 Sr. 3 H. 3 H- 3 He et al.), denoted as D isotope ={i1,i2,…,i n}, where D isotope For multi-isotope datasets, i m represents the isotopic composition of the mth sampling point.
[0056] Step S2: data preprocessing.
[0057] The specific method of step S2 is:
[0058] Step S21: Data cleaning, check the integrity and consistency of the data, and handle missing values and outliers. For missing values, use the mean to fill in the missing values:
[0059]
[0060] Among them, i missing is the missing value, N is the number of non-missing values, i m represents the isotopic composition of the mth sampling point.
[0061] Step S22: Data standardization: performing standardization processing on each isotope signature (data) i to eliminate dimensional differences:
[0062]
[0063] where i' is the isotope signature (data) after dimensionless differences are removed, μ is the mean of the original isotope dataset, and σ is the standard deviation of the original isotope dataset.
[0064] Step S23, data space, using geographic information system (GIS) software, interpolate the point data into surface data to generate isotope composition spatial distribution map, hydrogeological parameter spatial distribution map, etc. According to Kriging interpolation (Kriging):
[0065]
[0066] in, is the predicted value of the point to be interpolated, z(s i ) is the observed value of the known point, λ i is the weight coefficient.
[0067] Step S3: construct the characteristic value of the isotope to establish a quantitative model: t=F(i).
[0068] The specific process of step S3 includes:
[0069] Step S31: Calculate the correlation coefficient between the characteristic (isotope characteristic i) and the target variable t (the pipeline characteristic of the actual site, etc.) and perform correlation coefficient analysis:
[0070]
[0071] Where r is the correlation coefficient, Cov(x,y) is the covariance of i and t, σ i ,σ t are the standard deviations of i and t respectively.
[0072] Step S32: retain the features whose absolute value of the correlation coefficient is greater than the threshold θ.
[0073] Step S33: Shrink the coefficients of unimportant features to 0, thereby achieving feature selection.
[0074] Step S34: Calculate d-excess based on the structural principle and characteristics of isotope fractionation:
[0075]
[0076] Where d-excess is the spatial distribution characteristic of the isotope, M is the molar mass of the isotope, is the isotope average of the surrounding k neighboring points.
[0077] Step S35: Calculate the isotope average value of k adjacent points around each sampling point:
[0078]
[0079] where i m represents the isotopic composition of the mth sampling point. is the isotope average of k adjacent points around each sampling point.
[0080] Furthermore, adjusting the weights of each isotope characteristic includes: calculating the sensitivity coefficient of the output data of the quantitative model to the input data, screening out indicators that affect the environmental process based on the sensitivity coefficient, and adjusting the weights of each isotope characteristic based on the indicators.
[0081] Specifically, step S4 calculates the sensitivity of the model output to the input parameters and optimizes the prediction accuracy of the model.
[0082] The specific process of step S4 includes:
[0083] S41. Calculate the sensitivity coefficient:
[0084]
[0085] Where S is the sensitivity coefficient, which indicates how sensitive the model output t is to changes in the input parameter i. A larger absolute value of S indicates a more sensitive model output to changes in the input parameter. t is the model's output variable (target variable), and i is the model's input parameter (isotope signature variable).
[0086] S42. Screen out indicators that can respond to environmental processes.
[0087] Calculate φ i value:
[0088]
[0089] Among them, φ i is the Shapley value, which represents the contribution of each feature to the model prediction result. F is the set of all features, and S is a subset of F, which does not include feature x. i The feature combination of f(S) represents the predicted value of the model when the feature subset S is used. f(S∪{i}) represents the addition of feature x i The predicted value of the posterior model, |S| is the size of the subset S, and |F| is the number of all features. φ i A positive value indicates that the feature has a positive contribution to the prediction result, i A negative value indicates that the feature has a negative contribution to the prediction result.
[0090] Specifically, a positive SHAP value indicates that the feature contributes positively to the model's predictions, meaning that the presence or value of the feature increases the model's predictions. A negative SHAP value indicates that the feature contributes negatively to the model's predictions, meaning that the presence or value of the feature decreases the model's predictions.
[0091] S43. Adjust high-contribution features to optimize the model's predictive performance, while simplifying low-contribution features to reduce model computational requirements.
[0092] Furthermore, obtaining multi-isotope data in the area to be detected includes: obtaining hydrogeological characteristics of the area to be detected, inputting the hydrogeological characteristics into a random forest model, and obtaining multi-isotope data; the random forest model is trained using a second training set, and the second training set includes: original hydrogeological characteristics.
[0093] Furthermore, training model parameters includes: inputting multi-isotope initial data into the quantitative model to obtain prediction results, obtaining the isotope composition or migration path of the surface water body based on the prediction results, using k-fold cross-validation to verify the migration path of unknown water samples, and training model parameters.
[0094] Specifically, step S5, machine learning training quantitative model:
[0095] The specific method of step S5 is:
[0096] S51. Construct a quantitative model for machine learning training. The machine learning model may be a random forest model, etc.
[0097] Build a random forest model to process high-dimensional data and reduce data noise. The multiple decision trees of the random forest model are as follows:
[0098]
[0099] Where T is the number of trees, f t (x) is the prediction result of the t-th tree.
[0100] S52. Take the groundwater at each point as the source of the surface river, and determine the initial isotope value of the unknown water sample based on the measured isotope values and the source isotope values of the nearby basins. Solve the predicted value to obtain the isotope composition of the unknown groundwater or the migration path of the surface water body.
[0101] S53, model performance evaluation is performed using k-fold cross validation:
[0102]
[0103] Among them, MSE i is the mean square error of the i-th fold. The model training process includes data set partitioning: the data set is divided into a training set Xtrain and a test set Xtest with a ratio of 7:3, and cross-validation is used to evaluate the accuracy of the model.
[0104] Step S6: The quantitative model predicts the migration path of groundwater, selects points with similar target variables predicted based on isotope characteristics, connects them in spatial distribution, and infers the karst pipeline in the karst area.
[0105] The specific method of step S6 is:
[0106] S61. Construct a groundwater flow model, link different sampling points in space (prioritize points with strong correlation), and predict the spatial distribution characteristics of groundwater isotope composition. The model can be used to predict the source of unknown water samples:
[0107]
[0108] Among them, X new is the feature vector of the new sample, according to the prediction result.
[0109] S62. Establish a groundwater flow model based on numerical simulation to simulate the migration path. Optional flow field models include MODFLOW.
[0110] S63. Compare the flow model predictions and field survey test values with the machine learning prediction results, and calculate the mean square error (MSE) between the simulation results and the prediction results:
[0111]
[0112] If the MSE is within the preset range, the result is available, t m is the actual target variable value at the mth sampling point, If the MSE is within the preset range (<0.01), the result is acceptable.
[0113] An experiment was conducted on groundwater in a karst area affected by leachate from mining waste. Groundwater was sampled, runoff data from the basin where the waste was located was recorded, watershed-scale hydrochemical data and partial isotope data were monitored, and effective isotope indicators were selected through machine learning algorithms to characterize groundwater migration channels. Figure 1 As shown, the groundwater in the study area is characterized by high Ca 2+ -Mg 2+ and SO4 2- -F - The isotope results show that there is an exchange of isotopes between the water in the Wujiang River Basin and CO2, such as Figure 2 As shown in Figure 2, the sulfur isotopes in the water indicate that inorganic sulfide is the main source, such as Figure 3 As shown in Figure 2, strontium isotopes in water are mainly derived from carbonate weathering, such as Figure 4 As shown in Figure 2, the phosphorus isotope information of the spring water sampling points indicates its correlation with the slag filtrate, e.g. Figure 5 As shown in Figure 2, the existing information was integrated to obtain the migration path of groundwater, as shown in Figure 2. Figure 6 As shown in the figure, numerical simulations have shown that the method of combining isotopes with machine learning to judge groundwater migration is effective, as shown in the figure. Figure 7 As shown in the figure, the groundwater transport pipelines and recharge methods in the karst area are as follows: Figure 8 shown.
[0114] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A method for identifying karst conduits in karst areas based on machine learning and groundwater multi-isotope tracing, characterized in that: include: Acquire multi-isotope data in a region to be detected, input the multi-isotope data into a groundwater flow model, obtain spatial distribution characteristics of groundwater isotope composition, and use the spatial distribution characteristics to obtain karst conduits in the karst region; the groundwater flow model is obtained by training a quantitative model using a first training set, the first training set including: multi-isotope initial data; During the training of the quantitative model, the sensitivity coefficient of the output data of the quantitative model to the multi-isotope initial data is calculated, the weights of the individual isotope characteristics are adjusted, and the model parameters are trained.
2. The method for identifying karst conduits in karst areas based on machine learning and groundwater multi-isotope tracing according to claim 1, characterized in that: Acquiring multi-isotope data in the area to be detected includes: Obtaining hydrogeological characteristics of the area to be detected, inputting the hydrogeological characteristics into a random forest model to obtain the multi-isotope data; the random forest model is trained using a second training set, wherein the second training set includes: original hydrogeological characteristics; The multiple decision trees of the random forest model are: Where T is the number of trees, f t (x) is the prediction result of the t-th tree.
3. The method for identifying karst conduits in karst areas based on machine learning and groundwater multi-isotope tracing according to claim 1, characterized in that: Constructing the quantitative model includes: The multi-isotope initial data is preprocessed to obtain isotope characteristics, and the correlation coefficient between the isotope characteristics and the target variable is calculated. Feature selection is performed based on the correlation coefficient, and the quantitative model is selected based on the isotope characteristics after feature selection and the characteristics of the isotope fractionation structure principle.
4. The method for identifying karst conduits in karst areas based on machine learning and groundwater multi-isotope tracing according to claim 3, characterized in that: Obtaining the isotopic signature includes: Preprocessing the multi-isotope initial data includes: data cleaning, missing value and outlier processing, data standardization and point data interpolation processing, to generate a spatial distribution map of isotope composition and a spatial distribution map of hydrogeological parameters; The isotope characteristics are obtained based on the spatial distribution map of the isotope composition and the spatial distribution map of the hydrogeological parameters.
5. The method for identifying karst conduits in karst areas based on machine learning and groundwater multi-isotope tracing according to claim 1, characterized in that: Adjusting the weights of individual isotopic features includes: The sensitivity coefficient of the output data of the quantitative model to the input data is calculated, and the indicators affecting the environmental process are screened out according to the sensitivity coefficient, and the weights of the various isotope characteristics are adjusted according to the indicators.
6. The method for identifying karst conduits in karst areas based on machine learning and groundwater multi-isotope tracing according to claim 5, characterized in that: Calculating the sensitivity coefficient of the output data of the quantitative model to the input data includes: Among them, S is the sensitivity coefficient, which indicates the sensitivity of the model output t to the change of the input parameter i. The larger the absolute value of S, the more sensitive the model output is to the change of the input parameter. t is the output variable of the model, and i is the input parameter of the model.
7. The method for identifying karst conduits in karst areas based on machine learning and groundwater multi-isotope tracing according to claim 5, characterized in that: The indicators that affect the environmental process are screened out, including: Among them, φ i is the Shapley value, which represents the contribution of each feature to the model prediction result. F is the set of all features, and S is a subset of F, which does not include feature x. i The feature combination of x, f(S) represents the predicted value of the model when the feature subset S is used, and f(S∪{i}) represents the addition of feature x i The predicted value of the posterior model, |S| is the size of the subset S, |F| is the number of all features, and ! represents the factorial.
8. The method for identifying karst conduits in karst areas based on machine learning and groundwater multi-isotope tracing according to claim 1, characterized in that: Training the model parameters includes: The multi-isotope initial data is input into the quantitative model to obtain a prediction result. Based on the prediction result, the isotope composition or the migration path of the surface water body is obtained. The migration path of the unknown water sample is validated by k-fold cross validation to train the model parameters.