Well field aquifer hydraulic connection analysis method based on system clustering and Bayesian discrimination
By combining systematic clustering with Bayesian discriminant analysis, this method solves the problem of strong subjectivity in traditional water chemistry analysis methods, realizes automated and quantitative analysis of hydraulic connections, improves the accuracy and efficiency of water source identification, and is suitable for rapid emergency decision-making in mine water hazard prevention and control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-27
AI Technical Summary
Traditional hydrochemical analysis methods are highly subjective when processing high-dimensional, large-sample hydrochemical data, making it difficult to objectively, quickly, and accurately identify mixed water sources in different aquifers. Furthermore, existing cluster analysis methods lack quantitative discrimination criteria, failing to meet the needs of rapid emergency decision-making in areas such as mine water hazard prevention and control.
By combining hierarchical clustering and Bayesian discriminant analysis, a standardized dataset is generated through the detection of water chemical indicators. Dissimilarity measurement and cluster analysis are then performed to construct a Bayesian discriminant model. A set of discriminant functions is used to quantitatively distinguish water samples, thereby achieving automated and quantitative analysis of hydraulic connections.
It enables the objective, quantitative, and automated analysis of hydraulic connections, improving the accuracy and efficiency of water source identification. It can quickly identify mixed water sources and is suitable for dynamic monitoring of mine hydrology and emergency decision-making in case of water inrush.
Smart Images

Figure CN121744047A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of mine water disaster prevention, and particularly relates to a minefield aquifer hydraulic connection analysis method based on system clustering and Bayesian discrimination. BACKGROUND
[0002] Identifying the hydraulic connection between aquifers is an important basis for groundwater resource evaluation, mine water disaster prevention and ecological environment protection. Traditional methods mainly rely on hydrogeological tests such as pumping test and tracer test. Although the results of these methods are relatively reliable, they generally have limitations such as high cost, long cycle, limited monitoring range, etc. In recent years, the water chemical method based on the principle of "hydrogeochemical fingerprint" has gradually become an economical and effective supplement or alternative means. This method traces the water flow path and source by analyzing the ion composition and content in water. However, traditional water chemical analysis methods such as Piper tri-graph and Gibbs diagram mainly rely on the experience of researchers for qualitative or semi-quantitative interpretation. The analysis process is highly subjective, and it is difficult to objectively and efficiently process high-dimensional and large-sample water chemical data, especially in identifying mixed water sources of different aquifers.
[0003] Although existing research has tried to introduce multivariate statistical methods such as cluster analysis to automatically classify water samples, to some extent, it has reduced subjectivity, but this method still has obvious defects: as an unsupervised learning method, the classification result of cluster analysis itself lacks quantitative discrimination criteria and prediction ability, and cannot quickly and accurately trace the source of unknown water samples. Therefore, the existing technology has not yet formed a complete and automated analysis process from data-driven objective classification to the establishment of a quantitative discrimination model, and finally to the accurate prediction of new samples, which restricts the in-depth application of the water chemical method in mine water disaster prevention and other scenarios that require rapid emergency decision-making. SUMMARY
[0004] To solve the above technical problems, the application provides a minefield aquifer hydraulic connection analysis method based on system clustering and Bayesian discrimination, which comprises:
[0005] Based on the water chemical index detection data of different aquifers in the target area, an original data matrix is generated, and the original data matrix is standardized to obtain a standardized data set;
[0006] Based on the standardized data set, the dissimilarity between water samples is calculated to construct a dissimilarity matrix, and system clustering analysis is performed according to the dissimilarity matrix to generate a cluster tree diagram to determine the water chemical type group division of the water samples;
[0007] The water chemical type groups are taken as prior classification labels, and the normalized water chemical indexes are taken as discriminant variables to construct a Bayesian discriminant model, so as to obtain a set of discriminant functions corresponding to each water chemical type group;
[0008] The known classified samples are back-substituted and discriminated by using the set of discriminant functions, and the Bayesian discriminant model is verified and optimized according to the discriminant results.
[0009] The normalized data of the water sample to be discriminated are input into the optimized Bayesian discriminant model, the posterior probability of belonging to each water chemical type group is calculated, and the category to which the water sample belongs is determined according to the maximum posterior probability, so as to analyze the hydraulic connection between different aquifers.
[0010] Optionally, the water chemical indexes at least include the concentration values of K⁺, Na⁺, Ca²⁺, Mg²⁺, Cl⁻, SO4²⁻, HCO3⁻ and CO3²⁻, and the pH value and the total dissolved solids value.
[0011] Optionally, when calculating the dissimilarity measure between water samples, the Euclidean distance or the Mahalanobis distance is used; and the system clustering analysis uses the Ward method, the longest distance method or the class average method.
[0012] Optionally, the general form of the discriminant function is:
[0013] F k = C k +Σⱼ (W k ⱼ·Xⱼ);
[0014] Wherein, F k is the discriminant score of the sample belonging to the kth category, C k is the constant term of the kth category, W k is the discriminant coefficient of the kth category on the jth water chemical index, and Xⱼ is the normalized value of the jth water chemical index.
[0015] Optionally, after determining the category to which the water sample belongs according to the maximum posterior probability, if the water samples from different position aquifers are continuously and with a posterior probability higher than 85% discriminated as the same water chemical type group, it is determined that there is a close hydraulic connection between them.
[0016] Optionally, after determining the category to which the water sample belongs according to the maximum posterior probability, if the water samples are discriminated as belonging to multiple water chemical type groups with similar and no significant difference in probability, it is determined that the water sample is a mixed water sample, indicating that there is hydraulic exchange between different aquifers.
[0017] Optionally, the model verification includes calculating the discriminant accuracy and the misjudgment rate, and adjusting and optimizing the discriminant variables or the prior probability based on the verification results.
[0018] In another aspect, the present application also provides an electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the method when executing the computing program.
[0019] In another aspect, the present application also provides a computer readable storage medium storing a computer program, wherein the computer program is executable on a processor to implement the method.
[0020] Compared with the prior art, the present application has the following advantages and technical effects:
[0021] The present application realizes the objectification, quantification and automation of the hydraulic connection analysis by combining system clustering with Bayesian discrimination. The method classifies water samples in an objective manner in a data-driven manner, and constructs a discriminant model that can be quantitatively predicted, thereby significantly reducing the dependence on subjective experience of the traditional method. The model can output a probabilistic discrimination result, which not only improves the accuracy of water source identification, but also effectively identifies and quantifies mixed water sources, solving the deficiency of the traditional graphic method in this regard. Once the model is established, unknown water samples can be quickly and batch discriminated, which is low in cost and high in efficiency, and is especially suitable for mine hydrological dynamic monitoring and rapid identification of emergency water sources, and has strong practicability and popularization value. BRIEF DESCRIPTION OF DRAWINGS
[0022] The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The embodiments of this application and of the drawings illustrate to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:
[0023] Fig. 1 The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The embodiments of this application and of the drawings illustrate to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:
[0024] Fig. 2 The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The embodiments of this application and of the drawings illustrate to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:
[0025] Fig. 3 The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The embodiments of this application and of the drawings illustrate to explain the present application, and do not constitute an improper limitation on the present application. In the drawings: DETAILED DESCRIPTION
[0026] It should be noted that the embodiments and features in the present application can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and embodiments.
[0027] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in a different order.
[0028] Example 1
[0029] like Figs. 1-3 As shown, this embodiment provides a method for analyzing the hydraulic connection of aquifers in well fields based on system clustering and Bayesian discrimination, including:
[0030] Based on the detection data of water chemical indicators of different aquifers in the target area, an original data matrix is generated, and the original data matrix is standardized to obtain a standardized dataset.
[0031] Based on the standardized dataset, the dissimilarity measure between water samples is calculated to construct a dissimilarity matrix, and a hierarchical cluster analysis is performed based on the dissimilarity matrix to generate a cluster dendrogram to determine the water sample hydrochemical type group classification.
[0032] Using the water chemistry type groups as prior classification labels and standardized water chemistry indicators as discriminant variables, a Bayesian discriminant model is constructed to obtain a set of discriminant functions corresponding to each water chemistry type group.
[0033] The discriminant function set is used to perform back-substitution discrimination on samples with known classifications, and the Bayesian discriminant model is verified and optimized based on the discrimination results;
[0034] The standardized data of the water sample to be judged is input into the optimized Bayesian discriminant model to calculate the posterior probability of its belonging to each hydrochemical type group, and the category is determined according to the maximum posterior probability, thereby analyzing the hydraulic connection between different aquifers.
[0035] 1. Data collection and standardization processing:
[0036] The system collected clean water samples representing different aquifers (such as porous aquifers, fractured aquifers, and karst aquifers) within the study area. Various hydrochemical indicators were tested in the laboratory, including major cations (K⁺, Na⁺, Ca²⁺, Mg²⁺), major anions (Cl⁻, SO₄²⁻, HCO₃⁻, CO₃²⁻), pH value, and TDS value, forming a raw "sample-indicator" data matrix. Due to the significant differences in the dimensions and orders of magnitude of the indicators, to ensure the fairness of subsequent analyses, the raw data was processed using the Z-score standardization method, transforming it into standard normally distributed data with a mean of 0 and a standard deviation of 1.
[0037] 2. Q-type hierarchical cluster analysis:
[0038] Based on the observed data of the samples, the distance coefficient, a commonly used statistic in Q-type clustering analysis, was used to classify the samples. The principle is as follows:
[0039] If we consider N samples as N points in a P-dimensional space, the similarity between two samples can be measured by the distance between the two points in the P-dimensional space. This distance is actually the Mahalanobis distance. However, if we standardize the variables... If they are uncorrelated, then the Mahalanobis distance is the same as the usual Euclidean distance. In this case, the samples... and The distance between them is:
[0040] ;
[0041] Where p is the number of water chemistry indicators, and xaj and xak represent the standardized values of the j-th and k-th samples on the a-th indicator, respectively.
[0042] After calculating the distances between all pairs of samples, they can be arranged into a distance coefficient matrix:
[0043] ;
[0044] in It is a real symmetric matrix, so we only need to calculate the upper triangle part. Based on D, we can classify the N points and obtain the grouping graph.
[0045] This step aims to classify water samples into different hydrochemical types in a data-driven, unsupervised manner, with each type initially corresponding to a water source or a hydraulic system. First, the Euclidean distance (when the correlation between indicators is weak) or Mahalanobis distance (when the correlation between indicators is strong) between all sample points in the standardized data is calculated to construct a dissimilarity matrix. Second, Ward's method is used as the connectivity criterion for hierarchical clustering. This method aims to minimize the sum of squared deviations within each cluster, easily forming clusters with compact intra-cluster samples and high inter-cluster separation, with clear geological significance. Finally, a cluster dendrogram is generated and analyzed. Based on the abrupt change points (i.e., nodes where the distance coefficient increases significantly) on the vertical axis (distance) of the dendrogram, and combined with the hydrogeological conceptual model of the study area, the optimal number of classifications (k) is comprehensively determined, thereby objectively dividing all water samples into k hydrochemical type groups.
[0046] 3. Construction of Bayesian discriminant model:
[0047] Based on the preliminary classification obtained from cluster analysis, a more accurate discriminant model is established to quickly and quantitatively classify new samples.
[0048] Based on the fundamental principles of Bayesian discrimination, the source attribution of water samples in the study area is determined using the following calculation steps:
[0049] ①Calculation This means that the water sample is determined to belong to a certain water source without calculating the water ion content. In this case, the probability value of the water sample belonging to each water source is the same.
[0050] ;
[0051] ②Calculation Here, the distance method is used, which calculates the value by taking the reciprocal of the absolute value of the distance between the detected water ion value and the standard value.
[0052] ;
[0053] In the formula, , ;
[0054] ③Calculation Calculate according to formula (2.3.4-18);
[0055] ④ Calculate the probability that a water sample belongs to a water source when considering multiple water ion combinations, where... Water ions The weight.
[0056] ;
[0057] ⑤ Determine the water sample ownership with the highest probability.
[0058] ;
[0059] Using the k water chemistry type groups obtained in step 2 as known classification labels and the standardized water chemistry indicators as independent variables, a supervised Bayesian discriminant model is constructed. This model will generate a linear discriminant function for each water chemistry type group, with the form: F k =C k +W k1 X1+W k2 X2+ ... +W kp X p Among them, C k W is the constant term of the k-th group. k Let ⱼ be the discriminant coefficient of the k-th group on the j-th indicator, and Xⱼ be the standardized indicator value. For a new sample, its data are substituted into k discriminant functions to calculate k discriminant scores F. k .
[0060] 4. Model validation and optimization;
[0061] To evaluate the reliability of the established discriminant model, internal validation is required. Training samples with known classifications are substituted back into the discriminant function to calculate their discrimination scores, and then reclassified according to the maximum score principle. The overall discrimination accuracy and cross-validation accuracy of the model are calculated by comparing the original classifications with the discriminant classifications. Typically, the discrimination accuracy of the training set is required to be higher than 85%. If the misclassification rate is too high, the rationality of the clustering grouping needs to be checked, or dimensionality reduction (such as using principal component analysis - PCA) should be considered to optimize the model.
[0062] 5. Identification of unknown water samples and hydraulic relationship analysis;
[0063] For newly collected water samples from unknown sources (such as samples from sudden water inrush points or mixed samples from observation wells), the same hydrochemical testing and standardization processes are first performed. Then, the data are substituted into a validated Bayesian discriminant model to calculate the posterior probability (and discriminant score F) of belonging to each hydrochemical type group. k Correspondingly).
[0064] Based on the discrimination results, hydraulic connection is determined:
[0065] (1) Strong hydraulic connection: If an unknown water sample located in area A is identified with a high probability (>85%) as belonging to the same hydrochemical type group as the aquifer in area B, it strongly indicates that there is a close hydraulic connection between the aquifers in area A and area B.
[0066] (2) Mixed water source indication: If a water sample is judged to belong to two or more groups with similar probabilities and no absolutely dominant type, it indicates that the water sample is a mixture of water from different aquifers, which directly proves that there is a hydraulic exchange channel between these aquifers.
[0067] (3) Hydraulic isolation evidence: If water samples from different tectonic units are clearly identified as different hydrochemical types with a high probability of identification, then geochemical evidence is provided that there is a hydraulic barrier (such as a fault or dike) between them.
[0068] The beneficial effects of this embodiment are:
[0069] Objectivity and automation: The entire process from classification to judgment is data-driven, which greatly reduces the interference of subjective experience and judgment. The analysis process is standardized and highly repeatable.
[0070] Quantification and Precision: Bayesian discrimination provides probabilistic output results, making the judgment of hydraulic connection no longer a simple "yes" or "no", but "how likely" to be true, making decision support more scientific and precise.
[0071] Powerful mixed water identification capability: It can effectively identify and quantify the proportion of mixed water sources, which is difficult to achieve with traditional graphic methods, providing a key tool for identifying complex hydrogeological processes such as cross-flow replenishment.
[0072] High efficiency and economy: Once the model is established, it can quickly and cost-effectively determine the source of a large number of subsequent water samples. It is especially suitable for daily hydrological monitoring and emergency decision-making for water inrush in production mines, and has high value for promotion and application.
[0073] On the other hand, this embodiment also provides an electronic device, including a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the method when executing the computing program.
[0074] On the other hand, this embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method.
[0075] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for analyzing the hydraulic connection of aquifers in well fields based on system clustering and Bayesian discrimination, characterized in that, include: Based on the detection data of water chemical indicators of different aquifers in the target area, an original data matrix is generated, and the original data matrix is standardized to obtain a standardized dataset. Based on the standardized dataset, the dissimilarity measure between water samples is calculated to construct a dissimilarity matrix, and a hierarchical cluster analysis is performed based on the dissimilarity matrix to generate a cluster dendrogram to determine the water sample hydrochemical type group classification. Using the water chemistry type groups as prior classification labels and standardized water chemistry indicators as discriminant variables, a Bayesian discriminant model is constructed to obtain a set of discriminant functions corresponding to each water chemistry type group. The discriminant function set is used to perform back-substitution discrimination on samples with known classifications, and the Bayesian discriminant model is verified and optimized based on the discrimination results; The standardized data of the water sample to be judged is input into the optimized Bayesian discriminant model to calculate the posterior probability of its belonging to each hydrochemical type group, and the category is determined according to the maximum posterior probability, thereby analyzing the hydraulic connection between different aquifers.
2. The method according to claim 1, characterized in that, The water chemistry indicators include at least the concentrations of K⁺, Na⁺, Ca²⁺, Mg²⁺, Cl⁻, SO₄²⁻, HCO₃⁻, and CO₃²⁻, as well as the pH value and total dissolved solids value.
3. The method according to claim 1, characterized in that, When calculating the dissimilarity measure between water samples, Euclidean distance or Mahalanobis distance is used; the system cluster analysis adopts the Ward method, the longest distance method or the class average method.
4. The method according to claim 1, characterized in that, The general form of the discriminant function is: F k = C k +Σⱼ (W k ⱼ·Xⱼ); Among them, F k C is the discrimination score for a sample belonging to the k-th class. k W is the constant term of the k-th class. k Xⱼ is the discrimination coefficient of the kth category on the jth water chemical index, and Xⱼ is the standardized value of the jth water chemical index.
5. The method according to claim 1, characterized in that, After determining the category of a water sample based on the maximum posterior probability, if water samples from aquifers at different locations are consistently classified as belonging to the same hydrochemical type group with a posterior probability higher than 85%, then it is determined that there is a close hydraulic relationship between them.
6. The method according to claim 1, characterized in that, After determining the category of a water sample based on the maximum a posteriori probability, if the probability of a water sample being classified as belonging to multiple hydrochemical type groups is similar and there is no significant difference, it is determined to be a mixed water sample, indicating that there is hydraulic exchange between different aquifers.
7. The method according to claim 1, characterized in that, The model validation includes calculating the discrimination accuracy and the false positive rate, and adjusting and optimizing the discrimination variables or prior probabilities based on the validation results.
8. An electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, characterized in that, When the processor executes the computing program, it implements the method of any one of claims 1-7.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1-7.