Water resource monitoring data analysis method based on data mining

Through wavelet threshold denoising, directional Kriging interpolation, improved k-means++ spatiotemporal clustering and Stacking integration model, combined with reverse Darcy flow tracking algorithm and Bayesian hyperparameter optimization, an intelligent decision-making system was built, solving the complexity of data processing and pollution traceability problems in water resource monitoring, and achieving efficient and accurate water quality detection and pollution source positioning.

CN120470039AInactive Publication Date: 2025-08-12兴义民族师范学院
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510516017.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for existing water resource monitoring technologies to obtain accurate, timely and comprehensive water quality detection data in the basin, and traditional methods are complex in operation, high in cost, and are prone to secondary pollution, making it difficult to achieve accurate pollution traceability and efficient detection.

Method used

Wavelet threshold denoising and directional Kriging interpolation are used to improve data quality, combined with improved k-means++ spatiotemporal clustering and Stacking integrated model, precise pollution traceability is achieved through reverse Darcy flow tracking algorithm, and an intelligent decision-making system is built based on Bayesian hyperparameter optimization and knowledge graph.

Benefits of technology

It significantly improves the efficiency of water quality abnormality detection, pollution source positioning accuracy and system adaptability, reduces the difficulty of system construction and management, and achieves efficient water resource monitoring and pollution source positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470039A_ABST
    Figure CN120470039A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of water resource monitoring, in particular to a water resource monitoring data analysis method based on data mining, which improves the data quality through wavelet threshold denoising and directional Kriging interpolation, enhances the prediction precision in combination with improved k-means + + spatial-temporal clustering and a Stacking integrated model, and improves the prediction accuracy. A reverse Darcy flow tracking algorithm is adopted to realize accurate pollution traceability, and an intelligent decision-making system is constructed and formed based on Bayesian hyper-parameter optimization and a knowledge graph, so that the water quality anomaly detection efficiency, the pollution source positioning accuracy and the system adaptive capacity are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water resource monitoring, and in particular to a water resource monitoring data analysis method based on data mining. Background Art

[0002] Water resources, a fundamental and strategic resource, face increasingly severe challenges worldwide. The combined impacts of population growth, economic expansion, increasing water pollution, water environment damage, and climate change have exacerbated the vulnerability of water resources. The "dual hydrological structure" creates a direct connection between groundwater and surface water. Intense corrosion leads to severe surface water infiltration, widespread desertification, and a fragile ecological environment. Coupled with relatively underdeveloped economic and social development, water resource issues are particularly prominent.

[0003] The water resource characteristics of karst mountainous areas are primarily characterized by the close connection between surface and groundwater, the impact of land use changes on water quality, and the complexity of hydrochemical characteristics. Karst areas are highly developed with dissolution fissures and conduits, forming a dual-layer spatial structure of underground and surface water. The underground space becomes a key site for groundwater storage and migration. However, with the degradation of the ecological environment, the self-purification and regulation capacity of the groundwater system has gradually declined, resulting in uneven spatial and temporal distribution of groundwater resources, a decrease in total resource availability, and deterioration in water quality. For example, in the karst mountainous areas of Qianxinan Prefecture, land use patterns have gradually evolved from broad-leaved woodlands to shrublands, shrub grasslands, and cultivated land. Pollutants in surface runoff (such as SO42-, NH4+, PO43-, COD, etc.) have increased significantly, and water quality has shown a downward trend. In addition, the distribution of carbonate rock layers with different chemical properties largely controls the main components, types, and spatial distribution characteristics of hydrochemical characteristics, making water quality in karst areas susceptible to karstification.

[0004] To address the severe challenges facing water resources in Qianxinan Prefecture, in-depth research and development of an online water resource monitoring system and the use of monitoring data to study the temporal and spatial distribution characteristics of water resources are of great practical significance. Water resource monitoring methods include traditional chemical methods, electrochemical methods, atomic spectroscopy, molecular spectroscopy, and biosensors. While traditional chemical methods are mature and reliable, they are complex, time-consuming, and costly, and prone to secondary pollution. Electrochemical methods are pollution-free and require simple equipment, but they suffer from short electrode life and high maintenance costs. Atomic spectroscopy is suitable for detecting trace metal elements, but requires pretreatment and is primarily focused on single-element detection, limiting its application. Biosensors offer high precision and are suitable for real-time online monitoring, but the stability and lifespan of the biosensor elements are technical challenges. Molecular spectroscopy, particularly UV-visible absorption spectroscopy, offers the advantages of rapidity, pollution-free performance, and multi-parameter detection, and is considered a key development direction for online water quality monitoring.

[0005] Extensive research has been conducted both domestically and internationally on water resource monitoring systems. Foreign research primarily focuses on water quality monitoring systems based on Zigbee wireless sensor networks, the integration of remote sensing and GIS, and GPRS modules, enabling real-time monitoring of water quality within a watershed. For example, Zulhani Rasin et al. designed a water quality monitoring system based on a Zigbee wireless sensor network. Through distributed sensor nodes and remote base station monitoring modules, it transmits water quality information in real time, monitoring parameters such as pH, turbidity, and temperature. Domestic research focuses on distributed monitoring systems for large-scale water bodies, such as lakes and river basins, leveraging ARM embedded technology and GPRS wireless network communication to achieve real-time online water quality monitoring. Hohai University has developed a distributed eutrophication water quality early warning online monitoring system for lakes. This system utilizes ARM embedded technology and GPRS wireless network communication to achieve real-time online monitoring of eutrophication-related data parameters. Zhao Guangyu's research team at Zhejiang University has applied distributed database technology to a water quality monitoring system, integrating a variety of water quality sensors on an embedded platform to enable the collection, transmission, and management of water quality information across the watershed. Duan Qichuang's research team at Chongqing University has developed a water quality detection and monitoring system for the Three Gorges Reservoir area, which enables remote monitoring and control of the water quality detection and monitoring process through the Internet.

[0006] In summary, distributed water quality detection and monitoring systems are an important direction for the development of water quality detection technology. Especially when combined with wireless sensor network technology and continuous automatic online detection and monitoring technology, they can break through the limitations of traditional environmental monitoring stations, change the current backward situation where it is difficult to accurately, timely and comprehensively obtain water quality detection and monitoring data within the basin, and reduce the investment cost and management difficulty of system construction. Summary of the Invention

[0007] The purpose of the present invention is to provide a water resources monitoring data analysis method and system based on data mining, improve data quality through wavelet threshold denoising and directional kriging interpolation, enhance prediction accuracy by combining improved k-means++ spatiotemporal clustering and Stacking integrated model, realize accurate pollution source tracing by adopting reverse Darcy flow tracking algorithm, and form an intelligent decision-making system based on Bayesian hyperparameter optimization and knowledge graph construction, which significantly improves the efficiency of water quality anomaly detection, the accuracy of pollution source location and the system adaptability.

[0008] In order to achieve the above technical objectives and the above technical effects, the present invention is implemented through the following technical solutions:

[0009] A water resources monitoring data analysis method based on data mining includes the following steps:

[0010] S1: Data preprocessing and feature construction: A spatiotemporal data cube (dimensions: longitude × latitude × time × spectral parameter) was constructed. The monitoring data were processed through wavelet threshold denoising and directional kriging interpolation. The seasonal trend terms of water quality parameters were extracted using STL decomposition. A lagged feature matrix was constructed to capture the delayed effect of groundwater migration. The Savitzky-Golay derivative transformation was performed on the ultraviolet spectral data.

[0011] S2: Multi-model collaborative analysis: A stacking ensemble model was constructed to handle classification prediction tasks, an improved k-means++ algorithm was used for spatiotemporal clustering analysis, and a SARIMAX model was applied for time series prediction. The improved k-means++ algorithm introduced dynamic time warping distance and spatial constraint terms to calculate intra-class dissimilarity.

[0012] S3: Association rule mining and pollution source tracing: Extracting composite rules through a two-level association rule mining framework, designing spatiotemporal association rule templates to trigger the reverse Darcy flow tracing algorithm to locate pollution sources;

[0013] S4: Model Optimization and Knowledge Management: Use the Bayesian hyperparameter optimization framework to adjust model parameters, identify key driving factors based on Shapley values, and construct a knowledge graph containing the entity relationships between monitoring points, geological units, and pollution sources;

[0014] S5: Validation and Deployment: Implement geographic partition cross-validation to evaluate model performance, achieve lightweight model deployment through channel pruning and state space transformation, and deploy a real-time anomaly detection module at the edge.

[0015] Furthermore, the directional kriging interpolation in step S1 adopts a semivariogram bandwidth α=1.2 km and an anisotropy ratio β=1.5; the lag feature matrix includes three time delay features of t-3, t-7, and t-14 days; the Savitzky-Golay transform is a second-order derivative transform with a window width of 21 nm.

[0016] Furthermore, the base learners of the Stacking ensemble model in step S2 include L2 regularized Logistic regression (λ=0.1), improved ID3 decision tree (information gain ratio threshold>0.3) and 5-layer BP neural network (hidden layer nodes decrease in descending order of 64-32-16-8), and the meta-learner adopts XGBoost (max depth =6, learning rate =0.05).

[0017] Furthermore, the dissimilarity calculation formula of the improved k-means++ algorithm in step S2 is:

[0018] D total =γ·DTW(Ti ,T j )+(1-γ)·Haversine(loc i ,loc j )

[0019] Where γ = 0.6: spatiotemporal weight adjustment factor, reflecting the contribution ratio of dynamic time warping distance to spatial distance. Grid search has verified that this value can balance the spatiotemporal heterogeneity of pipeline flow in karst area hydrological data. DTW (T i , T j ): Dynamic time warping distance, which solves the problem of inconsistent sampling frequencies at monitoring sites. The calculation formula is:

[0020]

[0021] Where W is the set of all legal curved paths; Haversine (loc i ,loc j ): Geospatial distance calculation:

[0022]

[0023] R = 6371 km (average radius of the Earth);

[0024] Furthermore, the spatiotemporal association rule template in step S3 is expressed as:

[0025]

[0026] ΔT = 3d: Time propagation delay, verified by tracer experiments to satisfy:

[0027]

[0028] ΔD=2km: spatial action radius, and crack density ρ fracture Related:

[0029]

[0030] sup=0.12: support threshold, determined by significance test:

[0031] P(sup random ≥0.12)<0.05(permutation test)

[0032] When the rule is activated, the reverse Darcy flow tracking algorithm based on the MODPATH particle tracking model is started with a grid resolution of 100m×100m.

[0033] Furthermore, the Bayesian hyperparameter optimization described in step S4 uses a TPE sampler to iterate 50 rounds, and the objective function is a weighted combination of F1-score (weight 0.6) and inference delay (weight 0.4); the knowledge graph implements a chain query of [sudden event]-[rift zone]-[fertilization area] in the Neo4j graph database.

[0034] Furthermore, the model lightweighting in step S5 includes: implementing a sensitivity threshold of 1×10 -4 Channel pruning is performed to convert the SARIMAX model into a state space form implemented by Kalman filtering; the edge module adopts the TensorFlowLite format, the model size is <2MB and the response time is ≤15ms.

[0035] Furthermore, in step S2, outlier detection uses a joint criterion of the LOF algorithm (k=15 nearest neighbors) and spatial density correction, and an early warning is triggered when the LOF is greater than 2.5 and the density of outliers within the 500m buffer is greater than 30%.

[0036]

[0037] k=15: Determined based on coverage analysis:

[0038]

[0039] B 500m (p): Spatial buffer zone, radius is calculated based on:

[0040]

[0041] Furthermore, in step S4, model drift detection triggers online learning based on a PSI index > 0.25, and the FTRL-Proximal algorithm is used to update the logistic regression weight matrix.

[0042] Furthermore, the verification phase of step S5 adopts a three-element evaluation index system of spatial specific recall rate > 85%, warning lead time ≥ 6 hours, and traceability positioning error < 500m, and uses the TimeSeriesSplit method (n splits =3) Maintain temporal continuity.

[0043] Beneficial effects of the present invention:

[0044] The present invention effectively improves data processing efficiency and integrity through wavelet threshold denoising technology and directional kriging interpolation method. Wavelet threshold denoising technology uses Daubechies basis functions to perform multi-scale analysis and filtering on data, eliminating the noise generated by turbidity and scattering effects in the monitoring data. It not only improves the signal-to-noise ratio of the signal, but also maintains the integrity of the data characteristics, making subsequent analysis more accurate. Directional kriging interpolation interpolates geographic spatial data by introducing spatial variation function bandwidth and anisotropy ratio. Interpolation is performed based on the natural correlation of spatial data, which not only fills the data gaps, but also ensures that the interpolation results are consistent with the real geographic spatial environment, thereby improving the reliability of spatial data.

[0045] The present invention achieves the improvement of classification prediction tasks through the application of Stacking model architecture. The base learner combines L2 regularized Logistic regression, improved ID3 decision tree and BP neural network, which are respectively used to process sparse data, avoid overfitting and identify nonlinear data. In particular, the BP neural network uses the strategy of decreasing the number of hidden layer nodes to improve the generalization ability of the model. The meta-learner uses XGBoost to further reorganize the feature data, fully utilizes the prediction ability of each base learner, and achieves highly accurate classification results. In addition, in the spatiotemporal clustering analysis, the improved k-means++ algorithm is combined with the dynamic time warping distance (DTW) and geographic space constraints, so that the clustering analysis can take into account the heterogeneity of time and space, effectively solves the problem of monitoring frequency differences, and improves the accuracy and robustness of the analysis results.

[0046] The present invention uses a two-level association rule mining framework to deeply explore the complex association patterns of water resource pollution, designs a spatiotemporal association rule template, and realizes accurate tracing of pollution sources. This framework first uses the improved Apriori algorithm to preliminarily screen the pollution source data, and dynamically adjusts the support to adapt to the characteristics of different regions. Subsequently, the FP-Growth algorithm is used to perform deep association analysis to extract complex rules, such as the association pattern between high nitrate concentration and low pH value. This multi-level rule mining makes pollution source identification more accurate. When the rule is activated, the reverse Darcy flow tracking algorithm is used to achieve accurate positioning of pollution sources under complex terrain. This tracking algorithm combines the MODPATH particle tracking model to perform detailed spatial analysis to ensure the accuracy and real-time nature of tracing, providing key support for environmental protection decision-making.

[0047] The Bayesian hyperparameter optimization framework of the present invention is iteratively optimized through the TPE sampler, combined with the weighted combination of F1-score and inference delay, to achieve precise adjustment of model parameters. This optimization not only improves the accuracy of model predictions, but also reduces the consumption of computing resources. In terms of knowledge management, key driving factors are identified through Shapley value analysis, and a knowledge graph is constructed to integrate the spatiotemporal relationships of monitoring points, geological units and pollution sources. Complex relationship queries are implemented using the Neo4j graph database. It not only improves data processing and query efficiency, but also provides deep knowledge background support, providing a detailed foundation for scientific research and policy making. In addition, the model drift detection mechanism ensures the continued stable operation of the system under changing environmental conditions, and timely weight adjustments are made through online learning to maintain the high adaptability of the system.

[0048] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0050] Figure 1 It is a schematic diagram of the overall technical route of the present invention. DETAILED DESCRIPTION

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0052] Example 1

[0053] The water resources monitoring data analysis method based on data mining described in this embodiment includes the following steps:

[0054] Data preprocessing and feature construction stage

[0055] To address the high-dimensional, heterogeneous nature of monitoring data in karst areas, a spatiotemporal data cube structure (dimensions: longitude × latitude × time × spectral parameter) was first constructed. Wavelet threshold denoising (Daubechies 9 basis functions) was used to eliminate turbidity-induced scattering noise. Directional kriging interpolation (semivariogram bandwidth α = 1.2 km, anisotropy ratio β = 1.5) was performed for missing data. Seasonal trend terms (period = 365 days) of water quality parameters were extracted using STL decomposition, and lagged characteristic matrices (t-3, t-7, and t-14 days) were constructed to capture the delayed effects of groundwater transport in karst areas. A Savitzky-Golay second-order derivative transform (window width 21 nm) was applied to the UV spectral data to enhance the resolution of the nitrate peak at 220 nm. The dynamic correlation factor between COD and TOC was calculated using the moving window correlation coefficient method (window size 30 days).

[0056] Multi-model collaborative analysis stage

[0057] Adopt hybrid integration strategy to fuse spatiotemporal features:

[0058] For classification prediction tasks, a Stacking model architecture was constructed: the base learner used L2 regularized Logistic regression (λ = 0.1) to process sparse data, an improved ID3 decision tree (information gain ratio threshold > 0.3) to prevent overfitting, and a 5-layer BP neural network (the number of hidden layer nodes decreased in geometric progression 64-32-16-8); the meta-learner used XGBoost (max depth =6, learning rate =0.05) for feature recombination.

[0059] The spatiotemporal cluster analysis uses an improved k-means++ algorithm and introduces the dynamic time warping distance (DTW) to solve the problem of monitoring frequency differences. A spatial constraint term (weight γ = 0.6) is added when calculating the intra-cluster dissimilarity:

[0060] D total =γ·DTW(T i ,T j )+(1-γ)·Haversine(loc i ,loc j )

[0061] The outlier detection module integrates the LOF algorithm (k=15 nearest neighbors) and the spatial density correction based on crack density, triggering an early warning when the local outlier factor LOF>2.5 and the density of outliers within a 500m buffer zone>30%.

[0062] The SARIMAX (1, 1, 1) (1, 0, 1) 365 model is used for time series forecasting. Exogenous variables include rainfall (lagged 7 days) and agricultural fertilization intensity index. The seasonal difference order is determined by the Canova-Hansen test.

[0063] Association rule mining and pattern discovery stage

[0064] Design a two-level association rule mining framework:

[0065] Primary mining uses the improved Apriori algorithm, and the support is dynamically adjusted (urban area sup min =0.1, agricultural area sup min =0.05), confidence threshold conf min =0.7;

[0066] Deep association analysis uses the FP-Growth algorithm to extract frequent pattern trees, focusing on mining compound rules such as [high nitrate concentration → low pH value → sudden increase in COD]. Based on the characteristics of karst pipeline flow, a spatiotemporal association rule template is constructed:

[0067]

[0068] When the rule is activated, the reverse Darcy flow tracking algorithm (grid size 100m×100m) is started and combined with the MODPATH particle tracking model to locate the pollution source.

[0069] Model optimization and knowledge discovery stage

[0070] A Bayesian hyperparameter optimization framework (TPE sampler, 50 iterations) was established, with the objective function integrating prediction accuracy (F1-score weighted 0.6) and computational efficiency (inference latency weighted 0.4). Shapley value decomposition of the contribution of each feature identified fracture density (SHAP = 0.34) and rainfall accumulation (SHAP = 0.28) as key drivers. A knowledge graph was constructed to store three types of entities (monitoring sites, geological units, and pollution sources) and their spatiotemporal relationships. A Neo4j graph database was used to implement chained queries from [surge event] to [fracture zone] to [fertilization area]. A model drift detection module was developed, triggering online learning based on the PSI index (population stability index > 0.25). The FTRL-Proximal algorithm was used to update the logistic regression weight matrix.

[0071] Verification and deployment phase

[0072] Implement geographical division cross validation (dividing the study area into 5 hydrogeological units) and use the TimeSeriesSplit method (n splits=3) Maintaining temporal continuity. Performance evaluation indicators include: spatial specificity recall rate (>85%), early warning lead time (≥6 hours), and traceability positioning error (<500m). During deployment, model lightweight technology is used - channel pruning is performed on the BP neural network (sensitivity threshold 1×10 -4 ), reducing the number of parameters by 68%; converting the ARIMA model to a state-space form (implemented by Kalman filtering), reducing computational time by 42%. Deploying a lightweight anomaly detection module (TensorFlow Lite format, model size <2MB) at the edge achieves real-time response at 15ms.

[0073] Example 2

[0074] The water resources monitoring data analysis method based on data mining described in this embodiment includes the following steps:

[0075] S1: Data preprocessing and feature construction: A spatiotemporal data cube (dimensions: longitude × latitude × time × spectral parameter) was constructed. The monitoring data were processed through wavelet threshold denoising and directional kriging interpolation. The seasonal trend terms of water quality parameters were extracted using STL decomposition. A lagged feature matrix was constructed to capture the delayed effect of groundwater migration. The Savitzky-Golay derivative transformation was performed on the ultraviolet spectral data.

[0076] S2: Multi-model collaborative analysis: A stacking ensemble model was constructed to handle classification prediction tasks, an improved k-means++ algorithm was used for spatiotemporal clustering analysis, and a SARIMAX model was applied for time series prediction. The improved k-means++ algorithm introduced dynamic time warping distance and spatial constraint terms to calculate intra-class dissimilarity.

[0077] S3: Association rule mining and pollution source tracing: Extracting composite rules through a two-level association rule mining framework, designing spatiotemporal association rule templates to trigger the reverse Darcy flow tracing algorithm to locate pollution sources;

[0078] S4: Model Optimization and Knowledge Management: Use the Bayesian hyperparameter optimization framework to adjust model parameters, identify key driving factors based on Shapley values, and construct a knowledge graph containing the entity relationships between monitoring points, geological units, and pollution sources;

[0079] S5: Validation and Deployment: Implement geographic partition cross-validation to evaluate model performance, achieve lightweight model deployment through channel pruning and state space transformation, and deploy a real-time anomaly detection module at the edge.

[0080] In this embodiment, the directional kriging interpolation in step S1 adopts a semivariogram bandwidth α=1.2 km and an anisotropy ratio β=1.5; the lag feature matrix includes three time delay features of t-3, t-7, and t-14 days; the Savitzky-Golay transform is a second-order derivative transform with a window width of 21 nm.

[0081] In this embodiment, the base learners of the Stacking ensemble model in step S2 include L2 regularized Logistic regression (λ=0.1), improved ID3 decision tree (information gain ratio threshold>0.3) and 5-layer BP neural network (hidden layer nodes are reduced in descending order of 64-32-16-8), and the meta-learner adopts XGBoost (max depth =6, learning rate =0.05).

[0082] In this embodiment, the dissimilarity calculation formula of the improved k-means++ algorithm in step S2 is:

[0083] D total =γ·DTW(T i ,T j )+(1-γ)·Haversine(loc i ,loc j )

[0084] Where γ = 0.6: spatiotemporal weight adjustment factor, reflecting the contribution ratio of dynamic time warping distance to spatial distance. Grid search has verified that this value can balance the spatiotemporal heterogeneity of pipeline flow in karst area hydrological data. DTW (T i , T j ): Dynamic time warping distance, which solves the problem of inconsistent sampling frequencies of monitoring sites. The calculation formula is:

[0085]

[0086] Where W is the set of all legal curved paths; Haversine (loc i ,loc j ): Geospatial distance calculation:

[0087]

[0088] R = 6371 km (average radius of the Earth);

[0089] In this embodiment, the spatiotemporal association rule template in step S3 is expressed as:

[0090]

[0091] ΔT = 3d: Time propagation delay, verified by tracer experiments to satisfy:

[0092]

[0093] ΔD=2km: spatial action radius, and crack density ρ fracture Related:

[0094]

[0095] sup=0.12: support threshold, determined by significance test:

[0096] P(sup random ≥0.12)<0.05(permutation test)

[0097] When the rule is activated, the reverse Darcy flow tracking algorithm based on the MODPATH particle tracking model is started with a grid resolution of 100m×100m.

[0098] In this embodiment, the Bayesian hyperparameter optimization in step S4 uses a TPE sampler to iterate 50 rounds, and the objective function is a weighted combination of F1-score (weight 0.6) and inference delay (weight 0.4); the knowledge graph implements a chain query of [sudden event]-[rift zone]-[fertilization area] in the Neo4j graph database.

[0099] In this embodiment, the model lightweighting in step S5 includes: implementing a sensitivity threshold of 1×10 -4 Channel pruning is performed to convert the SARIMAX model into a state space form implemented by Kalman filtering; the edge module adopts the TensorFlow Lite format, the model size is <2MB and the response time is ≤15ms.

[0100] In this embodiment, the outlier detection in step S2 adopts the joint criterion of LOF algorithm (k=15 nearest neighbors) and spatial density correction, and triggers an early warning when LOF>2.5 and the density of outliers in the 500m buffer zone is>30%.

[0101]

[0102] k=15: Determined based on coverage analysis:

[0103]

[0104] B 500m (p): Spatial buffer zone, radius is calculated based on:

[0105]

[0106] In this embodiment, the model drift detection in step S4 triggers online learning based on a PSI index greater than 0.25, and the FTRL-Proximal algorithm is used to update the logistic regression weight matrix.

[0107] In this embodiment, the verification phase of step S5 adopts a three-element evaluation index system of spatial specific recall rate>85%, warning lead time ≥6 hours, and traceability positioning error<500m, and uses the TimeSeriesSplit method (n splits =3) Maintain temporal continuity.

[0108] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A water resources monitoring data analysis method based on data mining, characterized in that: The following steps are involved: S1: Establish a spatiotemporal data cube, process the monitoring data through wavelet threshold denoising and directional kriging interpolation, use STL decomposition to extract seasonal trend terms of water quality parameters, construct a lag characteristic matrix to capture the delayed effect of groundwater migration, and implement Savitzky-Golay derivative transformation on ultraviolet spectral data; S2: Build a Stacking ensemble model to handle classification prediction tasks, use an improved k-means++ algorithm for spatiotemporal clustering analysis, and apply a SARIMAX model for time series prediction. The improved k-means++ algorithm introduces dynamic time warping distance and spatial constraints to calculate intra-class dissimilarity. S3: Extract composite rules through a two-level association rule mining framework, design spatiotemporal association rule templates to trigger the reverse Darcy flow tracking algorithm to locate pollution sources; S4: Use the Bayesian hyperparameter optimization framework to adjust model parameters, identify key driving factors based on Shapley values, and construct a knowledge graph containing the entity relationships between monitoring points, geological units, and pollution sources; S5: Implement geographic partition cross-validation to evaluate model performance, achieve lightweight model deployment through channel pruning and state space transformation, and deploy real-time anomaly detection modules at the edge.

2. The water resources monitoring data analysis method based on data mining according to claim 1, characterized in that: The directional kriging interpolation in step S1 adopts a semivariogram bandwidth α = 1.2 km and an anisotropy ratio β = 1.5; the lag feature matrix includes three time delay features of t-3, t-7, and t-14 days; the Savitzky-Golay transform is a second-order derivative transform with a window width of 21 nm.

3. The water resources monitoring data analysis method based on data mining according to claim 1, characterized in that: The base learners of the Stacking ensemble model described in step S2 include L2 regularized Logistic regression, improved ID3 decision tree and 5-layer BP neural network, and the meta-learner adopts XGBoost.

4. The water resources monitoring data analysis method based on data mining according to claim 1, characterized in that: The dissimilarity calculation formula of the improved k-means++ algorithm in step S2 is: <h2 style=";text-align:left;direction:ltr">D<h2 style=";text-align:left;direction:ltr"> total <h2 style=";text-align:left;direction:ltr"> =γ·DTW(T<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> ,T<h2 style=";text-align:left;direction:ltr"> j <h2 style=";text-align:left;direction:ltr"> )+(1-γ)·Haversine(loc<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> ,loc<h2 style=";text-align:left;direction:ltr"> j <h2 style=";text-align:left;direction:ltr"> ) Where γ = 0.6: spatiotemporal weight adjustment factor, reflecting the contribution ratio of dynamic time warping distance to spatial distance. Grid search has verified that this value can balance the spatiotemporal heterogeneity of pipeline flow in karst area hydrological data. DTW (T i , T j ): Dynamic time warping distance, which solves the problem of inconsistent sampling frequencies at monitoring sites. The calculation formula is: Where W is the set of all legal curved paths; Haversine (loc i ,loc j ): Geospatial distance calculation: R=6371km.

5. The water resources monitoring data analysis method based on data mining according to claim 1, characterized in that: The spatiotemporal association rule template in step S3 is expressed as: ΔT = 3d: Time propagation delay, verified by tracer experiments to satisfy: u pipe ∈[0.02, 0.05]m / s ΔD=2km: spatial action radius, and crack density ρ fracture Related: sup=0.12: support threshold, determined by significance test: P(sup random ≥0.12)<0.05(permutation test) When the rule is activated, the reverse Darcy flow tracking algorithm based on the MODPATH particle tracking model is started with a grid resolution of 100m×100m.

6. The water resources monitoring data analysis method based on data mining according to claim 1, characterized in that: The Bayesian hyperparameter optimization described in step S4 uses a TPE sampler to iterate 50 rounds, and the objective function is a weighted combination of F1-score and inference delay; the knowledge graph implements a chain query of [surge event]-[rift zone]-[fertilization area] in the Neo4j graph database.

7. The water resources monitoring data analysis method based on data mining according to claim 1, characterized in that: The model lightweighting in step S5 includes: implementing a sensitivity threshold of 1×10 -4 The channel pruning is performed to convert the SARIMAX model into the state space form implemented by Kalman filtering; the edge module adopts the TensorFlowLite format, the model size is <2MB and the response time is ≤15ms.

8. The water resources monitoring data analysis method based on data mining according to claim 1, characterized in that: In step S2, outlier detection uses a combined criterion of the LOF algorithm and spatial density correction. When the LOF is greater than 2.5 and the density of outliers within the 500m buffer zone is greater than 30%, an early warning is triggered. k=15: Determined based on coverage analysis: B 500m (p): Spatial buffer zone, radius is calculated based on:

9. The water resources monitoring data analysis method based on data mining according to claim 1, characterized in that: In step S4, model drift detection triggers online learning based on a PSI index > 0.25, and the FTRL-Proximal algorithm is used to update the logistic regression weight matrix.

10. The water resources monitoring data analysis method based on data mining according to claim 1, characterized in that: The verification phase of step S5 adopts a three-element evaluation indicator system of spatial specificity recall rate > 85%, warning lead time ≥ 6 hours, and traceability positioning error < 500m, and maintains temporal continuity through the TimeSeriesSplit method.

Citation Information

Cited By

  • Water resource quality detection method and device applied to water diversion project and electronic equipment

    CN121431798A