Hydrogeological analysis system and method based on big data

By constructing a big data hydrogeological analysis system, dynamic monitoring and adaptive optimization of the hydrogeological environment were achieved, solving the problems of poor model adaptability and delayed decision response, and improving the accuracy and efficiency of hydrogeological analysis.

CN120851388BActive Publication Date: 2025-12-30SHANDONG ZHENGYUAN CONSTR ENG +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511339746.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-12-30
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

Existing hydrogeological analysis models lack dynamic performance monitoring and self-optimization mechanisms, making them unable to adapt to environmental changes, resulting in decreased prediction accuracy, difficulty in intuitively displaying three-dimensional structure and dynamic evolution, and delayed decision response.

Method used

We construct a hydrogeological analysis system based on big data. Through intelligent sensing and multi-source data fusion, model self-evolution, result visualization, and intelligent decision-making, we achieve dynamic threshold monitoring, KS test for outliers, and smooth iteration of A/B testing. We combine mechanism and machine learning with coupled models and parallel computing.

Benefits of technology

It improves the accuracy and efficiency of hydrogeological analysis predictions, supports the prevention and control of groundwater over-extraction and emergency decision-making on pollution, enhances data security and visualization capabilities, and provides scientific decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120851388B_ABST
    Figure CN120851388B_ABST
Patent Text Reader

Abstract

The application is specifically a hydrogeological analysis system and method based on big data, and relates to the technical field of hydrogeological analysis, comprising: putting new and old models into an A / B test stage and processing real-time data in parallel; automatically judging according to preset performance indexes and deciding whether to upgrade the candidate model to a new production model; and using the selected model after the decision for hydrogeological analysis to solve the long-term prediction accuracy problem under the dynamic change of the hydrogeological system. In the application, a dynamic threshold monitoring, K-S test outlier sample, model optimization and A / B test smooth iteration process are constructed, and a mechanism and machine learning coupled model and parallel computing are combined, so that model drift recognition delay is reduced, prediction accuracy is improved, underground water overexploitation prevention and control, pollution emergency and other scenes are effectively supported, and accurate and efficient scientific decision basis is provided for water resource management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hydrogeological analysis technology, and in particular to a hydrogeological analysis system and method based on big data. Background Technology

[0002] The following technical bottlenecks currently exist in hydrogeological analysis:

[0003] The analytical models are mostly single-mechanism models or machine learning models, lacking coupling capabilities. Furthermore, after deployment, the models lack dynamic performance monitoring and self-optimization mechanisms. When changes occur in the hydrogeological environment (such as sudden changes in rainfall or groundwater over-extraction), the models are prone to performance degradation, requiring manual readjustment and failing to meet long-term prediction needs. Fourth, the results are mostly presented in two-dimensional charts, making it difficult to intuitively display the three-dimensional hydrogeological structure and dynamic evolution process (such as pollutant transport). Moreover, decision support is limited to data output, lacking early warning linkage and scheme simulation capabilities, resulting in delayed decision response.

[0004] To address the aforementioned issues, there is an urgent need for a hydrogeological analysis system capable of deep fusion of multi-source data, self-evolution of models, visualization of results, and intelligent decision-making. This system aims to improve the accuracy and efficiency of hydrogeological analysis, provide scientific support for water resource management, and thus propose a hydrogeological analysis system and method based on big data. Summary of the Invention

[0005] The purpose of this invention is to propose a hydrogeological analysis system and method based on big data in order to solve the above-mentioned problems.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A big data-based hydrogeological analysis system includes:

[0008] Select the appropriate model to construct the production model and deploy it;

[0009] Continuously monitor the predictive performance of all deployed models and diagnose whether the models have failed due to environmental changes;

[0010] Once performance degradation is detected, the model update process is automatically triggered;

[0011] By using the latest data, the original model is fine-tuned to generate a candidate model with better performance;

[0012] The old and new models are put into the A / B testing phase to process real-time data in parallel.

[0013] The system automatically evaluates performance based on preset performance indicators and decides whether to upgrade the candidate model to a new production model.

[0014] The selected model after decision-making is used for hydrogeological analysis to address the issue of long-term prediction accuracy under dynamic changes in the hydrogeological system.

[0015] Preferably, the intelligent sensing and multi-source data fusion module specifically includes:

[0016] Acquire multi-source data, specifically including real-time access data from IoT devices, remote sensing data, business system data, and manually reported data;

[0017] At the data acquisition end, preliminary data processing is performed, including data cleaning, format standardization, simple filtering, and compression.

[0018] Data with different spatiotemporal resolutions are unified onto a standard grid and timeline to form a dataset.

[0019] Preferably, the big data governance and elastic storage module specifically includes:

[0020] Determine the time-series database layer, spatial database layer, and data lake layer of the hybrid storage architecture;

[0021] It manages data quality, provides a data quality dashboard, and monitors the integrity, consistency, and accuracy of data; and integrates AI-based anomaly detection algorithms to automatically identify and mark suspected abnormal data.

[0022] Build data security and permissions, provide role-based data access control, and anonymize sensitive data to meet security audit requirements.

[0023] Preferably, the step of selecting the appropriate model to construct the production model and deploying it specifically includes:

[0024] The mechanism model library integrates MODFLOW, MT3DMS, and SWAT, supports parametric configuration, has a built-in parameter template library, and has model coupling functionality.

[0025] The machine learning model library includes time series prediction models, spatial analysis models, and hybrid models.

[0026] Preferably, the core indicators of the continuously monitored model prediction performance include root mean square error and mean absolute error.

[0027] A normal distribution model is constructed based on historical performance data, and a dynamic early warning threshold is set. When the real-time root mean square error exceeds the threshold, it is determined to be model drift, and model self-optimization is initiated.

[0028] Preferably, the model self-optimization includes:

[0029] Outlier samples were extracted from the newly collected data, and the differences were quantified using the KS test. ,in For the new data distribution function, For historical data distribution function, when Samples exceeding a preset threshold are marked as differential samples.

[0030] For machine learning models, incremental gradient descent is used to update parameters;

[0031] For the mechanistic model, key parameters are re-inverted based on new data, and the objective function is optimized.

[0032] Preferably, the method further includes a model arena and smooth iteration:

[0033] The candidate model runs in parallel with the current production model, synchronously outputting prediction results based on real-time data for a period of T days, and then the overall score of the candidate model is calculated.

[0034]

[0035] in, , The standard deviation coefficient of the predicted values. The percentage increase in computing speed;

[0036] When the candidate model score hour, The production model is scored, and the iteration is started automatically.

[0037] Preferably, the 3D visualization and interaction module specifically includes:

[0038] Based on geological survey data and simulation results, a three-dimensional hydrogeological virtual model is constructed.

[0039] It visualizes the groundwater flow field and pollutant transport process, and supports dragging and playing along the timeline to show the evolution process;

[0040] Users can click on any location on the model to obtain all attributes and historical data at the clicked location in real time; they can customize profile lines to instantly generate geological profile maps and hydraulic gradient maps; and the model deeply integrates visualization and analysis, supporting real-time visual feedback for what-if scenarios.

[0041] Preferably, the intelligent decision support and automated action module specifically includes:

[0042] Users can customize early warning rules, start monitoring and issue alarms, and automatically trigger emergency plans;

[0043] It provides a strategy simulation tool, allowing users to set up different management plans, quickly simulate future trends, and provide optimal solution suggestions for decision-making through comparison of multiple plans;

[0044] Based on predefined templates and the latest data and analysis charts, generate compliant dynamic analysis reports, briefings, and thematic charts with a single click.

[0045] Hydrogeological analysis methods based on big data include:

[0046] Intelligent sensing of multi-source hydrogeological data, after preliminary processing at the edge, forms a high-quality consistent dataset through a preset algorithm;

[0047] Construct a hybrid storage architecture consisting of a time-series database layer, a spatial database layer, and a data lake layer, and simultaneously carry out data quality control, data traceability, and hierarchical security and disaster recovery protection.

[0048] Select suitable models from the mechanistic model library and the machine learning model library to construct a production model and deploy it;

[0049] The predictive performance of the production model is continuously monitored using RMSE and MAE as core indicators. Dynamic thresholds are constructed based on historical data to determine model drift, and self-optimization is completed for outlier samples based on model type.

[0050] The optimized model is used to conduct hydrogeological analysis, and the results are presented in a three-dimensional visualization interactive format.

[0051] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0052] 1. This invention constructs a dynamic threshold monitoring, KS test for outlier samples, sub-model optimization, and A / B testing smooth iteration process. By combining mechanism and machine learning coupled models with parallel computing, it reduces the latency of model drift identification and improves prediction accuracy. It effectively supports scenarios such as groundwater over-extraction prevention and control and pollution emergency response, and provides accurate and efficient scientific decision-making basis for water resource management.

[0053] 2. This invention utilizes an intelligent sensing and multi-source data fusion module to cover real-time IoT data, remote sensing data, business system data, and manually reported data. It combines algorithms such as Gauss-Kruger projection and Kriging interpolation to achieve spatiotemporal alignment, and then enhances the data through Bayesian fusion, significantly improving the accuracy of small-scale water volume estimation. Coupled with a three-level hybrid architecture of big data governance and elastic storage modules, and AI anomaly detection, this invention not only achieves preset requirements for data missing rate and real-time data entry delay, but also ensures the security of sensitive data and the recoverability of data under extreme conditions. Attached Figure Description

[0054] Further details, features, and advantages of this application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0055] Figure 1 This is a structural diagram of the core analysis and self-evolutionary modeling module of the present invention;

[0056] Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation

[0057] Several embodiments of this application will now be described in more detail with reference to the accompanying drawings to enable those skilled in the art to implement this application. This application may be embodied in many different forms and for various purposes and should not be limited to the embodiments set forth herein. These embodiments are provided to make this application thorough and complete, and to fully convey the scope of this application to those skilled in the art. The embodiments described do not limit this application.

[0058] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It will be further understood that terms such as those defined in commonly used dictionaries shall be interpreted as having a meaning consistent with their meaning in the relevant field and / or the context of this specification, and shall not be interpreted in an idealized or overly formal sense unless expressly defined herein.

[0059] Example 1

[0060] Its specific implementation method is combined with the appendix Figure 1 and attached Figure 2 Please provide a detailed explanation.

[0061] Appendix Figure 1 The block diagram of the hydrogeological analysis system based on big data provided in the embodiments of the present invention shows the connection relationship between the intelligent sensing and multi-source data fusion module and the intelligent decision support and automated action module, and marks the main functional interaction flow of each module.

[0062] Appendix Figure 2 The flowchart of the hydrogeological analysis method based on big data provided in the embodiments of the present invention shows the complete steps from intelligently sensing multi-source hydrogeological data to conducting hydrogeological analysis using the optimized model.

[0063] In this embodiment, it includes:

[0064] The intelligent sensing and multi-source data fusion module specifically includes:

[0065] Acquire multi-source data, specifically including real-time access data from IoT devices, remote sensing data, business system data, and manually reported data;

[0066] Real-time data access for IoT devices: Supports plug-and-play functionality for groundwater monitoring (water level / temperature / conductivity sensors), surface water monitoring (flow meters / turbidity meters), and meteorological monitoring (rain gauges / evaporation pans), compatible with communication protocols such as LoRa, NB-IoT, and 4G / 5G. Automatically identifies device models and matches data parsing rules (e.g., for TDR soil moisture sensors, automatically converts the raw voltage signal to volumetric water content). Built-in retransmission mechanism: When a device loses connection due to weak signal in the field, data is cached locally (up to 72 hours), and retransmitted according to the timestamp after connection is restored, avoiding data loss.

[0067] Remote sensing data: Automatically acquire inversion data from satellite remote sensing (such as GRACE gravity satellite, optical / radar satellite) via API to monitor land subsidence, soil moisture, etc.

[0068] Business system: Connects with databases of departments such as meteorology, water resources, and environmental protection to obtain data such as rainfall, evaporation, and water withdrawal permits.

[0069] Manual data reporting: Standardize and facilitate the entry of on-site survey data through a mobile app.

[0070] At the data acquisition end (edge ​​computing gateway), the data undergoes preliminary processing, including data cleaning, format unification, simple filtering, and compression, to reduce the load on the network and central system.

[0071] By utilizing spatial interpolation and time series alignment algorithms, data with different spatiotemporal resolutions are unified onto a standard grid and time axis, forming a high-quality, consistent dataset. This lays a solid foundation for subsequent analysis, including:

[0072] Spatial alignment: Based on the Gauss-Krüger projection, all data are unified to the 2000 National Geodetic Coordinate System. For discrete point data (such as monitoring well water levels), Kriging interpolation (suitable for areas with strong spatial correlation) or inverse distance weighted interpolation (suitable for sparse data areas) is used to generate a spatially continuous field (such as groundwater level isosurfaces) with a 10m×10m / 50m×50m grid. Topological relationship verification is performed on vector data (such as fault lines and rivers) (such as ensuring that fault lines do not cross well points), and the data is overlaid and aligned with raster data (such as remote sensing imagery).

[0073] Time alignment: All device clocks are calibrated based on an NTP time server, with the error controlled within ±1 second;

[0074] For data with different sampling frequencies (such as rainfall data at 1 hour / time and water level data at 10 minutes / time), linear interpolation is used to unify them to a standard time axis (such as 30 minutes / interval), while retaining the original sampling point markings.

[0075] Multi-source fusion enhancement: For example, regional groundwater storage change data retrieved from GRACE satellites can be fused with water level data from ground monitoring wells, and regional biases in satellite data can be corrected using Bayesian algorithms to improve the accuracy of water quantity estimation at small scales (such as counties).

[0076] The big data governance and elastic storage module specifically includes:

[0077] Determine the time-series database layer, spatial database layer, and data lake layer of the hybrid storage architecture;

[0078] Time-series database layer: The sensor time-series database is built using InfluxDB / TimescaleDB, and time and space composite indexes are designed for high-frequency hydrological monitoring data (such as water level and water temperature) (e.g., fast query by monitoring well ID and time range). Real-time data for the past year is stored on SSD (supporting high-concurrency writes), while historical data older than one year is automatically archived to HDD (reducing storage costs), and a hot query interface is retained.

[0079] Spatial Database Layer: Based on PostGIS, a spatial database is built to store vector data such as geological structures (faults, folds), hydrological units (aquifer boundaries, aquitard distribution), and infrastructure (monitoring wells, water intakes). It supports OGC standard spatial queries (e.g., querying all monitoring wells within a 5km radius of a fault). A built-in spatial index (R-tree index) optimizes the efficiency of operations such as buffer analysis and overlay analysis.

[0080] Data Lake Layer: An unstructured data lake is built based on Hadoop HDFS to store remote sensing images (TIFF format), geophysical reports (PDF), borehole videos (MP4), and intermediate numerical simulation results (binary files), with total capacity supporting petabyte-level expansion. It adopts a management model combining metadata and data blocks. Metadata (such as image capture time and resolution) is stored in a relational database, while data blocks are distributed for storage, supporting retrieval by metadata (e.g., searching for SAR radar images of a certain area in the summer of 2023).

[0081] It manages data quality, provides a data quality dashboard, and monitors the integrity, consistency, and accuracy of data; and integrates AI-based anomaly detection algorithms to automatically identify and mark suspected abnormal data.

[0082] The data quality indicator system defines four core indicators: completeness (e.g., monthly data missing rate of a monitoring well ≤5%), consistency (e.g., uniform units for permeability parameters of the same aquifer), accuracy (e.g., deviation between water level measurement and manual verification value ≤0.1m), and timeliness (e.g., delay from data collection to storage ≤30 seconds). These indicators are visualized in real time through a quality dashboard.

[0083] Intelligent anomaly detection and repair: Based on the isolated forest algorithm, outliers are identified (such as a sudden change in the pH value of a well to 12, which is far beyond the normal range). Combined with historical data distribution (95% confidence interval), the anomaly level (low / medium / high) is automatically marked. For moderately abnormal data (such as minor deviations caused by sensor drift), a spatiotemporal collaborative repair method is adopted: horizontally, the data of the same period of three surrounding wells are referenced, and vertically, the historical trend of the well is combined. The repair value is predicted by the LSTM model, and the original data is retained for manual review.

[0084] Data lineage tracing: Records the entire chain of each data source, processing, and flow (such as water level data, edge gateway filtering, spatiotemporal alignment, storage, and use in model training). It supports tracing back to the original acquisition device and processing logs through data ID, meeting audit and traceability requirements.

[0085] Build data security and access control, provide role-based data access control, and anonymize sensitive data (such as the location of water sources) to meet security audit requirements.

[0086] Tiered security strategy: Sensitive data (such as precise coordinates of drinking water sources and raw data on pollution exceeding standards) are stored using AES-256 encryption and transmitted using TLS1.3 encryption; Non-sensitive data (such as publicly available regional rainfall statistics) are anonymized (e.g., specific locations are obscured down to the township level).

[0087] Fine-grained access control: Based on the RBAC (role-based access control) model, roles such as administrator, analyst, and public user are defined: administrators have full data read and write permissions and system configuration permissions; analysts have data query and model operation permissions for specified areas; public users can only view anonymized statistical results (such as the annual average groundwater level change trend of a certain city).

[0088] Disaster recovery and mitigation: The production center is synchronized to the local disaster recovery center in real time, and incremental backups are made to the off-site disaster recovery center daily. It supports recovery at any point in time within 15 days after accidental data deletion, with RPO (Recovery Point Objective) ≤ 1 hour and RTO (Recovery Time Objective) ≤ 4 hours.

[0089] The core analysis and self-evolutionary modeling module specifically includes:

[0090] Select the appropriate model to construct the production model and deploy it;

[0091] Continuously monitor the predictive performance of all deployed models and automatically diagnose whether the models have failed due to environmental changes using a dynamic threshold algorithm;

[0092] Once performance degradation is detected, the model update process is automatically triggered;

[0093] By utilizing the latest data and fine-tuning the original model through online learning technology, a better-performing candidate model is generated.

[0094] The old and new models are put into the A / B testing phase to process real-time data in parallel.

[0095] The system automatically evaluates and decides whether to upgrade the candidate model to a new production model based on preset performance indicators (such as prediction accuracy and stability), thereby achieving smooth iteration and self-evolution of the model.

[0096] The selected model after decision-making is used for hydrogeological analysis to address the issue of long-term prediction accuracy under dynamic changes in the hydrogeological system.

[0097] Selecting the appropriate model to construct the production model and deploying it specifically includes:

[0098] The mechanism model library includes classic models such as MODFLOW (groundwater flow simulation), MT3DMS (solute transport simulation), and SWAT (watershed hydrological simulation), supports parameterized configuration (such as aquifer permeability coefficient and porosity), and has a built-in parameter template library (preset initial parameters according to regions such as the North China Plain and the Yangtze River Delta).

[0099] It also has model coupling capabilities: for example, it can couple MODFLOW with SEAWAT (seawater intrusion model) to simulate the seawater intrusion process caused by groundwater over-extraction in coastal areas and automatically transfer boundary conditions (such as the interaction between groundwater level and seawater level).

[0100] The machine learning model library includes time series prediction models, spatial analysis models, and hybrid models;

[0101] Temporal forecasting models: LSTM (suitable for long-term time series forecasting such as water level and rainfall), TCN (temporal convolutional network, which improves the prediction accuracy of short-term sudden increase / decrease events).

[0102] Spatial analysis models: U-Net (based on remote sensing image segmentation of aquifer distribution), Random Forest (inverting the correlation between permeability parameters and geological lithology);

[0103] Hybrid models: The output of the mechanistic model is used as a feature of the machine learning model (such as using the flow field features simulated by MODFLOW to assist LSTM in predicting water level), which improves the generalization ability in complex scenarios.

[0104] The model automatically records every parameter adjustment and training data update, generates a version number, supports version rollback (such as rolling back to the best-performing model version 3 months ago), and marks the reason for version iteration (such as data supplementation and update due to the rainy season).

[0105] By building a computing cluster based on Spark and MPI, large simulation tasks (such as groundwater simulation covering 100,000 square kilometers) can be computed in parallel by spatial partitioning (such as dividing into 100 10km×10km sub-regions) or time segmentation (such as splitting by quarters), improving computing efficiency by 10-50 times (a task that would take 72 hours on a traditional single machine can be completed in 2 hours by the cluster).

[0106] The computing nodes (CPU / GPU) are automatically allocated based on the complexity of the task (such as simulation time step and grid accuracy), with priority given to ensuring the resource needs of urgent tasks (such as pollution diffusion early warning simulation).

[0107] The grid is automatically densified in key areas (such as the area around pollution sources and areas with a high density of water wells) (from 100m×100m to 10m×10m), while the grid is simplified in non-key areas, reducing the amount of computation while ensuring accuracy (the total number of grids is reduced by more than 60%).

[0108] Built-in verification indicators such as Nash-Sutcliffe efficiency coefficient (NSE) and root mean square error (RMSE) automatically compare simulated values ​​with measured values ​​(such as simulated water level and monitoring well data), generate error heat maps, and locate areas with large simulation deviations (such as error areas caused by unreasonable fault parameter settings).

[0109] Based on observational data (such as well water level and water quality), unknown parameters (such as permeability coefficient and recharge rate) are inverted. The core objective function is: ,in Let be the parameter vector to be inverted. Let be the measured water level at the i-th observation point. To simulate water levels in the model, The standard deviation of the observation error;

[0110] The PEST tool is used to achieve efficient inversion, and the gradient descent algorithm is used to iteratively optimize the parameters until... ( (Preset precision threshold).

[0111] Continuously monitor the model's predictive performance, with key metrics including root mean square error (RMSE). ) and mean absolute error (MAE: );

[0112] in These are measured values. These are the model's predicted values. For sample size;

[0113] A normal distribution model is constructed based on historical performance data (such as the root mean square error (RMSE) over the past 6 months). Set the dynamic early warning threshold as (Upper limit of 95% confidence interval) When the real-time root mean square error (RMSE) exceeds this threshold, it is determined to be model drift, and model self-optimization is initiated.

[0114] By clearly defining core monitoring indicators and dynamic threshold algorithms, a precise quantitative basis for judging the performance stability of hydrogeological models is provided. It uses root mean square error (RMSE) and mean absolute error (MAE) as core indicators, which can intuitively and comprehensively reflect the degree of deviation between model predictions and measured values ​​(such as well water levels and rainfall). RMSE is more sensitive to extreme errors and can capture prediction anomalies caused by sudden environmental changes (such as rainstorms and pollution events), while MAE reflects the overall error level. The combination of the two avoids the limitations of a single indicator, transforming the model performance degradation from a vague perception to precise quantification, providing clear data support for subsequent interventions.

[0115] The dynamic early warning threshold mechanism effectively solves the problems of high false alarm rate and poor adaptability of traditional fixed thresholds. It constructs a normal distribution model based on RMSE data from the past 6 months and sets the upper limit of the 95% confidence interval (…). The threshold is set as an early warning threshold, which takes into account the natural fluctuation characteristics of hydrogeological data (such as normal error fluctuations caused by seasonal water level changes) and can accurately identify model drift that exceeds the normal range. Compared with a fixed threshold (such as using the same error standard regardless of season or region), this dynamic threshold can adapt to different hydrological scenarios, reduce invalid optimizations triggered by normal fluctuations, and avoid the risk of missing real model failures due to excessively high thresholds, thus ensuring efficient use of system resources.

[0116] The performance monitoring and drift judgment logic provides real-time assurance for the accuracy of long-term hydrogeological predictions. Hydrogeological systems are dynamic and changeable (e.g., groundwater over-extraction, changes in surface vegetation, and climate change can all affect model adaptability). Traditional models often fail to detect performance degradation in a timely manner, leading to increased deviations in long-term prediction results (e.g., misjudging the downward trend of groundwater levels, affecting water resource allocation decisions).

[0117] Through continuous monitoring and dynamic threshold diagnosis, the system can quickly identify and initiate self-optimization when model performance first becomes abnormal, preventing error accumulation from the source and ensuring that the model always adapts to the current hydrogeological environment, thus laying the foundation for the accuracy of long-term predictions (such as annual groundwater storage assessment and watershed water resources planning).

[0118] Model self-optimization includes:

[0119] Outlier samples (samples whose distribution differs significantly from historical data, such as data from extreme rainfall or sudden pollution events) are extracted from newly collected data, and the differences are quantified using the KS test: ,in For the new data distribution function, For historical data distribution function, when Samples exceeding a preset threshold are marked as differential samples.

[0120] For the machine learning model, the incremental gradient descent method is used to update the parameters. The parameter update formula for the t-th iteration is: ,in For model parameters, For learning rate, For loss function, For new samples;

[0121] For the mechanistic model, key parameters (such as replenishment amount) are re-inverted based on new data. ), optimize the objective function (Inversion formula with the same parameters).

[0122] Accurately extracting outlier samples and quantifying data discrepancies provides a targeted data foundation for model self-optimization, effectively solving the model adaptability problem caused by extreme events in hydrogeological scenarios. Hydrological systems are susceptible to unconventional events such as extreme rainfall and sudden pollution. The distribution of such data differs significantly from historical regular data, and traditional models often lead to deviations in optimization direction due to ignoring or misjudging these samples. By employing the KS test (calculating the maximum difference D between the distribution functions of new and old data) to quantify outlier sample characteristics, we can accurately identify anomalous data that is crucial to model performance. This ensures that subsequent optimization targets only the core samples that truly alter hydrological patterns, avoiding interference from invalid data and making model optimization more targeted.

[0123] Differentiated update strategies were designed for machine learning models and mechanistic models respectively, enabling on-demand optimization that ensures both optimization effectiveness and efficiency, breaking the limitations of traditional single update methods.

[0124] For machine learning models, incremental gradient descent updates parameters only based on new samples (the t-th parameter update depends on the gradient of the previous parameters and the loss function of the new sample), without the need to retrain on all data, which greatly reduces the consumption of computing resources and is suitable for the need for real-time updates of high-frequency hydrological data.

[0125] For mechanistic models, key parameters (such as supply amount) are re-inverted and the objective function is optimized. This ensures that the core parameters of the model are accurately matched with the current hydrogeological conditions (such as changes in groundwater extraction and adjustments in surface runoff), avoiding simulation deviations caused by parameter fixation. The two strategies complement each other, balancing update efficiency and model accuracy.

[0126] Model Arena and Smooth Iteration:

[0127] The candidate model (optimized) runs in parallel with the current production model, synchronously outputting prediction results on real-time data for a period of T days (T is set according to the data update frequency, such as 7-30 days). The overall score of the candidate model is then calculated.

[0128]

[0129] in, (Range represents the range of data values) The standard deviation coefficient of the predicted values. The percentage increase in computing speed;

[0130] When the candidate model score hour, To score the production model, an automatic iteration is initiated: first, 10% of real-time requests are switched to the candidate model, and after no anomalies are found, the switch is gradually expanded to 100%, while the original model is retained for 30 days as an emergency rollback version.

[0131] By employing a parallel testing and quantitative comprehensive scoring mechanism for candidate and production models, a scientific and rigorous decision-making basis is provided for the iteration of hydrogeological models, effectively avoiding the risks of blind updates. Its comprehensive scoring formula, based on accuracy (60% weighting, the core weight), stability, and efficiency, focuses on the most critical prediction accuracy requirement in hydrological analysis while also considering the stability of model operation (avoiding excessive fluctuations in predicted values) and computational efficiency (adapting to real-time analysis scenarios). Simultaneously, candidate and production models process real-time data concurrently and are continuously tested for T days, verifying the adaptability of the new model in real-world hydrological scenarios, rather than theoretical verification in a laboratory environment, ensuring that the upgraded model truly meets practical application needs.

[0132] The phased switchover and emergency rollback design ensures the continuous stability of the hydrogeological analysis system and resolves the contradiction between model iteration and business continuity. The hydrological system needs to support critical decisions such as water resource allocation and pollution early warning 24 / 7 and cannot tolerate interruptions or errors caused by model replacement. By first switching 10% of real-time requests to the candidate model, and then gradually expanding to 100% if no anomalies are found, small-scale adaptation issues (such as localized data deviations) can be identified and corrected in a timely manner, avoiding system risks caused by a full switchover. Simultaneously, the original model is retained for 30 days as an emergency version, allowing for rapid rollback and recovery even if the new model experiences a sudden failure, providing crucial assurance for the high reliability requirements of hydrogeological analysis.

[0133] The 3D visualization and interaction module specifically includes:

[0134] Based on geological survey data and simulation results, a three-dimensional hydrogeological virtual model is constructed, which can be made transparent, cut, and rotated to intuitively display the aquifer structure and fault distribution.

[0135] Data such as groundwater flow field and pollutant transport process are visualized in the form of dynamic particle flow, isosurface cloud map, heat map, etc., and the timeline can be dragged and played to show the evolution process.

[0136] Users can click on any location on the model to obtain all attributes and historical data for that location in real time; they can customize profile lines to instantly generate geological profile maps and hydraulic gradient maps; and the model deeply integrates visualization and analysis, supporting real-time visual feedback for what-if scenarios.

[0137] Real-time visual feedback for what-if scenarios: Users manually adjust parameters in a 3D scene, and the system calls the background model to quickly calculate (within 10 seconds) and update the 3D scene in real time: such as the range of water level rise caused by reinjection (dynamically expanded with a blue halo), and changes in water flow direction (particle flow direction deflection); synchronously generate comparison views (left and right split screens display the differences before and after adjustment) to assist in decision-making and evaluation;

[0138] The intelligent decision support and automated action module specifically includes:

[0139] Users can customize early warning rules (such as "when the water level in a certain area is lower than 50 meters above sea level"), start 24 / 7 automatic monitoring, and send alerts via SMS, email, App push, etc., and can automatically trigger emergency plans.

[0140] It provides a strategy simulation tool, allowing users to set different management plans (such as reducing mining volume by 10% or increasing artificial reinjection) to quickly simulate future trends and provide optimal solution suggestions for decision-making through comparison of multiple plans;

[0141] Based on predefined templates and the latest data and analysis charts, it can generate dynamic analysis reports, briefings, and thematic charts that meet the standards with one click, greatly improving work efficiency and the scientific nature of reports.

[0142] The intelligent early warning and emergency triggering mechanism provides all-weather, automated protection for hydrogeological risk prevention and control, solving the problems of low efficiency and delayed response of traditional manual monitoring. Users can customize early warning rules according to actual needs (such as water level in a certain area being lower than a specific altitude or water quality indicators exceeding standards). The system continuously monitors data 24 / 7, and once the threshold is triggered, it immediately alerts through multiple channels such as SMS, email, and App push notifications, ensuring that relevant personnel are aware of the risk as soon as possible. More importantly, the early warning can automatically trigger emergency plans (such as activating backup water sources or shutting down wells in polluted areas), avoiding response delays caused by manual judgment and operation, and effectively reducing losses caused by groundwater over-extraction and pollution spread.

[0143] The strategy simulation and automated reporting functions significantly improve the scientific rigor and efficiency of hydrogeological decision-making. Through the strategy simulation tool, users can set different management plans (such as reducing extraction by 10% or increasing artificial recharge). The system quickly simulates future hydrological trends (such as water level changes and reserve assessments) and intuitively presents the optimal solution through multi-plan comparisons, avoiding the subjectivity of traditional experience-based decision-making. Simultaneously, based on the latest data and analysis charts, it generates standardized dynamic reports, briefings, and thematic maps with a single click, eliminating the need for manual data processing and chart drawing. This not only significantly saves manpower and time costs but also ensures the accuracy and standardized format of report data, providing scientific and efficient output for water resource planning, geological exploration, and other work.

[0144] Example 2

[0145] Please see Figure 2 The hydrogeological analysis method based on big data includes the following parts:

[0146] Intelligent sensing of multi-source hydrogeological data (including real-time data from IoT devices, remote sensing data, business system data, and manually reported data) is initially processed at the edge and then formed into a high-quality consistent dataset through preset algorithms (spatial interpolation, temporal alignment, and Bayesian fusion, etc.).

[0147] Construct a hybrid storage architecture consisting of a time-series database layer, a spatial database layer, and a data lake layer, and simultaneously carry out data quality control (AI anomaly detection and repair), data traceability, and graded security and disaster recovery protection.

[0148] The production model is constructed and deployed by selecting suitable models from the mechanistic model library and machine learning model library. Parallel computing is achieved based on the Spark / MPI computing cluster. The model accuracy is ensured by parameter inversion and NSE and RMSE verification.

[0149] The production model prediction performance is continuously monitored using RMSE and MAE as core indicators. The model drift is determined by a dynamic threshold constructed based on historical data. Then, self-optimization is completed for outlier samples based on model type (incremental gradient descent for machine learning and parameter inversion for mechanistic models).

[0150] The optimized model is used to conduct hydrogeological analysis, and the results are presented in combination with 3D visualization interaction (dynamic particle flow, what-if scenario feedback). At the same time, custom early warning rules are set up to achieve 24 / 7 monitoring and alarm, and intelligent decision-making is assisted by strategy simulation and one-click generation of standardized reports.

[0151] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0152] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

[0153] It should be noted that, in this document, the use of relational terms such as "first" and "second" is merely for distinguishing one entity or operation from another, and does not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0154] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0155] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0156] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0157] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0158] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0159] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0160] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A hydrogeological analysis system based on big data, characterized by, The core analysis and self-evolution modeling module comprises the following: Selecting a corresponding model to constitute a production model and deploying the same; Continuously monitoring the prediction performance of all deployed models to diagnose whether the models have failed due to environmental changes; Once performance degradation is detected, automatically triggering a model updating process; Using the latest data to fine-tune the original model to generate a candidate model with better performance; Placing the new and old models into an A / B testing phase and processing real-time data in parallel; Automatically judging according to preset performance indicators and deciding whether to upgrade the candidate model to a new production model; Using the selected model after decision-making for hydrogeological analysis to solve the problem of long-term prediction accuracy under the dynamic changes of the hydrogeological system; The selection of a corresponding model to constitute a production model and deploy the same comprises the following: The mechanism model library comprises integrated MODFLOW, MT3DMS and SWAT, supports parameterized configuration, has a built-in parameter template library and has a model coupling function; The machine learning model library comprises a time series prediction model, a spatial analysis model and a hybrid model; The core analysis and self-evolution modeling module further comprises the following: An intelligent perception and multi-source data fusion module comprises the following: Obtaining multi-source data, specifically including real-time access data of Internet of Things devices, remote sensing data, business system data and manually reported data; Preliminarily processing the data at the data collection end, including data cleaning, format unification, simple filtering and compression; Uniformly converting data of different temporal and spatial resolutions to standard grids and time axes to form a data set; A big data management and elastic storage module comprises the following: Determining a hybrid storage architecture of a time series database layer, a spatial database layer and a data lake layer; Controlling data quality, providing a data quality board to monitor the integrity, consistency and accuracy of the data and integrating an AI-based anomaly detection algorithm to automatically identify and mark suspected abnormal data; Building data security and permissions, providing role-based data access control and desensitizing sensitive data to meet security audit requirements; A three-dimensional visualization and interaction module comprises the following: Based on geological survey data and simulation results, a three-dimensional hydrogeological virtual model is constructed; The groundwater flow field and the pollutant transport process are visualized, and time axis dragging and playback are supported to demonstrate the evolution process; The user can click on any position of the model to obtain all attributes and historical data at the clicked position in real time, can customize a profile line to instantly generate a geological profile graph and a hydraulic gradient graph, and can deeply integrate visualization and analysis to support instant visualization feedback of what-if scenarios; An intelligent decision support and automated action module comprises the following: The user can customize early warning rules, start monitoring and perform alarm and automatically trigger an emergency plan; A strategy simulation tool is provided, the user can set different management schemes, quickly simulate future trends, compare multiple schemes and provide an optimal solution suggestion for decision-making; Based on predefined templates and the latest data and analysis charts, a dynamic analysis report, a briefing and a thematic map can be generated in one click to meet the requirements of specifications.

2. The big data based hydrogeological analysis system of claim 1, wherein, The core indicators for continuously monitoring the prediction performance of the model include root mean square error and mean absolute error. Based on historical performance data, a normal distribution model is constructed to set a dynamic early warning threshold. When the real-time root mean square error exceeds the threshold, it is determined that the model has drifted, and the model self-optimization is started.

3. The big data based hydrogeological analysis system of claim 2, wherein, Model self-optimization includes: Extract outliers from new data, quantify the difference by K-S test: where is the new data distribution function, is the historical data distribution function, when is greater than the preset threshold, marked as difference sample; For machine learning models, use incremental gradient descent to update parameters; For mechanism models, re-invert key parameters based on new data to optimize the objective function.

4. The big data based hydrogeological analysis system of claim 3, wherein, It also includes model arena and smoothing iteration: Candidate models run in parallel with the current production model, output prediction results for real-time data, and run for T days to calculate the comprehensive score of the candidate model: wherein, , is a standard deviation coefficient of the predicted value, is a calculation speed improvement ratio; When the candidate model score is, To produce the model score, iterations are automatically initiated.

5. The hydrogeological analysis method based on big data according to any one of claims 1 to 4, characterized in that, Including: Intelligent sensing of multi-source hydrogeological data, after preliminary processing on the edge, high-quality consistent data sets are formed through pre-set algorithms; Build a hybrid storage architecture of time series database layer, spatial database layer and data lake layer, and simultaneously carry out data quality control, data traceability tracking and hierarchical security and disaster protection; Select suitable models from the mechanism model library and machine learning model library to form the production model and deploy it; Continuously monitor the prediction performance of the production model with RMSE and MAE as the core indicators, determine model drift based on the dynamic threshold constructed from historical data, and then complete self-optimization for out-of-sample models by type; Use the optimized model to analyze hydrogeology and present the results interactively with three-dimensional visualization.

Citation Information

Patent Citations

  • Dynamic optimization and adaptive learning method of water conservancy system large model and storage medium

    CN120218183A

  • Wide-area landslide rapid identification method based on interpretable intelligent algorithm

    CN120524834A