Physical logic driven groundwater storage intelligent prediction method and system

CN122088798BActive Publication Date: 2026-07-03SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG UNIV
Filing Date
2026-04-24
Publication Date
2026-07-03

Smart Images

  • Figure CN122088798B_ABST
    Figure CN122088798B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of groundwater storage prediction technology, and particularly relates to a physical logic-driven intelligent prediction method and system for groundwater storage. It includes: completing and physically verifying multi-source heterogeneous observation data related to groundwater storage in a target area to obtain a physically verified complete dataset; performing spatiotemporally decoupled causal contribution analysis on the physically verified complete dataset to generate a causal contribution matrix; constructing and training a physically constrained emergent spatiotemporal prediction model based on the causal contribution matrix; applying a physical anchoring adversarial migration strategy to adapt the trained emergent prediction model to the new target area to obtain a regionally adapted prediction model; and inputting future scenario conditions into the regionally adapted prediction model for simulation to generate groundwater storage prediction results. This invention can achieve high-precision, highly robust, and physically interpretable intelligent prediction of groundwater storage, and provides scientific decision support for water resource management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of groundwater storage prediction technology, and particularly relates to a physical logic-driven intelligent prediction method and system for groundwater storage. Background Technology

[0002] Groundwater, as a core water resource, directly impacts ecological security and social development due to changes in its reserves. Accurate monitoring and intelligent prediction of groundwater reserves are crucial for the scientific management of water resources, which typically requires comprehensive analysis of data from multiple sources, including rainfall, evaporation, and artificial pumping.

[0003] Currently, mainstream technical solutions are mainly divided into two categories:

[0004] One is the numerical simulation method based on physical processes. Although this method has a clear physical meaning, it is extremely dependent on high-precision geological parameters, which makes it difficult to construct and has huge computational costs, making it difficult to meet the needs of large-scale real-time prediction.

[0005] Second, there is the pure data-driven deep learning method. Although this method can uncover complex patterns, its process is like a "black box," lacking physical mechanism constraints. The prediction results often violate basic common sense such as water balance, and the performance will be severely degraded when moving to a new region or encountering extreme climate. The reliability of the results is difficult to gain the trust of experts. Summary of the Invention

[0006] To overcome the shortcomings of the existing technologies, this invention provides a physical logic-driven intelligent prediction method and system for groundwater reserves. It adopts a technical solution that deeply integrates physical mechanisms and data-driven approaches. Through steps such as adversarial physical verification reconstruction, spatiotemporal causal decoupling, physical constraint emergent modeling, and physical anchoring transfer learning, it can achieve intelligent prediction of groundwater reserves with high accuracy, high robustness, and physical interpretability, and provide scientific decision support for water resource management.

[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0008] The first aspect of this invention provides a physical logic-driven intelligent prediction method for groundwater reserves.

[0009] A physical logic-driven intelligent prediction method for groundwater storage includes the following steps:

[0010] The multi-source heterogeneous observation data related to groundwater storage in the target area were supplemented and physically verified to obtain a complete physical verification dataset that conforms to physical logic.

[0011] A spatiotemporally decoupled causal contribution analysis was performed on the complete dataset of physical verification to separate the independent influence of driving factors in the temporal and spatial dimensions and generate a causal contribution matrix.

[0012] Based on the causal contribution matrix, a physically constrained emergent spatiotemporal prediction model is constructed and trained. During the training process, a reward mechanism for discovering physical laws is introduced to enable the internal state of the model to spontaneously align with physical laws, thus obtaining a well-trained emergent prediction model.

[0013] By applying a physical anchoring adversarial migration strategy, the trained emergent prediction model is adapted to the target new area to obtain a region-adapted prediction model.

[0014] Future scenario conditions are input into the regional adaptation prediction model for simulation, generating groundwater storage prediction results and corresponding decision-making suggestions.

[0015] The second aspect of this invention provides a physical logic-driven intelligent prediction system for groundwater reserves.

[0016] A physical logic-driven intelligent prediction system for groundwater storage includes:

[0017] The multi-source data acquisition and physical logic verification module is configured to: complete and physically verify the multi-source heterogeneous observation data related to groundwater storage in the target area, and obtain a complete physical verification dataset that conforms to physical logic;

[0018] The driving factor analysis module is configured to perform spatiotemporally decoupled causal contribution analysis on the complete physical verification dataset, separate the independent effects of driving factors in the time and space dimensions, and generate a causal contribution matrix.

[0019] The self-learning modeling module is configured to: construct and train a physical constraint emergent spatiotemporal prediction model based on the causal contribution matrix; during the training process, a physical law discovery reward mechanism is introduced to enable the internal state of the model to spontaneously align with the physical law, thus obtaining a well-trained emergent prediction model.

[0020] The cross-regional knowledge transfer module is configured to: apply a physical anchoring adversarial transfer strategy to adapt the trained emergent prediction model to the target new area, thereby obtaining a regionally adapted prediction model;

[0021] The scenario simulation and decision-making module is configured to input future scenario conditions into the regional adaptive prediction model for simulation, and generate groundwater storage prediction results and corresponding decision-making suggestions.

[0022] The above one or more technical solutions have the following beneficial effects:

[0023] This invention provides a physical logic-driven intelligent prediction method and system for groundwater reserves. Through the organic combination of a series of techniques, including adversarial physical verification reconstruction, spatiotemporal causal decoupling, emergent modeling of physical constraints, and physical anchoring migration, it achieves a balance between prediction accuracy and physical plausibility. This method ensures the physical consistency of input data from the data source through adversarial training, and allows physical laws to be actively learned by the model as an intrinsic mechanism during model training. This fundamentally eliminates prediction results that violate natural laws, ensuring high accuracy and reliability even under extreme hydrological events.

[0024] This invention significantly improves the generalization ability and applicability of models across different regions through a physically anchored adversarial transfer learning strategy. By freezing the network layers representing universal physical mechanisms during model transfer and only fine-tuning the region-specific parameter layers, it effectively solves the technical bottleneck of traditional deep learning models' over-reliance on massive amounts of labeled data and the sharp performance drop when applied to new target regions. This method enables the model to quickly adapt with minimal target region data, greatly expanding the practical application value of this technology in data-scarce areas.

[0025] This invention tightly integrates the prediction process with decision support, constructing a complete intelligent closed loop from data to knowledge to decision-making. This method not only provides high-precision predicted values ​​but also clearly reveals the contribution of different driving factors to changes in groundwater reserves through simulation and other means, enhancing the model's interpretability. The final generated decision recommendation data, including risk assessment and multi-scenario analysis, can directly serve water resource management departments, transforming complex model outputs into intuitive and actionable scientific scheduling guidelines, achieving a deep integration of technological prediction and industry application.

[0026] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0027] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0028] Figure 1 This is a flowchart of the method in Example 1.

[0029] Figure 2 This is a mapping diagram connecting the physical logic and model of lateral groundwater flow.

[0030] Figure 3 The spatiotemporal causal contribution matrix generated for this embodiment.

[0031] Figure 4 Reconstructing a logic graph for physical equilibrium data based on adversarial learning.

[0032] Figure 5 This is a trajectory diagram for multi-scenario simulation. Detailed Implementation

[0033] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0034] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0035] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0036] Example 1

[0037] This embodiment discloses a physical logic-driven intelligent prediction method for groundwater reserves, aiming to develop an intelligent prediction system that can automatically extract big data patterns, embed physical logic constraints, and has the ability to quickly adapt across regions, so as to balance computational efficiency and physical consistency, and solve the pain points of current groundwater management such as difficulty in data acquisition, simulation that does not conform to common sense, and poor generalization ability.

[0038] like Figure 1 As shown, the physical logic-driven intelligent prediction method for groundwater storage includes the following steps:

[0039] The multi-source heterogeneous observation data related to groundwater storage in the target area were supplemented and physically verified to obtain a complete physical verification dataset that conforms to physical logic.

[0040] A spatiotemporally decoupled causal contribution analysis was performed on the complete dataset of physical verification to separate the independent influence of driving factors in the temporal and spatial dimensions and generate a causal contribution matrix.

[0041] Based on the causal contribution matrix, a physically constrained emergent spatiotemporal prediction model is constructed and trained. During the training process, a reward mechanism for discovering physical laws is introduced to enable the internal state of the model to spontaneously align with physical laws, thus obtaining a well-trained emergent prediction model.

[0042] By applying a physical anchoring adversarial migration strategy, the trained emergent prediction model is adapted to the target new area to obtain a region-adapted prediction model.

[0043] Future scenario conditions are input into the regional adaptation prediction model for simulation, generating groundwater storage prediction results and corresponding decision-making suggestions.

[0044] This embodiment first uses an adversarial network with embedded physical equilibrium equations to complete and verify multi-source heterogeneous data, generating a physically consistent complete dataset. Next, it employs spatiotemporal causal analysis to quantify the contribution of each driving factor to groundwater storage changes, generating a causal contribution matrix. Based on this, it constructs an emergent spatiotemporal prediction model capable of actively learning physical laws. Through a migration strategy that anchors to general laws and fine-tunes regional characteristics, the model quickly adapts to new areas with scarce data, achieving accurate cross-regional prediction of groundwater storage. Finally, it simulates future scenarios to generate prediction results and tiered decision-making suggestions. This invention achieves high-precision prediction with physical interpretability, significantly improves the model's generalization ability in heterogeneous regions, and provides a scientific tool for the sustainable management of groundwater resources.

[0045] The technical solution of this embodiment will now be explained in detail with reference to the accompanying drawings.

[0046] The overall technical solution of this embodiment includes the following steps:

[0047] First, acquire multi-source heterogeneous observation data of the target area. The multi-source heterogeneous observation data includes spatiotemporal field data reflecting meteorological, surface, hydrological and geological characteristics.

[0048] Data completion and physical logic verification: Multi-source data is input into a two-way verification network. The completion unit repairs missing observations, and the physical discrimination unit removes contradictory data based on the water balance equation, outputting a complete dataset that conforms to physical logic.

[0049] Based on the complete dataset, spatiotemporal causal analysis was performed to separate the independent temporal and spatial effects of rainfall replenishment and artificial pumping, quantify the contribution of each factor to water level fluctuations, and generate a driving influence matrix.

[0050] Self-learning modeling of physical laws: By combining the influence matrix with the spatial correlation of hydrology, a spatiotemporal prediction model is trained; a physical deviation correction mechanism is introduced during training, and by calculating the deviation between the model's predicted value and the physical law, the model is guided to automatically master the law of water level evolution.

[0051] Cross-regional model adaptation: Apply the knowledge transfer strategy to the target new area; by locking the common physical law layer in the model and only fine-tuning the parameter layer representing the geological characteristics of the new area, a prediction model adapted to the new environment is obtained.

[0052] The scenario conditions are input into the regional adaptation prediction model to generate groundwater storage prediction results and corresponding decision-making suggestions.

[0053] Reference Figure 1This embodiment proposes a physical logic-driven intelligent prediction method for groundwater reserves. It employs techniques such as adversarial physical verification reconstruction, spatiotemporal decoupling causal contribution analysis, physical constraint emergent modeling, and physical anchoring adversarial migration. This method enables closed-loop intelligent analysis throughout the entire process, from reconstruction of physical consistency of multi-source data to causal identification of driving factors, and then to cross-regional knowledge transfer and counterfactual scenario simulation. This improves the physical rationality of groundwater reserve prediction, the depth of analysis, and the system's adaptive evolution capability to complex working conditions.

[0054] Furthermore, the method in this embodiment specifically includes:

[0055] By constructing an adversarial training network to complete and verify multi-source heterogeneous observation data, and using a discriminator with embedded physical equilibrium equations to force constraints to generate data, the complex original observation state is abstracted into a physically self-consistent verification complete dataset.

[0056] Based on this dataset, a spatiotemporally decoupled causal contribution analysis is performed. By utilizing multi-scale time lag and spatial adjacency topology, the independent influence of driving factors in the spatiotemporal dimensions is separated and quantified, generating a causal contribution matrix to guide model learning.

[0057] Based on the matrix, a physical constraint emergent prediction model is constructed, and a physical law discovery reward mechanism is introduced during the training process, so that the internal state of the model can spontaneously align with the real physical laws when dealing with complex hydrological sequences.

[0058] By applying a physical anchoring adversarial migration strategy, the network layer representing a universal physical mechanism is frozen and the region-specific parameters are fine-tuned, enabling the trained model to quickly adapt to a new target region with scarce data.

[0059] Future scenario conditions are input into the adapted model for simulation, and attribution analysis and probability sampling techniques are used to generate groundwater storage prediction results and corresponding hierarchical decision-making suggestions.

[0060] Optionally, in the step of inputting multi-source heterogeneous observation data into a network that is adversarially trained by using a generator for data completion and a discriminator with embedded physical equilibrium equations for verification, to obtain a physically verified complete dataset, the multi-source heterogeneous observation data is spatiotemporally gridded to generate an initial gridded dataset; the initial gridded dataset is input into the generator of the network to generate preliminary complete spatiotemporal field data; the preliminary complete spatiotemporal field data is input into the discriminator of the network, and the discriminator calculates the physical equilibrium error of the preliminary complete spatiotemporal field data; the physical equilibrium error is fed back to the generator as an adversarial signal, and through iterative adversarial training, the data output by the generator is made to minimize the physical equilibrium error while filling in the missing data, and finally the physically verified complete dataset is output.

[0061] In the above process: the multi-source heterogeneous observation data is mapped onto a unified spatiotemporal grid to generate a standardized basic dataset; the initial gridded data is intelligently filled in with missing parts of the basic dataset using a completion unit to generate a preliminary data field covering the entire region; the preliminary data field is input into a verification unit, which detects whether there are any inconsistencies between the data that do not conform to physical laws based on the principle of water balance. The detected inconsistencies and deviations are fed back to the completion unit for dynamic adjustment. Through repeated corrections, the output data not only fills in missing values ​​but also eliminates logical conflicts to the greatest extent possible, forming an accurate and complete dataset.

[0062] Specifically, the engineering objective of this step is to process the raw, multi-source heterogeneous observation data, which contains spatiotemporal gaps and whose physical relationships are not verified, into a high-quality dataset that is spatiotemporally complete and internally consistent in its physical relationships. First, the system performs spatiotemporal gridding processing, resampling all multi-source heterogeneous observation data from different sources and in different formats, including site monitoring data, remote sensing image data, and geological map data, into a standardized spatiotemporal grid, for example, with a spatial resolution of 1 kilometer and a time step of one day, thus generating an initial gridded dataset containing null labels. Subsequently, this initial gridded dataset is used as input to the data completion generator in the adversarial physics verification and reconstruction network. This generator typically employs a convolutional autoencoder structure with multi-scale feature capture capabilities, and its task is to fill in the null values ​​in the initial gridded dataset, generating a preliminary complete spatiotemporal field data that is data-level complete.

[0063] Next, to ensure the physical validity of the generated data, this preliminary complete spatiotemporal field data is input into the network's physical law discriminator. This discriminator is not a traditional classifier; it integrates multiple physical calculation modules capable of gradient backpropagation. For example, its water balance calculation module extracts precipitation, evapotranspiration, surface runoff, and changes in water storage from the preliminary complete spatiotemporal field data for each grid cell, and calculates the physical balance error based on the water balance equation. The formula is

[0064]

[0065] Where P represents precipitation, E represents evapotranspiration, and R represents surface runoff. These four symbols represent the total change in water storage within the grid cell. The physical quantities represented by these symbols are all obtained directly from the preliminary complete spatiotemporal field data or calculated through embedded physical functions.

[0066] Similarly, the energy balance calculation module will also calculate the energy balance error. .

[0067] Ultimately, the system weights and combines these physical balance errors to form a total physical balance error. This physical balance error, as a key adversarial signal, together with the traditional reconstruction loss function, constitutes the total loss and is fed back to the data completion generator. Through hundreds of rounds of iterative adversarial training, the system continuously optimizes the network parameters of the data completion generator, forcing it to generate spatiotemporal field data that not only accurately reconstructs the data at known observation points but also brings the physical balance error close to zero. When the total loss converges to below a preset threshold, training stops, and the final data output from the data completion generator at this point becomes the complete physical verification dataset.

[0068] After the above processing steps, the complete physical verification dataset output by the system possesses two key characteristics. First, it is complete and continuous in both spatiotemporal dimensions, completely eliminating the problem of missing values ​​in the original data. Second, the relationships between the physical variables in the dataset are inherently self-consistent, strictly adhering to core physical laws such as water balance and energy balance built into the physical law discriminator. This high-quality dataset fundamentally guarantees the reliability and physical authenticity of the data source for subsequent causal analysis, model training, and other steps, laying a solid foundation for the final accuracy and interpretability of the entire prediction method.

[0069] Optionally, the step of performing spatiotemporally decoupled causal contribution analysis and generating a causal contribution matrix based on the physical verification complete dataset includes: extracting time series data from the physical verification complete dataset, analyzing it using a multi-scale time lag and memory effect identification algorithm to generate time-dimensional causal contribution components; extracting spatial distribution data from the physical verification complete dataset, and analyzing it in conjunction with the hydrological spatial adjacency relationship to generate spatial-dimensional causal contribution components; coupling the time-dimensional causal contribution components and the spatial-dimensional causal contribution components to separate the direct contributions of natural recharge factors and anthropogenic interference factors to the groundwater storage changes in each spatial grid, thereby forming the causal contribution matrix.

[0070] Furthermore, in the above steps, the changing patterns of various data over time are analyzed, the delayed effects of rainfall or pumping on water levels are identified, and the influence weights of different factors over time are quantified. Combining hydrological and geographical location relationships, the transmission characteristics of water level fluctuations between different regions are analyzed, and the degree of influence of the surrounding environment on specific grid spaces is quantified. By integrating the influence patterns of time and space, the contribution ratios of natural recharge (such as rainfall) and human disturbance (such as pumping) to groundwater changes at each location are accurately separated, ultimately forming a quantified driving influence matrix.

[0071] Specifically, the engineering objective of this step is to accurately quantify and separate the independent and direct impacts of different hydrological driving factors, such as rainfall and irrigation pumping, on groundwater storage changes in each geographic grid unit, based on the complete physical verification dataset generated in the previous step. This allows for the construction of a causal contribution matrix that clearly describes the driver-response relationship. First, the system extracts time-series data of all relevant physical quantities for each grid unit from the complete physical verification dataset. Next, the system activates a pre-defined multi-scale time lag and memory effect identification algorithm. For example, it uses a combination of cross-correlation analysis and convergent cross-mapping to calculate the time series of driving factors (such as rainfall) and response time series (such as groundwater level), identifying the optimal time lag period, typically between 1 and 90 days, and quantifying the causal correlation strength between the two, thereby generating time-dimensional causal contribution components.

[0072] Simultaneously, the system performs analysis in the spatial dimension. It retrieves spatial distribution data from the complete physical verification dataset and integrates external geological parameter distribution maps and digital elevation model (DEM) data to construct a dynamic hydrological spatial adjacency graph. The construction logic of this graph is as follows: nodes represent each grid cell, directed edges between nodes are established based on the hydraulic gradient determined by the DEM, and edge weights are determined by parameters such as the permeability coefficient in the geological parameter distribution map. After construction, the system uses a graph network propagation algorithm or a spatial autoregressive model to analyze how hydrological disturbances (such as upstream reservoir water release) propagate and influence downstream grid cells through this adjacency graph, thereby quantifying the spatial causal contribution components. Finally, the system performs a coupling operation to integrate the analysis results from the temporal and spatial dimensions. For each grid cell i and each driving factor j, its direct contribution can be determined by a coupling function, expressed as follows:

[0073]

[0074] in The final contribution value. The spatial influence coefficient represented by the spatial causal contribution component. To achieve the optimal time lag The strength of the driving factor j under the condition, and Let F be the spatial and temporal weighting coefficients obtained through model calibration, and let F be the nonlinear mapping function integrating the two. By performing this calculation on all grids and all driving factors (distinguished between natural supply factors and anthropogenic interference factors), a complete causal contribution matrix is ​​finally formed.

[0075] The final output of this step is a structured causal contribution matrix. Each row of this matrix corresponds to a spatial grid cell, and each column corresponds to a specific driving factor. The values ​​in the matrix precisely quantify the direct contribution of that driving factor to the change in groundwater storage in that grid cell after considering spatiotemporal transmission effects. This matrix is ​​not a simple list of raw data, but a high-level feature set with clear physical meaning and interpretability obtained after deep causal mining. It provides crucial, decoupled input for subsequent construction of physically constrained emergent spatiotemporal prediction models, ensuring that the model can learn based on real causal relationships rather than spurious statistical correlations, thereby significantly improving the accuracy and robustness of the prediction model.

[0076] Optionally, the step of constructing and training a physically constrained emergent spatiotemporal prediction model based on the causal contribution matrix and the adjacency relationship of the hydrological space to obtain the trained emergent prediction model includes: initializing the network structure and connection weights of the physically constrained emergent spatiotemporal prediction model using the causal contribution matrix and the adjacency relationship of the hydrological space; iteratively training the physically constrained emergent spatiotemporal prediction model using the complete physical verification dataset, outputting the predicted value and simulated intermediate state variables in each training iteration; calculating a first loss based on the error between the predicted value and the true value, and checking whether the simulated intermediate state variables show a trend consistent with the physical equilibrium relationship to generate a physical law discovery reward signal; updating the parameters of the physically constrained emergent spatiotemporal prediction model by combining the first loss and the physical law discovery reward signal until the model converges to obtain the trained emergent prediction model.

[0077] Furthermore, in the above steps, the initial state and connection logic of the prediction model are set using the driving influence matrix obtained in the previous step and the regional geographical correlation, thus building a basic framework for simulating hydrological evolution. The complete dataset is input into the model for iterative training, so that the model can simultaneously simulate intermediate process variables such as water flow and storage changes while outputting water level prediction values. The error between the predicted values ​​and the actual observed values ​​is compared, and the intermediate processes simulated by the model are checked to see if they conform to physical common sense such as "water balance". If the simulation process is closer to the actual physical laws, a positive "logical reward signal" is given to the model. Combining the prediction error and the logical reward signal, the internal parameters of the model are continuously and automatically adjusted. When the model can accurately predict the water level and fully conform to the physical laws, the training is completed and the model parameters are fixed.

[0078] Specifically, the engineering objective of this step is to explicitly and structurally integrate the causal relationship knowledge obtained from previous steps into a deep learning model, and to employ an innovative training mechanism that enables the model not only to fit data but also to actively learn and internalize physical laws. First, the system performs model initialization. Based on the causal contribution matrix and hydrological spatial adjacency relationships, it constructs the initial network structure of a physically constrained emergent spatiotemporal prediction model. Specifically, this model typically employs a graph neural network architecture, where the hydrological spatial adjacency relationships directly define the graph's topology, i.e., the connection methods between nodes; while the values ​​in the causal contribution matrix are used to initialize the initial weights or attention scores for message passing between nodes, ensuring that information flow tends to follow the discovered strong causal paths from the outset. The parameters of the model's specialized units, such as recurrent neural network units simulating hydrological storage and graph convolutional layers simulating hydrological flow, are also assigned initial values ​​at this stage.

[0079] Subsequently, the system enters the model training iteration phase. It uses the complete physical verification dataset as input, feeding it into the initialized physical constraint emergent spatiotemporal prediction model. In each training iteration, after processing the input data, the model outputs two types of information: the predicted value for the next time step, and the intermediate state variables simulated by each specialized unit within the model, such as the simulated water flow between nodes calculated by the graph convolutional layer, and the simulated water storage updated within the recurrent unit. Next, the system calculates a composite loss function. The function consists of two parts. The first part is the predicted loss. The first part is the root mean square error between the predicted value and the true value in the complete physical verification dataset. The second part is the physical constraint loss. The system periodically checks the intermediate state variables of the simulation, calculates whether they satisfy preset physical equilibrium relationships such as mass conservation through a differentiable physics engine, and uses the deviation as... The total loss function is expressed as:

[0080]

[0081] in, This is a hyperparameter, typically dynamically adjusted between 0.1 and 1.0, used to balance the learning weights of prediction accuracy and physical consistency. The design of this composite loss function is essentially a reward mechanism for discovering physical laws; when... Decreasing this loss is equivalent to providing a positive stimulus to the model. Finally, the system utilizes this composite loss. Perform gradient backpropagation and update all trainable parameters of the model using an optimizer such as Adam. Repeat this process until... If the model converges on the validation set, the resulting model is the trained emergent prediction model.

[0082] The final output of this step is a well-trained, current-mode predictive model. This model differs from traditional black-box models in that, through structural initialization and an innovative training mechanism, physical laws are no longer external, rigid constraints, but have been internalized into the model's own behavioral preferences. It not only demonstrates high predictive accuracy on known data, but more importantly, its internal computational logic has spontaneously aligned with hydrophysical processes. This makes the model's predictions and inferences inherently physically plausible and interpretable, providing a solid model foundation for subsequent reliable cross-regional transfer and scenario analysis.

[0083] Optionally, the step of applying a physical anchoring adversarial migration strategy to adapt the trained emergent prediction model to the target new area to obtain a region-adapted prediction model includes:

[0084] In the trained emergent prediction model, a network layer representing a universal physical mechanism is identified as a frozen layer, and a neural network layer representing region-specific parameters is identified as an adjustable layer. A small amount of observation data of the target new region is acquired and input into the trained emergent prediction model to extract feature representations from the intermediate layers of the model. The feature representations are then input into a domain discriminator, which is used to distinguish the source of the features. With the goal of deceiving the domain discriminator while keeping the parameters of the frozen layer unchanged, the parameters of the adjustable layer are adaptively adjusted using only the small amount of observation data to complete the model transfer and obtain the region-adaptive prediction model.

[0085] Furthermore, in the above steps, the trained prediction model is split into two parts: one is a general rule layer representing universal hydrological common sense, which is locked (frozen) and no longer changed; the other is a regional feature layer representing regional differences, which serves as a variable that can be fine-tuned later. A small amount of actual observation data of the target new area is collected and input into the model. After that, the intermediate feature information generated by the model when processing this new data is extracted. A difference discrimination module is introduced, which is specifically responsible for determining whether the features currently output by the model belong to the experience of the source area or the actual situation of the new area. Under the premise of keeping the general rule layer unchanged, by making the features output by the model continuously move closer to the actual situation of the new area, the difference discrimination module cannot distinguish between the two, thereby completing the parameter adjustment of the "regional feature layer" and obtaining a prediction model adapted to the new area.

[0086] Specifically, the engineering objective of this step is to efficiently and reliably transfer and adapt an emergent prediction model trained in a data-rich source region to a data-scarce target area, ensuring that the model quickly learns the unique hydrogeological characteristics of the target area while maintaining its core physical logic. First, the system performs a hierarchical model identification operation. It analyzes the network structure of the trained emergent prediction model and classifies its parameters into two categories based on the function of each layer. The first category consists of network layers representing universal physical mechanisms, such as computational layers simulating fundamental laws like mass conservation and energy balance; these layers are labeled as frozen layers, and their parameters remain unchanged during the transfer process. The second category consists of neural network layers representing region-specific parameters, such as network input layers or partial attention weight layers associated with specific permeability coefficients and soil types; these layers are labeled as adjustable layers.

[0087] Next, the system enters the adversarial transfer training phase. It first acquires a very limited amount of observational data from the target new region, which may only be 5% to 10% of the data from the source region. This limited amount of observational data is input into a partially frozen emergent prediction model, which performs forward propagation and generates feature representations for this data in its intermediate layers. Simultaneously, a pre-defined domain discriminator, typically a lightweight binary classifier, is used to receive these feature representations. The domain discriminator's task is to distinguish whether the input feature representations originate from the source region data or the target new region data. The core of the transfer training is an adversarial game process, and its total loss function can be expressed as…

[0088]

[0089] in, It is migration loss. It is the prediction loss of the model on a small amount of observation data in the target new area, used to drive the model to learn the specific data patterns of the target new area. This is the adversarial loss, calculated from the classification error of the neighborhood discriminator. The goal of updating the adjustable layer parameters is to minimize... , while maximizing This involves "deceiving" the domain discriminator, making it unable to distinguish the source of features. Gamma is a weighting coefficient used to balance these two objectives. By alternately optimizing the domain discriminator and the adjustable layers of the model, the system ultimately aligns the feature representation generated by the adjustable layers with the feature distribution of the source region, while minimizing the prediction error in the target new region. When the migration loss... When convergence occurs, training is complete, and the resulting model is the region-adaptive prediction model.

[0090] After this step, the system outputs a finely tuned regional adaptation prediction model. This model successfully combines the universal hydrophysical knowledge learned from the source region (encapsulated in the frozen layer) with the unique geological and hydrological characteristics of the target area (learned into the adjustable layer). This "physical anchoring" transfer method avoids the physical logic drift or overfitting to small sample data that may occur with traditional fine-tuning methods. The final model possesses both strong generalization ability and accurate capture of local details in the target area, enabling it to make high-precision and high-reliability groundwater storage predictions even under conditions of extremely scarce data.

[0091] Optionally, the step of inputting future scenario conditions into the regional adaptive prediction model for simulation to generate groundwater storage prediction results and corresponding decision suggestion data includes:

[0092] Multiple sets of future scenario conditions, including different climate change parameters and different water resource management parameters, are constructed. Each set of future scenario conditions is input into the regional adaptive prediction model to simulate the corresponding spatiotemporal evolution trajectory data of groundwater storage. Based on the spatiotemporal evolution trajectory data of groundwater storage, the contribution data of the core driving factors causing trajectory changes are analyzed and output. The Monte Carlo method is used to perform probability sampling and simulation on multiple sets of future scenario conditions to generate a comparative analysis chart of future multi-scenario projections of groundwater storage. By combining the spatiotemporal evolution trajectory data of groundwater storage, the contribution data of the core driving factors, and the comparative analysis chart of future multi-scenario projections, the hierarchical decision-making recommendation data, including specific control measures and expected effects, is generated.

[0093] Furthermore, in the above steps, multiple sets of future simulation scenarios (such as continuous drought, water-saving irrigation, etc.) containing different rainfall and climate changes and different water management schemes are constructed as input conditions for the model; each set of simulation scenarios is input into the adapted model to simulate and record the dynamic evolution path of groundwater storage over time in the future; the water level changes under different simulation paths are analyzed, and the contribution ratio of factors such as natural climate and artificial pumping to future water level fluctuations is quantified and output to identify the key points affecting the water level; through multiple random sampling simulations, the probability of future groundwater storage is calculated to quantify the risk level of water level over-extraction in different regions; by combining the water level evolution path, the contribution ratio of factors, and risk zoning, a hierarchical management recommendation including specific water-saving measures, expected recovery effects, and priority treatment areas is automatically generated.

[0094] Specifically, the engineering objective of this step is to utilize the trained regional adaptation prediction model to conduct multi-scenario simulations and interpretability analyses, ultimately generating a set of decision support information that directly serves water resource management, including quantitative predictions and qualitative recommendations. First, the system performs multi-scenario construction and simulation operations. Based on the output of future climate change prediction models (such as CMIP6) and regional water resource management plans, it constructs multiple sets of future scenario conditions with clearly defined boundary conditions. For example, it constructs specific combination scenarios such as "under a high-emission scenario, agricultural irrigation water consumption increases by 15%" or "under a medium-emission scenario, water-saving measures are implemented to reduce pumping volume by 10%." Subsequently, the system uses each set of future scenario conditions as input to drive the regional adaptation prediction model to perform long-term forward simulations, simulating groundwater storage changes over the next 5 to 30 years, generating a set of spatiotemporal evolution trajectory data for groundwater storage.

[0095] Next, the system initiates attribution and probabilistic analysis. For key evolution trajectories, such as predicting a severe drop in water level in a certain area, the system utilizes the model's internal interpretability tools to attribute the result. That is, while keeping other conditions constant, it adjusts each driving factor (such as precipitation or pumping volume) one by one, quantifying the contribution of the change in this factor to the final water level drop, and ultimately outputting a core driving factor contribution data map that clearly defines the influence weight of each factor. Simultaneously, to quantify the uncertainty of the prediction, the system uses the Monte Carlo method to conduct thousands of simulations within the parameter space of future scenario conditions (such as random sampling of precipitation within 20% above and below the predicted mean). Based on the results of these thousands of simulations, the system generates a multi-scenario comparative analysis map of future scenarios. Finally, the system generates decision recommendation data. It integrates all the above analysis results, including the spatiotemporal evolution trajectory data of groundwater storage under different management strategies, the contribution data of core driving factors, and the multi-scenario comparative analysis map of future scenarios, automatically generating hierarchical decision recommendation data that includes specific control measures, implementation areas, recommended time windows, and expected effect assessments.

[0096] The final output of this step is a highly integrated set of decision support products, rather than a single predicted value. It includes quantitative simulation results of multiple future possibilities, in-depth interpretability analysis of the causes of change, comprehensive quantification of predictive uncertainty, and finally, actionable, tiered management recommendations. This output transforms complex model predictions into intelligence that water resource managers can directly understand and use, achieving a complete closed loop from data to knowledge to action, and greatly enhancing the application value of scientific prediction in practical engineering management.

[0097] Optionally, the step of the discriminator calculating the physical equilibrium error of the preliminary complete spatiotemporal field data includes: receiving input data, output data, and stored change data from the preliminary complete spatiotemporal field data; calculating the closure error between the input data, output data, and stored change data through differentiable operations; and outputting the closure error as part of the physical equilibrium error.

[0098] Specifically, the engineering objective of this step is to precisely define and implement the water balance calculation module in the physical law discriminator, ensuring that it can quantitatively verify the physical consistency of the input data and generate gradient signals that can be used for neural network training. First, this module is designed as a differentiable computational unit that receives multiple physical quantities corresponding to a single spatiotemporal grid cell from the preliminary complete spatiotemporal field data as input. Specifically, it slices the input tensor to obtain four key subsets: input data representing inflow, mainly precipitation and lateral inflow; output data representing outflow, including evapotranspiration, surface runoff, and lateral outflow; and storage change data representing changes in system state, namely, changes in groundwater storage.

[0099] Subsequently, the module internally performs the core closure error calculation. It performs algebraic operations on the extracted physical quantities based on the classic water balance equation. All input physical quantities are pre-standardized to equivalent water depth (e.g., millimeters per day) to ensure the correctness of the physical logic of the calculation. The calculation process follows a specific formula to quantify the degree of water imbalance in the grid cell within a specific time step; this closure error is the key factor in the calculation. It can be represented as

[0100]

[0101] in, This indicates taking the absolute value, where P is the precipitation. E is the lateral inflow, R is the evapotranspiration, and E is the surface runoff. This refers to the lateral outflow. This is to store the variation data. These variables are all values ​​directly extracted from the preliminary complete spatiotemporal field data. The calculated closure error... , is a scalar value that directly reflects the degree to which the input data deviates from the law of conservation of mass at that point in time and space. Finally, this closure error This is used as the output of this module and, as part of the physical equilibrium error, integrated into the total loss function of the adversarial physics verification reconstruction network. Since the entire computation process is based on differentiable fundamental operations (addition, subtraction, absolute value), the closure error... The gradient can be automatically calculated and backpropagated to the data completion generator, thereby guiding its parameter updates.

[0102] The direct output of this step is that the physical law discriminator possesses the core capability to quantify the physical consistency of the data. Through this differentiable water balance calculation module, the system can transform abstract physical laws into concrete, calculable loss terms. This allows the adversarial training process to no longer merely pursue the similarity of data in statistical distribution, but introduces strict physical constraints. Ultimately, the existence of this module ensures that the complete physical verification dataset output by the entire adversarial physical verification reconstruction network highly conforms to the fundamental principle of mass conservation at every spatiotemporal point, providing a physically reliable data foundation for all subsequent analyses.

[0103] Optionally, the network structure of the physical constraint emergent spatiotemporal prediction model includes a memory unit simulating water storage state and a graph convolutional connection unit simulating water flow transmission, wherein the initial connection weights of the graph convolutional connection unit are determined by the hydrological spatial adjacency relationship.

[0104] Furthermore, the memory unit, which simulates the water storage state, is used to record and store the historical state and trend of water level changes in the target area over time; the graph convolution connection unit, which simulates water flow transmission, is used to characterize the mutual infiltration and transport process of water between different geographical locations; it also includes an association module, the initial connection relationship of which is determined by the regional hydrological spatial location map, ensuring that the model conforms to the real geographical laws of "great influence from nearby areas and small influence from distant areas" and "flow with terrain" in the initial state.

[0105] Specifically, the engineering objective of this step is to directly map the concepts of hydrophysical processes into the computational architecture of the model by designing specific neural network structural units, thereby enabling the model's internal operating mechanism to naturally simulate real-world hydrological dynamics. In the physically constrained emergent spatiotemporal prediction model, the system constructs two core dedicated units. The first is a memory unit that simulates the water storage state, typically implemented using a variant of Long Short-Term Memory (LSTM) or Gated Recurrent Unit (GRU). Each spatial grid node is associated with such a memory unit, whose internal state vector is physically designed to characterize the current groundwater storage of that grid. This unit dynamically controls the inflow, maintenance, and outflow of information through its input gate, forget gate, and output gate structure, directly simulating the water storage, consumption, and release processes of aquifers in engineering terms.

[0106] The second type is the graph convolutional connection unit that simulates water flow. This unit is the core of the graph neural network and is responsible for handling the interactions between different grid nodes. Unlike standard graph convolutional networks, its connection weights are not randomly initialized, but are initialized based on the hydrological spatial adjacency relationships determined in the previous steps.

[0107] Specifically, such as Figure 2 As shown, the system constructs an adjacency matrix A, where if grid node j has a hydraulic connection to node i (e.g., j is upstream of i), then the matrix elements... It is assigned a non-zero initial value, the magnitude of which is proportional to the hydraulic gradient and permeability between the two; otherwise, it is zero. During model training, information (such as water volume) is weighted and aggregated from the feature vector of a node through this graph convolution kernel defined by A, and then passed to its neighboring nodes. This process directly simulates the lateral flow of groundwater between different plots in engineering. Each layer of the model contains these graph convolutional connection units, allowing information to propagate and update along physically plausible paths throughout the study area.

[0108] Figure 2 The diagram, from left to right, illustrates three processes: physical hydrogeological network (regional flow simulation), adjacency matrix construction, and graph convolutional network hierarchical linking.

[0109] In the physical hydrogeological network (regional flow simulation) stage, for the given topography, the lateral flow is determined by the local topography and geology. The topography is divided into 9 units according to the set rules, which are represented by the numbers 1, 2, 3, 4, 5, 6, 7, 8, and 9 respectively. The lateral flow is represented by the gray arrows pointing between the grid units, and its specific flow direction is dominated by the hydraulic gradient between the units. The meaning represented is the connection weight between lattice cells i and j, and its magnitude is proportional to the hydraulic gradient and permeability, i.e. This intuitively reflects the ability and trend of water flow transmission between units.

[0110] In the adjacency matrix construction phase, a 9x9 matrix is ​​used to represent the adjacency matrix, where white cells represent zero initial weights and dark cells represent non-zero initial weights. The initial weights are determined as follows: if there is no physical flow path between cells, the corresponding connection weight in the adjacency matrix is ​​zero; if there is a physical flow path between cells, the corresponding connection weight in the adjacency matrix is ​​not zero. For example, in the matrix... , The location has a non-zero weight, corresponding to the existence of lateral flow paths between units 4 and 8 and the central unit 5 in the physical hydrological network. The magnitude of the connection weight is determined based on the flow potential.

[0111] In the layer connection stage of the graph convolutional network, the simplified aggregation formula of the graph convolutional neural network is first shown:

[0112]

[0113] in, For the first Grid unit Feature information, For physically anchored connection weights, For activation function, For the first Grid unit The aggregation update feature.

[0114] During the network initialization phase, weights are directly based on the physical adjacency matrix. The system is designed so that subsequent information propagates only along valid physical paths, and the weights are learnable during training but are always strictly limited by physical constraints, ultimately achieving compliant propagation and feature updating of hydrological information in graph convolutional networks.

[0115] Through this structured design, the introduction of specialized units gives the internal computation of the physically constrained emergent spatiotemporal prediction model a clear physical correspondence. The memory unit simulating water storage states enables the model to capture the temporal dynamics and memory effects of the groundwater system, while the graph convolutional connection unit initialized by hydrological spatial adjacency relationships ensures the physical rationality of spatial interactions. This "hard-coded" physical prior in the architecture, compared to simply adding constraints to the loss function, guides the model to learn the correct physical processes at a deeper level. This results in a well-trained emergent prediction model that not only provides accurate predictions but also has a more interpretable internal state, providing a transparent model foundation for subsequent mechanism analysis and decision support.

[0116] Optionally, the domain discriminator is a convolutional neural network classifier, and in the physical anchoring adversarial transfer strategy, the domain discriminator and the adjustable layer in the trained emergent prediction model are trained alternately in adversarial mode.

[0117] In this embodiment, an image feature recognition logic is used to construct a difference discriminator, which is used to compare the feature differences of the model when processing new and old regional data. During the model adaptation process, the difference discriminator and the adjustable parameter layer of the model conduct reciprocal exercises (adversarial training): the discriminator tries to find differences, and the parameter layer tries to eliminate differences, until the model can perfectly capture the unique patterns of the new region.

[0118] Specifically, the engineering objective of this step is to precisely guide the parameter updates of the adjustable layers in the emergent prediction model by introducing a dedicated domain discriminator and establishing an adversarial training mechanism. This ensures that the model adapts to the data distribution of the new target region without losing the general features learned from the source region. First, the system-defined domain discriminator is an independent, lightweight convolutional neural network classifier. This classifier typically consists of 2 to 3 convolutional layers and fully connected layers. Its input is the feature representation from the intermediate layers of the emergent prediction model, and its output is a single probability value representing the confidence that the feature representation originates from the source region. Its design aims to efficiently capture subtle differences in the feature distribution.

[0119] In the execution of the physical anchoring adversarial transfer strategy, the system employs an alternating adversarial training mode. This process is divided into two phases and repeated cyclically. In the first phase, the system fixes the parameters of all layers in the emergent prediction model, including the frozen layer and the adjustable layer. Then, it extracts intermediate layer feature representations from a small amount of observation data from the source region and the new target region, respectively, and assigns them corresponding domain labels (e.g., 0 for the source region and 1 for the new target region). These labeled feature representations are used to train the domain discriminator, which aims to minimize the classification cross-entropy loss, i.e., improve its ability to distinguish different domain features. In the second phase, the system, in turn, fixes the parameters of the domain discriminator and unfreezes the adjustable layer in the emergent prediction model. At this point, the system only inputs a small amount of observation data from the new target region into the model and extracts intermediate layer feature representations. The training objective is to update the parameters of the adjustable layer so that the domain labels output by the generated feature representations, after passing through the domain discriminator, are misclassified as being from the source region's label as much as possible. This process is achieved by maximizing the classification loss of the domain discriminator, essentially "deceiving" the discriminator. These two phases are performed alternately, with the batch size for each training round typically set between 16 and 64, and the learning rate using a low initial value to ensure the stability of fine-tuning.

[0120] Through this alternating adversarial training mechanism, the system establishes a dynamic game process. The presence of the domain discriminator provides a clear, feature distribution-aligned optimization direction for the parameter updates of the adjustable layers. Compared to fine-tuning solely based on task loss, this approach more effectively prevents the model from overfitting to small sample data of the new target region and encourages the model to learn a domain-invariant feature representation. Ultimately, this adversarial training process ensures that the region-adaptive prediction model maintains its core feature representation capabilities consistent with those trained in the data-rich source region while learning the specificity of the new target region, thus achieving efficient and robust knowledge transfer.

[0121] To verify the feasibility and practical effect of this invention, it was applied to the dynamic prediction and management of groundwater storage in a large irrigation area in the North China Plain (hereinafter referred to as the "target area"). This area has a high proportion of agricultural water use and severe groundwater over-extraction. Existing monitoring data suffers from uneven station distribution, missing data for certain time periods, and inconsistent physical calibers among multiple data sources. Traditional prediction models struggle to accurately depict its complex hydrological processes.

[0122] The engineering objectives of this embodiment are: (1) to integrate multi-source heterogeneous data within the target area to generate a high-quality dataset that is spatiotemporally complete and physically self-consistent; (2) to train a model based on the dataset that can accurately predict future changes in groundwater storage and can be transferred to adjacent data-scarce areas; and (3) to produce decision recommendations that can guide the adjustment of irrigation strategies.

[0123] Step 1: Physically verify the generation of the complete dataset.

[0124] This embodiment acquired observational data of the target area from 2010 to 2020, including daily data from 35 groundwater level monitoring stations, MODIS remote sensing evapotranspiration and surface temperature products (8-day composite) covering the entire area, TRMM precipitation data (daily, 0.25 degree resolution), regional geological maps (including permeability coefficient distribution), and a 30-meter resolution digital elevation model (DEM).

[0125] First, spatiotemporal gridding was performed, resampling all data to a 1 km × 1 km grid with a time step of 1 day, generating an initial gridded dataset containing approximately 25% null labels. This dataset was then fed into an adversarial physics verification and reconstruction network. The network's generator employs a convolutional autoencoder structure, and the discriminator integrates a differentiable physics computation module based on the water balance equation.

[0126] In a specific training iteration, for the numbered... For the data point on June 15, 2015, the preliminary complete spatiotemporal field data generated by the generator includes the following values: precipitation P = 15.2 mm, evapotranspiration E = 5.8 mm, surface runoff R = 1.1 mm, and changes in groundwater storage. = 7.5 mm. These values ​​are fed into the water balance calculation module of the discriminator, which calculates the physical closure error. = abs(P - E - R - dS) = abs(15.2 - 5.8 - 1.1 - 7.5) = abs(0.8) = 0.8 mm. This error As an adversarial signal, it, together with the reconstruction loss, constitutes the total loss and is backpropagated to update the generator's network parameters, forcing the generator to produce a value closer to physical equilibrium in the next iteration. value.

[0127] After 500 rounds of iterative adversarial training, the total loss function converged to below the preset threshold 1e-4. The final output physical verification complete dataset had an average physical closure error of less than 0.1 mm / day across all grid cells, achieving 100% data integrity. Figure 4 As shown.

[0128] Figure 4This paper fully demonstrates the entire process of transforming multi-source heterogeneous observational data from raw gridded data with missing data labels into a physically self-consistent complete dataset. The process comprises four core stages: input, reconstruction network layer, discriminator verification of water balance, and output. In the input stage, grid cells are numbered 1 through 9, representing the raw gridded dataset with missing data labels. In the reconstruction network layer, the input data undergoes missing value imputation and preliminary spatiotemporal field reconstruction via a convolutional autoencoder generator. The discriminator calculates the physical balance error based on the water balance equation, determines whether the data is autonomous, and backpropagates the error to update the generator. Finally, the output converges to a physically verified complete dataset with 100% data integrity and a physical error of less than 0.1 mm / day, providing high-quality physically self-consistent data for subsequent analysis.

[0129] Table 1 shows a comparison of the specific data quality of the original observation data, the traditional interpolation dataset, and the complete physical verification dataset of this embodiment.

[0130] Table 1. Comparison of Dataset Quality

[0131]

[0132] Step 2: Generation of the causal contribution matrix.

[0133] Based on the generated complete physical verification dataset, the system performs a spatiotemporally decoupled causal contribution analysis. In the time dimension, a convergent cross-mapping method is used to analyze the relationship between rainfall and grid data. The relationship with groundwater levels identified an optimal time lag period of 18 days and a causal correlation strength of 0.72. Spatially, a hydrological spatial adjacency map constructed based on the DEM and geological maps shows that the upstream reservoir (grid) The water release behavior of the grid is achieved through a highly permeable aquifer. There is a hydraulic connection, and the spatial influence coefficient is... It was quantified as 0.45.

[0134] The quantization diagram of the spatiotemporal causal contribution matrix generated for the embodiments of the present invention is as follows: Figure 3 As shown in the figure, the horizontal axis (X-axis) represents the four core driving factors: precipitation (P), evapotranspiration (E), pumping (W), and runoff (R). The vertical axis (Y-axis) represents typical geographic monitoring grids within the target area, labeled as grids G101, G102, G103, G104, and G105. Each cell in the matrix is ​​labeled with a specific value, representing the contribution weight of the corresponding driving factor to the change in groundwater storage in that grid. The values ​​are normalized to a score of 0–1, with larger values ​​indicating a stronger influence of the driving factor.

[0135] The matrix numerical distribution shows that grids G101 and G102 are most significantly influenced by precipitation, with scores of 0.78 and 0.65 respectively; grid G103 is dominated by pumping, with a score as high as 0.75; grids G104 and G105 exhibit the characteristics of combined effects of natural and anthropogenic factors. By quantitatively comparing the contributions of each factor, the dominant driving factors of groundwater storage changes in different regions can be intuitively distinguished, clarifying the differences in contributions between natural recharge and anthropogenic disturbance. This provides interpretable and strongly causal feature inputs for subsequent physically constrained emergent spatiotemporal prediction models.

[0136] The system uses coupling functions The direct contribution of each driving factor to each grid cell was calculated. A partial example of the final causal contribution matrix is ​​shown below. Detailed data is shown in Table 2.

[0137] Table 2 Causal Contribution Matrix

[0138]

[0139] Step 3: Construction, training, and transfer of the emergent prediction model.

[0140] The system constructs a physically constrained emergent spatiotemporal prediction model using a graph neural network (GNN) as its framework. Each grid node is associated with a memory unit based on a GRU variant to simulate water storage conditions. The graph topology of the GNN is defined by hydrological spatial adjacency relationships, while the initial weights of the graph convolutional connection units simulating water flow transmission are directly initialized using the values ​​in the aforementioned causal contribution matrix.

[0141] During training, a composite loss function is used. ,in To predict the root mean square error, The model identifies reward signals based on the deviation between simulated water flow and the law of conservation of mass. Using the Adam optimizer with an initial learning rate of 1e-4, it is trained in the data-rich region A (70% of the target region) to obtain a well-trained emergent prediction model.

[0142] To validate the physical anchoring adversarial transfer strategy, the model was applied to region B, which has scarce data (accounting for 30% of the target area, with only 10% of the monitoring data from region A). In the model, the graph convolutional layer simulating mass conservation was labeled as a frozen layer, while the input layer related to region-specific soil parameters was labeled as an adjustable layer. A domain discriminator with two convolutional layers was introduced and trained adversarially against the adjustable layers of the model. The training batch size was 32, and the fine-tuning learning rate for the adjustable layers was 1e-5. After adversarial training, the domain discriminator could not effectively distinguish whether features came from region A or region B (classification accuracy approached 50%), indicating that the model successfully learned domain-invariant features and completed transfer adaptation.

[0143] Step 4: Scenario simulation and decision recommendation generation.

[0144] We will use a region-adapted prediction model that has performed well in region B to simulate scenarios for the next 5 years. Two scenarios will be set up:

[0145] Scenario 1 (Continued Status Quo): Rainfall is predicted based on the CMIP6 climate model SSP2-4.5 path, and agricultural irrigation pumping remains at the current level.

[0146] Scenario 2 (Water Conservation Management): The climate conditions are the same as above, but the pumping volume is reduced by 20% by optimizing irrigation technology.

[0147] Simulation results show that under Scenario 1, the average groundwater level in Area B will drop by 3.2 meters; under Scenario 2, the water level will only drop by 0.9 meters. SHAP value attribution analysis confirms that agricultural pumping is the core driving factor for water level changes, contributing 78%. The system further generated risk analysis data through 1000 Monte Carlo simulations, showing that under Scenario 1, the probability of groundwater levels dropping to the ecological red line in more than 40% of Area B is higher than 80%.

[0148] Figure 5The figure shows a comparative analysis of the future evolution trajectory of groundwater reserves generated based on the simulation function of this invention. The horizontal axis represents the future prediction time step (quarter), and the vertical axis represents the Groundwater Reserve Index (GSI), a dimensionless normalized index used to characterize the relative abundance or scarcity of groundwater reserves; a higher value indicates more abundant reserves. The solid line represents Scenario I: the change in the reserve index under the current situation, showing a trend of continuous decline in groundwater reserves without intervention; the dashed line represents Scenario II: the expected evolution trajectory under water-saving management, reflecting the slowing decline and gradual stabilization of groundwater reserves after water-saving intervention measures are implemented. The diagonal lines surrounding the curves in the figure represent uncertainty intervals, where uncertainty interval I corresponds to the 95% confidence interval of Scenario I, and uncertainty interval II corresponds to the 95% confidence interval of Scenario II, used to intuitively quantify the range of fluctuations and risk levels of results under different prediction scenarios. Through multi-scenario comparison and uncertainty analysis, this figure clearly presents the differences in the impact of different management strategies on the long-term changes in groundwater reserves, providing key visual evidence and scientific support for the system to generate hierarchical decision-making suggestions and formulate water resource regulation plans.

[0149] Based on the above analysis, the system generates the following decision recommendation: "It is recommended to prioritize the promotion of water-saving technologies such as drip irrigation in the southeastern part of Zone B (highest risk level). It is expected that the risk of water level over-extraction in this sub-region can be reduced by 65% ​​within 5 years, achieving basic stability of the groundwater level." Detailed data is shown in Table 3.

[0150] Table 3. Comparison of prediction accuracy (RMSE) of different models in regions A and B.

[0151]

[0152] Tables 1-3 above record the actual application data of this invention in a large irrigation area in the North China Plain, and demonstrate in detail the comprehensive performance of the system in data reconstruction, causal analysis, model prediction and migration application.

[0153] Table 1 clearly shows that, compared with traditional interpolation methods, the physical verification reconstruction network of the present invention can not only generate spatiotemporally complete data, but more importantly, it significantly reduces the physical inconsistency of the data (physical closure error is reduced from 1.25 to 0.08 mm / day), thereby reducing the root mean square error of the prediction basis by more than 70%.

[0154] The example causal contribution matrix in Table 2 quantifies and decouples the influence of fuzzy driving factors, providing direct, data-driven prior knowledge for the subsequent construction of physically interpretable model structures.

[0155] The comparison results in Table 3 are particularly crucial, demonstrating the superiority of the model presented in this invention. In the data-rich region A, its accuracy (RMSE 0.29 m) far surpasses that of the traditional model (0.98 m). More importantly, in the data-scarce region B, through the physical anchoring adversarial migration strategy, the model's prediction accuracy (0.42 m) is fundamentally improved compared to both the conventional fine-tuning (2.15 m) and the pre-migration (1.88 m) results, validating the powerful ability of this invention to solve small sample problems. This fully demonstrates the effectiveness, robustness, and practical value of the full-chain method of this invention.

[0156] It should be noted that the electrical connections between the various units described above do not necessarily represent direct or indirect connections. Any indirect connection method can be applied to the embodiments of the present invention as long as it achieves the purpose of the present invention. The above descriptions are merely exemplary embodiments of the present invention and should not be construed as limiting the scope of the present invention.

[0157] Example 2

[0158] This embodiment discloses a physical logic-driven intelligent prediction system for groundwater reserves.

[0159] A physical logic-driven intelligent prediction system for groundwater storage includes:

[0160] The multi-source data acquisition and physical logic verification module is configured to: complete and physically verify the multi-source heterogeneous observation data related to groundwater storage in the target area, and obtain a complete physical verification dataset that conforms to physical logic;

[0161] The driving factor analysis module is configured to perform spatiotemporally decoupled causal contribution analysis on the complete physical verification dataset, separate the independent effects of driving factors in the time and space dimensions, and generate a causal contribution matrix.

[0162] The self-learning modeling module is configured to: construct and train a physical constraint emergent spatiotemporal prediction model based on the causal contribution matrix; during the training process, a physical law discovery reward mechanism is introduced to enable the internal state of the model to spontaneously align with the physical law, thus obtaining a well-trained emergent prediction model.

[0163] The cross-regional knowledge transfer module is configured to: apply a physical anchoring adversarial transfer strategy to adapt the trained emergent prediction model to the target new area, thereby obtaining a regionally adapted prediction model;

[0164] The scenario simulation and decision-making module is configured to input future scenario conditions into the regional adaptive prediction model for simulation, and generate groundwater storage prediction results and corresponding decision-making suggestions.

[0165] The aforementioned multi-source data acquisition and physical logic verification module actually includes a multi-source data acquisition module and a physical logic verification module, wherein:

[0166] The multi-source data acquisition module is used to acquire multi-source heterogeneous observation data of the target area;

[0167] The physical logic verification module includes a network trained adversarially by using a generator for data completion and a discriminator for verification using an embedded physical equilibrium equation. This network receives the multi-source heterogeneous observation data and outputs a complete physical verification dataset.

[0168] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0169] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A physical logic-driven intelligent prediction method for groundwater reserves, characterized in that, Includes the following steps: The multi-source heterogeneous observation data related to groundwater storage in the target area were supplemented and physically verified to obtain a complete physical verification dataset that conforms to physical logic. A spatiotemporally decoupled causal contribution analysis was performed on the complete dataset of physical verification to separate the independent influence of driving factors in the temporal and spatial dimensions and generate a causal contribution matrix. Based on the causal contribution matrix, a physically constrained emergent spatiotemporal prediction model is constructed and trained. During the training process, a reward mechanism for discovering physical laws is introduced to enable the internal state of the model to spontaneously align with physical laws, thus obtaining a well-trained emergent prediction model. By applying a physical anchoring adversarial migration strategy, the trained emergent prediction model is adapted to the target new area to obtain a region-adapted prediction model. Future scenario conditions are input into the regional adaptive prediction model for simulation, generating groundwater storage prediction results and corresponding decision-making suggestions. The process of training a physically constrained emergent spatiotemporal prediction model to obtain a well-trained emergent prediction model includes: Based on the causal contribution matrix and combined with the hydrological spatial adjacency relationship, the network structure and connection weights of the physical constraint emergent spatiotemporal prediction model are initialized. The physical constraint emergent spatiotemporal prediction model is iteratively trained using a complete physical verification dataset. In each training iteration, the predicted value and the simulated intermediate state variables are output. The first loss is calculated based on the error between the predicted and actual values, and the intermediate state variables in the simulation are checked to see if they show a trend consistent with the physical equilibrium relationship, so as to generate a reward signal for the discovery of physical laws. By combining the first loss with physical laws to discover reward signals, the parameters of the physical constraint emergent spatiotemporal prediction model are updated until the model converges, thus obtaining a well-trained emergent prediction model. The network structure of the physical constraint emergent spatiotemporal prediction model includes a memory unit that simulates the water storage state and a graph convolutional connection unit that simulates water flow transmission. The initial connection weights of the graph convolutional connection unit are determined by the hydrological spatial adjacency relationship. The intermediate state variables include the simulated water flow between nodes calculated by the graph convolutional layer and the simulated water storage updated within the cyclic unit. The hydrological spatial adjacency relationships are generated from geological parameter distribution maps and digital elevation model data.

2. The physical logic-driven intelligent prediction method for groundwater reserves as described in claim 1, characterized in that, Multi-source heterogeneous observation data related to groundwater storage in the target area include spatiotemporally distributed precipitation data, evapotranspiration data, surface temperature data, groundwater level monitoring data, geological parameter distribution maps, and digital elevation model data.

3. The physical logic-driven intelligent prediction method for groundwater reserves as described in claim 1, characterized in that, Multi-source heterogeneous observation data is input into a network that incorporates adversarial training between a generator and a discriminator. The generator performs data completion, and the network is validated using a discriminator embedded with physical equilibrium equations. This results in a complete physical verification dataset, which includes: Spatiotemporal gridding is performed on multi-source heterogeneous observation data to generate an initial gridded dataset; The initial gridded dataset is input into the generator to generate preliminary complete spatiotemporal field data; The preliminary complete spatiotemporal field data is input into the discriminator, which calculates the physical equilibrium error of the preliminary complete spatiotemporal field data. The physical balance error is fed back to the generator as an adversarial signal. Through iterative adversarial training, the generator outputs data to minimize the physical balance error while filling in the missing data, and finally outputs a complete physical verification dataset. The discriminator calculates the physical equilibrium error of the preliminary complete spatiotemporal field data, specifically including: Input data, output data, and stored change data are received from the preliminary complete spatiotemporal field data. The input data includes precipitation and lateral inflow. The output data includes evapotranspiration, surface runoff, and lateral outflow. The stored change data includes changes in groundwater storage. The closure error between the input data, output data, and stored change data is calculated using differentiable operations. The closure error is output as part of the physical equilibrium error.

4. The physical logic-driven intelligent prediction method for groundwater reserves as described in claim 1, characterized in that, A spatiotemporally decoupled causal contribution analysis was performed on the complete physical verification dataset to separate the independent influences of driving factors in the temporal and spatial dimensions, generating a causal contribution matrix, specifically including: Time series data are extracted from the complete dataset of physical verification to identify the time lag period and causal correlation strength between driving factors and water level response, and to generate time-dimensional causal contribution components. Spatial distribution data are extracted from the complete dataset of physical verification and analyzed in conjunction with hydrological spatial adjacency relationships to generate spatial causal contribution components. The temporal causal contribution component is coupled with the spatial causal contribution component to separate the direct contribution of natural recharge factors and anthropogenic interference factors to the groundwater storage change of each spatial grid, thus forming a causal contribution matrix.

5. The physical logic-driven intelligent prediction method for groundwater reserves as described in claim 1, characterized in that, By applying a physical anchoring adversarial migration strategy, the trained emergent prediction model is adapted to the target new area to obtain a region-adapted prediction model, which specifically includes: In the well-trained emergent prediction model, the network layer representing the universal physical mechanism is identified as the frozen layer, and the neural network layer representing the region-specific parameters is identified as the adjustable layer. Acquire observation data of the target new area and input the observation data into the trained emergent prediction model to extract the feature representation of the intermediate layer of the model; The feature representations of the intermediate layers of the model are input into a domain discriminator, which is used to distinguish the source of the features. With the goal of deceiving the domain discriminator, while keeping the parameters of the frozen layer unchanged, the parameters of the adjustable layer are adaptively adjusted using only the observation data of the new target area to complete the model transfer and obtain the region-adaptive prediction model. in: The network layer representing universal physical mechanisms includes a computational layer that simulates the fundamental laws of mass conservation and energy balance. Neural network layers representing region-specific parameters include network input layers or partial attention weight layers associated with specific permeability coefficients and soil types.

6. The physical logic-driven intelligent prediction method for groundwater reserves as described in claim 5, characterized in that, The domain discriminator is a convolutional neural network classifier. In the physical anchoring adversarial transfer strategy, the domain discriminator and the adjustable layers in the pre-trained emergent prediction model are trained in an adversarial manner. The task of the domain discriminator is to distinguish whether the input feature representation originates from source region data or target new region data.

7. The physical logic-driven intelligent prediction method for groundwater reserves as described in claim 1, characterized in that, Future scenario conditions are input into a regional adaptive prediction model for simulation, generating groundwater storage prediction results and corresponding decision-making recommendations, specifically including: Construct multiple sets of future scenario conditions that include different climate change parameters and different water resource management parameters; Each set of future scenario conditions is input into the regional adaptation prediction model to simulate the corresponding spatiotemporal evolution trajectory data of groundwater storage. Based on the spatiotemporal evolution trajectory data of groundwater storage, we analyze and output the contribution data of the core driving factors that lead to the trajectory changes; The Monte Carlo method was used to perform probability sampling and simulation of multiple future scenarios to generate a comparative analysis chart of future multi-scenario projections of groundwater storage. By integrating spatiotemporal evolution data of groundwater reserves, contribution data of core driving factors, and comparative analysis maps of future multi-scenario projections, hierarchical decision-making recommendations data containing specific control measures and expected effects are generated.

8. The physical logic-driven intelligent prediction method for groundwater reserves as described in claim 3, characterized in that, The closure error between the input data, output data, and stored change data is calculated using differentiable operations, specifically including: ; in, This indicates taking the absolute value, where P is the precipitation. E is the lateral inflow, R is the evapotranspiration, and E is the surface runoff. This refers to the lateral outflow. To store change data; This represents the closure error.

9. A physical logic-driven intelligent prediction system for groundwater reserves, characterized in that, include: The multi-source data acquisition and physical logic verification module is configured to: complete and physically verify the multi-source heterogeneous observation data related to groundwater storage in the target area, and obtain a complete physical verification dataset that conforms to physical logic; The driving factor analysis module is configured to perform spatiotemporally decoupled causal contribution analysis on the complete physical verification dataset, separate the independent effects of driving factors in the time and space dimensions, and generate a causal contribution matrix. The self-learning modeling module is configured to: construct and train a physical constraint emergent spatiotemporal prediction model based on the causal contribution matrix; during the training process, a physical law discovery reward mechanism is introduced to enable the internal state of the model to spontaneously align with the physical law, thus obtaining a well-trained emergent prediction model. The cross-regional knowledge transfer module is configured to: apply a physical anchoring adversarial transfer strategy to adapt the trained emergent prediction model to the target new area, thereby obtaining a regionally adapted prediction model; The scenario simulation and decision-making module is configured to: input future scenario conditions into the regional adaptive prediction model for simulation, and generate groundwater storage prediction results and corresponding decision-making suggestion data; The process of training a physically constrained emergent spatiotemporal prediction model to obtain a well-trained emergent prediction model includes: Based on the causal contribution matrix and combined with the hydrological spatial adjacency relationship, the network structure and connection weights of the physical constraint emergent spatiotemporal prediction model are initialized. The physical constraint emergent spatiotemporal prediction model is iteratively trained using a complete physical verification dataset. In each training iteration, the predicted value and the simulated intermediate state variables are output. The first loss is calculated based on the error between the predicted and actual values, and the intermediate state variables in the simulation are checked to see if they show a trend consistent with the physical equilibrium relationship, so as to generate a reward signal for the discovery of physical laws. By combining the first loss with physical laws to discover reward signals, the parameters of the physical constraint emergent spatiotemporal prediction model are updated until the model converges, thus obtaining a well-trained emergent prediction model. The network structure of the physical constraint emergent spatiotemporal prediction model includes a memory unit that simulates the water storage state and a graph convolutional connection unit that simulates water flow transmission. The initial connection weights of the graph convolutional connection unit are determined by the hydrological spatial adjacency relationship. The intermediate state variables include the simulated water flow between nodes calculated by the graph convolutional layer and the simulated water storage updated within the cyclic unit. The hydrological spatial adjacency relationships are generated from geological parameter distribution maps and digital elevation model data.

Citation Information

Patent Citations

  • Tunnel unfavorable geology physical field-hydrological field fusion holographic detection method and system

    CN121091397A

  • Water level prediction method and system for power grid disaster prevention

    CN121434653A