A method for analyzing dependence relationship between groundwater level and driving factors based on multi-source fusion

By using a multi-source fusion analysis method, the causal contribution of driving factors in groundwater level changes is quantified, which solves the problem that existing technologies cannot accurately identify the net impact of driving events and provides a precise basis for water resource management.

CN120910475BActive Publication Date: 2025-12-26河北省地质环境监测院 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511394815.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-12-26
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately quantify the net impact of a single driving event in groundwater level changes, and cannot effectively identify and separate the causal contribution of specific precipitation or extraction activities to water levels, leading to ambiguity and uncertainty in water resource management decisions.

Method used

By using a multi-source fusion analysis method, groundwater level and driving factor data are obtained, peak events and inflection points of change are detected, the causal contribution scores of driving factor events to water level changes and response efficiency indicators are calculated, and the dependence of driving factors on water level changes is quantified.

Benefits of technology

It enables precise quantification of groundwater level changes, providing a scientific basis for differentiated and refined water resource management decisions, and overcoming the macroscopic and parameter-dependent shortcomings of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910475B_ABST
    Figure CN120910475B_ABST
Patent Text Reader

Abstract

The application provides a kind of groundwater level and driving factor dependent relationship analysis method based on multi-source fusion, it is related to groundwater level and driving factor dependent relationship technical field, the application decomposes groundwater level and driving factor time series data into independent events by event processing, then builds candidate event pair and calculates local groundwater level trend before driving event, establishes expected water level baseline by linear extrapolation, calculates the deviation absolute value integral of actual water level relative to baseline, defines as causal contribution score, to quantify the net effect of single driving event, finally, through significance test, match the optimal response event, and calculate response efficiency index combined with event intensity, to accurately characterize the influence efficiency and lag characteristics of different driving factors.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of dependence relationship between underground water level and driving factors, in particular to an analysis method for dependence relationship between underground water level and driving factors based on multi-source fusion. BACKGROUND

[0002] Accurate analysis of the dependence relationship between the dynamic of underground water level and its driving factors, such as precipitation and exploitation, is the scientific basis for accurate evaluation, reasonable development and effective management of water resources. The complexity of this problem lies in that the underground water system is a typical nonlinear time-varying system, and its response to external driving has significant time lag effect. In addition, the action mechanism and efficiency of different driving factors are different. When facing multi-source and heterogeneous driving data, the traditional analysis method often has difficulty in stripping the independent contribution of each factor, and cannot quantify the time-effect characteristics of its influence, thereby restricting the mechanism-based prediction and management decision.

[0003] The existing technology mainly solves this problem through two types of ways: one is based on mathematical statistics method, such as multiple regression, correlation analysis, etc., trying to establish statistical correlation between water level and driving factors; the other is based on physical mechanism numerical simulation, trying to invert the driving relationship by constructing and calibrating complex hydrogeological model; however, the statistical method is difficult to capture the complex nonlinear time series relationship and lag effect, and the result is macro and averaged, which cannot reveal the causal mechanism under specific events. The numerical simulation method has clear physical meaning, but its construction is seriously dependent on the hydrogeological parameters which are difficult to obtain accurately, and the model calibration process is complex and the calculation cost is high, so its reliability has great uncertainty in different regions.

[0004] These existing technologies have a common essential deficiency: they tend to approximate the behavior of the system from a global or statistical average perspective, while ignoring the fine description of individual events. Specifically, the existing methods cannot effectively identify and separate the net influence of a specific precipitation event or exploitation activity on the underground water level. They lack a mechanism to subtract the trend of water level itself to construct a counterfactual benchmark, so they cannot accurately calculate the extent to which an external driving factor leads to the observed water level change. It is this lack of quantitative ability of event-level causal contribution that leads to the fundamental difficulty of existing technologies in revealing the dynamic influence process of driving factors and accurately evaluating their efficiency.

[0005] The above information disclosed in the background section is only used to strengthen the understanding of the background of the present disclosure, and therefore it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0006] The present application aims to provide a groundwater level and driving factor dependency relationship analysis method based on multi-source fusion to solve the problems in the background art.

[0007] To achieve the above-mentioned purpose, the present application provides the following technical solutions.

[0008] A groundwater level and driving factor dependency relationship analysis method based on multi-source fusion, the specific steps comprising:

[0009] Step 1: Obtain the groundwater level and driving factor data of the monitoring well in the target area, detect the peak value event of the driving factor data, extract multiple driving factor events, detect the change inflection point of the groundwater level data, and extract the water level change event;

[0010] Step 2: For each driving factor event, construct it and all water level change events in the preset candidate period as a candidate event pair, for each candidate event pair, calculate the function expression of the groundwater level in the historical period before the driving factor event occurs, extrapolate the function expression to the entire duration of the paired water level change event to obtain the expected water level sequence;

[0011] Step 3: Calculate the absolute area of the residual sequence between the actual observed water level sequence and the expected water level sequence in the entire duration of the water level change event, and record the value of the absolute area as the causal contribution score of the driving factor event to the water level change event. For each driving factor event, select the water level change event with the largest causal contribution score from all candidate event pairs as the optimal response event;

[0012] Step 4: Set the lag time as the difference between the start time of the water level change event and the start time of the driving factor event, combine the causal contribution score and the driving factor event, calculate the response efficiency index, and count the lag time and the response efficiency index of the optimal response event to quantify the dependency relationship between the driving factor and the groundwater level change.

[0013] Further, the peak value event of the driving factor data is detected, and multiple driving factor events are extracted. The logic of detecting the change inflection point of the groundwater level data and extracting the water level change event is as follows:

[0014] Each event in the driving factor event contains a start time, an end time and an event cumulative intensity, and each event in the water level change event contains a start time, an end time and an event duration;

[0015] The driving factors include precipitation, surface water level and groundwater exploitation in the target area; the event cumulative intensity refers to the cumulative value of the driving factor data within the duration from the start time to the end time of the event;

[0016] Peak detection is performed on the driving factor data, specifically including: smoothing the driving factor data, identifying local peak points that exceed a preset intensity threshold, and taking each peak point as the core, tracing back to the moment when the data value first falls back to the preset threshold or encounters a local minimum as the event start time, and tracing back to the moment when the data value first falls back to the preset threshold or encounters the next driving factor event start time as the event end time. The preset threshold is determined based on the peak value and a preset proportional coefficient.

[0017] The groundwater level data is subjected to change inflection point detection, and the extracted water level change events are specifically water level rise events: the groundwater level data is smoothed, and the first-order difference sequence of the groundwater level data is calculated. The zero-crossing point of the first-order difference sequence from negative to positive is detected, and this is taken as the start time of the water level rise event. The zero-crossing point of the first-order difference sequence from positive to negative is detected, and this is taken as the end time of the water level rise event.

[0018] Furthermore, the logic for constructing candidate event pairs is as follows: based on the hydrogeological characteristics of the target area, the maximum reasonable response lag time is preset to be... For any driving factor event, a preset candidate time period is set as... ,in Given the start time of the driving factor event, the driving factor event is paired with all water level rise events whose start times fall within the preset candidate time period to form candidate event pairs.

[0019] Furthermore, the logic for determining the expected water level is as follows:

[0020] The specific method for calculating the functional expression of groundwater level during a historical period prior to the occurrence of the driving factor event is as follows: Select the start time of the driving factor event. As the endpoint, confirm a segment of length to the left. Historical time window By fitting the groundwater level data within the historical time window using the least squares method, the time-groundwater level function expression within the historical time window is obtained.

[0021] The function expression is further extrapolated to calculate its duration during the water level rise event paired with the driving factor event. The function values ​​at each time point within the time frame constitute the expected water level sequence;

[0022] The absolute area of ​​the difference between the actual observed water level series and the expected water level series is calculated. This value is used as the causal contribution score of the driving factor event to the water level change event. The calculation formula is as follows:

[0023] ;

[0024] in, Contribute scores to causality. is a time variable, is an actual observed water level value, is an expected water level sequence.

[0025] Further, for each driving factor event, from all candidate event pairs, its optimal response event is determined according to the following rules:

[0026] Significance test is performed to exclude accidental matches: when the number of candidate events is greater than or equal to 2, the first quartile of causal contribution scores of all candidate event pairs is calculated and the third quartile , and the significance threshold is defined as , where is a preset scaling coefficient; all water level change events with causal contribution scores are filtered out to form a set of significant response event candidates;

[0027] If the set of significant response event candidates is empty, it is determined that the current driving factor event does not cause a significant water level change event response in this analysis, and if the set of significant response event candidates is not empty, the water level change event with the largest causal contribution score is selected from the candidate set as the optimal response event of the driving factor event, thereby forming a successfully matched driving-response event pair.

[0028] Further, for each successfully matched driving-response event pair, its lag time is calculated and defined as the difference between the start time of the water level rise event and the start time of the driving factor event, i.e. ;

[0029] The response performance indicator is calculated according to the following formula:

[0030] ;

[0031] wherein is the response performance indicator of the driving-response event pair, is the causal contribution score of the driving-response event pair, is the event cumulative intensity of the driving factor event in the driving-response event pair.

[0032] Further, for any driving factor, its corresponding complete driving-response event pairs are aggregated to form a set, and the following statistics are performed respectively:

[0033] 1) Obtain the lag time distribution of each driving-response event pair in the set, draw a frequency distribution histogram, and determine the mode of the distribution as the typical lag time of the driving factor affecting the groundwater level;

[0034] 2) Extract the response performance indicator of each driver-response event pair in the set And descriptive statistics, calculate its mean, median, standard deviation and quartile, to evaluate the efficiency and stability of the impact of the driving factor.

[0035] Compared with the prior art, the beneficial effects of the present application are:

[0036] The present application effectively solves the core deficiency that the existing method cannot accurately quantify the net influence of a single driving event by its unique event analysis and causal contribution quantification framework. By linear extrapolation of the local historical trend before the driving event, a counterfactual reference frame excluding the interference of the driving event is cleverly established for each water level change event. This technical feature makes it possible to highlight the net contribution of external driving events by stripping the background changes of water level, thereby overcoming the defects of general results of traditional statistical methods and dependence on parameters of physical models;

[0037] The present application realizes the accurate quantification of the cumulative energy of water level change caused by a specific driving event by integrating the absolute deviation area of the actual water level and the expected benchmark. This index, combined with significance test and response performance indicator, forms a complete event-level causal correlation judgment and performance evaluation system, which enables the analyst to determine whether a rainfall or exploitation event has truly triggered a significant water level response, and further quantitatively evaluate the impact efficiency of different driving factors such as precipitation and exploitation unit intensity;

[0038] The present application provides an analysis tool with clear physical meaning, efficient calculation and no need for complex geological parameters. The output results are no longer macro statistical correlations, but present the typical lag time distribution of driving factors and the specific statistical characteristics of impact efficiency, thereby providing direct and reliable scientific basis for differentiated and refined water resource management decisions, and fundamentally solving the ambiguity and uncertainty problem in water level attribution analysis of the prior art. BRIEF DESCRIPTION OF DRAWINGS

[0039] Figure 1 The present application is a schematic diagram of the overall method flow;

[0040] Figure 2 The present application is an absolute deviation density plot;

[0041] Figure 3 The present application is an absolute deviation-cause contribution score line chart;

[0042] Figure 4 The present application is an absolute deviation column chart;

[0043] Figure 5Fitting curve diagram of actual observation water level-absolute deviation of the application. DETAILED DESCRIPTION

[0044] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to specific examples.

[0045] It should be noted that, unless otherwise defined, technical terms or scientific terms used in the present application should be understood as their common meanings by those skilled in the art to which the present application belongs. The terms "first", "second" and similar words used in the present application do not represent any order, number or importance, but are only used to distinguish different components. The terms "include" or "contain" and similar words mean that the elements or objects before the words cover the elements or objects listed after the words and their equivalents, without excluding other elements or objects. The terms "connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "up", "down", "left", "right" and the like are only used to represent relative positional relationships, which may change accordingly when the absolute positions of the described objects change.

[0046] Embodiment:

[0047] Please refer to Figures 1-5 The present application provides a technical solution:

[0048] A method for analyzing the dependence relationship between groundwater level and driving factors based on multi-source fusion, comprising the following specific steps:

[0049] Step 1: Obtain the groundwater level and driving factor data of the monitoring well in the target area, detect the peak value events of the driving factor data, extract a plurality of driving factor events, detect the change inflection points of the groundwater level data, and extract the water level change events;

[0050] The "target area" defines the spatial scale of the analysis, such as a specific hydrogeological unit or administrative area; the "monitoring well" is the specific location for data acquisition, and its selection directly determines the representativeness of the analysis results; the "groundwater level data" is the core response variable, which is usually time series data recorded by a pressure sensor at equal time intervals (such as daily); and the "driving factor data" is an input variable for explaining water level changes, which usually includes:

[0051] The driving factors include precipitation, surface water level and groundwater exploitation in the target area, and the data is obtained by fusing meteorological stations, surface water monitoring stations, exploitation amount statistical records and multi-source remote sensing data;

[0052] Precipitation is the most important natural recharge source of groundwater, and the direct infiltration of rainwater into the aquifer is the key driving factor for analyzing water level recovery. The water level of surface water bodies refers to the water level of nearby rivers, lakes, reservoirs, etc. When the water level is higher than the groundwater level, it will seep laterally to recharge the groundwater. Conversely, when the groundwater level is higher, it may drain to the surface water, which is an important boundary condition. Groundwater exploitation refers to the extraction of groundwater through water wells, which is the most important human driving factor for water level decline. These three categories basically cover the most important and most common factors affecting the dynamics of shallow groundwater such as pore water and fissure water. Meteorological stations provide precipitation data such as hourly rainfall and daily rainfall, which are point data. Surface water monitoring stations provide time series data of surface water body water level or flow, which are also point data. The extraction amount statistics are usually from the reports of water companies, government management departments or farmers, providing data on groundwater extraction, which may be daily or monthly statistics;

[0053] Multi-source remote sensing data can obtain areal rainfall data to make up for the lack of meteorological stations in sparse areas. For surface water bodies, optical satellites such as Landsat and Sentinel-2 or radar satellites such as Sentinel-1 can be used to monitor water body area changes and indirectly reflect water level information or verify monitoring station data. For exploitation, although it is difficult to directly monitor, remote sensing such as InSAR ground deformation monitoring can be used to retrieve land subsidence caused by large-scale exploitation as indirect evidence of exploitation intensity;

[0054] Fusion means not simply using a single data source, but combining station data (high precision, point) with remote sensing data (wide coverage, areal) to generate a set of high-quality driving factor data sets that can completely cover the target area in time and space through spatial interpolation, assimilation and other techniques;

[0055] Peak event detection is performed on the driving factor data to extract multiple driving factor events, and change inflection point detection is performed on the groundwater level data to extract water level change events. The logic for extracting water level change events is as follows:

[0056] Each event in the driving factor event contains a start time, an end time, and an event cumulative intensity. Each event in the water level change event contains a start time, an end time, and an event duration;

[0057] Event cumulative intensity refers to the cumulative value of driving factor data from the start time to the end time of the event;

[0058] For precipitation, the cumulative intensity is the total rainfall during the event duration, obtained by integrating the rainfall intensity data over time; for surface water body water level, the cumulative intensity is the average water level exceeding the threshold height or the water level increment integral during the event duration, to represent the average state of the surface water level or its cumulative deviation from the baseline water level; for groundwater exploitation, the cumulative intensity is the total exploitation water volume during the event duration, obtained by integrating the exploitation rate data over time or accumulating the discrete exploitation volume records;

[0059] The event cumulative intensity is an index aiming to quantify the overall scale or total amount of the event, rather than describing its instantaneous intensity, which answers the question of "how much does this event bring in total?" Due to the different physical meanings and original data dimensions of different driving factors, the cumulative way of "intensity" must be defined respectively to ensure that the physical meaning of each index is clear and can be used for subsequent response efficiency index calculation;

[0060] Precipitation (flux): its impact depends on how much rain falls, not how strong the rain is but how fleeting it is, so its cumulative intensity is the integral (i.e. sum) of rainfall intensity over time; surface water body water level (state variable) impact depends on how high the water level is and how long it stays high, so its cumulative intensity can have two definitions: 1) average water level exceeding the threshold height, calculating the average water level exceeding a certain baseline such as the average of the surrounding groundwater level, which measures the "potential" of recharge; 2) water level increment integral, calculating the increment (positive value) of water level at each time during the event period relative to the baseline, and then integrating these increments over time, which measures both the "potential" and "duration" of recharge, a more comprehensive index;

[0061] Groundwater exploitation (flux) impact depends on how much water is pumped, so its cumulative intensity is the integral of exploitation rate over time (for continuous data) or the direct accumulation of total exploitation volume (for discrete daily / monthly report data);

[0062] The start / end time defines the boundaries of the event, which is the basis for calculating key indicators such as event duration and lag time, and the event cumulative intensity is the core attribute of the driving factor event, for precipitation event it is the total rainfall, for groundwater exploitation event it is the total exploitation volume, using cumulative intensity rather than peak intensity can better represent the total energy input of an event, which is more matched with the response mechanism of the groundwater system (often related to total volume);

[0063] Change duration For water level response events, duration is an important characteristic, which helps to distinguish the response speed under different hydrogeological conditions, such as phreatic aquifer response fast, duration short; confined aquifer response slow, duration may be longer;

[0064] The peak detection is performed on the driving factor data, specifically including: smoothing the driving factor data, identifying a local peak point exceeding a preset intensity threshold, and taking each peak point as a core, tracing back to a time when the data value first falls below the preset threshold or encounters a local minimum value as an event start time, and tracing back to a time when the data value first falls below the preset threshold or encounters a next driving factor event start time as an event end time, the preset threshold being determined based on the peak value and a preset proportion coefficient;

[0065] The smoothing usually uses algorithms such as moving average and Savitzky-Golay filter to remove high-frequency noise, and short-term fluctuations, measurement errors and other “spikes” that may exist in the original data, which are not meaningful hydrological events. The smoothing process can filter out these noises, making the subsequent peak detection more accurate and robust, and avoiding detecting a large number of meaningless small events;

[0066] The peak finding algorithm, such as find_peaks, is used to preliminarily screen significant events, and an absolute threshold is set, such as precipitation greater than 5 mm, which can ignore those insignificant and unlikely to cause water level changes small events, and focus on analyzing those driving events with potential important influence;

[0067] Taking each peak point as a core is the core logic of defining event boundaries. The preset threshold here is a dynamic threshold, and the calculation formula is usually: the threshold is the peak intensity multiplied by the proportion coefficient, for example, the proportion coefficient is set to 0.1, that is, the 10% level of the peak value; Different event peak values, it is unscientific to use a fixed absolute threshold to determine the boundaries of all events, and the dynamic threshold can adaptively determine the boundaries of each event according to its own intensity, so that the boundaries of strong events are wide, and the boundaries of weak events are narrow, which is more consistent with physical intuition; Defining the event boundary at a certain proportion of the intensity falling back to the peak value means that we only care about the impact of the main part of the event, and ignore the weak fluctuations at the tail, which makes the definition of the event more consistent and comparable;

[0068] The preset threshold calculation formula is: , Wherein, is the intensity value of the local peak point identified in the driving factor event currently being processed, for precipitation events, is the peak rainfall intensity, and for exploitation events, is the peak exploitation rate;

[0069] is a preset proportion coefficient, which is a constant, and the value range is usually between 0.1 and 0.5 (i.e. 10% to 50%), which needs to be preset according to the data situation and application scenario; The greater the value, the higher the threshold, the narrower the event boundary, and the shorter the event duration, which is suitable for capturing those core highlights with clear boundaries; The smaller the value, the lower the threshold, the wider the event boundary, and the longer the event duration, which is suitable for capturing those events with less intensity but wider impact;

[0070] Local minimum is a protective condition. Sometimes data may fluctuate during the rebound process, and the dynamic threshold line may cross the data line multiple times. By taking the local minimum as an alternative termination condition, the end of the event can be determined more reasonably, and two independent events can be avoided from being mistakenly connected into one;

[0071] Event end time is a logical termination condition. When two events are very close in time, this rule can prevent overlapping between events and ensure that each event is independent, which is particularly important when dealing with continuous rainfall events;

[0072] The change inflection point detection is performed on the groundwater level data, and the extracted water level change event is a water level rise event: the groundwater level data is smoothed, and the first-order difference sequence of the groundwater level data is calculated. The zero-crossing point where the first-order difference sequence changes from negative to positive is taken as the start time of the water level rise event, and the zero-crossing point where the first-order difference sequence changes from positive to negative is taken as the end time of the water level rise event;

[0073] The smoothing principle is the same as above, which is to remove high-frequency noise and make the water level trend clearer, avoiding a large number of false inflection points caused by data fluctuations; the first-order difference is to calculate the difference between adjacent time points, which converts the water level value sequence into a water level change rate sequence. The sign (positive / negative) of the difference sequence directly represents whether the water level is rising or falling, and the numerical value represents the speed of rising / falling. This step converts the problem from identifying water level to identifying change direction turning point, which is easier and more accurate;

[0074] The zero-crossing point where the first-order difference sequence changes from negative to positive is the start time (valley point) of a water level rise event, and the zero-crossing point where the first-order difference sequence changes from positive to negative is the end time (peak point) of the water level rise event;

[0075] This method directly captures the turning point of the water dynamics process, and has extremely clear physical meaning. It is more accurate than simply setting a water level rise threshold, such as rising more than 0.2 meters, because it can identify any small but turning upward process, regardless of its final amplitude.

[0076] Step 2: For each driving factor event, construct all candidate event pairs with all water level change events within the pre-set candidate time period, for each candidate event pair, calculate the function expression of the groundwater level in the historical period before the driving factor event, extrapolate the function expression to the entire duration of the paired water level change event to obtain the expected water level sequence;

[0077] The logic of constructing candidate event pairs is: according to the hydrogeological characteristics of the target area, the maximum reasonable response lag time is pre-set as For any driving factor event, set its pre-set candidate time period as , where is the start time of the driving factor event, and the driving factor event is constructed as a candidate event pair with all water level rise events whose start time falls within the pre-set candidate time period;

[0078] Based on the prior knowledge of causality in hydrogeology to narrow down the analysis range, improve the calculation efficiency and reduce blind matching, it recognizes a basic physical fact: a driving event, such as rainfall, needs a certain propagation time from its occurrence to the full emergence of its potential impact on groundwater level, which depends on the properties of the aquifer such as the specific yield, transmissivity, is the maximum time span allowed physically;

[0079] It is not an arbitrary guess value, but a physical time scale estimated based on hydrogeological parameters such as aquifer lithology, thickness, permeability coefficient, etc. For example, for a highly permeable sand and gravel aquifer, precipitation may reach the groundwater body in one or two days, so may be set to 3-5 days, while for clay or fractured bedrock aquifers with poor permeability, the response may be slow and long, may need to be set to 30 days or even longer. This parameter ensures that the algorithm is consistent with the physical reality of the specific study area;

[0080] Reflects the speed of the target aquifer system in response to external driving, The larger the value, the slower the response speed of the aquifer system, and the longer the time for water to migrate in the vadose zone or aquifer, for example, for areas with thick clay layers and poor permeability, may need to be set large, such as tens of days; for sand and gravel aquifers, the response is very fast, may be shorter, such as a few days;

[0081] It is a pre-set parameter itself, but its value is determined by the inherent properties of the aquifer, such as permeability coefficient, storage rate, the larger these independent variables, i.e. the better the permeability and the stronger the water release capacity, Generally, the smaller the value, the faster the response, which is an inverse relationship;

[0082] reflects the full potential time range that a driving event may trigger a water level response, the larger the time window, the wider the analysis range, the more possible response events considered, but also the more calculation and the possible introduction of irrelevant noise events; a "search window" is dynamically defined for each driving factor event, which starts from the moment of the driving event, and the time window lasts for a maximum reasonable lag time, the starting point of the window is defined as the start time of the driving event rather than the end time, because the impact of driving factors, such as rainfall infiltration, often occurs at the beginning, rather than after it ends;

[0083] The logic for determining the expected water level is:

[0084] Calculate the average change trend of the groundwater level in a historical period before the driving factor event, which is done by selecting the driving factor event start time as the end point, a historical time window of length is identified, and a time-groundwater level function expression is obtained in the historical time window based on least squares fitting of the groundwater level data in the historical time window;

[0085] Least squares fitting is a standard and efficient method for fitting data trends, its goal is to find a straight line (or curve) that minimizes the sum of the squared vertical distances (residuals) from all actual data points, which means it can maximize the elimination of random fluctuations (noise) and extract the most likely internal linear trend;

[0086] Any water level rise event, as long as its start time is within the search window of the driving event, it is eligible to be a candidate response for the driving event, selecting the start time (rather than the peak time or end time) of the water level event as the matching anchor point is to capture the "response start" time, which forms the most direct causal time chain with the "drive start" time;

[0087] The length of the historical time window is determined comprehensively according to the hydrogeological characteristics of the target aquifer system and the data time resolution, and its value should be sufficient to reflect the stable change trend of the groundwater level before the driving factor event, usually set to 1 to 3 times the typical response time of the regional groundwater system to the driving factor;

[0088] The typical response time is a physical quantity that can be estimated according to hydrogeological parameters such as aquifer lithology and thickness, which The determination of the subjective preset becomes a rule, and 1 to 3 times provides a reasonable and operable experience range, which ensures the stability of the trend and avoids the introduction of irrelevant historical information by too long window;

[0089] Capture and quantify the natural or background trend of the groundwater level before the driving event occurs, which may be caused by other slow-changing factors not considered, such as evaporation, lateral runoff, regional slow recharge or exploitation, etc. By analyzing the local historical data before the event, the future instantaneous change rate of the water level can be best estimated if there is no driving event;

[0090] Further extrapolate the above function expression to calculate the function value at each time point within the duration of the water level rise event paired with the driving factor event, which is the expected water level sequence;

[0091] Assuming that "if there is no driving event, how will the groundwater level change?" The answer is: it will most likely continue the previous trend, so use the historical data before the driving event to fit a trend line, and extrapolate this trend to the time period after the event. The expected water level sequence obtained in this way represents the water level situation under the assumption of "no driving event"; Compare the actual observed water level with the expected water level sequence, and the difference between the two can be more purely attributed to the driving event;

[0092] Extrapolation is a hypothetical operation that assumes that in the future, the groundwater level will continue to strictly follow the linear change rule it has shown in the historical window For each time point within the duration of the water level rise event, substitute the function expression to calculate the expected water level sequence, and arrange the function values of all time points in chronological order to obtain a continuous "should have" water level line without the influence of the driving event;

[0093] This expected water level sequence is the baseline for measuring the real impact of the driving event. By calculating the difference (residual) between the actual observed water level and the expected water level, and taking the area of the difference as the causal contribution score, the net impact of the driving event can be clearly and quantitatively isolated, rather than simply analyzing the correlation.

[0094] Step 3: Calculate the absolute area of the residual sequence between the actual observed water level sequence and the expected water level sequence within the entire duration of the water level change event, and record the value of the absolute area as the causal contribution score of the driving factor event to the water level change event. For each driving factor event, select the water level change event with the largest causal contribution score from all its candidate events as the optimal response event.​

[0095] is the observed reality, is the "parallel reality" that would have occurred without the driving event, the difference between the two (i.e. the residual) represents the net impact of the driving event; the absolute area because we are only interested in the magnitude of the impact (whether positive or negative), and integrating (summing) the magnitude over time gives us the total impact, which is more comprehensive and stable than looking at the maximum deviation at a single point in time;

[0096] The absolute area of the difference between the actual observed water level sequence and the expected water level sequence is calculated, and this value is taken as the causal contribution score of the driving factor event to the water level change event, and the formula is:

[0097]

[0098] wherein, is the causal contribution score, is the time variable, is the actual observed water level value, is the expected water level sequence;

[0099] quantitatively reflects the total amount of water level deviation from its expected development trend caused by the driving event in the entire process of the water level change event, and indicates the strength or scale of the driving factor impact, which is an energy-based and comparable index; The larger the value of , the more intense and significant the impact of the driving event on the water level change, for example, the contribution of a heavy rain to the water level rise

[0100] instantaneous deviation absolute value is the core independent variable of the formula, the larger this difference is, the larger the value of , which directly measures the "instantaneous impact" of the driving event at that moment; the longer the duration, the larger the value of , the integral operation (sum in discrete data) accumulates the impact of each moment, obtaining the total impact, which enables

[0101] to capture sustained and moderate impacts, rather than just sharp peaks; the specific data of the partial event number and the causal contribution score is shown in Table 1.

[0102]

[0103] Through the analysis of the data, it is observed that there is a clear response relationship between the groundwater level change and the expected benchmark. From the data, it can be seen that the expected water level benchmark shows a stable downward trend, from 100.000 to 99.860, while the actual observed water level shows a characteristic of first falling, then rising, and then stabilizing. In the first 6 monitoring points, the deviation between the actual water level and the expected benchmark is small (0.008-0.016), and then the deviation gradually expands, reaching a peak (0.131) at the 10th monitoring point, and then gradually narrows;

[0104] This change pattern shows that in the early stage of monitoring, the groundwater level system is in a relatively balanced state, and the actual water level basically coincides with the expected benchmark. With the passage of time, external driving factors, such as precipitation recharge or groundwater exploitation, begin to appear, leading to an increase in the deviation of the actual water level from the expected benchmark. In particular, during the 8th-15th monitoring points, the deviation value remains at a high level (0.042-0.449), indicating that there are strong external interference factors at this stage;

[0105] The cumulative change trend of the causal contribution score further confirms this law. The cumulative speed is slow in the early stage, significantly faster in the middle stage, and gradually tends to be flat in the later stage, which shows that the response of the groundwater system to external driving factors has obvious time characteristics. The impact is small in the early stage, significant in the middle stage, and gradually returns to balance in the later stage. This change law is of great significance for understanding the dynamic response mechanism of the groundwater system and can provide scientific basis for water resource management and geological disaster warning.

[0106] For each driving factor event, from all candidate event pairs, the optimal response event is determined according to the following rules:

[0107] Significance test is performed to exclude accidental matches: when the number of candidate events is greater than or equal to 2, the first quartile of the causal contribution score of all candidate event pairs is calculated and the third quartile , and the significance threshold is defined as , where is a preset scaling coefficient; all water level change events with causal contribution scores are screened out to form a set of significant response event candidates;

[0108] A non-parametric significance test method based on the data distribution itself is adopted, using quartiles and interquartile ranges. This method is more robust than the normal distribution-based test, such as Z-score, and is not sensitive to outliers, which is very suitable for hydrological data with non-normal distribution;

[0109] The upper boundary of the normal fluctuation range of the causal contribution score of all candidate response events reflecting a driving event, events exceeding this boundary can be considered statistically "abnormally significant"; it indicates the threshold value for judging whether a response is "significant", which is not a fixed value, but dynamically calculated according to the specific circumstances of all candidate responses of this driving event, with adaptability; The greater, the higher the set significance threshold, the more stringent the screening conditions, only very strong response events can be selected;

[0110] The greater, The greater, The upper middle level of all candidate event values, its size determines the basic height of the threshold; The interquartile range, that is, IQR, the greater the IQR, The greater, the IQR measures the dispersion of all candidate event values, the greater the IQR, the more dispersed the data, the greater the fluctuation, so a higher threshold needs to be set to exclude these fluctuations; IQR is not sensitive to outliers, while standard deviation is greatly affected by extreme values, using the IQR-based method can ensure that your significance test is not distorted by one or two abnormally high values, so it is more reliable;

[0111] The greater, The greater, this is a groundwater preset parameter, used to adjust the strictness of the test, The greater, the threshold is more relaxed, easier to match the response, The smaller, the threshold is more stringent;

[0112] If the significant response event candidate set is empty, it is determined that the current driving factor event does not cause significant water level change event response in this analysis, if the significant response event candidate set is not empty, the water level change event with the largest causal contribution score is selected from the candidate set, and the optimal response event of the driving factor event is finally confirmed, thereby forming a successful matching driving-response event pair;

[0113] The significant response event candidate set reflects which events of all candidate water level events of the current driving event have abnormally prominent causal contribution scores that are unlikely to be caused by random background fluctuations, it indicates the list of water level events that are truly qualified as "suspects" of the driving event after statistical filtering; usually this set will only contain 0 or 1 events, if it contains multiple, it means that there may be multiple driving sources acting simultaneously, or the driving event itself is complex, such as multiple peaks in a rainfall, but only the most prominent one will be selected Taking the largest (the one that is most large) as the "optimal response event" is a very rigorous logic;

[0114] The optimal response event reflects the most important water level response event that is ultimately identified as the driving event after rigorous screening through two layers of "contribution quantification" and "statistical significance" from all possible candidates. It indicates the most reliable correspondence between the "cause" (driving event) and the "effect" (water level event). Its existence means that the analyst can have a high degree of confidence that this water level change is mainly caused by the driving event.

[0115] Step 4: Let the lag time be the difference between the start time of the water level change event and the start time of the driving factor event. Combine the causal contribution score and the driving factor event to calculate the response efficiency index. Statistically analyze the lag time and response efficiency index of the optimal response event to quantify the dependency between the driving factor and the groundwater level change.

[0116] For each successfully matched driver-response event pair, calculate its latency. , defined as the difference between the start time of the water level rise event and the start time of the driving factor event, i.e. ;

[0117] The response performance index is defined by the following formula:

[0118] ;

[0119] in, This is a performance metric for the response of the driver-response event pair. The causal contribution score for this driver-response event pair. The cumulative intensity of the driving factor events in this driver-response event pair;

[0120] It quantitatively reflects the "efficiency" of the driving factor, measuring the total water level response that a unit intensity driving event can cause. It indicates the conversion efficiency of the driving factors on the groundwater system; it is an "input-output" efficiency indicator. The larger the value, the higher the efficiency of the driving factor. This means that a small driving intensity can produce a large water level response. For example, a 50mm rainfall falling on an exposed, highly permeable gravel layer could produce a large water level response. In contrast, the same rainfall falling in urbanized areas (where large areas of the ground are covered by impermeable layers) mostly forms surface runoff, with little infiltration. The value will be very low;

[0121] A rainfall of 50mm ( ), if it falls on bare, very permeable gravel layer (good surface infiltration conditions), will produce a very large , so the calculation of a large value, the same rainfall in urban areas, a large number of ground covered by impervious layer, poor infiltration conditions, most of the surface runoff, recharge, small, , the value will be very low;

[0122] The same total amount of exploitation ( ), in the elastic water storage coefficient of large confined aquifer, caused by water level drawdown (reflected in ) is small, , the value is low; in the large specific yield of phreatic aquifer, the water level drawdown caused by large, , the value is high;

[0123] For a rare heavy rainstorm ( great), due to the heavy soil, poor infiltration conditions, most of the rainwater formed surface runoff, only a small amount of recharge groundwater, the final water level rise is very limited small, very small, the value is very low; a small rainfall ( very small), but because of the fall on bare, very permeable gravel layer, almost all the rainwater quickly infiltrated, effectively recharge groundwater, caused a very significant water level rise ( great), great, the value is tens of times the former, which perfectly reflects the good permeability of aquifer, hydrogeological characteristics of good recharge conditions;

[0124] Lag time Quantitatively reflects the speed of the hydrological system response to driving factors, it measures the time from the "cause" to the "effect" between the appearance, directly indicates a core physical transport properties of groundwater system, it is the length of the flow path, aquifer permeability, water storage coefficient and other hydrogeological parameters of the comprehensive embodiment; The greater the value, the slower the response speed of the system, which may mean: 1) driving factor occurs far from the monitoring well, water needs a long time to flow, 2) poor permeability of the aquifer rock, such as clay, dense fractured rock, water flow slowly, 3) the specific yield of the aquifer is large, it needs longer time to fill the huge "capacity" to show obvious water level change;

[0125] ​The smaller the value, the faster the system response speed. This usually means that the monitoring well is located near the recharge area or the aquifer has very high permeability, such as a large karst conduit or a coarse sand and gravel layer. The size is not directly determined by the timestamp itself, but is an external manifestation of the inherent properties of the system determined by the driving event, hydrogeological conditions, and the location of the monitoring well.

[0126] For any given driving factor, we summarize the set of all driving-response event pairs corresponding to it, and then perform the following statistics:

[0127] 1) Obtain the latency of each driver-response event pair in the set. The distribution of the frequency distribution histogram was plotted, and the mode of the distribution was determined as the typical lag time of the driving factor's influence on the groundwater level.

[0128] 2) Extract the response performance metrics for each driver-response event pair in the set. Descriptive statistics were performed to calculate the mean, median, standard deviation, and quartiles to assess the efficiency and stability of the driving factor's influence.

[0129] The mode is defined as the typical lag time, and the peak of the histogram corresponds to... The value represents the most common and typical response speed, which is more representative than the average value because it is not affected by a few extreme outliers, such as a special event with an extremely slow response. The distribution pattern shows that the shape of the histogram itself also contains a lot of information. A thin distribution indicates that the system response speed is very stable. A flat distribution indicates that the response speed varies greatly, which may suggest that the aqueous system has strong spatial heterogeneity.

[0130] The mean / median represents the average efficiency level of the driving factor's influence. The median is more robust than the mean, being less affected by events of extreme efficiency or inefficiency. It indicates the average efficiency of this type of driving factor, such as precipitation, in the region. A higher value indicates a higher overall efficiency of the driving factor.

[0131] Standard deviation is a measure of the dispersion of data; a small standard deviation indicates efficiency for most events. The fact that the values ​​are all concentrated around the average indicates that the influence of the driving factors is very stable, which may mean that the replenishment mechanism is simple and stable; a large standard deviation indicates efficiency. The values ​​are very dispersed, with some events being highly efficient and others being very inefficient, indicating that the efficiency is unstable. This may be due to a variety of factors, such as good infiltration conditions in some areas and poor conditions in others, or the variable nature of the driving events themselves, such as different rainfall intensities and durations leading to different infiltration rates. Analyzing the reasons behind these efficient and inefficient events can lead to deeper scientific discoveries.

[0132] The IQR also reflects the dispersion and distribution shape of the data, but is more robust than the standard deviation, and can be calculated as the interquartile range The greater the IQR, the more dispersed the distribution of the middle 50% of the data, and the greater the stability of the median, The greater the value, the more events are high-efficiency events.

[0133] The above formulas are all dimensionless values calculated, and the formula is obtained by software simulation of a large amount of data to obtain a formula of the most recent real situation, and the preset parameters in the formula are set by a person skilled in the art according to the actual situation.

[0134] The above embodiments can be realized wholly or partially by software, hardware, firmware or any combination thereof. When realized by software, the above embodiments can be realized wholly or partially in the form of a computer program product. Those skilled in the art can realize that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solutions.

[0135] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, which can be located in one place or distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0136] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application.

Claims

1. A method for analyzing the dependence relationship between groundwater level and driving factors based on multi-source fusion, characterized in that, The specific steps include: Step 1: Obtain the groundwater level and driving factor data of the monitoring well in the target area, perform peak event detection on the driving factor data, extract a plurality of driving factor events, perform change inflection point detection on the groundwater level data, and extract water level change events; Step 2: For each driving factor event, construct a candidate event pair with all water level change events in a preset candidate period, calculate a function expression of the groundwater level in a historical period before the occurrence of the driving factor event, and extrapolate the function expression to the entire duration of the paired water level change event to obtain an expected water level sequence; Step 3: Calculate the absolute area of the residual sequence between the actual observed water level sequence and the expected water level sequence during the entire duration of the water level change event, and record the value of the absolute area as the causal contribution score of the driving factor event to the water level change event. For each driving factor event, select the water level change event with the maximum causal contribution score from all candidate event pairs as the optimal response event; Step 4: Set the lag time as the difference between the start time of the water level change event and the start time of the driving factor event, combine the causal contribution score and the driving factor event, calculate the response efficiency index, and statistically analyze the lag time and the response efficiency index of the optimal response event to quantify the dependence between the driving factor and the groundwater level change. 2.The method of claim 1, wherein the method comprises: The logic for performing peak event detection on the driving factor data and extracting a plurality of driving factor events and performing change inflection point detection on the groundwater level data and extracting water level change events is as follows: Each event in the driving factor event includes a start time, an end time, and an event cumulative intensity. Each event in the water level change event includes a start time, an end time, and an event duration; The driving factors include precipitation, surface water level, and groundwater exploitation in the target area. The event cumulative intensity refers to the cumulative value of the driving factor data during the duration from the start time to the end time of the event; The peak detection on the driving factor data specifically includes: smoothing the driving factor data, identifying local peak points that exceed a preset intensity threshold, and tracing back to the time when the data value first falls below the preset threshold or encounters a local minimum value as the event start time, and tracing back to the time when the data value first falls below the preset threshold or encounters the start time of the next driving factor event as the event end time, wherein the preset threshold is determined based on the peak value and a preset proportion coefficient; The change inflection point detection on the groundwater level data extracts water level rise events: smooth the groundwater level data, calculate the first-order difference sequence of the groundwater level data, detect the zero-crossing point where the first-order difference sequence changes from negative to positive as the water level rise event start time, and detect the zero-crossing point where the first-order difference sequence changes from positive to negative as the water level rise event end time.

3. The method according to claim 2, wherein the method is characterized by: The logic of constructing the candidate event pair is: according to the hydrogeological characteristics of the target area, the maximum reasonable response lag time is preset as , for any driving factor event, the preset candidate period corresponding thereto is set as , wherein is the start time of the driving factor event, and the driving factor event and all water level rising events with start times falling within the preset candidate period are constructed as a candidate event pair.

4. The method according to claim 3, wherein the method is characterized by: The logic for determining the expected water level is as follows: The function expression of the groundwater level in a historical period before the driving factor event occurs is calculated in the following way: the starting time of the driving factor event is selected As the end point, a historical time window with a length of is confirmed in the forward direction The groundwater level data in the historical time window is fitted based on the least square method to obtain the time-groundwater level function expression in the historical time window. The function expression is further extrapolated to calculate its duration during the water level rise event paired with the driving factor event. The function values ​​at each time point within the time frame constitute the expected water level sequence; Calculate the absolute area of the difference between the actual observed water level sequence and the expected water level sequence. The value is used as the causal contribution score of the driving factor event to the water level change event, and the calculation formula is as follows: ; wherein, is the causal contribution score, is the time variable, is the actual observed water level value, is the expected water level sequence.

5. The method according to claim 4, wherein the method is characterized by: For each driving factor event, from all candidate event pairs, its optimal response event is determined according to the following rules: Significance test is performed to exclude chance matches: when the number of candidate events is greater than or equal to 2, the first quartile of causal contribution scores of all pairs of candidate events is calculated with the third quartile and a significance threshold is defined as where is a preset scaling factor; all causal contribution scores of water level change events are screened out to form a set of significant response event candidates; If the set of significant response event candidates is empty, it is determined that the current driving factor event does not cause a significant water level change event response in this analysis. If the set of significant response event candidates is not empty, the water level change event with the largest causal contribution score is selected from the candidate set as the optimal response event of the driving factor event, thereby forming a successful matching driving-response event pair.

6. The method according to claim 5, wherein the method is characterized by: For each successfully matched driver-response event pair, calculate its lag time , defined as the difference between the start time of the water level rise event and the start time of the driver event, i.e. ; The response effectiveness index is defined by the following formula: ; wherein, is a response efficacy indicator for the driver-response event pair, is a causal contribution score for the driver-response event pair, is an event cumulative strength of a driver factor event in the driver-response event pair.

7. The method according to claim 6, wherein the method is characterized by: For any driving factor, all its corresponding driving-response event pairs are aggregated to form a set, and the following statistics are performed respectively: 1) Obtain the lag times for each driver-response event pair in the set Distribute, plot a frequency distribution histogram, and determine the mode of the distribution as the typical lag time for the driver to affect the groundwater level; 2) Extracting the response performance indicators of each driver-response event pair in the set And descriptive statistics, calculate its mean, median, standard deviation and quartile, to assess the efficiency and stability of the impact of the driver factor.

Citation Information

Patent Citations

  • Groundwater level analysis method and system based on event text data mining

    CN108182178A

  • Rainfall prediction method and system for Beijing, Tianjin and Hebei regions

    CN119556377A