Closed watershed lake water level evolution prediction method based on multi-source data

By constructing an XGBoost-SHAP white-box prediction model and combining the Mantel test and SHAP threshold, the problems of parameter uncertainty and interpretability in the prediction of water level in closed watershed lakes are solved, achieving high-precision and interpretable water level prediction and improving the robustness and decision support capability of the model.

CN121745366APending Publication Date: 2026-03-27INST OF GEOGRAPHY HENAN ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies for predicting water levels in closed watershed lakes suffer from problems such as difficulty in obtaining parameters, high model uncertainty, lack of interpretability of machine learning models, and difficulty in distinguishing between spurious correlations and real causal driving mechanisms, resulting in insufficient robustness and decision support value of the prediction models.

Method used

We employ an XGBoost-SHAP white-box prediction model based on multi-source data. By processing multi-source feature inputs through spatiotemporal alignment and standardization, we construct a SHAP dependency graph, perform Mantel statistical tests and nonlinear response threshold extraction, screen causal driving feature sets, and reconstruct the prediction model to achieve high-precision and interpretable predictions.

Benefits of technology

It achieves high-precision nonlinear prediction, quantitatively analyzes the physical response threshold of driving factors, improves the physical interpretability and robustness of the prediction model in complex environments, and provides a reliable basis for watershed water resources management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121745366A_ABST
    Figure CN121745366A_ABST
Patent Text Reader

Abstract

The invention provides a closed watershed lake water level evolution prediction method based on multi-source data, and relates to the technical field of hydrological resources. The method comprises the following steps: firstly, collecting natural environment and social economic multi-source data of a closed drainage basin, and constructing a normalized multi-source feature input matrix; carrying out total element analysis on the features and extracting a nonlinear response threshold value; and meanwhile, Mantel statistical test is executed on the feature matrix to obtain a statistical correlation coefficient. And judging the characteristics which simultaneously meet the statistical significance and mechanism threshold conditions as real causal driven factors, and eliminating false related factors which are statistical significance but have no physical threshold. And finally, reconstructing a prediction model based on the screened causal-driven feature set, and outputting a water level evolution prediction value. Through'statistics-mechanism 'dual verification, interference of false correlation is effectively eliminated, the problem that a traditional black box model cannot distinguish causal and coincident trends is solved, and physical interpretability and robustness of water level prediction are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hydrological resources technology, and in particular to a method for predicting the water level evolution of closed watershed lakes based on multi-source data. Background Technology

[0002] As sensitive indicators of regional climate change, the water level evolution of closed watershed lakes is crucial for maintaining ecological balance and ensuring the safety of surrounding human activities. Currently, methods for predicting lake water levels mainly fall into two categories: one is hydrological models based on physical processes, which calculate water levels by simulating physical processes such as precipitation, evaporation, infiltration, and runoff confluence within the watershed; the other is data-driven statistical or machine learning models, which predict water levels by mining the nonlinear patterns in historical hydrological and meteorological data. With the development of remote sensing technology and big data, ensemble learning models that integrate multi-source data from meteorology, hydrology, and human activities are gradually becoming a research hotspot.

[0003] Current technological trends are shifting from single-dimensional natural factor-driven analysis to comprehensive prediction based on the coupling of the "nature-society" dual water cycle. At the algorithmic level, to address the weak generalization ability of traditional single models, ensemble learning algorithms such as Extreme Gradient Boosting (XGBoost) are widely used in hydrological prediction due to their advantages in handling high-dimensional nonlinear data. Simultaneously, research is beginning to focus on how to utilize multi-source heterogeneous data to quantify the complex interference of human activities on hydrological processes, attempting to construct comprehensive prediction systems that better reflect real-world environments.

[0004] However, existing technologies still have significant limitations in practical applications: First, physical models often face difficulties in obtaining parameters and high structural uncertainty in closed watersheds with scarce data, such as the Qinghai-Tibet Plateau, resulting in limited simulation accuracy; second, most mainstream machine learning models are "black box" structures, which, although highly accurate in prediction, cannot quantify the specific contribution of each input feature and the threshold of nonlinear response, and lack interpretability at the physical mechanism level; most importantly, existing data-driven methods have difficulty distinguishing between "statistical correlation" and "physical causality," and are prone to misjudging socioeconomic indicators that coincidentally synchronize with water level changes in time as driving factors, thus masking the truly effective human factors, which greatly reduces the robustness and decision support value of the prediction model. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a method for predicting the water level evolution of lakes in closed watersheds based on multi-source data. This invention solves the problems of existing technologies, such as the large uncertainty of parameters in physical models in areas without data, and the lack of interpretability of machine learning models and their inability to effectively distinguish between spurious correlations and real causal driving mechanisms.

[0006] To achieve the above objectives, the present invention provides the following solution: A method for predicting the water level evolution of lakes in a closed watershed based on multi-source data includes: The original multi-source dataset of the closed watershed is collected, and the original multi-source dataset is spatiotemporally aligned and standardized to obtain the normalized multi-source feature input matrix and the corresponding target water level vector. An XGBoost-SHAP white-box prediction model is constructed, and the normalized multi-source feature input matrix is ​​trained and analyzed with all elements to obtain the SHAP dependency map of each feature dimension. Extract the nonlinear response threshold based on the SHAP dependency graph; Perform Mantel statistical tests on the normalized multi-source feature input matrix to obtain the statistical correlation coefficients between each feature and the target water level vector; Based on the statistical correlation coefficient and the nonlinear response threshold, the true and false driving factors are identified, and the causal driving feature set is obtained by screening. Based on the causal-driven feature set, the XGBoost-SHAP white-box prediction model is reconstructed to obtain the reconstructed prediction model. The original multi-source dataset for the period to be tested is obtained and input into the reconstructed prediction model to obtain the predicted value of water level evolution.

[0007] The present invention discloses the following technical effects: This invention provides a method for predicting the water level evolution of lakes in closed watersheds based on multi-source data. By constructing an XGBoost-SHAP white-box architecture, this invention not only achieves high-precision nonlinear prediction but also quantitatively analyzes the physical response thresholds of driving factors, overcoming the shortcomings of traditional "black box" models. More importantly, this invention innovatively introduces a dual "statistical-mechanism" verification logic, using Mantel tests and SHAP thresholds to collaboratively eliminate spurious factors (such as economic indicators) that are merely statistical coincidences, while retaining true physical driving factors. This mechanism significantly improves the physical interpretability and robustness of the prediction model under complex environmental changes, providing a reliable basis for the scientific management of watershed water resources. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1A flowchart illustrating a method for predicting the water level evolution of lakes in a closed watershed based on multi-source data, provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the interannual variations and long-term trends of key meteorological elements in a lake basin, provided as an embodiment of the present invention. Figure 2 (a) is a schematic diagram showing the interannual and long-term trends of annual precipitation in a certain lake basin; Figure 2 (b) is a schematic diagram showing the interannual and long-term trends of the annual average temperature in a certain lake basin; Figure 2 (c) is a schematic diagram showing the interannual and long-term trends of annual evaporation in a certain lake basin; Figure 2 (d) is a schematic diagram of the interannual and long-term trends of annual sunshine hours in a certain lake basin; Figure 3 This is a schematic diagram illustrating the changes in water consumption by sector and its long-term evolution characteristics in a certain lake basin, provided as an embodiment of the present invention. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0011] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0012] like Figure 1 As shown, this invention provides a method for predicting the water level evolution of lakes in a closed watershed based on multi-source data, including: Step 100: Collect the original multi-source dataset of the closed watershed, and perform spatiotemporal alignment and standardization on the original multi-source dataset to obtain the normalized multi-source feature input matrix and the corresponding target water level vector. Step 200: Construct an XGBoost-SHAP white-box prediction model, and perform full-element training and analysis on the normalized multi-source feature input matrix to obtain the SHAP dependency map of each feature dimension; Step 300: Extract the nonlinear response threshold based on the SHAP dependency graph; Step 400: Perform Mantel statistical tests on the normalized multi-source feature input matrix to obtain the statistical correlation coefficients between each feature and the target water level vector; Step 500: Based on the statistical correlation coefficient and the nonlinear response threshold, perform true and false driving factor identification to obtain the causal driving feature set; Step 600: Reconstruct the XGBoost-SHAP white-box prediction model based on the causal-driven feature set to obtain the reconstructed prediction model; Step 700: Obtain the original multi-source dataset for the period to be tested and input it into the reconstructed prediction model to obtain the predicted value of water level evolution.

[0013] like Figure 2 As shown, the interannual variations and long-term trends of key meteorological elements in a certain lake basin from 2000 to 2024 were assessed. The results indicate that despite significant interannual fluctuations, annual precipitation generally showed a significant increasing trend, suggesting that the overall climate of the basin has become significantly wetter over the past approximately 25 years. Figure 2 a). In contrast, although the annual average temperature fluctuates within a certain range, its long-term warming trend is not obvious, and it does not show a meaningful signal of sustained increase. Figure 2 b). Annual evaporation also exhibits significant interannual variation, with a long-term trend of slow decline, but this has not reached a statistically significant level. Figure 2 c). Unlike precipitation changes, the annual sunshine duration showed a strong and highly significant downward trend, indicating that the lake basin has experienced a significant darkening process in recent years. This change may have a significant impact on the lake's water balance by affecting the underlying surface energy budget and evapotranspiration processes. Figure 2 d).

[0014] like Figure 3 As shown, in order to understand the direct impact of human activities on water resources, this embodiment analyzes the changes in water consumption of various sectors in a certain lake basin from 2000 to 2024. Figure 3The results show that irrigation water for cultivated land, which accounted for approximately 87.3% of total water consumption in 2000, decreased significantly by about 59.1% during this period, from approximately 178 million cubic meters in 2000 to approximately 72.9 million cubic meters in 2024, a net reduction of approximately 105.5 million cubic meters. This significant decline is highly consistent with the substantial reduction in cultivated land area (approximately 70.8%) recorded in the land use / cover change analysis, indicating that land use change is the main driver of changes in watershed water consumption. The continuous reduction in agricultural irrigation water has dominated the evolution trajectory of total water consumption in the watershed, giving it a distinct "V"-shaped reversal characteristic: total water consumption peaked at approximately 264 million cubic meters around 2010, and then declined rapidly, reaching approximately 115.62 million cubic meters by 2024, a decrease of approximately 43.4% compared to 2000 and a decrease of approximately 56.2% compared to the 2010 peak. It is noteworthy that although indicators such as nighttime light pollution suggest a continued increase in urban construction and activity, industrial water consumption exhibits a pattern of initial increase followed by a sharp decline. After peaking at approximately 18.5 million cubic meters around 2006, it has fallen to less than 1 million cubic meters annually since 2012, reflecting a significant adjustment in the regional industrial structure. During the same period, residential and urban public water consumption remained at a low level of approximately 3 million cubic meters annually, contributing minimally to the overall water resource balance of the basin. Urban intensive development has not significantly increased total water consumption. In summary, water consumption analysis provides quantitative evidence for the intensity and structural retreat of human activities in the basin: the approximately 43.4% decrease in total water consumption was not a synchronous and uniform reduction across all sectors, but rather almost entirely driven by a significant contraction in agricultural irrigation water use. This result is highly consistent with the substantial shrinkage of arable land observed in land use / cover change analysis.

[0015] Furthermore, the specific implementation process of step 100 is as follows: This embodiment first constructs a multidimensional heterogeneous raw observation database for a closed watershed. The data collection scope covers meteorological forcing factors, hydrophysical factors, land use cover factors, and anthropogenic disturbances and socioeconomic factors. Among them, meteorological forcing factors include daily precipitation, temperature, evaporation, and sunshine duration records from meteorological stations within and around the watershed; hydrophysical factors include monthly data on water level, water area, water storage, inflow into the lake, and lake surface temperature; land use cover factors are derived from annual data on cultivated land, forest land, grassland, and wetland area from remote sensing interpretation products; and anthropogenic disturbances and socioeconomic factors cover annual data on human footprint index, nighttime light index, regional population, regional GDP, and water consumption by industry.

[0016] This embodiment uses the temporal resolution of monthly water level data in hydrophysical factors as a benchmark to perform multi-scale temporal resampling and alignment processing on the aforementioned multi-source heterogeneous data. For daily-scale meteorological forcing factors, this embodiment uses monthly cumulative or monthly average algorithms to aggregate them into monthly-scale sequences to match the temporal granularity of the target variable. For annual-scale land use cover factors and socioeconomic factors, this embodiment uses linear interpolation or spline function interpolation algorithms to upsample annual-scale records into continuous monthly-scale sequences, and then matches and truncates each processed data sequence on the time axis to generate a synchronous multi-source time series set with consistent time spans.

[0017] This embodiment integrates synchronous multi-source time series sets to construct an original feature matrix and extracts contemporaneous water level data as target water level vectors. To eliminate scale differences between data of different dimensions and improve the convergence speed of subsequent models, this embodiment performs standardization processing on each feature column in the original feature matrix. Specifically, this embodiment calculates the arithmetic mean and standard deviation of each feature column, subtracts the arithmetic mean from each value in the feature column, and then divides it by the standard deviation, thereby obtaining a normalized multi-source feature input matrix with uniform numerical distribution, providing standardized data input for subsequent nonlinear model training.

[0018] Furthermore, the specific implementation process of step 200 is as follows: This embodiment first establishes an extreme gradient boosting model architecture based on decision tree ensembles, and divides the normalized multi-source feature input matrix into training and validation datasets. On the training dataset, this embodiment performs full-factor iterative training based on the extreme gradient boosting model architecture, generating a boosting tree ensemble by constructing a new regression decision tree to fit the cumulative prediction residuals of the preceding decision trees. During iterative training, this embodiment uses the validation dataset to perform grid search and cross-validation to optimize the key hyperparameters of the extreme gradient boosting model architecture. These key hyperparameters include the learning rate, the maximum tree depth, and the subsample sampling ratio, thereby determining the boosting tree ensemble with the optimal generalization performance and identifying it as the trained nonlinear water level prediction model.

[0019] This embodiment then introduces a Shapley additive interpreter to encapsulate the nonlinear water level prediction model, constructing an XGBoost-SHAP white-box prediction model. Based on cooperative game theory principles, this embodiment uses the white-box prediction model to calculate the marginal contribution value of each sample in each feature dimension of the normalized multi-source feature input matrix. This calculation process decomposes the model's predicted output value for a specific sample into the sum of the baseline value and the contribution values ​​of each feature, quantitatively assessing the positive or negative promoting effect of each feature on the deviation of the water level from the baseline value, and integrating all calculated marginal contribution values ​​into a SHAP value matrix with the same dimensions as the input matrix.

[0020] Finally, this embodiment traverses each feature dimension contained in the normalized multi-source feature input matrix, extracts the original feature value sequence under that feature dimension as the x-axis data, and extracts the SHAP value sequence corresponding to that feature dimension from the SHAP value matrix as the y-axis data. This embodiment uses the original feature value sequence and the SHAP value sequence to construct a scatter plot reflecting the nonlinear mapping relationship between feature value changes and model prediction contributions, and uses local weighted regression or multinomial fitting algorithms to perform nonlinear fitting on the scatter plot. This fitting process aims to smooth random fluctuations and highlight feature response trends, thereby obtaining SHAP dependency maps for each feature dimension that can intuitively display the changes in nonlinear response thresholds and marginal effects.

[0021] Furthermore, the expression for the nonlinear water level prediction model is: ; in, For a moment Predicted lake water levels; For a moment Normalized multi-source feature input vector; This is the overall nonlinear water level prediction function trained using the extreme gradient boosting algorithm; For the integration of the first Sub-prediction functions corresponding to each regression decision tree; This represents the total number of regression decision trees integrated into the nonlinear water level prediction model.

[0022] Specifically, the nonlinear water level prediction model constructed in this embodiment is an ensemble model, composed of multiple regression decision trees. The model's prediction function utilizes the sub-prediction function of each regression decision tree in the ensemble to perform feature space mapping on the input normalized multi-source feature vector and outputs the weight values ​​corresponding to the leaf nodes of that decision tree. The output weights of all decision trees are scaled by the learning rate parameter and then weighted and accumulated to finally obtain the model's predicted lake water level. The physical meaning of the model is that the final water level prediction value is obtained by jointly correcting and accumulating the prediction results of all ensemble decision trees; each tree attempts to compensate for the prediction errors left by the previous tree combinations.

[0023] This embodiment provides specific parameter examples for logical deduction. Assume that the total number of ensemble regression decision trees determined through cross-validation during the model training phase is 500; to ensure the model's generalization ability, the learning rate parameter (shrinkage step size) is set to 0.1. During the model's inference calculations, assume the input normalized multi-source feature vector contains data such as precipitation and water consumption measured at a certain time point. The model inputs this vector to the first ensemble regression decision tree, obtaining its predicted value (e.g., 3.5). Similarly, the second tree outputs a predicted value (e.g., -0.2), and so on, until the 500th tree outputs a predicted value (e.g., 0.05). At this point, the model performs a weighted summation calculation according to the formula: 0.1 multiplied by 3.5 plus 0.1 multiplied by -0.2, continuing until 0.1 multiplied by 0.05. The final calculated value is the predicted lake water level at the corresponding time point output by the model. In this way, the model can aggregate the small contributions of hundreds or thousands of weak predictors into a high-precision final water level prediction.

[0024] Furthermore, the expression for the XGBoost-SHAP white-box prediction model is: ; in, For a moment Lake water level predictions given by the XGBoost-SHAP white-box prediction model; To develop a nonlinear water level prediction function within the SHAP framework The white-boxed prediction mapping obtained after additive decomposition; This is the expected value of the benchmark water level calculated over the entire sample space; For the sample time Above, the first Marginal contribution of dimensional features to predicted water level (SHAP value); Normalized multi-source feature input vector The total number of feature dimensions included.

[0025] Specifically, after obtaining the trained nonlinear water level prediction model, this embodiment introduces the Shapley additive interpreter to encapsulate it, thereby constructing the XGBoost-SHAP white-box prediction model. The core objective of this model is to decompose the complex prediction values ​​output by the traditional black-box model into a highly interpretable linear additive structure. This structure follows the additive axiom in cooperative game theory, that is, the predicted lake water level at any given time is strictly equal to the sum of the marginal contributions of all features to a global baseline expectation value.

[0026] The white-box model decomposed in this embodiment consists of three core elements: First, the baseline water level expectation value, which is the average predicted value calculated over the entire sample space, representing the expected water level starting point given by the model without considering any specific feature information input. Second, the total number of feature dimensions, which determines the number of features involved in the prediction, corresponding to the sum of effective features contained in the normalized multi-source feature input vector. Finally, the feature marginal contribution, which measures the strength and direction of the influence of a certain feature dimension on the deviation of the predicted water level from the baseline expectation value at a specific sample time; positive values ​​represent a promoting effect, and negative values ​​represent an inhibiting effect. The physical function of this white-box prediction model is to transform the complex nonlinear mapping into a simple linear summation of the contributions of each driving factor, thereby achieving a transparent analysis of the driving mechanism of water level evolution.

[0027] Assuming the expected baseline water level, calculated across the entire sample space, is 3200 meters, this serves as the starting point for the model's prediction. Assuming the normalized multi-source feature input vector contains a total of 5 feature dimensions. At a specific moment, the marginal contribution of precipitation is +50, temperature is -5, agricultural water use is -10, nighttime light index is +0.5, and the sum of other feature contributions is +1. In this embodiment, an addition operation is performed: the expected baseline water level of 3200 meters is added to all positive and negative contributions, resulting in a final predicted lake level of 3236.5 meters. This final value is mathematically identical to the predicted value output by the nonlinear water level prediction model, but it also reveals the specific impact of each driving factor on the prediction result.

[0028] Furthermore, the specific implementation process of step 300 is as follows: In this embodiment, when performing attribution analysis on driving factors, a zero-contribution baseline is first introduced based on the curve shape of the SHAP dependency plot to visually distinguish the direction of the driving factors' influence on water level evolution. The zero-contribution baseline represents the position where the marginal contribution value is zero, used to delineate the promoting and inhibiting regions of the characteristic effect. This embodiment aims to identify key critical points in the nonlinear response of driving factors by analyzing the geometric morphological characteristics of the fitted curves in the plot relative to the zero-contribution baseline. This analysis is a prerequisite for subsequent identification of genuine and false driving factors and verification of physical mechanisms in this embodiment.

[0029] This embodiment then extracts the positive-negative effect reversal threshold. This threshold characterizes the critical point at which the driving factor's influence on water level fundamentally changes direction. This embodiment identifies the intersection points between the fitted curve in the SHAP dependency plot and the zero-contribution baseline, and extracts the original feature values ​​corresponding to these intersection points on the horizontal axis. These original feature values ​​are then determined as the positive-negative effect reversal threshold. For example, when precipitation is below this threshold, it may lead to a drop in water level (inhibition), while when precipitation is above this threshold, it may lead to a rise in water level (promotion). This threshold clearly defines the nature of the influence.

[0030] Building upon this, this embodiment further extracts the action intensity mutation threshold. This threshold characterizes the critical point where the efficiency or intensity of the feature contribution rate changes significantly, aiming to capture nonlinear mutation mechanisms. This embodiment identifies inflection points where the tangent slope changes significantly or extreme points in the curve shape within the fitted curve, and extracts the original feature values ​​corresponding to these inflection points or extreme points on the horizontal axis. These original feature values ​​are then determined as the action intensity mutation threshold. This action intensity mutation threshold is used to distinguish between mild and high-efficiency regions of feature action. Finally, this embodiment summarizes all extracted positive and negative effect reversal thresholds and action intensity mutation thresholds as the nonlinear response threshold output for this feature dimension, used for subsequent causal logic verification.

[0031] Furthermore, the specific implementation process of step 400 is as follows: This embodiment first prepares the input data matrix for the Mantel statistical test. It iterates through each feature dimension of the normalized multi-source feature input matrix, calculating the numerical distance between any two time-step samples within the current feature dimension, thereby constructing a feature time-series difference matrix reflecting the similarity relationship of the feature over time. Simultaneously, this embodiment calculates the numerical distance between any two time-step samples in the target water level vector, constructing a target time-series difference matrix. Both the target time-series difference matrix and the feature time-series difference matrix are symmetric matrices describing the similarity relationship between samples, used for subsequent inter-matrix correlation assessment.

[0032] This embodiment then calculates the Mantel statistic based on the constructed feature time-series difference matrix and the target time-series difference matrix. Specifically, this embodiment expands the effective elements of the feature time-series difference matrix and the target time-series difference matrix into two one-dimensional vectors, and calculates the linear correlation coefficient between these two one-dimensional vectors. This embodiment uses the calculated linear correlation coefficient as the initial Mantel statistic. This statistic intuitively reflects whether there is synchronicity or antisynchronicity between the time evolution pattern of this feature dimension and the time evolution pattern of lake water level, providing a preliminary basis for distinguishing between coincidental trends and potential causal relationships.

[0033] To verify the reliability of the initial Mantel statistic, this embodiment performs a random permutation significance verification. This embodiment performs multiple synchronous random shuffling and rearrangement of the rows and columns of the feature time-series difference matrix, and recalculates the linear correlation coefficient after each rearrangement to generate a random correlation coefficient distribution. This embodiment calculates the statistical significance probability value by statistically analyzing the position of the initial Mantel statistic within the random correlation coefficient distribution. Finally, this embodiment determines whether the statistical significance probability value meets a preset confidence threshold: when the confidence threshold is met, this embodiment outputs the initial Mantel statistic as the statistical correlation coefficient for that feature dimension; when the confidence threshold is not met, this embodiment marks the statistical correlation coefficient as invalid or zero, thereby eliminating accidental statistical correlations before identifying true and false driving factors.

[0034] Furthermore, the specific implementation process of step 500 is as follows: This embodiment performs a true / false driving factor identification for each feature dimension to ensure that subsequent prediction models rely only on driving factors with physical control mechanisms, thus resolving the problem of confusion between statistical correlation and physical causality. This embodiment first checks whether the feature dimension simultaneously satisfies both the statistical significance condition and the mechanism threshold condition. The statistical significance condition means that the Mantel statistical correlation coefficient corresponding to the feature dimension must meet a preset confidence level requirement to exclude accidental statistical correlations. The mechanism threshold condition means that the feature dimension must be extracted to at least one valid nonlinear response threshold to prove that the feature's influence on water level has a clear physical control node or turning point mechanism.

[0035] This embodiment then classifies and determines each feature dimension based on the results of the dual conditions. Features that simultaneously satisfy both the statistical significance condition and the mechanism threshold condition are identified as true causal drivers. These true causal drivers indicate that the feature directly influences the water level evolution process through specific physical control nodes or nonlinear mutation mechanisms. Conversely, features that satisfy the statistical significance condition but fail to satisfy the mechanism threshold condition are identified as spurious correlation coincidence factors. These spurious correlation coincidence factors indicate that the feature exhibits a monotonous change without an inflection point or a simple synchronous trend in the SHAP dependency graph, showing only a statistically significant numerical coincidence with water level evolution in the time dimension without any physical causal relationship.

[0036] Based on the classification results described above, this embodiment constructs a causal driving feature set. From the feature list corresponding to the normalized multi-source feature input matrix, this embodiment removes all features identified as spurious correlation coincidence factors, completely eliminating their interference with the model. Subsequently, this embodiment aggregates all features identified as genuine causal driving factors to construct the causal driving feature set. This causal driving feature set contains only driving factors that have passed both statistical and physical verification, providing a pure and highly reliable input set for subsequent model reconstruction and robust prediction.

[0037] Furthermore, the specific implementation process of step 600 is as follows: This embodiment first performs dimension masking based on the results of the true / false driving factor identification to reconstruct the input data. Based on the causal driving feature set, this embodiment performs a filtering operation on the normalized multi-source feature input matrix, retaining only the column data corresponding to the feature dimensions belonging to the true causal driving factors, and removing all feature dimensions belonging to the false correlation coincidence factors and irrelevant factors. Through this operation, this embodiment generates a causal constraint feature input matrix with smaller dimensions, containing only the true physical driving elements that have passed dual verification.

[0038] This embodiment then initializes the extreme gradient boosting model architecture within the XGBoost-SHAP white-box prediction model to prepare for model reconstruction training. This embodiment loads the newly generated causal constraint feature input matrix as the new model training input set and loads the target water level vector as the model training target set. This ensures that the model fitting process completely eliminates the interference of all spurious correlation factors, allowing the model's learning to focus on resolving the true physical mechanisms, thereby improving the model's causal robustness and generalization ability.

[0039] In this embodiment, full-element iterative training is performed on the causal constraint feature input matrix to initiate the model reconstruction process. This embodiment constructs a new decision tree ensemble to fit the target water level vector, re-establishing a nonlinear mapping relationship controlled only by the true causal driving factors. After training and convergence, this embodiment generates a reconstructed prediction model. This reconstructed prediction model has higher physical reliability and can accurately reflect the actual response of the closed watershed water level to core climate and anthropogenic driving factors, avoiding prediction biases caused by spurious correlations in the original model.

[0040] Furthermore, the expression for the reconstructed prediction model is: ; in, For at any time The lake water level predictions given by the reconstructed prediction model are only controlled by the real causal driving factors; For a moment The causal constraint feature input vector obtained after causal-driven feature set masking; This is the reconstructed nonlinear water level prediction function obtained by retraining on the causal constraint feature input matrix; To reconstruct the model of the first Sub-prediction functions corresponding to each regression decision tree; To reconstruct the number of regression decision trees integrated in the predictive model.

[0041] Furthermore, the specific implementation process of step 700 is as follows: This embodiment then performs preprocessing operations on the original multi-source dataset for the test period before prediction. This operation strictly follows the rules of the model training phase: First, spatiotemporal alignment and standardization are performed, but here the feature mean and standard deviation parameters stored during model training must be used to ensure that the test data and training data are on the same numerical scale. Second, dimensionality masking is performed, retaining only the feature dimensions contained in the causal driving feature set obtained after identifying true and false driving factors, and removing all irrelevant features. After these two steps, this embodiment obtains a causal constraint feature input matrix that can be directly input into the reconstructed prediction model.

[0042] In this embodiment, the causal constraint feature input matrix is ​​input into the reconstructed prediction model. Based on its internally established nonlinear mapping relationship, controlled only by the true causal driving factors, the reconstructed prediction model performs inference calculations on each sample point in the causal constraint feature input matrix to obtain the predicted lake water level at that moment. This embodiment ultimately summarizes the prediction results for all moments to form a complete time series of predicted water level evolution values, thus achieving an accurate prediction of the future water level change trend of lakes in a closed watershed.

[0043] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0044] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for predicting the water level evolution of lakes in a closed watershed based on multi-source data, characterized in that, include: The original multi-source dataset of the closed watershed is collected, and the original multi-source dataset is spatiotemporally aligned and standardized to obtain the normalized multi-source feature input matrix and the corresponding target water level vector. An XGBoost-SHAP white-box prediction model is constructed, and the normalized multi-source feature input matrix is ​​trained and analyzed with all elements to obtain the SHAP dependency map of each feature dimension. Extract the nonlinear response threshold based on the SHAP dependency graph; Perform Mantel statistical tests on the normalized multi-source feature input matrix to obtain the statistical correlation coefficients between each feature and the target water level vector; Based on the statistical correlation coefficient and the nonlinear response threshold, the true and false driving factors are identified, and the causal driving feature set is obtained by screening. Based on the causal-driven feature set, the XGBoost-SHAP white-box prediction model is reconstructed to obtain the reconstructed prediction model. The original multi-source dataset for the period to be tested is obtained and input into the reconstructed prediction model to obtain the predicted value of water level evolution.

2. The method for predicting the water level evolution of a closed watershed lake based on multi-source data according to claim 1, characterized in that, The original multi-source dataset includes: Meteorological forcing factors, hydrophysical factors, land use cover factors, anthropogenic disturbance factors, and socioeconomic factors.

3. The method for predicting the water level evolution of lakes in a closed watershed based on multi-source data according to claim 1, characterized in that, The XGBoost-SHAP white-box prediction model is constructed, and the normalized multi-source feature input matrix is ​​trained and analyzed using full-element methods to obtain SHAP dependency maps for each feature dimension, including: An extreme gradient boosting model architecture based on the first decision tree ensemble is established, and the normalized multi-source feature input matrix is ​​divided into a training dataset and a validation dataset. Iterative training is performed on the training dataset, and a boosting tree ensemble is generated by constructing a second decision tree to fit the prediction residuals produced by the first decision tree. Based on boosting tree ensemble, during the iterative training process, grid search and cross-validation are performed using the validation dataset to optimize the key hyperparameters of the extreme gradient boosting model architecture, generating a trained nonlinear water level prediction model; the key hyperparameters of the model include learning rate, maximum tree depth, and subsample sampling ratio. The nonlinear water level prediction model is encapsulated by introducing a Shapley additive interpreter to construct an XGBoost-SHAP white-box prediction model. Based on the principle of cooperative game theory, the marginal contribution value of each sample in the normalized multi-source feature input matrix on each feature dimension is calculated using the XGBoost-SHAP white-box prediction model, and all the calculated marginal contribution values ​​are integrated into a SHAP value matrix. Traverse each feature dimension contained in the normalized multi-source feature input matrix, extract the original feature value sequence under each feature dimension as the horizontal axis data, and extract the SHAP value sequence corresponding to each feature dimension from the SHAP value matrix as the vertical axis data. A scatter plot reflecting the mapping relationship between feature value changes and the prediction contribution of the XGBoost-SHAP white-box prediction model is constructed using the original feature value sequence and the SHAP value sequence; The SHAP dependency map is obtained by fitting the scatter plot.

4. The method for predicting the water level evolution of lakes in a closed watershed based on multi-source data according to claim 3, characterized in that, The expression for the nonlinear water level prediction model is: ; in, For a moment Predicted lake water levels; For a moment Normalized multi-source feature input vector; This is the overall nonlinear water level prediction function trained using the extreme gradient boosting algorithm; For the integration of the first Sub-prediction functions corresponding to each regression decision tree; This represents the total number of regression decision trees integrated into the nonlinear water level prediction model.

5. The method for predicting the water level evolution of a closed watershed lake based on multi-source data according to claim 3, characterized in that, The expression for the XGBoost-SHAP white-box prediction model is: ; in, For a moment Lake water level predictions given by the XGBoost-SHAP white-box prediction model; To develop a nonlinear water level prediction function within the SHAP framework The white-boxed prediction mapping obtained after additive decomposition; This is the expected value of the benchmark water level calculated over the entire sample space; For the sample time Above, the first Marginal contribution of dimensional features to predicted water level (SHAP value); Normalized multi-source feature input vector The total number of feature dimensions included.

6. The method for predicting the water level evolution of a closed watershed lake based on multi-source data according to claim 1, characterized in that, The extraction of the nonlinear response threshold based on the SHAP dependency graph includes: A zero-contribution baseline is introduced into the SHAP dependency graph, and the geometric morphological characteristics of the fitted curve in the SHAP dependency graph relative to the zero-contribution baseline are analyzed. Based on geometric morphological features, the intersection points between the fitted curve and the zero contribution baseline are identified, the original feature values ​​corresponding to the intersection points on the horizontal axis are extracted, and the original feature values ​​are determined as the positive and negative effect flipping thresholds. Identify the inflection points or extreme points of the curve shape in the fitted curve where the slope of the tangent changes significantly, extract the original feature values ​​corresponding to the inflection points or extreme points on the horizontal axis, and determine the current original feature values ​​as the threshold for sudden change in the intensity of action. The nonlinear response threshold is obtained by summing all the extracted positive and negative effect reversal thresholds and the effect intensity mutation threshold.

7. The method for predicting the water level evolution of a closed watershed lake based on multi-source data according to claim 1, characterized in that, Perform Mantel statistical tests on the normalized multi-source feature input matrix to obtain the statistical correlation coefficients between each feature and the target water level vector, including: Based on each feature dimension of the normalized multi-source feature input matrix, calculate the numerical distance between any two time step samples under the current feature dimension, and construct the feature temporal difference matrix. Simultaneously calculate the numerical distance between any two time step samples in the target water level vector to construct the target time series difference matrix; Expand the feature time series difference matrix and the target time series difference matrix into one-dimensional vectors, and calculate the linear correlation coefficient between the two one-dimensional vectors to obtain the initial Mantel statistic. The rows and columns of the characteristic time series difference matrix are synchronously and randomly shuffled multiple times, and the linear correlation coefficient is recalculated after each rearrangement to generate a random correlation coefficient distribution. The initial Mantel statistic is located in the distribution of the random correlation coefficient, and the statistical significance probability value is calculated. Determine whether the statistical significance probability value meets a preset confidence threshold. If the confidence threshold is met, output the initial Mantel statistic as the statistical correlation coefficient of the feature dimension. If the confidence threshold is not met, mark the statistical correlation coefficient of the current feature dimension as invalid or zero.

8. The method for predicting the water level evolution of a closed watershed lake based on multi-source data according to claim 1, characterized in that, The step of distinguishing between true and false driving factors based on the statistical correlation coefficient and the nonlinear response threshold, and filtering to obtain a causal driving feature set, includes: For each feature dimension of the normalized multi-source feature input matrix, check whether the current feature dimension simultaneously satisfies the statistical significance condition and the mechanism threshold condition. The statistical significance condition is that the statistical correlation coefficient corresponding to the feature dimension meets the preset confidence level requirement. The mechanism threshold condition is that the feature dimension is extracted to at least one valid nonlinear response threshold. Features that simultaneously satisfy the statistical significance condition and the mechanism threshold condition are identified as true causal driving factors, wherein the true causal driving factor represents that the feature directly acts on the water level evolution process through a specific physical control node or a nonlinear mutation mechanism. Features that meet the statistical significance condition but fail to meet the mechanism threshold condition are identified as spurious correlation coincidence factors. The spurious correlation coincidence factor indicates that the feature shows a monotonous change without an inflection point or a simple synchronous trend in the SHAP dependency graph, and has only a statistical numerical coincidence with water level evolution in the time dimension without any physical causal relationship. Remove all features identified as false correlation coincidence factors from the feature list corresponding to the normalized multi-source feature input matrix, and collect all features identified as true causal driving factors to construct the causal driving feature set.

9. The method for predicting the water level evolution of a closed watershed lake based on multi-source data according to claim 8, characterized in that, The step of reconstructing the XGBoost-SHAP white-box prediction model based on the causal-driven feature set to obtain the reconstructed prediction model includes: Based on the causal driving feature set, the normalized multi-source feature input matrix is ​​subjected to dimension masking to filter, retaining only the column data corresponding to the feature dimensions belonging to the real causal driving factors, and removing the feature dimensions belonging to the false correlation coincidence factors and irrelevant factors, thereby generating a causal constraint feature input matrix that contains only real physical driving elements. Initialize the extreme gradient boosting model architecture inside the XGBoost-SHAP white-box prediction model, load the causal constraint feature input matrix as a new model training input set, and load the target water level vector as the model training target set; Full-factor iterative training is performed on the causal constraint feature input matrix. A new decision tree ensemble is constructed to fit the target water level vector, and a nonlinear mapping relationship controlled only by the true causal driving factor is established to generate a reconstructed prediction model.

10. The method for predicting the water level evolution of a closed watershed lake based on multi-source data according to claim 1, characterized in that, The expression for the reconstructed prediction model is: ; in, For at any time The lake water level predictions given by the reconstructed prediction model are only controlled by the real causal driving factors; For a moment The causal constraint feature input vector obtained after causal-driven feature set masking; This is the reconstructed nonlinear water level prediction function obtained by retraining on the causal constraint feature input matrix; To reconstruct the model of the first Sub-prediction functions corresponding to each regression decision tree; To reconstruct the number of regression decision trees integrated in the predictive model.

Citation Information

Cited By

  • A three-source characteristic similarity transfer hydrological simulation method for a data-deficient river basin

    CN122221702A