Water resource safety assessment and future change trend prediction method
By combining the DPSR model and random forest with Bayesian optimization algorithm and SHAP algorithm, a water resource security evaluation index system is constructed, which solves the problems of insufficient accuracy and interpretability of traditional methods and realizes efficient water resource security assessment and future trend prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIANJIN UNIV
- Filing Date
- 2025-12-10
- Publication Date
- 2026-04-17
AI Technical Summary
Traditional water resource security assessment methods are insufficient in terms of accuracy and interpretability, making it difficult to effectively respond to emergencies and nonlinear responses. Furthermore, machine learning models rely on historical data and lack interpretability.
A four-dimensional water resources security evaluation index system was constructed using the DPSR model. By combining random forest and Bayesian optimization algorithms, the contribution of each factor was explained through the SHAP algorithm, a water resources security prediction model was constructed, and data standardization and catastrophe series method were used for evaluation.
It improves the accuracy and interpretability of water resource security assessments, enabling prediction of trends over the next 10 years, adaptability to assessments in multiple regions, and provision of scientific management support.
Smart Images

Figure CN121882684A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of safety prediction technology, specifically relating to a method for water resource safety assessment and prediction of future change trends. Background Technology
[0002] Water resources are crucial for human social development and ecological balance. Global population growth and accelerated industrialization have exacerbated the supply-demand imbalance. Some regions in my country also face shortages and uneven distribution, necessitating scientific assessment and protection to achieve sustainable utilization. Traditional quantitative models and multivariate statistical analysis are not accurate enough. Although entropy method, analytic hierarchy process (AHP) and fuzzy comprehensive evaluation method have been improved in recent years, they still have limitations in dealing with sudden events and nonlinear responses. The catastrophe series method, by analyzing the catastrophe characteristics of the system, can more effectively identify potential risks and has a higher evaluation effectiveness. In water resource security prediction, traditional mathematical and statistical methods have limited accuracy, while machine learning methods, although offering improvements, suffer from reliance on historical data and insufficient interpretability. The Shapley Additive Interpretation Algorithm (SHAP), based on machine learning, can quantify the contribution of various factors, addressing the problem of model interpretability and providing new insights for management. Random forests (RF) are widely used for water resource security prediction. They can handle high-dimensional nonlinear data and are highly fault-tolerant, but they are computationally complex and have weak interpretability. Bayesian optimization (BO) can optimize RF hyperparameters without increasing computational complexity, thereby enhancing prediction accuracy and robustness. The combination of the two can maintain high accuracy even when data is missing or noisy, and it also breaks through the dependence of traditional methods on linear relationships. Currently, while evaluation methods based on machine learning and optimization algorithms are developing rapidly, traditional methods (such as linear programming and heuristic scheduling rules) still face challenges such as poor interpretability and difficulty in considering the complexities of reality. To address this, this paper focuses on the "Drive-Pressure-State-Response" (DPSR) model indicators, combining the RF+BO model and the SHAP algorithm to develop an evaluation indicator system. This system aims to improve the interpretability of the model and the reliability of the results. Furthermore, it will predict water resource changes over the next 10 years and interpolate unknown indicators, providing scientific support for sustainable water resource management. Summary of the Invention
[0003] The purpose of this invention is to provide a method for water resource security assessment and prediction of future trends, addressing the shortcomings and deficiencies of existing technologies.
[0004] The technical problem solved by this invention is achieved through the following technical solution: A method for assessing water resource security and predicting future trends, the method comprising the following steps: S1. Construct a water resource security evaluation index system based on the DPSR model, covering four dimensions: resources, society, economy, and ecology. Establish a water resource carrying capacity evaluation index system with four baseline layers: driving force system, pressure system, state system, and response system, comprising a total of 30 secondary indicators. Construct the water resource security evaluation index system, introduce the entropy method to calculate the indicator weights, and rank the indicators according to the weight calculation results. Establish a hierarchical catastrophe model to evaluate water resource carrying capacity. S2. The SHAP additive interpretation algorithm is used to comprehensively analyze the contribution of each water resource indicator to the assessment results, calculate the impact of the contribution value of each feature on water resource security, identify key factors, optimize feature selection, screen out the core features most important to the assessment results, remove redundant features, and simplify the model. S3. By constructing an appropriate water resource security prediction model combining random forest and Bayesian optimization algorithm, rapid prediction and decision-making are achieved, and different benchmark levels of water resource security assessment are classified and evaluated.
[0005] Moreover, S1 specifically refers to: S1.1 Data standardization: Standardize the sample data before measurement; ; in: The standardized value of the j-th indicator for the i-th sample; Let be the value of the j-th indicator for the i-th sample; S1.2 Calculate the weight of the index value of the i-th sample under the j-th index: ; S1.3 Calculate the entropy value of the j-th index: ; In the formula: k = -1 / lnn; S1.4 Calculate the weight of the j-th indicator: ; S1.5 Calculate the indicator layer weight values according to S1.1~S1.5, and then add the indicator layer weight values to obtain the system layer weight values; S1.6 The specific calculation steps of the mutation series method are as follows: a. Establish a multi-level indicator system. Based on the evaluation purpose, decompose the indicator system into a multi-level subsystem composed of several evaluation indicators, and rank the importance of each indicator in each level according to the entropy value calculation results. b. Determine the type of mutation system. Commonly used mutation system types include folding mutations, cusp mutations, swallowtail mutations, and butterfly mutations. Their system models and normalization formulas are shown below: ; In the formula: The potential function is the abrupt change of the fold; It is the potential function for a cusp abrupt change; Let be the potential function for the swallowtail mutation; The potential function for the butterfly mutation is given by: x is the state variable; a, b, c, and d are all control variables of x. x a 、x b、 x c and x d This corresponds to the normalization formula; c. Standardize the raw data of the indicators and convert the raw data of each indicator into dimensionless values between [0,1]. d. Use the normalization formula to perform comprehensive quantitative calculations to obtain the total mutation membership value, and conduct comparative analysis based on this; use the normalization formula to calculate the corresponding state variables for each control variable. x There are two principles for determining the value: "taking the smaller value from the larger, medium, and small" and the average value. If the control variables are not complementary, the "taking the smaller value from the larger, medium, and small" principle is used; otherwise, the average value principle is used.
[0006] Moreover, S2 specifically refers to: S2.1 Determination of Baseline Values: Baseline Values It is the model's average prediction under the "reference distribution," typically calculated based on training data: ; In the formula: This represents the mean of the model predictions for all samples in the training set, ensuring that the starting point of the additive interpretation model conforms to the overall distribution characteristics of the data; S2.2, Define the Shapley value: The Shapley value is used to measure the contribution of each feature to the model's prediction. The formula is: ; In the formula: Let S be the Shapley value of feature i, where S is a subset of features and N is the set of all features. It is the model output based on the feature subset S; S2.3 Weighted Interpretation: The model's predictions can be obtained by weighted summation of features, using the following formula: ; In the formula: It is a baseline prediction. It represents the contribution of each feature to the model output; S2.4 Consistency of Local Explanation For a single sample x, its predicted value must satisfy the condition that the predicted result of the sample is equal to the baseline value plus the sum of the SHAP values of all features for that sample, ensuring the consistency between the local explanation and the model output.
[0007] Moreover, S3 specifically refers to: S3.1 Construct a water resources security assessment framework based on the DPSR model. Through sample experiments, couple the four subsystems of driving force, pressure, state, and response to establish a water resources carrying capacity assessment index system covering 30 secondary indicators in Hubei Province. On this basis, use the entropy-mutation series method to complete the construction of the water resources security assessment model. S3.2 Perform hierarchical analysis of the DPSR model to complete the water resource security level assessment and risk identification; classify the 30 secondary indicators into four baseline layers: driving force, pressure, state and response; and combine random forest and Bayesian optimization algorithms to finally construct a complete water resource security evaluation index system. S3.3, the dataset is divided into training and test sets. The correlation between the indicators is analyzed by mutual information (MIR) and F test (FR). The model is trained based on random forest and Bayesian optimization algorithm to finally obtain the water resource security prediction model. S3.4, Model training and optimization: Using a tree assigner, the dimensions of input factors are optimized to predict the four major subsystems of driving force, pressure, state, and response, as well as the overall level of water resource security for the next ten years, and to determine the optimal combination of hyperparameters; at the same time, missing data values in the baseline layer are imputed.
[0008] The advantages and beneficial effects of this invention are as follows: This invention constructs an evaluation index system encompassing four dimensions: resources, society, economy, and ecology, using the entropy-mutation series method. It employs the entropy-variance ordinal method to assess water resource carrying capacity and combines it with the interpretable algorithm Shapley Weighted Approach (SHAP) to predict water resource trends. Future trend prediction utilizes random forest and Bayesian optimization models, possessing both the ability to accurately predict future security status and effectively imput missing data. This method is applicable to water resource security assessments in multiple regions, exhibiting good adaptability and strong practicality. It can effectively formulate effective protection measures by quantitatively assessing the current state of water security and diagnosing hindering factors, thus promoting the sustainable and high-quality development of water resource security. Attached Figure Description
[0009] Figure 1 This is a schematic diagram of the process of the present invention; Figure 2 This is a schematic diagram of the coupling coordination degree of the Hubei Province water resources security subsystem from 2007 to 2024 according to the present invention; Figure 3This is a schematic diagram of the obstacle degree of the water resources security assessment criteria layer in Hubei Province from 2007 to 2024, as presented in this invention. Figure 4 This is a schematic diagram of the obstacle degree of water resource security obstacle factors in Hubei Province from 2007 to 2024, based on the present invention. Figure 5 This is a schematic diagram of the bee colony graph of the SHAP values of each feature of the model of the present invention; Figure 6 This is a schematic diagram illustrating the changing trend of the water resources security assessment index in Hubei Province from 2007 to 2024. Figure 7 This is a schematic diagram illustrating the results of the automatic optimization of hyperparameters according to the present invention; Figure 8 This is a schematic diagram comparing the invention before and after optimization. Figure 9 This is a schematic diagram illustrating the comparison of prediction errors in this invention; Figure 10 These are the predicted values for each indicator from 2025 to 2034 of this invention; Figure 11 This is a schematic diagram illustrating the prediction values and prediction errors of 10 random insertions repeated 500 times according to the present invention. Detailed Implementation
[0010] The present invention will be further described in detail below through specific embodiments. The following embodiments are merely descriptive and not limiting, and should not be used to limit the scope of protection of the present invention.
[0011] like Figure 1 As shown, this invention provides a method for water resource security assessment and prediction of future trends. Taking Hubei Province's water resources as an example, a layer-by-layer decomposition method is adopted. Combined with the current status of Hubei's water resources, and based on principles of scientific rigor, representativeness, and data availability, a water resource carrying capacity evaluation index system for Hubei Province is established from four baseline layers: driving force system, pressure system, state system, and so on, comprising a total of 30 secondary indicators. The specific water resource security evaluation index system is shown in Table 1. Here, "positive" means that the larger the data, the more conducive it is to improving water resource security, and vice versa.
[0012] Table 1. Water Resources Security Evaluation Index System of Hubei Province
[0013] Step 1: Construct a water resources security evaluation index system based on the DPSR model. Considering the current water resources situation in Hubei Province, and adhering to principles of scientific rigor, representativeness, and data availability, a coupling coordination degree analysis of 30 secondary indicators at the criterion level was conducted. The coupling coordination degree of the water resources security subsystem in Hubei Province from 2007 to 2024 is shown below. Figure 2 The obstacle degree of the water resources security assessment criteria in Hubei Province from 2007 to 2024 is as follows: Figure 3The degree of water resource security barrier factors in Hubei Province from 2007 to 2024 is as follows: Figure 4 A water resources security assessment model for Hubei Province based on the entropy-mutation series method was constructed. The weights of the calculated indicators and the types of system mutations are shown in Table 2.
[0014] Table 2. Indicator Weights and System Mutation Types
[0015] The prediction method of the present invention includes the following steps: Step 1.1, Data Standardization. Since the units of measurement for each indicator are not uniform, the sample data needs to be standardized before measurement.
[0016] ; ; In the formula: The standardized value of the j-th indicator for the i-th sample; a i,j For the i-th sample j The numerical value of each indicator.
[0017] Step 1.2, calculate the weight of the index value of the i-th sample under the j-th index: ; Step 1.3, calculate the first... j Entropy value of the item indicator: ; In the formula: k = -1 / lnn; Step 1.4, calculate the weight of the j-th indicator. ; Step 1.5: Calculate the indicator layer weight values according to steps 1-4, and then add the indicator layer weight values together to obtain the system layer weight values.
[0018] Step 1.6, the specific calculation steps of the mutation series method are as follows.
[0019] a. Establish a multi-level indicator system. Based on the evaluation objectives, decompose the indicator system into a multi-level subsystem composed of several evaluation indicators, and rank the importance of each indicator in each level according to the entropy value method. b. Determine the type of mutation system. Commonly used mutation system types include fold mutations, cusp mutations, swallowtail mutations, and butterfly mutations. Their system models and normalization formulas are shown below.
[0020] ; In the formula: The potential function of the folded abrupt change Potential function of cusp abrupt change The potential function of the swallowtail mutation The potential function for the butterfly mutation is: x x is a state variable; a, b, c, and d are all control variables of x. , , and This is the corresponding normalization formula.
[0021] c. Standardization of raw indicator data. The raw data of each indicator are converted into dimensionless values between [0,1].
[0022] d. Use the normalization formula to perform comprehensive quantitative calculations to obtain the total mutation membership value, and then conduct comparative analysis based on this. The corresponding state variable x values calculated using the normalization formula for each control variable have two selection principles: "taking the smaller of the larger and medium values" and the average value. If the control variables are "non-complementary," the "taking the smaller of the larger and medium values" principle is used; otherwise, the average value principle is used.
[0023] Step 2 involves constructing an additive interpretation algorithm for SHAP feature selection to help analyze the contribution of various water resource indicators to the assessment results and provide a scientific basis. Through SHAP feature selection, key factors affecting water resource security can be identified, thereby enabling more precise formulation of protection and management strategies, and improving the reliability and transparency of model predictions. The beehive diagram of the SHAP values of each feature in the model is shown below. Figure 5 As shown.
[0024] Step 2.1, Determination of baseline values: Baseline values It is the model's average prediction under the "reference distribution," typically calculated based on training data: ; In the formula: This represents the mean of the model predictions for all samples in the training set, ensuring that the starting point of the additive interpretation model conforms to the overall distribution characteristics of the data.
[0025] Step 2.2, Define the Shapley value: The Shapley value measures the contribution of each feature to the model's predictions. The formula is: ; In the formula: Let S be the Shapley value of feature i, where S is a subset of features and N is the set of all features. It is the model output based on the feature subset S.
[0026] Step 2.3, Explanation of Addition: The model's predictions can be obtained by weighted summation of features, using the following formula: ; In the formula: It is a baseline prediction. It represents the contribution of each feature to the model output.
[0027] Step 2.4, Consistency of Local Explanation: For a single sample x, its predicted value must satisfy the condition that the predicted result of the sample is equal to the baseline value plus the sum of the SHAP values of all features for that sample, ensuring consistency between the local explanation and the model output.
[0028] Step 3, by classifying the results of the water resources security evaluation index, and referring to the research results at home and abroad, we set 1 to 5 levels, namely I, II, III, IV and V, and the corresponding conventional security evaluation value range is: extremely unsafe (0~0.2), relatively unsafe (0.2~0.4), basically safe (0.4~0.6), safe (0.6~0.8), and very safe (0.8~1). The evaluation value of the mutation series method is relatively large, and the evaluation value under the mutation series method needs to be corrected
[23] . The values of each secondary index are set to 0.2, 0.4, 0.6 and 0.8 respectively. The mutation series method is used to normalize each value, and finally the safety rating level range under the mutation series method is obtained. The final calculation results are shown in Table 3. The trend of water resources security evaluation index change in Hubei Province from 2007 to 2024 is as follows. Figure 6 As shown.
[0029] Table 3 Classification of Water Resources Security Assessment Levels in Hubei Province
[0030] Step 4: By constructing an appropriate water resource security prediction model based on random forest combined with Bayesian optimization algorithm, rapid prediction and decision-making can be achieved. The water resource prediction model based on random forest (RF) and Bayesian optimization (BO) enhances the fitting ability to nonlinear relationships through ensemble learning and realizes adaptive optimization of hyperparameters by using Bayesian optimization, thereby significantly improving prediction accuracy and model robustness.
[0031] Step 4.1, Dataset Partitioning: Using the Bootstrap method, split the original dataset... n samples are randomly drawn with replacement to form multiple training subsets. Each subset is used to train a decision tree.
[0032] Decision tree generation: Each decision tree uses random feature selection during training to avoid overfitting. For example, suppose the feature subset of a certain node is... When partitioning nodes, a subset of features is randomly selected. When performing splitting, features that maximize information gain are typically selected. The formula for calculating information gain is: ; In the formula: It is a dataset entropy, It is the i-th subset.
[0033] Node splitting criteria: The splitting of each node in a decision tree depends on minimizing the Gini index or maximizing the information gain. The formula for the Gini index is: ; in, It is the probability of the i-th class in the dataset R.
[0034] Tree training and prediction: Each decision tree makes predictions independently, and the final result is determined by voting or averaging. For example, for classification problems, the voting formula is: ; For regression problems, the prediction results are obtained by taking the mean: ; In the formula: T is the number of trees, This is the prediction result for the t-th tree.
[0035] Feature importance evaluation: During training, the importance of each feature can be evaluated for each tree. Feature importance can be calculated by randomly permuting features using the following formula: ; In the formula: It represents the change in the impact of feature f on model performance.
[0036] Step 4.2, Selecting a surrogate model: First, Bayesian optimization requires constructing a surrogate model for the objective function. A Gaussian process (GP) is typically chosen as the surrogate model. The Gaussian process describes the relationship between different input points x using a mean function and a covariance function. The Gaussian process model is as follows: ; In the formula: It is a mean function. Let be the covariance function.
[0037] Choosing the acquisition function: Bayesian optimization determines the point for the next evaluation by using the expected improvement function. The formula for the expected improvement (E) is: ; In the formula: It is the currently known optimal value. It is the predicted value of the objective function at point x.
[0038] Choosing a new evaluation point: In each iteration, Bayesian optimization selects a new point by maximizing the acquisition function to choose the optimal evaluation point. That is: After this point is evaluated, the value of the objective function is... It is calculated and added to the model.
[0039] Updating the surrogate model: Bayesian optimization updates the surrogate model after each evaluation to include the new objective function evaluation results. The posterior distribution of the Gaussian process is updated using Bayes' theorem: ; In this way, the model can better predict the results for other unevaluated points.
[0040] Convergence Criterion: After multiple iterations, Bayesian optimization will gradually converge to the optimal or near-optimal solution of the objective function. The criterion for ending the iteration is usually when the improvement is less than a set threshold or when the maximum number of iterations is reached.
[0041] Step 4.3: Construct a water resources security assessment framework based on the DPSR model. Through sample experiments coupling four subsystems—driving force, pressure, state, and response—a water resources carrying capacity assessment index system for Hubei Province, encompassing 30 secondary indicators, is established. Based on this, the entropy-catastrophic series method is used to construct the water resources security assessment model. Hierarchical analysis of the DPSR model is performed to complete the water resources security level assessment and risk identification. The 30 secondary indicators are categorized into four baseline layers: driving force, pressure, state, and response. Combining random forest and Bayesian optimization algorithms, a complete water resources security assessment index system is finally constructed. Because these indicators are strongly correlated with the current water resources situation and management decisions, they are identified as the core input factors of the model.
[0042] Step 4.4 involves dividing the dataset into training and testing sets, and analyzing the correlation between various indicators using mutual information (MIR) and F-test (FR). A water resource security prediction model is then trained based on random forest and Bayesian optimization algorithms. Model training and optimization are then performed. The input factor dimensions are optimized using a tree assigner to predict the four subsystems of driving forces, pressures, states, and responses, as well as the overall level of water resource security over the next ten years, and to determine the optimal hyperparameter combination. Simultaneously, missing data values in the baseline layer are imputed. This model provides a scientific basis for the dynamic adjustment of water resource management strategies and helps maintain the ecological balance of the water resource system.
[0043] Step 4.5: Evaluate the accuracy of the water resources security assessment model predictions. Step 4.5.1, Automatic Hyperparameter Optimization Evaluation: Automatic hyperparameter optimization can improve the robustness of multi-energy complementary models, reduce human error, increase accuracy, avoid overfitting, and enhance generalization and prediction stability, outperforming traditional manual adjustments that rely on experience. Taking the Hubei water resources security assessment in Figure 7 as an example, combining random forest and Bayesian optimization, the optimal parameters are NumTrees=50, MinLeafSize=3, with an objective function value of 0.8106, and a time consumption of 68.08 seconds. This method has broad application prospects in many fields and can improve model reliability.
[0044] Step 4.5.2, Comparison before and after optimization: In the water resource prediction model, combining Random Forest (RF) and Bayesian optimization (BO) can significantly improve prediction accuracy and model stability. To verify the effectiveness of this method, this paper uses the index values of five systems—driving system, pressure system, state system, response system, and water resource security—from 2007 to 2019 as a test set, and predicts the index values for 2020 to 2024 based on these values. By comparing the results before and after optimization (e.g., ... Figure 8 As shown in the figure, the prediction results using a combination of random forest and Bayesian optimization are significantly improved compared to traditional random forest prediction.
[0045] First, the difference between the optimized predictions and the actual values was significantly reduced. Specifically, the prediction accuracy for water resource security was the highest, reaching 97.57%, an improvement of 0.06% compared to the unoptimized random forest prediction. This indicates that Bayesian optimization further improved the accuracy of the predictions by adjusting the model's hyperparameters during the optimization process. Second, the prediction results for the response system also showed significant improvement, with an accuracy of 93.98%, an improvement of 1.24% compared to the random forest prediction. Although the improvement in water resource security was relatively limited, the improvement in the state system demonstrated the adaptability of Bayesian optimization to different systems. The prediction model using random forest + Bayesian optimization improved the average accuracy by 0.84% compared to the traditional random forest model. This improvement not only reflects the advantages of Bayesian optimization in model parameter tuning but also demonstrates that Bayesian optimization can effectively avoid overfitting and improve the model's generalization ability, thus enabling more stable and accurate prediction results when dealing with complex data.
[0046] Bayesian optimization plays a crucial role in improving the accuracy and stability of water resource prediction models. In particular, when dealing with different systems and data types, Bayesian optimization can significantly improve the overall performance of the model, providing more reliable data support for the scientific management and decision-making of water resources.
[0047] Step 4.5.3, Comparison of Prediction Errors: Comparing prediction errors directly reflects the differences in accuracy and stability between the two methods, revealing the crucial role of Bayesian optimization in improving model performance. Random forests, as an ensemble learning method, can typically reduce model variance to some extent. However, because it cannot automatically adjust parameters, it may introduce systematic errors. On some datasets, especially when the data contains complex nonlinear relationships or is noisy, random forests may exhibit some bias. Although random forests reduce overfitting through the ensemble of multiple decision trees, the variance can still be large when the training data fluctuates significantly or is noisy, leading to unstable prediction results. Therefore, random forest models typically exhibit high variance, especially when the dataset varies greatly, where prediction errors may be more significant. Bayesian optimization, by automatically adjusting the model's hyperparameters (such as the number and depth of decision trees), can reduce bias and variance to some extent, thereby significantly optimizing prediction results. By comparing the prediction performance of random forests and random forest + Bayesian optimization, the comparison of prediction errors is shown below. Figure 9 As shown in the figure, the results indicate that the water resource security system had the smallest prediction error, with an average error of 0.0229 and a minimum error of 0.0020, representing a reduction of 0.0006 in average error compared to random forest prediction. The second smallest error was observed in the "Response system," with an average error of 0.0574 and a minimum error of 0.0260, a reduction of 0.0118 in average error compared to random forest prediction. This demonstrates that Bayesian optimization can effectively reduce the bias and variance of the random forest model, significantly improving prediction accuracy and stability.
[0048] The optimized model exhibited lower prediction errors when faced with complex datasets, demonstrating the powerful role of Bayesian optimization in improving model performance. Therefore, the combination of random forests and Bayesian optimization has become an important means of improving the accuracy of complex tasks, especially in fields such as water resource prediction.
[0049] Step 4.5.4, Forecasting of Indicators: Predicting the indicator values of the five major systems—driving system, pressure system, state system, response system, and water resource security—over the next 10 years plays a crucial role in water resource management. These forecasts not only help address the ever-changing environment and technological advancements but also provide data support for policymakers to optimize resource allocation, thereby enhancing societal adaptability. Especially against the backdrop of global climate change and increasing resource scarcity, accurate forecasting of system indicators has become a key tool for improving social management efficiency and promoting sustainable development.
[0050] Through forecasting and analyzing various indicators from 2025 to 2034 (such as...) Figure 10As shown in the figure, the predicted values and trends of different systems can be obtained. From the prediction results, the prediction results of the water resource security and response system are relatively stable with small fluctuations, and the security assessment level is III (basic safety) for all systems. This indicates that in the next 10 years, the water resource security and response system will maintain a relatively stable operating state in most cases and can effectively cope with current environmental changes; the prediction results of the pressure system are relatively more volatile. The index value predicted by random forest ranges from 0.8929 to 0.9131, while the prediction range after combining random forest with Bayesian optimization is 0.9041 to 0.9061. According to the security assessment level classification, the security level of the pressure system is between III (basic safety) and IV (safe). Although the security situation is relatively stable, attention still needs to be paid to the potential impact of external pressure factors on the system, leading to some uncertainty. In the prediction of the state system, the prediction index value of random forest is 0.8580 to 0.8880, while the prediction value after combining random forest with Bayesian optimization is 0.8617 to 0.8691. According to the safety rating classification, the state system's safety assessment level falls between II (relatively unsafe) and III (basically safe), indicating that the system faces certain risks and requires close monitoring to avoid unforeseen problems. The predicted values for the driving system range from 0.8293 to 0.8408 in the random forest model and from 0.8350 to 0.8362 after Bayesian optimization. Based on the safety rating classification, this system's safety assessment level is II (relatively unsafe). This suggests that the driving system may face significant uncertainties and risks, and its future operation may require more technical support and optimization to ensure its stability.
[0051] In conclusion, forecasting indicators from these five major systems allows for a better understanding of the potential risks and changing trends of each system, thus providing policymakers with more scientific data support. These forecasts not only help in developing more effective water resource management and emergency response strategies but also provide strong decision-making support for society in addressing complex environmental challenges.
[0052] Step 4.5.5: Comparison of 10 random data insertion predictions before and after optimization: Water resource security assessment is the core basis for policy making and resource management, and the completeness and accuracy of data directly determine the stability of system operation and the scientific nature of decision-making. However, due to the influence of multiple factors, data gaps are common and can easily interfere with the rationality of water resource security assessments. If not handled properly, it may lead to decision-makers making misjudgments based on incomplete analysis, especially in key scenarios such as resource allocation and emergency response, which can easily cause resource imbalances or management delays. Therefore, data insertion prediction is crucial, as it can reduce decision-making risks and provide comprehensive and reliable data support for decision-making. In the five major systems of water resource security assessment—driving system, pressure system, state system, response system, and water resource security—there are 10 data gaps in various indicators from 2007 to 2024. Using the random forest coupled Bayesian optimization algorithm for data insertion prediction can effectively improve the reliability of assessment results: this algorithm can accurately fill in missing values, ensure data continuity and consistency, and thus improve the system's prediction accuracy for future changes.
[0053] The predictions and prediction errors of 10 random insertions were repeated 500 times. (a) Random Forest prediction; (b) Random Forest + Bayesian optimization prediction; (c) Random Forest prediction error; (b) Random Forest + Bayesian optimization prediction error is shown below. Figure 11 As shown, the optimization effect of the Random Forest + Bayesian optimization algorithm is significant. Comparing the predicted values of missing data with the actual values reveals a substantial reduction in prediction error: the overall average error after optimization is 0.04582, the maximum error is 0.25391, and the minimum error is 0.00002, representing reductions of 0.00140, 0.00808, and 0.00001 respectively compared to before optimization. Among these, the optimization effect on water resource security indicators is particularly outstanding, with the average error after optimization being 0.02560, the maximum error 0.14477, and the minimum error 0.00002, representing reductions of 0.00217, 0.00215, and 0.00001 respectively compared to before optimization. The optimization of the driving system indicator error is also significant, with the average error decreasing to 0.03845, the maximum error to 0.13360, and the minimum error to 0.00002.
[0054] In summary, the water resource security prediction model based on random forests and Bayesian optimization possesses the dual capabilities of accurately predicting future security status and effectively imputing missing data, ensuring the scientific rigor and accuracy of the assessment. The application of data imputation in water resource management not only optimizes the decision-making process but also provides solid data support for sustainable development planning.
[0055] Although embodiments and drawings of the present invention have been disclosed for illustrative purposes, those skilled in the art will understand that various substitutions, variations and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the scope of the present invention is not limited to the contents disclosed in the embodiments and drawings.
Claims
1. A method for assessing water resource security and predicting future trends, characterized in that: The steps of the method are as follows: S1. Construct a water resource security evaluation index system based on the DPSR model, covering four dimensions: resources, society, economy, and ecology. Establish a water resource carrying capacity evaluation index system with four aspects: driving force system, pressure system, state system, and response system, with a total of 30 secondary indicators. A water resources security evaluation index system was constructed, the entropy method was introduced to calculate the index weight values, and the indicators were ranked according to the weight calculation results. A hierarchical catastrophe model was established to evaluate the water resources carrying capacity. S2. The SHAP additive interpretation algorithm is used to comprehensively analyze the contribution of each water resource indicator to the assessment results, calculate the impact of the contribution value of each feature on water resource security, identify key factors, optimize feature selection, screen out the core features most important to the assessment results, remove redundant features, and simplify the model. S3. By constructing an appropriate water resource security prediction model combining random forest and Bayesian optimization algorithm, rapid prediction and decision-making are achieved, and different benchmark levels of water resource security assessment are classified and evaluated.
2. The method for water resource security assessment and future trend prediction according to claim 1, characterized in that: Specifically, S1 is: S1.1 Data standardization: Standardize the sample data before measurement; ; in: The standardized value of the j-th indicator for the i-th sample; Let be the value of the j-th indicator for the i-th sample; S1.2 Calculate the weight of the index value of the i-th sample under the j-th index: ; S1.3 Calculate the entropy value of the j-th index: ; In the formula: k = -1 / lnn; S1.4 Calculate the weight of the j-th indicator: ; S1.5 Calculate the indicator layer weight values according to S1.1~S1.5, and then add the indicator layer weight values to obtain the system layer weight values; S1.6 The specific calculation steps of the mutation series method are as follows: a. Establish a multi-level indicator system. Based on the evaluation purpose, decompose the indicator system into a multi-level subsystem composed of several evaluation indicators, and rank the importance of each indicator in each level according to the entropy value calculation results. b. Determine the type of mutation system. Commonly used mutation system types include folding mutations, cusp mutations, swallowtail mutations, and butterfly mutations. Their system models and normalization formulas are shown below: ; In the formula: The potential function is the abrupt change in folding. It is the potential function for a cusp abrupt change; Let be the potential function for the swallowtail mutation; The potential function for the butterfly mutation is given by: x is the state variable; a, b, c, and d are all control variables of x. x a 、x b、 x c and x d This corresponds to the normalization formula; c. Standardize the raw data of the indicators and convert the raw data of each indicator into dimensionless values between [0,1]. d. Use the normalization formula to perform comprehensive quantitative calculations to obtain the total mutation membership value, and conduct comparative analysis based on this; use the normalization formula to calculate the corresponding state variables for each control variable. x There are two principles for determining the value: "taking the smaller value from the larger, medium, and small" and the average value. If the control variables are not complementary, the "taking the smaller value from the larger, medium, and small" principle is used; otherwise, the average value principle is used.
3. The method for water resource security assessment and future trend prediction according to claim 1, characterized in that: Specifically, S2 is: S2.1 Determination of Baseline Values: Baseline Values It is the model's average prediction under the "reference distribution," typically calculated based on training data: ; In the formula: This represents the mean of the model predictions for all samples in the training set, ensuring that the starting point of the additive interpretation model conforms to the overall distribution characteristics of the data; S2.2, Define the Shapley value: The Shapley value is used to measure the contribution of each feature to the model's prediction. The formula is: ; In the formula: Let S be the Shapley value of feature i, where S is a subset of features and N is the set of all features. It is the model output based on the feature subset S; S2.3 Weighted Interpretation: The model's predictions can be obtained by weighted summation of features, using the following formula: ; In the formula: It is a baseline prediction. It represents the contribution of each feature to the model output; S2.4 Consistency of Local Explanation For a single sample x, its predicted value must satisfy the condition that the predicted result of the sample is equal to the baseline value plus the sum of the SHAP values of all features for that sample, ensuring the consistency between the local explanation and the model output.
4. The method for water resource security assessment and future trend prediction according to claim 1, characterized in that: Specifically, S3 is: S3.1 Construct a water resources security assessment framework based on the DPSR model. Through sample experiments, couple the four subsystems of driving force, pressure, state, and response to establish a water resources carrying capacity assessment index system covering 30 secondary indicators in Hubei Province. On this basis, use the entropy-mutation series method to complete the construction of the water resources security assessment model. S3.2 Perform hierarchical analysis of the DPSR model to complete the water resource security level assessment and risk identification; classify the 30 secondary indicators into four baseline layers: driving force, pressure, state and response; and combine random forest and Bayesian optimization algorithms to finally construct a complete water resource security evaluation index system. S3.3, the dataset is divided into training and test sets. The correlation between the indicators is analyzed by mutual information (MIR) and F test (FR). The model is trained based on random forest and Bayesian optimization algorithm to finally obtain the water resource security prediction model. S3.4, Model training and optimization: Using a tree assigner, the dimensions of input factors are optimized to predict the four major subsystems of driving force, pressure, state, and response, as well as the overall level of water resource security for the next ten years, and to determine the optimal combination of hyperparameters; at the same time, missing data values in the baseline layer are imputed.