A socio-economic-ecological system water resource demand calculation and optimal allocation method and system
Patent Information
- Application Number
- CN202610671463.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-18
AI Technical Summary
[0005](1)部门与区域适配性不足:现有技术多采用统一模型对某一大型区域/流域整体需水进行预测,缺乏对于生活、农业、工业与生态不同部门需水驱动机制差异和区县等小区域需水异质性特征的考虑,无法反映个性化需水规律;
[0055] (1) Sectoral Differentiation Modeling and County-Level Adaptation: By constructing customized prediction models for the four sectors of life, agriculture, industry and ecology respectively, the differences in water demand driving mechanisms of each sector are fully considered. Through multi-source heterogeneous data processing, scenario data at the global and national scales are adapted to county-level administrative units to solve the spatial matching problem of existing technologies in small-scale regional applications and support refined management and precise allocation of water resources.
Smart Images

Figure CN122596484A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of water resource measurement and management technology, specifically involving a multi-scenario driven method for measuring and optimizing the allocation of socio-economic and ecosystem-differentiated water resource demands. It is applicable to the refined measurement of water demand at the county and district scale, providing decision support for water resource planning, management, and sustainable development. Background Technology
[0002] Water resources, as a vital strategic resource, are essential for human livelihoods and the continuation of civilization, a fundamental support for sustainable economic and social development, and a necessary guarantee for ecological civilization construction. Water scarcity has become a prominent bottleneck restricting high-quality regional development.
[0003] Under the combined impacts of global climate change and regional human activities, water resource systems in arid regions face increasingly severe challenges. On the one hand, global warming exacerbates variations in the hydrological cycle, leading to changes in precipitation patterns and increased uncertainty in glacial meltwater replenishment. On the other hand, the sustained and rapid economic and social development in Northwest China, along with population agglomeration, urbanization, and industrial upgrading, further drives up water demand, exacerbating the supply-demand imbalance.
[0004] Against this backdrop, the existing technology has the following objective shortcomings (Mu et al., 2020; Pichler et al., 2023; Zhu Wanyi et al., 2023; Tian Jintao et al., 2025):
[0005] (1) Insufficient adaptability to sectors and regions: Existing technologies mostly use a unified model to predict the overall water demand of a large region / basin, lacking consideration of the differences in water demand driving mechanisms among different sectors such as domestic, agricultural, industrial and ecological sectors, as well as the heterogeneous characteristics of water demand in small areas such as districts and counties, and thus failing to reflect individualized water demand patterns.
[0006] (2) Poor dynamic response capability: Traditional water demand forecasting methods based on empirical formulas and static parameterization schemes are difficult to effectively characterize the nonlinear, high-dimensional water demand response mechanism driven by climate change and human activities, which restricts the decision support capability for refined water resource management and scientific allocation.
[0007] (3) Poor interpretability of the model: Although existing machine learning algorithms have improved prediction performance, their black box mechanism has lost the transparency of the mechanism and cannot explain the action path of key factors, which affects the credibility and application value of the prediction results in actual water resource management decision-making.
[0008] (4) Insufficient scenario coverage: Existing technologies lack systematic consideration of the ongoing global climate change, increasingly stringent carbon emission constraints, and socio-economic transformation, making it difficult to support water resource planning toward carbon neutrality and sustainability goals;
[0009] (5) Single uncertainty handling method: Existing technologies handle the uncertainty of prediction results in a relatively crude manner, often using symmetrical interval estimation or failing to quantify probabilistic uncertainty, which cannot handle the skewed distribution characteristics of prediction errors and makes it difficult to provide decision-makers with accurate risk boundary basis. Summary of the Invention
[0010] The purpose of this invention is to provide a method for measuring and optimizing the allocation of socio-economic and ecosystem water resources demand, so as to realize multi-scenario driven multi-sectoral differentiated water resources demand prediction and optimization allocation at the county and district scale. On the other hand, it provides a system for measuring and optimizing the allocation of socio-economic and ecosystem water resources demand, providing an interpretable, verifiable and scalable scientific decision-making basis for regional water resources refined management, risk early warning and adaptive optimization allocation.
[0011] In a first aspect, the present invention provides a method for calculating and optimizing the allocation of socio-economic-ecosystem water resource demand, the method comprising:
[0012] Construct a basic database for regional socio-economic-ecosystem water demand measurement, which includes a historical basic database and a shared socio-economic pathway-representative concentration pathway (SSP-RCP) future scenario database;
[0013] The multi-source heterogeneous data in the basic database are preprocessed to construct an initial dataset for water resource demand prediction;
[0014] Based on preprocessed historical data, key driving factors of water demand in various sectors of life, agriculture, industry, and ecology are screened.
[0015] A multi-model set including traditional statistical models and machine learning algorithms is constructed. Based on the key driving factors, the optimal simulation models of each department are trained and screened. The Shapley additive interpretation method is used to conduct interpretability analysis on the optimal models and analyze the mechanism of action of each driving factor.
[0016] Based on the optimal simulation model, differentiated water resource demand calculations under multiple scenarios are realized. By combining kernel density estimation and Monte Carlo simulation to generate asymmetric confidence intervals, the uncertainty of the calculation results is quantified.
[0017] Conduct water supply and demand balance analysis, construct multi-objective optimized water resource allocation schemes, and provide decision support for regional sustainable water resource management.
[0018] In an optional implementation, the preprocessing of the multi-source heterogeneous data in the basic database includes:
[0019] Standardized format: Convert vector data, raster data, and text / table data into a standardized, computable format;
[0020] Scale transformation: Spatially aggregate and scale up the raster data of the shared socio-economic path-representative concentration path scenario database to match the spatial scale of district and county administrative units;
[0021] Missing value imputation: For missing values in time series data, linear interpolation is used to impute them;
[0022] Outlier identification: Outlier identification and replacement are performed using the interquartile range method;
[0023] Data normalization: All data are normalized using the feature scaling method to eliminate the dimensional differences between different indicators.
[0024] In an optional implementation, the scaling transformation includes scaling the urbanization rate data, specifically including:
[0025] The urbanization gap coefficient between counties and provinces is calculated based on historical data to obtain the multi-year average stable gap coefficient. Based on the stable gap coefficient and future provincial urbanization data, the future urbanization rate of counties is calculated. The calculated future urbanization rate of counties is subject to categorical upper limit constraints, and then smoothed using a local weighted regression scatter smoothing method to eliminate inter-year abrupt changes.
[0026] In an optional implementation, the scaling transformation includes scaling the temperature data, specifically including:
[0027] Daily near-surface temperature data were aggregated into monthly data, and the Kelvin units of near-surface temperature were converted to degrees Celsius. Six applicable global climate models were selected based on literature, and the future climate ensemble results were constructed using the equal-weighted averaging method. Using 1982-2024 as the historical baseline period, monthly multi-year averages were performed on the 1km high-resolution temperature and the multi-model ensemble results to construct a historical temperature baseline map and model historical baseline monthly data. Based on the calculation of the temperature anomaly field between the baseline period observation data and the future simulation data, the resolution of the anomaly field was improved to the target resolution through a spatial interpolation resampling algorithm and superimposed onto the historical baseline climate state raster surface to generate a high-resolution future temperature prediction raster map. Spatial regional statistical operations were performed on the raster map to extract the average temperature sequence corresponding to the administrative boundaries of districts and counties.
[0028] In an optional implementation, the step of screening key driving factors for water demand in various sectors, including domestic, agricultural, industrial, and ecological areas, based on preprocessed historical data, includes:
[0029] Construct an initial set of influencing factors on water demand for each department;
[0030] Pearson correlation analysis was used to calculate the correlation coefficients between each influencing factor and the water consumption of the corresponding department.
[0031] Factors with correlation coefficients not less than a specific threshold are selected as key drivers of water demand for the corresponding departments.
[0032] In an optional implementation, the construction of the multi-model set includes:
[0033] A multi-model set, including traditional statistical models and machine learning models, is constructed to adapt to the water use simulation needs of different departments and to achieve multi-model comparison and selection.
[0034] In an optional implementation, training and selecting the optimal simulation model for each department based on the key driving factors includes:
[0035] The key driving factors of each department are used as model inputs, and the historical water consumption of the corresponding department is used as model outputs. The models are divided into training and testing sets according to the proportions.
[0036] Cross-validation is used to optimize model parameters and avoid overfitting.
[0037] The accuracy of each model is evaluated based on mean square error, root mean square error, mean absolute error, and coefficient of determination, and the optimal simulation model for each department is selected.
[0038] In an optional implementation, after training and selecting the optimal simulation model for each department, the method further includes:
[0039] Cross-regional modeling validation was conducted by selecting comparative regions with different socio-economic and climatic characteristics to verify the regional transferability and application robustness of the model.
[0040] The Shapley additive interpretation (SHAP value) method was used to conduct interpretability analysis on the optimal model and to evaluate the degree of influence of each input variable on the prediction results.
[0041] By combining kernel density estimation and Monte Carlo simulation to generate asymmetric confidence intervals, the uncertainty of the calculation results is quantitatively analyzed.
[0042] Obtain multiple types of water supply datasets that match the coverage and time range of future scenarios, and perform spatial upscaling and temporal standardization on the water supply datasets to form a standardized water supply input dataset that corresponds one-to-one with the water resource demand prediction scenario.
[0043] Using the median of the water demand confidence interval as the demand-side input and standardized water supply data as the supply-side input, the annual supply and demand gap under each scenario is calculated.
[0044] Referring to the critical ratio method, and combining the ratio of the supply-demand gap to the water demand, the degree of water shortage is classified.
[0045] With the goals of maximizing economic benefits, ensuring social stability and ecological security, a multi-objective optimization model is constructed to comprehensively weigh the economic efficiency, fairness and ecological constraints of water resource allocation, and to formulate an optimal water resource allocation plan.
[0046] Secondly, this invention provides a socio-economic-ecosystem water resource demand measurement and optimization allocation system, the system comprising:
[0047] The data acquisition module is used to acquire and construct a basic database for regional socio-economic-ecosystem water demand measurement. The basic database includes a historical basic database and a shared socio-economic path-representative concentration path scenario database.
[0048] The data preprocessing module is used to preprocess the multi-source heterogeneous data in the basic database to construct an initial dataset for water resource demand forecasting.
[0049] The correlation analysis module is used to screen key driving factors of water demand in various sectors of life, agriculture, industry, and ecology based on preprocessed historical data.
[0050] The historical simulation module is used to construct a multi-model set, train and select the optimal simulation model for each department based on the key driving factors, and analyze the driving mechanism in conjunction with interpretability analysis.
[0051] The future prediction module is used to calculate differentiated water resource demand under multiple scenarios based on the optimal simulation model.
[0052] The optimization module is used to conduct water supply and demand balance analysis, construct multi-objective optimized water resource allocation schemes, and provide decision support for regional sustainable water resource management.
[0053] This invention provides a method and system for measuring and optimizing the allocation of socio-economic and ecosystem water resource demand. It acquires and constructs a basic database encompassing historical and future multi-scenario data, preprocesses heterogeneous multi-source data, and filters key driving factors for water demand in various sectors, including domestic, agricultural, industrial, and ecological areas. A multi-model ensemble is constructed to train and select the optimal simulation model for each sector, and interpretability analysis is used to analyze the driving mechanisms. Based on the optimal model, differentiated water resource demand measurement under multiple scenarios is achieved. Finally, supply and demand balance analysis is conducted, and a multi-objective optimized water resource allocation scheme is constructed. This scheme, through the integration of sector-specific differentiated modeling, a multi-scenario adaptation framework, and multi-objective optimized allocation, constructs a complete technical system from data preprocessing and water resource demand prediction to optimized allocation. This improves the accuracy and dynamic response capability of water demand measurement, enhances the interpretability and decision support value of prediction results, and provides a scientific basis for refined water resource measurement, planning management, and sustainable development decision-making at the county and district levels.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] (1) Sectoral Differentiation Modeling and County-Level Adaptation: By constructing customized prediction models for the four sectors of life, agriculture, industry and ecology respectively, the differences in water demand driving mechanisms of each sector are fully considered. Through multi-source heterogeneous data processing, scenario data at the global and national scales are adapted to county-level administrative units to solve the spatial matching problem of existing technologies in small-scale regional applications and support refined management and precise allocation of water resources.
[0056] (2) Selection and integration of machine learning algorithms: A model set of 4 traditional models and 8 machine learning algorithms was constructed. The parameters were optimized by cross-validation and grid search. The optimal model of each department was selected. Combined with cross-regional robustness verification, the advantages of machine learning in regional water resource demand prediction were verified, and the transferability and application value of the model were ensured.
[0057] (3) Model interpretability: The SHAP value is introduced to conduct interpretability analysis of the optimal model, revealing the contribution of key driving factors to water demand and nonlinear mechanisms, thereby enhancing the credibility of the prediction results and the reference value for decision-making.
[0058] (4) Multi-scenario coupling drive: Integrating four SSP-RCP scenarios under the framework of the sixth phase of the International Coupled Model Comparison Program (CMIP6), covering multiple development paths from low carbon to high greenhouse gas emissions, providing scenario support for addressing medium- and long-term water resources planning under the current policy background;
[0059] (5) Asymmetric uncertainty quantification: Kernel density estimation is used to fit the error distribution, Monte Carlo simulation is used to generate prediction confidence intervals, and the shortest confidence interval method is used to process the asymmetric error distribution, providing decision-makers with a more accurate basis for risk boundaries;
[0060] (6) Multi-objective optimization configuration and dynamic adaptation: Based on water demand forecasting, water supply data is coupled to construct a multi-objective optimization configuration module that takes into account economic, social and ecological benefits. This can effectively improve the efficiency of regional water resource utilization and configuration benefits, and support scientific water resource management and decision-making. Attached Figure Description
[0061] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0062] Figure 1A flowchart illustrating the method for calculating and optimizing the allocation of socio-economic-ecosystem water resources demand provided in this embodiment of the invention;
[0063] Figure 2 A schematic diagram illustrating water resource demand forecasting and optimal allocation provided in an embodiment of the present invention;
[0064] Figure 3 This is a functional block diagram of the socio-economic-ecosystem water resource demand measurement and optimization allocation system provided in an embodiment of the present invention. Detailed Implementation
[0065] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0066] This invention provides a method for calculating and optimizing the allocation of socio-economic-ecosystem water resource demand. Please refer to [link / reference]. Figure 1 The flowchart below shows a method for calculating and optimizing the allocation of socio-economic-ecosystem water resources demand according to an embodiment of the present invention. The method is as follows:
[0067] Step S1: Construct a historical basic database covering 2000-2023 and a database of 4 SSP-RCP future scenarios.
[0068] Step S2: Perform format unification, spatial scale conversion, missing value imputation, outlier identification and normalization on multi-source heterogeneous data.
[0069] Step S3: Use Pearson correlation analysis to identify the key drivers of water demand in the four sectors of domestic, agricultural, industrial and ecological water use.
[0070] Step S4: Construct a model set of 4 traditional models and 8 machine learning algorithms, and select the optimal prediction model for each department through cross-validation and evaluation by multiple statistical indicators.
[0071] Step S5: Couple the four scenarios under the CMIP6 framework to predict the evolution trend of water demand from 2030 to 2100, and combine Monte Carlo simulation to generate 95% asymmetric confidence intervals for water demand in each sector.
[0072] Step S6: Couple water supply capacity data to conduct supply and demand gap analysis, construct a multi-objective optimization model that takes into account economic, social and ecological benefits, and output the optimal configuration scheme.
[0073] The following will provide a detailed explanation of each step; please refer to [link / reference]. Figure 2.
[0074] The proposed method for calculating and optimizing regional socio-economic-ecosystem water resource demand mainly comprises four modules: a basic database module, a historical simulation module, a future prediction module, and a water resource optimization module, as well as a socio-economic-ecosystem water resource demand calculation and optimization system. The specific technical solution is as follows:
[0075] S1. Construct a basic database for calculating water demand in the regional socio-economic and ecosystem sectors.
[0076] The database comprises a basic database of 15 districts and counties covered by the Heihe River Basin, including a historical basic database and an SSP-RCP scenario database. The historical basic database, based on data availability and model requirements, covers a relatively long period, such as 2000-2023. Specifically, it includes socio-economic data (population, urbanization, GDP, disposable income, etc.); meteorological data (temperature, humidity, precipitation, evapotranspiration, etc.); water resources data (historical water consumption, water supply, runoff, etc.); land use data (cultivated land area, green area, irrigated area, construction land area, etc.); and elevation DEM data. The above data mainly comes from official data such as statistical yearbooks, water resources bulletins, and data from the China Meteorological Administration, Inner Mongolia Autonomous Region, Gansu Province, Qinghai Province, and their respective cities (prefectures), districts (counties), and the China Meteorological Administration. The SSP-RCP scenario database uses data from 2030 to 2100 covering four scenarios under the CMIP6 framework: SSP1-2.6 (low-carbon sustainable development), SSP2-4.5 (moderate development), SSP3-7.0 (regional competition and high emissions), and SSP5-8.5 (fossil fuel-dominated and high emissions). Future updates may extend the data or include scenarios with higher precision, depending on the database's development. Specific data includes socio-economic projections (population, GDP, etc.); meteorological projections (temperature, precipitation, etc.); and land use projections (green space area, arable land area, etc.). The data primarily originates from the CMIP6 website and the National Earth System Science Data Center.
[0077] S2. Data preprocessing: Based on the S1 database, the data is divided into a training set (2000-2017) and a test set (2018-2023) in a 7:3 ratio. The initial dataset for water resource demand prediction is constructed using the training set.
[0078] S21. Unified Format: Vector data (SHP format), raster data (NC / TIFF format), and text / table data (TXT / Excel format) are uniformly converted into a calculable numerical format. Raster data is uniformly converted to TIFF format, and statistical data is uniformly converted to Excel format to facilitate subsequent model calculations.
[0079] S22. Scale Transformation: Primarily applicable to the SSP-RCP scenario database. Using ArcGIS's SpatialAnalyst tool, the 1km resolution SSP-RCP raster scenario data is spatially aggregated and upscaled to the administrative units of the county / district under study, maintaining spatial scale consistency with historical baseline data and ensuring spatial matching. Due to the relatively small geographical area of the county / district studies, further calculations and processing are required for some future scenario data with resolutions below 1km, such as urbanization rates only obtained down to the provincial administrative unit level, and temperature data with a spatial resolution of 0.25° and a temporal resolution of daily / monthly. The specific calculation steps are as follows: S221. Urbanization Rate Calculation:
[0080] (1) Calculate the urbanization gap coefficient between counties and provinces each year using historical data. ;
[0081] =
[0082] (2) Based on the existing data, the county-level stability gap coefficient from 2000 to 2023 can be calculated. ;
[0083]
[0084] (3) Through Multiplying the future provincial urbanization data with the future urbanization rate data of the districts and counties yields the future urbanization rate data of the districts and counties. To avoid extreme values that deviate from the land spatial planning and urbanization patterns, the prediction results are subject to categorized upper limit constraints based on the proportion of agricultural population and urbanization development patterns within the region: the upper limit for fully urbanized municipal districts is 0.95, the upper limit for general municipal districts is 0.90, and the upper limit for counties / agricultural county-level cities is 0.85. The constraint values are based on domestic research on new urbanization, the actual distribution of county-level urbanization levels, and international urbanization convergence patterns, and are maintained consistently in subsequent analyses. For the districts and counties studied in this invention example, due to their location in the arid inland river basin of Northwest China, the overall regional development is relatively backward. Therefore, for highly urbanized areas such as Jiayuguan City and Ganzhou District of Zhangye City in this example, the constraint value is truncated to 0.95, while the constraint value for other districts and counties is 0.85.
[0085] (4) To address the potential interannual abrupt changes in the constrained data, a secondary processing method using local weighted regression scatter smoothing (LOWESS) was employed. This effectively eliminated interannual abrupt changes in the data sequence while preserving the long-term trend and upper limit of the constraints, ensuring that the temporal changes in the urbanization rate conformed to the objective law of Logistic growth and guaranteeing the usability of the data in spatial analysis and scenario simulation.
[0086] S222, Temperature Calculation:
[0087] (1) Aggregate daily near-surface air temperature (tas) data into monthly data;
[0088]
[0089] in, The average monthly temperature of a certain month. For the first time this month Daily temperature, This represents the number of days in the current month.
[0090] Since the official data uses Kelvin as the unit for tas, it needs to be converted to degrees Celsius.
[0091]
[0092] (2) Considering the structural biases inherent in single-model regional climate simulations, this invention, based on the validation results of existing literature (Liu et al., 2022; Song et al., 2024; Xu et al., 2024; Yang et al., 2024) on regional climate simulations in China, selected several applicable models, including BCC-CSM2-MR, FGOALS-g3, INM-CM4-8, INM-CM5-0, MRI-ESM2-0, and EC-Earth3, which exhibit low errors in temperature simulations in China, especially in arid and semi-arid regions. Therefore, the equal-weighted averaging method was used to construct the future climate ensemble results; for any year... ,month The results of the multi-mode set are calculated as follows:
[0093]
[0094] in, This is a result of a multi-mode set; For the first The pattern in the first Year Monthly temperature variable values; The number of patterns in this invention ;
[0095] (3) Using 1982-2024 as the historical baseline period, and combining the existing high-resolution historical data and meteorological station data, we will conduct multi-year averages of 1 km high-resolution temperature and precipitation for each month from 1982 to 2024 to construct 12 historical temperature baseline maps.
[0096]
[0097] in, Indicates the historical baseline period Monthly temperature data per 1 km; The base period is the number of years. It is the first Year Monthly temperature value;
[0098] (4) Similarly, the multi-model ensemble results were averaged over many years from 1982 to 2024 to obtain monthly data with a resolution of 0.25°:
[0099]
[0100] in, Representation of the historical baseline period of the pattern Monthly temperature data at 0.25 degrees Celsius; It is the model baseline period Year Monthly temperature value;
[0101] (5) The additive Delta method is used for downscaling, and finally a future temperature dataset with a resolution of 1km is obtained; for the future... Year In [month], calculate the outliers of the multi-model ensemble temperature relative to the baseline monthly mean:
[0102]
[0103] in, For the future Year The temperature anomaly field for the month, with a resolution of 0.25°.
[0104] Subsequently, the temperature anomaly field was resampled to 1 km using bilinear interpolation and superimposed onto the high-resolution historical baseline monthly climatology of the corresponding month to obtain the future 1 km monthly mean temperature:
[0105]
[0106] in, For the future Year Temperature at 1 km in the month; To resample the temperature anomaly field to 1 km;
[0107] (6) Calculate the monthly data into yearly data using Python software;
[0108] (7) Use the ArcGIS software Spatial Analyst tool to spatially aggregate the 1km resolution temperature data into the district / county administrative unit for subsequent calculations.
[0109] S23. Missing Value Imputation: For missing values in time series data, such as historical water consumption data, the ipolate command of the Stata software is called to perform linear interpolation, and the missing values are fitted using the values of adjacent time points.
[0110] S24. Outlier Identification: Outlier identification is performed using the interquartile range (IQR) method. The calculation formula is as follows:
[0111]
[0112] in, It is the lower quartile, or 25%. It is the upper quartile, or 75%. This represents the lower bound of the minimum value of the data. Represents the upper bound of the maximum value of the data;
[0113] The identified outliers may be due to statistical errors or specific events at a particular point in time. The study uses the median to replace all outliers.
[0114] S25. Data Normalization: To reduce the model's sensitivity to data, eliminate the dimensional differences between different indicators (e.g., precipitation is in mm, GDP is in 100 million yuan), and improve the model's convergence speed and measurement accuracy, a feature scaling method is used to normalize all data.
[0115]
[0116] in, For the first The original values of the explanatory variables. Let it be its maximum and minimum values.
[0117] S3. Correlation calculation to screen key driving factors of water demand in each department.
[0118] Based on preprocessed historical data and the driving mechanisms of water demand in each sector, Pearson correlation analysis was used to screen the key driving factors of water demand in each sector. The specific steps are as follows:
[0119] S31. Construct a preliminary set of factors influencing water demand in various sectors: Taking into account economic development, social conditions, natural resources and meteorological environment factors, and combining existing research with regional realities, construct preliminary sets of factors influencing water demand for domestic, agricultural, industrial and ecological purposes respectively.
[0120] S32. Pearson Correlation Analysis Calculation: Input all data from the initial set of influencing factors and the water consumption of each sector into the Pearson correlation coefficient calculation formula to obtain the correlation coefficient between each influencing factor and the water demand of each sector. .
[0121]
[0122] S33. Key Driving Factor Screening: In this invention, correlation coefficients were calculated for 24-year time series data from 15 different districts and counties. The degrees of freedom for each district and county were then determined. Consult the Pearson correlation coefficient critical value table to find the significance level. (Two-tailed test) Filter by absolute value of correlation coefficient, keeping one decimal place. Influence factors ≥ 0.5 were selected as key drivers of water demand in each sector, ensuring a strong linear correlation between the selected factors and water volume in each sector. The following key drivers for each sector were ultimately determined, providing fundamental data support for subsequent water resource demand forecasting and modeling.
[0123]
[0124] S4. Historical water consumption modeling, training and selecting the optimal simulation model for each department.
[0125] A multi-model ensemble encompassing four traditional models and eight machine learning algorithms was constructed. The models were trained, tested, validated, and selected based on the varying water demand characteristics of different departments. Interpretive analysis was also introduced to elucidate the internal mechanisms of machine learning. The specific steps are as follows:
[0126] S41. Construct a set of 12 models in 2 categories, including 4 traditional statistical models covering Linear Regression, Ridge Regression, Lasso Regression with Least Absolute Shrinkage and Lasso, and Autoregressive Integral Moving Average (ARIMA), and 8 machine learning algorithms including Support Vector Machine (SVM), Decision Tree (DT), K-Nearest Neighbors (KNN), Random Forest (RF), Gradient Boosting Machine (LightGBM), Extreme Boosting Machine (XGboost), Backpropagation Neural Network (BPNN), and Grey Wolf-Optimized Backpropagation Neural Network (GWO-BP). In particular, the Grey Wolf-Optimized Backpropagation Neural Network model is introduced. Its advantage lies in the fact that the Grey Wolf optimization algorithm is inspired by the hunting behavior of grey wolf packs, simulating the leadership hierarchy and hunting mechanism of grey wolves in nature. Four types of grey wolves, such as α, β, δ, and ω, are used to simulate the leadership hierarchy, and the three main steps of hunting are also implemented: finding prey, surrounding prey, and attacking prey. It features a simple structure, few adjustable parameters, and ease of implementation, achieving a balance between local optimization and global search. Therefore, it exhibits good performance in terms of solution accuracy and convergence speed, and can effectively handle the nonlinear and high-dimensional characteristics in water resource demand forecasting. In this invention example, the hidden layer node number is optimized using a grid search method, resulting in a hidden structure of (127, 60) for the BP neural network, and the final calculation result is obtained after 479 iterations.
[0127] S42. Model Training and Testing: The key driving factors of each department selected in S33 are used as the model input variables, and the historical water consumption of each department is used as the model output. The model parameters are optimized using the 5-fold time-series cross-validation method to avoid model overfitting. That is, the training set is randomly divided into 5 different subsets, each subset is called a fold. The algorithm model is used for 5 training and evaluation cycles. Each time, 4 folds are selected for training and the remaining 1 fold is evaluated. The average score and standard deviation of the prediction are calculated using the 5 scores output by cross-validation. Finally, the model accuracy is verified using the training set and the test set respectively.
[0128] S43. Optimal Model Selection: Each model is evaluated for accuracy using different statistical indicators, and these are compared to select the best model. These indicators include mean squared error (MSE), root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²). 2 Statistical indicators. First, use MSE, RMSE, and MAE to assess the degree of deviation between the actual and predicted values;
[0129]
[0130]
[0131]
[0132] in, Represents the true value of the explained variable. The values represent predicted values, and n represents the sample size. MSE (Mean Sequence Error) is the squared error, as its squared term is based on a larger penalty for larger errors, but it is also more sensitive to outliers. RMSE restores the dimensions consistent with the explanatory variables, making it easier to understand and interpret, but it is also sensitive to outliers in the data and needs to be compared with the magnitude of the true values. MAE addresses the sensitivity to outliers by using absolute values, but it is difficult to reflect the distribution of prediction errors. The closer these metrics are to 0, the smaller the error and the better the prediction performance. Additionally, R... 2 Evaluate the goodness of fit of the predicted straight line or curve to the data;
[0133]
[0134] in, This represents the mean of the explained variable. The closer its value is to 1, the higher the goodness of fit, and the better the predictive model can interpret the data.
[0135] S44, with lower error and higher To determine the optimal algorithm for each department, using historical water consumption simulation results of the social system as an example, the specific calculation indicators are as follows:
[0136]
[0137] Based on the calculation results, the optimal algorithm for the social system is determined to be Random Forest (RF), which has the smallest error index. The accuracy is improved by an average of 40% compared with traditional statistical methods, providing basic model support for predicting water resource demand under different future scenarios, and finally calculating the water resource demand of the entire regional system.
[0138] S45. To verify the regional transferability of the model, this invention further includes: selecting at least two comparison regions that differ from the modeling region in terms of socioeconomic level and climate conditions; using the same method as in S42-S43, i.e., retraining and evaluating all models in the multi-model set based on historical data from the comparison regions; and recording the performance metrics of each model on the test set of the comparison regions. This verification result can be used to evaluate the applicability of the models selected by this invention in different regions.
[0139] S46. Model interpretability analysis: The Python SHAP library is used to calculate the SHAP value for each input feature of the trained optimal model (such as a random forest model obtained through social system screening). This method is based on Shapley values in cooperative game theory, evaluating the importance of each input variable by calculating the marginal contribution of features when added to the model. The specific calculation formula is as follows: A SHAP summary plot is drawn to show the direction and magnitude of the influence of each feature on the model output; A SHAP dependency plot is drawn to analyze the changing trend of the SHAP value of a single feature within its range, thereby identifying the linear and nonlinear relationships between the feature and water demand.
[0140]
[0141] A larger absolute average SHAP value indicates a greater impact of the feature on the model's prediction results. The linear and nonlinear interaction effects between each feature and water demand are analyzed using a SHAP dependency graph: if the SHAP value changes monotonically with the feature value, a linear relationship exists; if the SHAP value fluctuates or changes non-monotonicly within the feature value range, a nonlinear mechanism exists. In the example of this invention, taking the calculation of the SHAP value of population size on the water demand of the social system as an example, a basically linear positive driving effect is observed, while the temperature factor has a nonlinear effect on ecological water demand (increasing first and then decreasing), verifying the necessity of using a nonlinear machine learning model in this invention.
[0142] S47. Through multi-model comparison and screening in S4, the optimal model for water demand in each sector is determined, and the historical simulation module is constructed to provide basic model support for future water resource demand forecasting.
[0143] S5. Future water demand forecasting, combined with SSP-RCP scenario to address future uncertainties.
[0144] Based on the optimal models for each sector obtained through S4 screening, four SSP-RCP scenarios (SSP1-2.6, SSP2-4.5, SSP3-7.0, and SSP5-8.5) from CMIP6 are coupled to cover different development paths from low-carbon to high-carbon, predicting the evolution trend of water demand under multiple scenarios from 2030 to 2100. The specific steps are as follows:
[0145] S51. Scenario data input: Input the future calculation data of key driving factors of various sectors such as socio-economic, meteorological, and land use, which were preprocessed in step S3, into the optimal model of the corresponding sector.
[0146] S52. Calculation of future water demand: By modeling, the water demand for domestic, agricultural, industrial and ecological purposes from 2030 to 2100 under four scenarios is obtained. The total amount of social, economic and ecological water demand can be obtained year by year in the future.
[0147]
[0148] S53. Uncertainty analysis: A method combining kernel density estimation (KDE) and Monte Carlo simulation is used to quantify the error range of the measurement results and generate asymmetric confidence intervals. The specific steps are as follows:
[0149] S531. Based on the simulation prediction results of the historical validation set (2018-2023), calculate the relative prediction error of water demand for each sector:
[0150]
[0151] in, These are the model's predicted values. This represents the actual water consumption. Statistical analysis was performed on the error sequences of each department, and kernel density estimation was used to estimate the probability density function of the fitted error, avoiding prior assumptions about the error distribution pattern.
[0152]
[0153] in, For Gaussian kernel function, The bandwidth parameter is optimized using leave-one-out cross-validation.
[0154] S532. For the water resource demand prediction models of each department, set the input parameters for the Monte Carlo simulation respectively:
[0155] (1) Number of simulations: N=10000 times to ensure a balance between statistical robustness of simulation results and computational efficiency;
[0156] (2) Input variable distribution: The error probability distribution obtained in S531 is used as the input distribution of the random disturbance term, and random disturbances are applied to the baseline forecast values of each department:
[0157]
[0158] in, For the first In the simulation, the first Annual water demand The baseline model prediction is used. The relative error is randomly sampled from the error distribution;
[0159] (3) Scenario coupling: Monte Carlo simulations were performed independently for the four SSP-RCP scenarios to reflect the range of uncertainty under different development paths.
[0160] S533. Perform statistical analysis on the N simulation results under different scenarios for each department, and generate asymmetric confidence intervals:
[0161] (1) Calculate the p-th percentile as the upper and lower bounds of the confidence interval, and set the confidence level α = 95%; the lower bound of the interval = upper bound of the interval =
[0162] (2) The shortest confidence interval method is used to handle asymmetric error distributions, and the inverse cumulative distribution function is used to solve for the error distribution that satisfies the condition. And the interval width is the smallest [ , ];
[0163] (3) Output a 95% confidence interval for the water demand of each department, and calculate the ratio of the predicted interval width as a relative uncertainty indicator.
[0164] S534. Sum the confidence intervals of each department according to step S52 to obtain the confidence interval of the total water demand:
[0165]
[0166] S535. The reliability of the uncertainty quantification results is evaluated using two metrics: Prediction Interval Coverage Probability (PICP) and Prediction Interval Normalized Average Width (PINAW).
[0167]
[0168]
[0169] in, To determine the number of samples in the validation set, The actual value range is used to ensure that PICP is close to the set confidence level (95%) and PINAW is within an acceptable range through validation set testing.
[0170] S536. Output the annual water demand forecasts and 95% confidence intervals for each scenario from 2030 to 2100, as the risk boundary basis for water resource planning decisions.
[0171] S54. Verify the calculation results of S5 in conjunction with the regional local government planning documents, including checking the consistency between the predicted trend and the planning target, analyzing the deviation between the predicted value and the planning target value in 2030, and the coverage of the confidence interval to the planning target value.
[0172] S6. Based on the future water demand forecast in S5, this module conducts a water supply and demand balance analysis and constructs a water resource allocation scheme based on multi-objective optimization to provide decision support for regional sustainable water resource management. The specific steps are as follows:
[0173] S61. Obtain publicly available water supply capacity product datasets officially released, matching the time coverage (2030-2100) and future scenario coverage (SSP1-2.6, SSP2-4.5, SSP3-7.0, SSP5-8.5), covering water supply potential data for various water sources such as surface water, groundwater, reclaimed water, externally diverted water, and unconventional water. Perform spatial upscaling and temporal standardization on the dataset. First, use the ArcGIS software Spatial Analyst tool to upscale the high-precision data to the county-level administrative unit. At the same time, convert the obtained daily / monthly water supply data into yearly sequences using Python software to ensure complete consistency with the spatiotemporal resolution and scenario system of the water demand forecast results, forming a standardized water supply capacity input dataset that corresponds one-to-one with the water resource demand forecast scenarios.
[0174] S62. Supply and demand gap analysis: Using the confidence intervals of water demand for each sector output from S5 as the demand-side input and the water supply capacity data obtained from S61 as the supply-side input, calculate the annual supply and demand gap under each scenario:
[0175]
[0176] in, For the supply and demand gap in year t and scenario s, a positive value indicates a surplus and a negative value indicates a shortage. For the water supply capacity in year t and scenario s, Let be the total water demand in year t and scenario s. Here, the total water demand is represented by the median of the water demand confidence interval output by S5, ensuring that the supply and demand analysis takes into account both uncertainty and decision robustness.
[0177] S63. Water scarcity severity classification, referencing the withdrawal-to-availability ratio method adopted by the Food and Agriculture Organization of the United Nations (FAO) AQUASTAT, combined with the ratio of supply-demand gap to water demand. The degree of water shortage is classified into the following levels:
[0178]
[0179] S64. With the goals of maximizing economic benefits, ensuring social stability, and protecting the ecological environment, a multi-objective optimization model is constructed to comprehensively weigh the economic efficiency, fairness, and ecological constraints of water resource allocation. The specific steps are as follows:
[0180] S641. Economic Benefit Objective Function: Maximizing economic benefits refers to maximizing water production benefits, that is, maximizing the net water supply benefit value of each water-using sector in each planning year, which serves as the economic benefit objective function.
[0181]
[0182] in, The water supply from water source i to water user h, in billions of m³ 3 ; The water supply efficiency coefficient from water source i to water user h, in billions of yuan; Priority coefficient for water-using departments; The priority water supply coefficient for water source i;
[0183] In the formula, The calculation is performed using a linearly decreasing weight allocation method, and the specific steps are as follows:
[0184] The priority water use (water supply) coefficient refers to the degree of importance of a user's access to water supply relative to other water-using sectors, while the priority water supply coefficient refers to the degree of importance of a water source's access to water relative to other water sources. According to Article 21 of the *Water Law of the People's Republic of China (2016 Amendment)*: The development and utilization of water resources shall first meet the domestic water needs of urban and rural residents, while also taking into account the needs of the ecological environment, agriculture, industry, water use, and navigation. In particular, the development and utilization of water resources in semi-arid and arid regions shall fully consider the water needs of the ecological environment.
[0185]
[0186] in, This is the priority water use (water supply) coefficient; Let j be the water usage sequence number of the water-using department (water supply department); This represents the maximum value of the serial number of the water-using (water supply) department.
[0187] S642. Social Benefit Objective Function: In the process of social development, with the goal of maintaining social stability, the objective function for regional social benefits is to minimize the total water shortage.
[0188]
[0189] in, The total water demand of water-using sectors is 100 million m³. 3 .
[0190] S643. Ecological and Environmental Benefit Objective Function: Actively responding to the concept of ecological and environmental protection, the objective function for regional ecological and environmental benefits is to minimize water shortage.
[0191]
[0192] in, For the ecological sector's water needs, 100 million cubic meters 3 .
[0193] S65. Constraints mainly include total water demand constraint, total water supply constraint, and non-negativity constraint. The total water demand constraint means that the water supply from each water source to each water user should not exceed the water demand of that user. The total water supply constraint means that the water allocated to each calculation zone should not exceed the total water supply of the water supply departments in that calculation zone. The non-negativity constraint means that the water allocated to each water user in each calculation zone should be greater than 0.
[0194] S651, Total Water Demand Constraints
[0195]
[0196] in, The water supply from each water source to department h, in billions of m³. 3 ; For the water demand of water-using sectors, 100 million m³ 3 .
[0197] S652, Total Water Supply Constraints
[0198]
[0199] in, The total water supply from each water source to each department, in billions of cubic meters. 3 ; The total available water volume is 100 million m³.3 .
[0200] S653, Nonnegativity Constraint
[0201]
[0202] S66. Solving the water resource optimization allocation model: The Non-Dominated Sorting Genetic Algorithm (NSGA-III) based on reference points is used to solve the multi-objective optimization model. NSGA-III introduces an adaptive reference point mechanism based on the NSGA-II framework, which can effectively handle multi-objective optimization problems in three-dimensional and higher-dimensional objective spaces, overcoming the deficiency of insufficient selection pressure in high-dimensional objective spaces in NSGA-II. The specific steps are as follows:
[0203] S661. Set parameters and initialize the population. Based on the recommended values from relevant literature on water resource optimization, the parameters are corrected and verified according to the dimensional characteristics of the decision variables in this model. The final population size is set to 200, which can be dynamically adjusted according to the number of water-using departments in the district / county. In this example, it is set to the number of water-using departments to ensure that the population size matches the number of reference points and maintains the diversity of the solution set. The number of generations is set to 500 to ensure a reasonable balance between convergence accuracy and computational cost. The crossover probability is set to 0.9 to enhance the global search capability and promote the transmission of excellent gene structures. The mutation probability is set to 1 / 20 according to the empirical formula to maintain moderate local perturbation and avoid premature convergence. The initial population is generated using Latin hypercube sampling.
[0204] S662. Calculate the values of the three objective functions: economic benefits, social benefits, and ecological benefits.
[0205] S663. Perform fast non-dominated sorting on the individuals in the population, divide the Pareto front, and associate the individuals with reference points generated in the three-dimensional target space; select individuals to enter the next generation based on the non-dominated level and the reference point association density; perform simulated binary crossover and polynomial mutation operations on the selected individuals to generate the offspring population.
[0206] S664. If the current iteration reaches the maximum iteration number of 500, or the rate of change of the Pareto front hypervolume is less than 0.1% for 50 consecutive iterations, then terminate the iteration and output the Pareto front solution set; otherwise, return to step S662 to continue the iteration.
[0207] S665. The TOPSIS method based on ideal points is used to select the comprehensive optimal compromise scheme from the Pareto front solution set: First, the Euclidean distance between each scheme and the positive ideal solution and the negative ideal solution is calculated, and then the relative proximity is calculated. The scheme with the largest proximity is selected as the recommended configuration scheme.
[0208] S67. For different water supply guarantee rates (50%, 75%) and different SSP-RCP scenarios (SSP1-2.6, SSP2-4.5, SSP3-7.0, SSP5-8.5), execute the above solution process independently.
[0209] S68. Integrate the water resource optimization allocation schemes under different scenarios and different water supply guarantee rates to form a dynamic water resource optimization allocation scheme set. Output the future annual water allocation for domestic, agricultural, industrial, and ecological use, supply and demand balance results, water shortage level, and comprehensive benefit indicators under different scenarios. Provide quantitative decision-making basis for long-term planning, dynamic regulation, and refined management of water resources in the study area. The allocation schemes obtained in this module and the water resource demand forecast results obtained in S5 are used together as core data and connected to the subsequent socio-economic-ecosystem water resource demand calculation and optimization allocation system for unified display, query, and interactive application.
[0210] S7. Based on the inventive concept and calculation process of S1-S6 above, the present invention also provides a socio-economic-ecosystem water resource demand measurement and optimization allocation system. This system focuses on the needs of water resource shortage and refined management, adopts a front-end and back-end separation architecture, and realizes full-process integration from the data layer, service layer to the presentation layer, providing scientific support and technical tools for regional water resource management and planning.
[0211] S71. System Overall Architecture: This system adopts a layered design architecture, including a three-layer structure: data layer, service layer, and presentation layer. Data interaction and collaborative operation between these layers are achieved through standardized interfaces.
[0212] Data layer: Provides multi-source heterogeneous data support for the system, including historical basic data, future scenario data, model calculation result data, etc., to realize unified access, storage and management of data;
[0213] Service layer: Encapsulates core functions such as data processing, model calculation, and business logic, including data preprocessing module, correlation analysis module, historical simulation module, future prediction module, and optimization configuration module;
[0214] Presentation layer: Provides an interactive visual interface to enable data query, model operation, result display and export, and other functions, providing users with an intuitive and convenient operating experience.
[0215] S72. Technical Implementation Scheme: This system adopts a front-end and back-end separation architecture, and the specific implementation method is as follows:
[0216] Backend framework: Spring Boot is used to build the backend service, which implements functions such as business logic processing, data interface service, and model calculation scheduling;
[0217] Front-end interactive interface: The front-end interactive interface is built based on Vue3 and TypeScript, realizing interactive operations such as user login, function menu, parameter configuration, and result display;
[0218] Spatial Analysis Engine: Integrates Cesium 3D Geographic Information Engine with ArcGIS spatial analysis tools to enable visualization, spatial overlay analysis, and scale conversion of regional spatial data;
[0219] Scientific computing engine: Combining the Python scientific computing ecosystem, it enables complex computational tasks such as multi-model training, water resource demand forecasting, Monte Carlo simulation, and multi-objective optimization.
[0220] S73. Core Module Functions: This system can be divided into functional modules based on the methods described in S1-S6 above for the socio-economic-ecosystem water resource demand measurement and optimization allocation system. Each function can be divided into its own module, or two or more functions can be integrated into a single processing module. These integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this invention is illustrative and represents only one logical functional division; other division methods may be used in actual implementation.
[0221] For example, when dividing functional modules according to their respective functions, Figure 3 The socio-economic-ecosystem water resource demand measurement and optimization system shown is only a schematic diagram of a device. This system may include a data acquisition module, a data preprocessing module, a correlation analysis module, a historical simulation module, a future prediction module, and an optimization allocation module. The functions of each module of this socio-economic-ecosystem water resource demand measurement and optimization system will be described in detail below.
[0222] The data acquisition module constructs a historical basic database covering 2000-2023 and a database of four SSP-RCP future scenarios;
[0223] The data preprocessing module performs format unification, spatial scale conversion, missing value imputation, outlier identification, and normalization on multi-source heterogeneous data.
[0224] The correlation analysis module uses Pearson correlation analysis to identify key driving factors of water demand in the four sectors of domestic, agricultural, industrial, and ecological water use.
[0225] The historical simulation module constructs a model set of 4 traditional models and 8 machine learning algorithms, and selects the optimal prediction model for each department through cross-validation and evaluation by multiple statistical indicators.
[0226] The future forecasting module couples four scenarios under the CMIP6 framework to predict the evolution trend of water demand from 2030 to 2100, and combines Monte Carlo simulation to generate 95% asymmetric confidence intervals for water demand in each sector.
[0227] The optimization module couples water supply capacity data to conduct supply and demand gap analysis, and constructs a multi-objective optimization model that takes into account economic, social and ecological benefits to output the optimal configuration scheme.
[0228] The socio-economic-ecosystem water resource demand calculation and optimization allocation system provided by the present invention can be used to perform socio-economic-ecosystem water resource demand calculation and optimization allocation under any of the above-described embodiments.
[0229] Comparison documents
[0230] [1]Mu, L., Zheng, F., Tao, R., Zhang, Q., & Kapelan, Z. (2020). Hourly and Daily Urban Water Demand Predictions Using a Long Short-TermMemory Based Model. Journal of Water Resources Planning and Management, 146(9), 05020017. https: / / doi.org / 10.1061 / (asce)wr.1943-5452.0001276.
[0231] [2]Pichler, M., & Hartig, F. (2023). Machine learning and deeplearning—A review for ecologists. Methods in Ecology and Evolution, 14(4). https: / / doi.org / 10.1111 / 2041-210x.14061.
[0232] [3] Zhu Wanyi, Zhang Zhenke, Guo Xinya, et al. Characteristics and estimation of vegetation ecological water demand in the Mara River Basin [J]. Acta Ecologica Sinica, 2023, 43(18):7523-7535. DOI:10.20103 / jstxb.202209182672.
[0233] [4] Tian Jintao, Zuo Qiting, Bayinji, et al. Analysis and prediction of economic, social and ecological water balance in Qinhe River Basin [J]. Advances in Water Resources and Hydropower Science and Technology, 2025, 45(02):9-16+45.
[0234] [5]Liu F, Xu C, Long Y et al, 2022. Assessment of CMIP6 ModelPerformance for Air Temperature in the Arid Region of Northwest China andSubregions. Atmosphere, 13(3):454.
[0235] [6]Song X, Xu M, Kang S et al, 2024. Evaluation and projection ofchanges in temperature and precipitation over Northwest China based on CMIP6models. International Journal of Climatology, 44(14):5039-5056.
[0236] [7]Xu M, Hou Z, Kang S et al, 2024. Projection of snowfall andprecipitation phase changes over the Northwest China based on CMIP6multimodels. Journal of Hydrology, 641:131743.
[0237] [8]Yang B, Wei L, Tang H et al, 2024. Future changes in extremesacross China based on NEX-GDDP-CMIP6 models. Climate Dynamics, 62(10):9587-96。
Claims
1. A method for measuring and optimizing the allocation of socio-economic-ecosystem water resource demand, characterized in that, The method includes: Construct a basic database for regional socio-economic-ecosystem water demand measurement, which includes a historical basic database and a shared socio-economic pathway-representative concentration pathway future scenario database. The multi-source heterogeneous data in the basic database are preprocessed to construct an initial dataset for water resource demand prediction; Based on preprocessed historical data, key driving factors of water demand in various sectors of life, agriculture, industry, and ecology are screened. A multi-model set including traditional statistical models and machine learning algorithms is constructed. Based on the key driving factors, the optimal simulation models of each department are trained and screened. The Shapley additive interpretation method is used to conduct interpretability analysis on the optimal models and analyze the mechanism of action of each driving factor. Based on the optimal simulation model, differentiated water resource demand calculations under multiple scenarios are realized. By combining kernel density estimation and Monte Carlo simulation to generate asymmetric confidence intervals, the uncertainty of the calculation results is quantified. Conduct water supply and demand balance analysis, construct multi-objective optimized water resource allocation schemes, and provide decision support for regional sustainable water resource management.
2. The method for calculating and optimizing the allocation of socio-economic-ecosystem water resource demand according to claim 1, characterized in that, The preprocessing of the multi-source heterogeneous data in the basic database includes: Standardized format: Convert vector data, raster data, and text / table data into a standardized, computable format; Scale transformation: Spatially aggregate and scale up the raster data of the shared socio-economic path-representative concentration path scenario database to match the spatial scale of district and county administrative units; Missing value imputation: For missing values in time series data, linear interpolation is used to impute them; Outlier identification: Outlier identification and replacement are performed using the interquartile range method; Data normalization: All data are normalized using the feature scaling method to eliminate the dimensional differences between different indicators.
3. The method for calculating and optimizing the allocation of socio-economic-ecosystem water resource demand according to claim 2, characterized in that, The scaling transformation includes scaling transformation of urbanization rate data, specifically including: The urbanization gap coefficient between counties and the province is calculated based on historical data to obtain the multi-year average stable gap coefficient. Based on the stable gap coefficient and future provincial urbanization data, the future urbanization rate of counties is calculated. The calculated future urbanization rate of counties is subject to categorical upper limit constraints. Finally, a local weighted regression scatter smoothing method is used to smooth the data and eliminate inter-year abrupt changes.
4. The method for calculating and optimizing the allocation of socio-economic-ecosystem water resource demand according to claim 2, characterized in that, The scaling transformation includes scaling of temperature data, specifically including: Daily near-surface temperature data were aggregated into monthly data, and the Kelvin units of near-surface temperature were converted to degrees Celsius. Six applicable global climate models were selected based on literature, and the future climate ensemble results were constructed using the equal-weighted averaging method. Using 1982-2024 as the historical baseline period, monthly multi-year averages were performed on the 1km high-resolution temperature and the multi-model ensemble results to construct a historical temperature baseline map and model historical baseline monthly data. Based on the calculation of the temperature anomaly field between the baseline period observation data and the future simulation data, the resolution of the anomaly field was improved to the target resolution through a spatial interpolation resampling algorithm and superimposed onto the historical baseline climate state raster surface to generate a high-resolution future temperature prediction raster map. Spatial regional statistical operations were performed on the raster map to extract the average temperature sequence corresponding to the administrative boundaries of districts and counties.
5. The method for calculating and optimizing the allocation of socio-economic-ecosystem water resource demand according to claim 1, characterized in that, Based on preprocessed historical data, key driving factors for water demand in various sectors, including domestic, agricultural, industrial, and ecological areas, are screened, including: Construct an initial set of influencing factors on water demand for each department; Pearson correlation analysis was used to calculate the correlation coefficients between each influencing factor and the water consumption of the corresponding department. Factors with correlation coefficients not less than a specific threshold are selected as key drivers of water demand for the corresponding departments.
6. The method for calculating and optimizing the allocation of socio-economic-ecosystem water resource demand according to claim 1, characterized in that, The construction of the multi-model set includes: A multi-model set, including traditional statistical models and machine learning models, is constructed to adapt to the water use simulation needs of different departments and to achieve multi-model comparison and selection.
7. The method for calculating and optimizing the allocation of socio-economic-ecosystem water resource demand according to claim 1, characterized in that, The process of training and selecting the optimal simulation model for each department based on the key driving factors includes: The key driving factors of each department are used as model inputs, and the historical water consumption of each department is used as model outputs. The models are divided into training and testing sets according to the proportions. Temporal cross-validation was used to optimize the model parameters to avoid overfitting. The accuracy of each model is evaluated based on mean square error, root mean square error, mean absolute error, and coefficient of determination, and the optimal simulation model for each department is selected.
8. The method for calculating and optimizing the allocation of socio-economic-ecosystem water resource demand according to claim 1, characterized in that, After training and selecting the optimal simulation model for each department, the process also includes: Cross-regional modeling validation was conducted by selecting comparative regions with different socio-economic and climatic characteristics to verify the regional transferability and application robustness of the model. The Shapley additive interpretation (SHAP value) method was used to conduct interpretability analysis on the optimal model and to evaluate the degree of influence of each input variable on the prediction results. By combining kernel density estimation and Monte Carlo simulation to generate asymmetric confidence intervals, the uncertainty of the calculation results is quantitatively analyzed. The water resource demand forecast results were verified by combining regional local government planning documents.
9. The method for calculating and optimizing the allocation of socio-economic-ecosystem water resource demand according to claim 1, characterized in that, The aforementioned water resource supply and demand balance analysis and the construction of a multi-objective optimized water resource allocation scheme specifically include: Obtain multiple types of water supply datasets that match the coverage and time range of future scenarios, and perform spatial upscaling and temporal standardization on the water supply datasets to form a standardized water supply input dataset that corresponds one-to-one with the water resource demand prediction scenario. Using the median of the water demand confidence interval as the demand-side input and standardized water supply data as the supply-side input, the annual supply and demand gap under each scenario is calculated. Referring to the critical ratio method, and combining the ratio of the supply-demand gap to the water demand, the degree of water shortage is classified. With the goals of maximizing economic benefits, ensuring social stability and ecological security, a multi-objective optimization model is constructed to comprehensively weigh the economic efficiency, fairness and ecological constraints of water resource allocation, and to formulate an optimal water resource allocation plan.
10. A socio-economic-ecosystem water resource demand measurement and optimization allocation system, characterized in that, The system includes: The data acquisition module is used to acquire and construct a basic database for regional socio-economic-ecosystem water demand measurement. The basic database includes a historical basic database and a shared socio-economic path-representative concentration path scenario database. The data preprocessing module is used to preprocess the multi-source heterogeneous data in the basic database to construct an initial dataset for water resource demand forecasting. The correlation analysis module is used to screen key driving factors of water demand in various sectors of life, agriculture, industry, and ecology based on preprocessed historical data. The historical simulation module is used to construct a multi-model set, train and select the optimal simulation model for each department based on the key driving factors, and analyze the driving mechanism in conjunction with interpretability analysis. The future prediction module is used to calculate differentiated water resource demand under multiple scenarios based on the optimal simulation model. The optimization module is used to conduct water supply and demand balance analysis, construct multi-objective optimized water resource allocation schemes, and provide decision support for regional sustainable water resource management.