Rural drought-caused drinking difficulty risk dynamic assessment and early warning method

Through multi-source data processing and model building, the passive response to the risk of drinking water difficulties in rural areas due to drought has been solved, and proactive early warning and precise intervention have been achieved. This has improved the timeliness and accuracy of risk identification and ensured the drinking water safety of rural residents.

CN120875129AActive Publication Date: 2025-10-31CHINA INST OF WATER RESOURCES & HYDROPOWER RES
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510912424.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-31
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

In the current technology, it is difficult to make early warning and timely intervention for the problem of drinking water difficulties caused by drought in rural areas. Traditional disaster statistics have problems of lag and inconsistent data. There is a lack of specialized risk assessment index and monitoring system, which makes it difficult to make drought relief and water supply decisions without scientific basis and to achieve proactive early warning and precise intervention.

Method used

By acquiring and preprocessing multi-source data, physical constraints are determined, data dimensionality reduction and variable structure analysis are performed, linear, nonlinear and machine learning models are constructed, and model parameters are calibrated with historical data to achieve precise identification and early warning of risk periods, forming a full-process early warning system.

Benefits of technology

This has enabled a shift from passive response to proactive early warning, improved the timeliness and accuracy of risk identification, provided a scientific basis for drought relief and water supply decisions, and enhanced the drinking water safety of rural residents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875129A_ABST
    Figure CN120875129A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic assessment and early warning method for the drinking difficulty risk of people due to drought in rural areas. The method comprises the steps of multi-source data acquisition and preprocessing, physical constraint relation determination, data dimension reduction and variable structure analysis, risk assessment model construction, model parameter calibration and risk period determination. The method comprises the following steps: carrying out dimensionality reduction processing on a collinear variable based on a principal component analysis method, constructing three risk assessment models including a linear model, a nonlinear model and a machine learning model in parallel, screening an optimal model by utilizing a dual-objective function, optimizing and determining a risk threshold value, and identifying and combining high-risk time periods by applying a run-length theory. The core innovation point of the method is that the daily scale risk model is calibrated by using the annual scale reported risk population data, and the problem that the historical data granularity is not matched with the early warning demand is effectively solved. By establishing a variable processing system and multi-model integrated assessment under physical constraint conditions, the timeliness and accuracy of risk assessment are significantly improved, and the transformation from passive disaster statistics to active risk early warning is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of earth science and disaster risk management, and specifically relates to a method for dynamic assessment and early warning of the risk of rural residents facing drinking water difficulties due to drought. Background Technology

[0002] In rural areas, the drinking water shortage caused by drought has seriously affected residents' basic living conditions and production activities.

[0003] Currently, the assessment of rural residents' drinking water shortages due to drought relies primarily on disaster statistics and reporting at the county and township levels—a typical passive response mechanism. This approach suffers from significant delays, often only becoming known after residents have already experienced drinking water difficulties, hindering early warning and timely intervention. Furthermore, the quality of reported data is affected by human factors, exhibiting inconsistencies and delays, making it difficult to provide a reliable basis for scientific decision-making.

[0004] Although standardized precipitation index (SPI) and standardized soil moisture index (SSI) are available in the field of drought monitoring, these indicators are mainly used for monitoring and assessing meteorological and agricultural droughts, and are not designed for the risk of rural drinking water shortages. In fact, rural drinking water shortages due to drought are a complex process involving natural conditions and social factors, and are not directly related to general drought indices.

[0005] More importantly, due to the lack of a specific assessment index and daily monitoring system for the risk of drinking water difficulties in rural areas, relevant departments struggle to promptly grasp the development trend and spatial distribution characteristics of the risk. Furthermore, the absence of a scientific method for determining risk thresholds makes it difficult to accurately reflect the actual risk situation in different regions.

[0006] These problems often lead to a "firefighting" response in drought relief and water supply decisions, making it difficult to achieve proactive early warning and precise intervention, which affects the efficiency of drought relief and the protection of drinking water safety for rural residents.

[0007] Therefore, developing a dynamic monitoring and early warning method for the risk of drinking water difficulties in rural areas due to drought is of great significance for improving the monitoring and early warning capabilities for drinking water safety and the level of disaster prevention and mitigation. Summary of the Invention

[0008] To address the problems existing in current technologies, this invention aims to solve the passive response problem caused by relying on disaster reporting in the risk management of rural drought-related drinking water difficulties. It provides a system solution integrating data fusion, model building, risk period identification, and judgment. By establishing a complete technical route from acquiring environmental variables and daily meteorological and hydrological data, to determining physical constraints, then to data dimensionality reduction and variable structure analysis, and subsequently constructing three types of risk assessment models—linear, nonlinear, and machine learning—and finally calibrating model parameters based on historical reported data to achieve precise identification of risk periods, this invention achieves a shift from passively receiving disaster reports to proactive monitoring and early warning, improving the timeliness and accuracy of risk identification, providing a scientific basis for drought relief and water supply decisions, and effectively ensuring the drinking water safety of rural residents.

[0009] The objective of this invention is achieved through the following technical solution:

[0010] This invention provides a method for dynamic assessment and early warning of the risk of drinking water shortages in rural areas due to drought, comprising the following steps:

[0011] Step 1: Multi-source data acquisition and preprocessing

[0012] Collect basic data, including environmental variable data and daily input variables; the environmental variable data includes population density, elevation, slope, historical average evapotranspiration and historical average precipitation; the daily input variables include maximum temperature, temperature anomaly, soil moisture content, number of consecutive high-temperature days, average wind speed, reference crop evapotranspiration (ET0) and number of consecutive rainless days, as well as the percentage of precipitation anomaly at multiple time scales.

[0013] After all the data undergoes imputation of missing values, outlier identification and removal, and standardization, a spatiotemporally consistent multi-source dataset is formed, providing data support for subsequent determination of physical constraints and data dimensionality reduction and variable structure analysis.

[0014] Step 2: Determine physical constraints

[0015] Based on the different impacts of various input variables on drought risk, the physical constraints between each input variable and the risk of rural residents facing drinking water difficulties due to drought are clearly defined. Based on hydrometeorological principles and actual observational analysis, the direction of each variable's impact on risk is determined. Specifically:

[0016] For factors including maximum temperature, temperature anomaly, and number of consecutive rainless days, an increase in their values ​​leads to an increased risk of drought, which is marked as a positive relationship (+).

[0017] For factors including soil moisture content, number of consecutive high-temperature days, and percentage of precipitation anomaly, an increase in their values ​​helps alleviate drought risk and is marked as a negative relationship (-);

[0018] For complex factors including mean wind speed and reference crop evapotranspiration, since they may exhibit different influencing mechanisms under different geographical environments and climatic conditions, they are marked as undetermined relationships (?).

[0019] The direction and intensity of its effect need to be clarified through subsequent parameter calibration and sensitivity analysis; the determination of the above physical constraint relationship provides a scientific basis for data dimensionality reduction and variable structure analysis, and also provides constraints for parameter setting in the model construction process to ensure the physical rationality of the model.

[0020] Step 3: Data Dimensionality Reduction and Variable Structure Analysis

[0021] After determining the physical constraints, data processing is performed on the acquired multi-source dataset;

[0022] Since collinearity exists among the input variables, it may lead to model redundancy and reduced computational efficiency. Therefore, principal component analysis is used for dimensionality reduction: a correlation matrix is ​​constructed from the standardized environmental variable data and daily input variables, eigenvalues ​​and eigenvectors are calculated, and principal components are extracted in order of eigenvalue magnitude.

[0023] The key formula for principal component extraction is as follows:

[0024]

[0025] Where: PC i Let a be the i-th principal component; ij X is the loading coefficient of the j-th original variable in the i-th principal component; j Let be the j-th original variable after standardization; m is the total number of original variables.

[0026] This method can effectively remove redundant information from the data and retain key feature variables, which reduces the amount of data processing and ensures the accuracy of the risk assessment model.

[0027] Step 4: Risk Assessment Model Construction

[0028] Three different types of models are constructed in parallel: linear model, nonlinear model and machine learning model, to comprehensively assess the risk of rural people having difficulty accessing drinking water due to drought; all three types of models adopt the idea of ​​dimensionality reduction, further reducing the dimensionality of multiple principal components extracted in step 3 to obtain a one-dimensional sequence as the final risk value;

[0029] (1) Linear Model

[0030] The linear model is constructed using a linear combination of principal components, and its mathematical expression is as follows:

[0031]

[0032] In the formula: Rlinear The risk value obtained from the linear model; w i PC represents the weight coefficient of the i-th principal component; i Let i be the i-th principal component; n is the number of principal components retained.

[0033] (2) Nonlinear Model

[0034] The nonlinear model adopts a quadratic function form and introduces the interaction between principal components; its expression is:

[0035]

[0036] In the formula: R nonlinear The risk value obtained from the nonlinear model; α i β is the coefficient of the linear term of the i-th principal component; ij PC represents the coefficient of the interaction term between the i-th principal component and the j-th principal component. i and PC j These are the sequences of the i-th and j-th principal components, respectively;

[0037] (3) Machine learning model

[0038] Machine learning models use clustering algorithms to cluster sample points in the principal component space and calculate the distance from each sample to the cluster center as a risk measure.

[0039] R ml =dist(PC sample C k (4)

[0040] In the formula: R ml The risk value obtained from the machine learning model; dist represents the distance function; PC sample C represents the coordinates of the sample points in the principal component space. k The corresponding cluster centers;

[0041] The three types of models reduce the principal component space to a one-dimensional risk sequence using different methods, providing multi-dimensional risk quantification results for subsequent comprehensive assessment.

[0042] Step 5: Model parameter calibration

[0043] Based on the training set data, the parameters and risk thresholds of the three types of models constructed in step 4 are calibrated to determine the optimal risk assessment criteria. A dual-objective function evaluation system is established to comprehensively consider the consistency between the model's prediction results and historical reported data, ensuring the accuracy and practicality of the evaluation results. Specifically, this includes:

[0044] S51: Model Parameter Initialization and Optimization

[0045] The initial weight coefficients of the linear model are set according to the proportion of the variance contribution rate of each principal component, and iterative optimization is performed by combining the least squares method with the gradient descent algorithm; the coefficients of the first term and the interaction term are determined by the stepwise regression method; the machine learning model determines the coordinates of the cluster center by K-means clustering, and the calculated risk value sequence is standardized and mapped to the [0,1] interval.

[0046] S52: Risk Distribution Analysis and Initial Threshold Setting

[0047] Probability distribution analysis was performed on the risk value sequences obtained from the three types of models, and an initial risk threshold was set based on historical data as the initial boundary for determining risk and no-risk states.

[0048] S53: Construction of Dual Objective Functions and Parameter Optimization

[0049] A dual-objective function evaluation system is established, considering two optimization objectives simultaneously: First, maximize the Pearson correlation coefficient between the annual high-risk days series and the reported high-risk population to ensure the consistency of their changing trends; Second, minimize the root mean square error between the annual high-risk days series and the reported high-risk population to ensure numerical similarity.

[0050] A comprehensive evaluation index J = correlation coefficient / root mean square error is constructed. By iteratively adjusting the model parameters and risk threshold, the optimal parameter combination for J is found. Through multiple rounds of optimization, the optimal model and its parameter settings are determined, providing a reliable basis for subsequent risk period identification.

[0051] S54: Model Validation

[0052] The optimized model was applied to the test set data to statistically analyze the correlation between the number of high-risk days and the actual number of people at risk, calculate the comprehensive evaluation index, verify the model's generalization ability and predictive stability, and ensure that the model can adapt to different climatic conditions and socio-economic backgrounds.

[0053] Step 6: Determine the risk period

[0054] Based on run theory, high-risk periods can be identified and merged to achieve accurate monitoring and early warning of the risk of people facing drinking water difficulties due to drought.

[0055] The advantages of this invention compared to the prior art are as follows:

[0056] 1. Filling technological gaps in the field and realizing the transformation from passive response to proactive early warning: Breaking through the traditional disaster statistics model: Through real-time acquisition and dynamic analysis of multi-source data, it achieves early identification and early warning of the risk of rural drinking water difficulties due to drought, changing the passive mechanism that relies on disaster reporting at the county and township levels, and solving the problem of the lag in existing technologies that "respond only after the disaster occurs." Constructing a full-process early warning system: Through a complete technical route of "data acquisition - physical constraint analysis - variable dimensionality reduction - multi-model construction and optimization - risk period identification," a systematic proactive early warning capability has been formed, which can provide technical support for rural drinking water safety.

[0057] 2. Innovative Solution to the Matching Challenges of Historical Data Granularity with Early Warning Needs: Cross-Scale Data Fusion Technology: This technology cleverly utilizes annually reported risk population data to calibrate daily-scale risk models. Through dual-objective function optimization, it organically combines data from different time granularities, overcoming the bottleneck of insufficient historical data temporal resolution in traditional methods. Improved Model Applicability: This technology enables models to inversely deduce daily-scale risk evolution processes based on limited annual-scale statistical data, solving the model construction challenges in rural areas caused by a lack of historical monitoring data. It is particularly suitable for rural, decentralized water supply areas with insufficient data accumulation.

[0058] 3. Enhanced Assessment Accuracy Through Multi-Model Integration and Physical Constraint System: Multi-dimensional Risk Quantification Capability: Parallel construction of linear, nonlinear, and machine learning models, cross-validation of risks using different algorithms, and dimensionality reduction of collinear variables using principal component analysis effectively avoids the limitations of a single model, resulting in more comprehensive risk assessment results. Physical Mechanism-Driven Variable Processing: Establishing physical constraints between input variables and risks (e.g., positive / negative impacts) provides scientific constraints for model construction, ensuring that assessment results conform to hydrological and meteorological principles.

[0059] 4. Dynamic Risk Period Identification Enhances Early Warning Timeliness: Refined Analysis Based on Runs Theory: High-risk periods are identified and merged based on runs theory, transforming daily risk assessment results into continuous risk cycles. This accurately pinpoints the duration of risks, providing a clear time window for drought relief decision-making. Early Warning: Dynamic monitoring enables timely identification of risk periods, allowing relevant departments to deploy water supply security measures in advance, effectively reducing the scope and severity of drinking water shortages. Attached Figure Description

[0060] The present invention will be further described below with reference to the accompanying drawings and embodiments:

[0061] Figure 1 A framework diagram for a risk assessment model of drinking water difficulties caused by drought;

[0062] Figure 2 A graph showing the frequency distribution of risk values ​​and threshold determination;

[0063] Figure 3 Example chart for identifying and merging risk periods (April-May 2013). Detailed Implementation

[0064] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0065] Example 1

[0066] like Figure 1 The process shown in this embodiment, taking the risk assessment of rural drinking water difficulties due to drought in Yanting County, Sichuan Province as an example, provides a method for dynamic assessment and early warning of rural drinking water difficulties due to drought, including the following steps: Step 1, multi-source data acquisition and preprocessing

[0067] First, obtain two types of basic data, including environmental variable data and daily input variables.

[0068] (1) Obtaining environmental variable data

[0069] Population density data: Using LandScan global population density data from 2000 to 2019, after cropping and statistical analysis, the average population density within the administrative region of Yanting County was determined to be 220.5 people / square kilometer;

[0070] Elevation data: Obtained from the SRTM 90m resolution digital elevation model, the elevation range of Yanting County is 350-650 meters;

[0071] Slope data: Based on the above DEM data, the area with slopes of 6-15 degrees (inclusive) accounts for 66.09% of the county.

[0072] Historical average evapotranspiration: Calculated using the Penman-Monteith formula based on observation data from the Yanting County Meteorological Station and five surrounding meteorological stations from 1980 to 2019. The annual average reference crop evapotranspiration is 800-1200 mm.

[0073] Historical average precipitation: Based on historical observations from the same meteorological station network, the annual average precipitation in the region was calculated to be 907 mm.

[0074] (2) Obtaining variables by daily input

[0075] Meteorological elements: Maximum temperature, temperature anomaly, average wind speed, and ETO data were extracted from daily observation records of the Yanting County Meteorological Station from 1990 to 2019.

[0076] Soil moisture content: obtained by combining MODIS satellite soil moisture products with data from 9 local automatic soil moisture monitoring stations, after bias correction;

[0077] Consecutive high-temperature days: The number of consecutive days with a maximum daily temperature of ≥35℃ is counted.

[0078] Number of consecutive rainless days: The standard for rainless days is daily precipitation <0.5mm, and the number of consecutive days is counted.

[0079] Precipitation anomaly percentage at multiple time scales: Calculate the cumulative precipitation over 30 days, 60 days, 90 days, 120 days, 180 days, and 300 days respectively, and compare it with the historical average for the same period to obtain the percentage value.

[0080] (3) Data preprocessing methods

[0081] Handling missing data: For meteorological station data with a missing data rate of less than 2%, linear interpolation is used to fill the missing data; for remote sensing data missing data, which is mainly caused by cloud cover, a combination of time series interpolation and spatial interpolation is used to fill the missing data.

[0082] Outlier identification and removal: Outliers are identified using the 3σ principle, combined with historical extreme value records, and outliers account for less than 0.5% of the total data.

[0083] Data standardization: Variables with different dimensions are processed using the Z-score standardization method to transform them into standard normal distribution data with a mean of 0 and a standard deviation of 1;

[0084] Spatiotemporal consistency processing: all data are unified to the administrative boundary of Yanting County, with a spatial resolution of 1km×1km and a temporal resolution of daily scale;

[0085] After all the above data undergoes missing value imputation, outlier identification and removal, and standardization, a spatiotemporally consistent multi-source dataset is formed.

[0086] Step 2: Determine physical constraints

[0087] Based on the geographical and climatic characteristics of Yanting County and historical data on rural drinking water shortages due to drought, this embodiment defines the physical constraints between each input variable and the risk. By analyzing the correspondence between historical data on rural drinking water shortages due to drought in Yanting County from 1990 to 2019 and each input variable, and combining the principles of regional hydrometeorology, the following physical constraints are determined:

[0088] (1) Positive relationship (+) variable

[0089] An increase in the following variables will lead to a higher risk of rural residents facing drinking water shortages due to drought, thus establishing a positive relationship:

[0090] Highest temperature: Historical data from Yanting County shows that when the highest temperature exceeds 32℃, the probability of rural residents experiencing drinking water shortages due to drought increases significantly.

[0091] Temperature anomaly: When the temperature is more than 3°C higher than the same period in previous years, the risk increases significantly;

[0092] Number of consecutive rainless days: This variable is significantly positively correlated with risk, especially when there are more than 25 consecutive rainless days, the risk increases sharply;

[0093] Consecutive days of high temperatures: Analysis shows that for every additional day of consecutive high temperatures, the risk increases by approximately 12%.

[0094] (2) Negative Relationship (-) Variable

[0095] The following variables, when their values ​​increase, help mitigate risk and are therefore identified as having a negative relationship:

[0096] Soil moisture content: When the soil moisture content in the top 20cm layer is below 18%, the risk increases significantly;

[0097] Precipitation anomaly percentage: The average precipitation anomaly at each time scale is negatively correlated with risk.

[0098] (3) Clarification of uncertain relationship (?) variables

[0099] Through historical data analysis and parameter calibration in Yanting County, the initially uncertain relationships between variables were clarified:

[0100] Average wind speed: Under specific climatic conditions in Yanting County, wind speed is positively correlated with risk (+), mainly because increased wind speed exacerbates evaporation;

[0101] ET0 (reference crop evapotranspiration): ET0 in this region shows a significant positive correlation with risk (+).

[0102] Step 3: Data Dimensionality Reduction and Variable Structure Analysis

[0103] Based on multi-source data from Yanting County between 1990 and 2019 and established physical constraints, this embodiment performs data dimensionality reduction and variable structure analysis on all 17 variables to reduce data dimensionality and extract key features. The specific implementation process is as follows:

[0104] (1) Construction of the correlation matrix

[0105] First, Pearson correlation coefficients were calculated for the 17 standardized variables, and a 17×17 correlation matrix was constructed. The analysis results showed significant collinearity among the variables. For example, the correlation coefficient between consecutive rainless days and soil moisture content reached -0.83, the correlation coefficient between temperature anomaly and maximum temperature was 0.76, and the correlation coefficients between the percentage of precipitation anomalies at each time scale ranged from 0.65 to 0.92. These high correlations confirmed the existence of redundant information among the variables, necessitating dimensionality reduction through principal component analysis.

[0106] (2) Eigenvalue calculation and principal component extraction

[0107] Based on the correlation matrix, eigenvalues ​​and eigenvectors were calculated. The results showed that the first four eigenvalues ​​were 5.24, 2.87, 1.53, and 1.42, with corresponding variance contribution rates of 40.3%, 22.1%, 11.8%, and 10.9%, respectively. The cumulative variance contribution rate was 85.1%, exceeding the preset 85% threshold. Therefore, the first four principal components were selected as the basic parameters for subsequent model construction.

[0108] (3) Principal component loading analysis

[0109] The formulas for calculating the loading coefficients of each principal component and the original variables are as follows:

[0110] |R-λI|=0 (1)

[0111] In the formula: R is the correlation matrix, λ is the eigenvalue, and I is the identity matrix.

[0112] Formula for calculating the variance contribution rate of each eigenvalue:

[0113]

[0114] In the formula: λ i Let be the i-th eigenvalue, p be the total number of variables (17), and j be the index or sequence number of each eigenvalue when summing all eigenvalues.

[0115] Formula for calculating cumulative variance contribution rate:

[0116]

[0117] In the formula: k is the number of principal components selected.

[0118] The calculation results show that the first four eigenvalues ​​are 5.24, 2.87, 1.53, and 1.42, with corresponding variance contribution rates of 40.3%, 22.1%, 11.8%, and 10.9%, respectively. The cumulative variance contribution rate is 85.1%, which has reached the preset threshold of 85%. Therefore, the first four principal components are selected as the basic parameters for subsequent model construction.

[0119] After maximum variance rotation, the loading coefficients of each principal component and the original 17 variables are as follows:

[0120] Principal Component 1 (PC1): Water Supply Index. This primarily reflects water availability, with soil moisture content (0.85), 90-day precipitation anomaly percentage (0.89), 120-day precipitation anomaly percentage (0.92), and 180-day precipitation anomaly percentage (0.87) as the main positive loading variables, and the number of consecutive rainless days (-0.82) as the main negative loading variable. This principal component explains 40.3% of the total variance; lower scores indicate insufficient water supply and a higher risk of drought.

[0121] Principal Component 2 (PC2): Heat Accumulation Index. Primarily reflects heat conditions, with maximum temperature (0.83), temperature anomaly (0.86), and number of consecutive high-temperature days (0.79) as the main positive loading variables. This principal component explains 22.1% of the total variance; higher scores indicate more severe heat accumulation and a higher risk of drought.

[0122] Principal Component 3 (PC3): Water Consumption Index. Primarily reflects evaporation intensity, with mean wind speed (0.75) and reference crop evapotranspiration ET0 (0.81) as the main positive loading variables. This principal component explains 11.8% of the total variance; higher scores indicate faster water consumption and a higher risk of drought.

[0123] Principal Component 4 (PC4): Temporal Scale Characteristic Index. It primarily reflects the differences in precipitation characteristics across different time scales, with the short-term precipitation anomaly percentage (30 days: 0.84, 60 days: 0.82) as the main positive loading variable and the long-term precipitation anomaly percentage (300 days: -0.77) as the main negative loading variable. This principal component explains 10.9% of the total variance, reflecting the unevenness of precipitation temporal distribution and is of great significance for judging drought development trends.

[0124] (4) Principal component score calculation

[0125] According to formula (1), the daily principal component scores are calculated using the standardized values ​​and loading coefficients of each variable. Specifically, for each day (7670 days) in Yanting County from 1990 to 2010, the scores of the four principal components are calculated, forming a 7670×4 principal component score matrix. This period covers multiple typical drought events, and the principal component score sequence can fully reflect the development pattern of drought in Yanting County and the impact of hydrological and meteorological conditions on drinking water risk. This matrix retains 85.1% of the information of the original 17 variables and significantly reduces the data dimensionality, providing a scientific basis for subsequent risk assessment model construction and parameter calibration.

[0126] Step 4: Risk Assessment Model Construction

[0127] Based on four principal components extracted from Yanting County, this embodiment constructs three different types of risk assessment models in parallel to achieve a comprehensive assessment of the risk of drinking water shortages due to drought. The specific implementation process of the three models is as follows:

[0128] (1) Linear Model Construction

[0129] The linear model is constructed using a weighted linear combination of principal components, and its mathematical expression is as follows:

[0130]

[0131] In the formula: R linear The risk value obtained from the linear model; w i PC represents the weight coefficient of the i-th principal component; i Let i be the i-th principal component.

[0132] (2) Construction of nonlinear models

[0133] The nonlinear model introduces the interaction between principal components and is constructed using a quadratic function form. Its mathematical expression is:

[0134]

[0135] In the formula: R nonlinear The risk value obtained from the nonlinear model; α i β is the coefficient of the linear term of the i-th principal component; ij PC represents the coefficient of the interaction term between the i-th principal component and the j-th principal component. i and PC j These are the sequences of the i-th and j-th principal components, respectively.

[0136] (3) Machine learning model construction

[0137] The machine learning model employs the K-means clustering algorithm. First, cluster analysis is performed on the four principal components of the training set from 1990 to 2010, determining the optimal number of clusters to be 2, representing risky and risk-free states respectively. The distance function used is Euclidean distance, calculating the distance from the daily principal component vector to the risk cluster center; a smaller distance value indicates higher risk. To ensure consistency with other models (larger values ​​indicate higher risk), the reciprocal of the distance is used as the risk indicator.

[0138]

[0139] In the formula: R ml The risk value obtained from the machine learning model; PC i C is the sequence of the i-th principal component of the sample point; 1iLet be the coordinates of the i-th dimension of the risk cluster center.

[0140] Step 5: Model parameter calibration

[0141] Based on training data from 1990 to 2010, this embodiment calibrates the parameters and risk thresholds of the three types of models to determine the optimal risk assessment criteria.

[0142] S51: Model Parameter Initialization and Optimization

[0143] (1) Calibration of linear model parameters

[0144] The initial weight coefficients of the linear model were set according to the proportion of the variance contribution rate of each principal component, and iterative optimization was performed using the least squares method combined with the gradient descent algorithm. After 500 iterations, the optimal weight coefficients were obtained: w1 = -0.45, w2 = 0.38, w3 = 0.32, w4 = 0.15. Among them, w1 is negative, which conforms to the physical meaning of PC1 (water supply index); w2, w3, and w4 are positive, which is consistent with the physical constraint relationship.

[0145] (2) Calibration of nonlinear model parameters

[0146] The nonlinear model determines the coefficients of the first-order and interaction terms using a stepwise regression method. After optimization calculations, the optimal parameter values ​​are obtained:

[0147] The coefficients of the linear terms are: α1 = -0.42, α2 = 0.35, α3 = 0.30, α4 = 0.14

[0148] Key interaction coefficient: β 12 = -0.18, β 13 = -0.15, β 23 =0.12, β 14 = -0.08, β 24 =0.06, and the absolute values ​​of the coefficients of the other interaction terms are all less than 0.05.

[0149] (3) Machine learning model parameter calibration

[0150] The coordinates of the cluster centers determined by the K-means clustering results are:

[0151] Cluster center C1 (at risk): [-1.25, 0.85, 0.65, 0.40];

[0152] Cluster center C2 (risk-free): [0.95, -0.45, -0.55, -0.35];

[0153] The risk value sequence calculated by the machine learning model is processed by Min-Max standardization and mapped to the [0,1] interval. The closer the value is to 1, the higher the risk.

[0154] S52: Risk Distribution Analysis and Initial Threshold Setting

[0155] (1) Perform probability distribution analysis on the risk value sequences obtained from the three types of models:

[0156] The risk values ​​of the linear model are approximately normally distributed, with a mean of 0.05 and a standard deviation of 0.73.

[0157] The risk values ​​of the nonlinear model are distributed to the right, with a mean of 0.08 and a standard deviation of 0.82.

[0158] The standardized risk values ​​of the machine learning model exhibit a bimodal distribution with a mean of 0.42 and a standard deviation of 0.18.

[0159] (2) Based on historical data of drought-related drinking water shortages in Yanting County from 1990 to 2010, an initial risk threshold was set:

[0160] Linear model: r0_linear = 0.56 (corresponding to approximately 85th percentile);

[0161] Nonlinear model: r0_nonlinear = 0.58 (corresponding to approximately 85th percentile);

[0162] Machine learning model: r0_ml = 0.66 (corresponding to approximately 85% quantile).

[0163] S53: Construction of Dual Objective Functions and Parameter Optimization

[0164] (1) Based on the initial threshold, the risk value series from 1990 to 2011 was statistically analyzed year by year to obtain the number of high-risk days in each year:

[0165] Table 1 Simulation results of various models and actual reported population at risk.

[0166]

[0167] (2) Based on the statistical results of the training set (1990-2010) above, calculate the correlation coefficient and root mean square error between the number of high-risk days in each model and the actual reported number of at-risk people:

[0168] The correlation coefficient of the linear model is 0.74, and the root mean square error is 0.70; the correlation coefficient of the nonlinear model is 0.88, and the root mean square error is 0.48; the correlation coefficient of the machine learning model is 0.84, and the root mean square error is 0.54. A comprehensive evaluation index J = correlation coefficient / root mean square error is constructed, and the calculated values ​​are: J = 1.06 for the linear model; J = 1.83 for the nonlinear model; and J = 1.565 for the machine learning model.

[0169] (3) Based on the nonlinear model (because it has the highest J value), perform parameter optimization and adjustment:

[0170] Threshold optimization: Adjusting the threshold r0 within the range of [0.55, 0.67], it was found that the J value reaches its maximum of 1.83 when r0 = 0.61 (e.g., Figure 2 (As shown).

[0171] Optimization of the coefficients of the first term: Adjust α1 to α4 within a range of ±10%. The optimized coefficients are α1 = -0.44, α2 = 0.37, α3 = 0.32, and α4 = 0.14.

[0172] Interaction term coefficient optimization: Adjust the coefficients of the main interaction terms; the optimized coefficients are β. 12 = -0.20, β 13 = -0.17, β 23 =0.14.

[0173] After 100 rounds of optimization, the comprehensive performance index J value of the nonlinear model stabilized at 1.83, and it was determined to be the optimal model.

[0174] S54: Model Validation

[0175] The optimized nonlinear model (threshold r0 = 0.61) was used to assess the risk of the test set from 2011 to 2019. The statistical results showed a significant correlation between the number of high-risk days and the actual number of people at risk (correlation coefficient 0.91, root mean square error 0.33, and comprehensive evaluation index J value of 2.76), indicating that the model has good generalization ability and predictive stability.

[0176] Step 6: Determine the risk period

[0177] Based on the optimized nonlinear model and its risk threshold from step 5, this embodiment applies the runs theory method to identify and merge the risk periods of drinking water shortages due to drought in rural areas of Yanting County, thereby achieving refined risk monitoring. (See appendix) Figure 3 ).

[0178] The specific implementation process is as follows:

[0179] (1) Identification of the initial risk day

[0180] First, a threshold of 0.61 was applied to the optimized nonlinear model risk value sequence for binary classification, resulting in an initial risk state sequence (1 representing risk, 0 representing no risk). Analysis was then conducted on the drought process in Yanting County from April to May 2013, and the initial identification results are shown by the blue line "original value" in the figure. It can be observed that the risk state fluctuated multiple times between April 8th and May 8th, exhibiting several short-term interruptions.

[0181] (2) Risk periods will be combined for implementation.

[0182] The initial risk state sequence was processed using the merging rules of runs theory: A period of three consecutive days with a risk value exceeding the threshold of 0.61 was identified as the initial high-risk group; the intervals between adjacent high-risk groups were analyzed, such as the one-day interval between April 13th and 15th, which met the merging condition of "interval ≤ 3 days"; the merging was iteratively performed according to this rule, ultimately combining all high-risk groups from April 8th to May 8th into a single complete risk period. The red line "correction value" in the figure illustrates the merged result. It is clearly visible that multiple previously discontinuous high-risk periods have been merged into a continuous and complete risk period.

[0183] (3) Verification of the conditions for terminating the merger

[0184] Boundary confirmation was performed on the merged risk period: the risk values ​​for the 7 days prior to the beginning of the risk period (April 8th) (April 1st-7th) did not exceed the threshold; the risk values ​​for the 7 days following the end of the risk period (May 8th) (May 9th-15th) also did not exceed the threshold. The termination condition of "no risk for 7 days before and after" was met, therefore, April 8th to May 8th was confirmed as a complete risk period, with a total duration of 31 days. This is basically consistent with the reported drinking water risk period in 2013.

[0185] (4) Verification of actual impact

[0186] A comparison was made between the risk period identified by the model and the actual reported data from Yanting County in 2013: The actual reports showed that the first report of localized rural drinking water difficulties occurred on April 10, 2013, affecting approximately 18,700 people; the number of reported cases peaked at 21,500 on April 30; and the drought was reported to have largely eased by May 10. The risk period identified by the model (April 8 to May 8) largely matches the actual reported data.

[0187] By using the above-mentioned method for determining risk periods, the daily risk assessment results were successfully transformed into practically operable risk periods, providing a scientific basis for monitoring and early warning of the risk of drinking water difficulties caused by drought in Yanting County, and significantly improving the timeliness and accuracy of risk management.

[0188] Finally, it should be noted that the above description is only used to illustrate the technical solutions of the present invention and is not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention.

Claims

1. A method for dynamic assessment and early warning of the risk of drinking water difficulties in rural areas due to drought, characterized in that, The method includes the following steps: Step 1: Multi-source data acquisition and preprocessing Collect basic data, including environmental variable data and daily input variables; the environmental variable data includes population density, elevation, slope, historical average evapotranspiration and historical average precipitation; the daily input variables include maximum temperature, temperature anomaly, soil moisture content, number of consecutive high-temperature days, average wind speed, reference crop evapotranspiration and number of consecutive rainless days, as well as the percentage of precipitation anomaly at multiple time scales. After all the data undergoes imputation of missing values, outlier identification and removal, and standardization, a spatiotemporally consistent multi-source dataset is formed. Step 2: Determine physical constraints Based on the different impacts of various input variables on drought risk, the physical constraints between each input variable and the risk of rural residents facing drinking water difficulties due to drought are clearly defined, specifically: For factors including maximum temperature, temperature anomaly, and number of consecutive rainless days, an increase in their values ​​leads to an increased risk of drought, which is marked as a positive relationship (+). For factors including soil moisture content, number of consecutive high-temperature days, and percentage of precipitation anomaly, an increase in their values ​​helps alleviate drought risk and is marked as a negative relationship (-); For complex factors including mean wind speed and reference crop evapotranspiration, since they may exhibit different influencing mechanisms under different geographical environments and climatic conditions, they are marked as undetermined relationships (?). Step 3: Data Dimensionality Reduction and Variable Structure Analysis After determining the physical constraints, data processing is performed on the acquired multi-source dataset; Principal component analysis was used for dimensionality reduction: a correlation matrix was constructed from the standardized environmental variable data and daily input variables, eigenvalues ​​and eigenvectors were calculated, and principal components were extracted in order of eigenvalue magnitude. The key formula for principal component extraction is as follows: Where: PC i Let a be the i-th principal component; ij X is the loading coefficient of the j-th original variable in the i-th principal component; j Let j be the j-th original variable after standardization; m is the total number of original variables; Step 4: Risk Assessment Model Construction Three different types of models are constructed in parallel: linear model, nonlinear model and machine learning model; all three types of models adopt the idea of ​​dimensionality reduction, and the multiple principal components extracted in step 3 are further reduced in dimensionality to obtain a one-dimensional sequence as the final risk value; (1) Linear Model The linear model is constructed using a linear combination of principal components, and its mathematical expression is as follows: In the formula: R linear The risk value obtained from the linear model; w i PC represents the weight coefficient of the i-th principal component. i Let i be the i-th principal component; n is the number of principal components retained. (2) Nonlinear Model The nonlinear model adopts a quadratic function form and introduces the interaction between principal components; its expression is: In the formula: R nonlinear The risk value obtained from the nonlinear model; α i β is the coefficient of the linear term of the i-th principal component; ij PC represents the coefficient of the interaction term between the i-th principal component and the j-th principal component. i and PC j These are the sequences of the i-th and j-th principal components, respectively; (3) Machine learning model Machine learning models use clustering algorithms to cluster sample points in the principal component space and calculate the distance from each sample to the cluster center as a risk measure. R ml =dist(PC sample ,C k ) (4) In the formula: R ml The risk value obtained from the machine learning model; dist represents the distance function; PC sample C represents the coordinates of the sample points in the principal component space. k The corresponding cluster centers; Step 5: Model parameter calibration Based on the training set data, the parameters and risk thresholds of the three types of models constructed in step 4 are calibrated to determine the optimal risk assessment criteria, specifically including: S51: Model Parameter Initialization and Optimization The initial weight coefficients of the linear model are set according to the proportion of the variance contribution rate of each principal component, and iterative optimization is performed by combining the least squares method with the gradient descent algorithm; the coefficients of the first term and the interaction term are determined by the stepwise regression method; the machine learning model determines the coordinates of the cluster center by K-means clustering, and the calculated risk value sequence is standardized and mapped to the [0,1] interval. S52: Risk Distribution Analysis and Initial Threshold Setting Probability distribution analysis was performed on the risk value sequences obtained from the three types of models, and an initial risk threshold was set based on historical data as the initial boundary for determining risk and no-risk states. S53: Construction of Dual Objective Functions and Parameter Optimization A dual-objective function evaluation system is established, considering two optimization objectives simultaneously: First, maximize the Pearson correlation coefficient between the annual high-risk days series and the reported high-risk population to ensure the consistency of their changing trends; Second, minimize the root mean square error between the annual high-risk days series and the reported high-risk population to ensure numerical similarity. A comprehensive evaluation index J = correlation coefficient / root mean square error is constructed. The optimal parameter combination for J is found by iteratively adjusting the model parameters and risk threshold. Through multiple rounds of optimization, the optimal model and its parameter settings are determined. S54: Model Validation The optimized model was applied to the test set data to analyze the correlation between the number of high-risk days and the actual number of people at risk, calculate the comprehensive evaluation index, and verify the model's generalization ability and predictive stability. Step 6: Determine the risk period Based on run theory, high-risk periods can be identified and merged to achieve accurate monitoring and early warning of the risk of people facing drinking water difficulties due to drought.

2. The method for dynamic assessment and early warning of rural drinking water difficulties due to drought according to claim 1, characterized in that, In step 1, elevation and slope data are extracted from the SRTM 90m resolution digital elevation model with an accuracy of no less than 90m; population density is obtained from LandScan global population density data; historical average evapotranspiration and precipitation are calculated based on more than 30 years of observation records from meteorological stations; the daily input variables are obtained through automatic weather stations, rain gauge networks and remote sensing monitoring.

3. The method for dynamic assessment and early warning of rural drinking water difficulties due to drought as described in claim 1, characterized in that, In step 3, when performing principal component analysis, start from principal component 1 and add up to principal component n sequentially until the cumulative variance contribution rate reaches 85%.

4. The method for dynamic assessment and early warning of rural drinking water difficulties due to drought as described in claim 1, characterized in that, In step 5, S52, the initial risk threshold corresponds to the high quantile of the risk value distribution.

5. The method for dynamic assessment and early warning of rural drinking water difficulties due to drought as described in claim 1, characterized in that, In step 6, the specific ideas behind the runs theory include: Analyze the risk value sequence obtained in step 5. If the risk value exceeds the set threshold for 3 consecutive days, it is identified as an initial high-risk day. If the interval between the number of risk-free days between two high-risk day groups is less than or equal to 3 days, the two high-risk day groups are merged into a longer high-risk period. The merging process was carried out iteratively until no high-risk days occurred within the 7 days before the start of the high-risk period and the 7 days after its end.

Citation Information

Patent Citations

  • Typhoon disaster risk assessment and dynamic forecasting method

    CN116757305A

  • Drought risk prediction method and system for changing climate

    CN117592663A

  • Data shortage estuary composite flood disaster research method considering climatic change

    CN118760899A

  • Soil water content prediction method and system based on canopy-atmospheric environment information

    CN120146328A

  • Dynamic risk control method and device, equipment and storage medium

    CN120163653A