A rural drought drinking difficulty risk dynamic assessment and early warning method

By processing multi-source data and building models, the passive response problem in risk assessment of rural drinking water difficulties caused by drought has been solved, and proactive early warning of the risk of rural drinking water difficulties caused by drought has been achieved. This has improved the timeliness and accuracy of risk identification and ensured the drinking water safety of rural residents.

CN120875129BActive Publication Date: 2026-02-27CHINA INST OF WATER RESOURCES & HYDROPOWER RES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510912424.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2026-02-27
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

In existing technologies, the assessment of the risk of drinking water difficulties in rural areas due to drought mainly relies on passive disaster reporting, which leads to delayed response and a lack of scientific risk assessment methods. This makes it difficult to achieve early warning and precise intervention, affecting the efficiency and effectiveness of drought relief and water supply decisions.

Method used

By acquiring and preprocessing multi-source data, establishing physical constraints, performing data dimensionality reduction and variable structure analysis, constructing linear, nonlinear, and machine learning models, and calibrating model parameters with historical data, dynamic monitoring and early warning of the risk of rural residents facing drinking water difficulties due to drought can be achieved.

Benefits of technology

This has enabled a shift from passive response to proactive early warning, improved the timeliness and accuracy of risk identification, provided a scientific basis for risk assessment, and ensured the safety of drinking water for rural residents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875129B_ABST
    Figure CN120875129B_ABST
Patent Text Reader

Abstract

The application provides a rural drought drinking difficulty risk dynamic assessment and early warning method, which comprises multi-source data acquisition and preprocessing, physical constraint relationship determination, data dimension reduction and variable structure analysis, risk assessment model construction, model parameter calibration and risk period determination. The collinear variables are processed by principal component analysis method, three types of risk assessment models, including linear model, nonlinear model and machine learning model, are constructed in parallel, the optimal model is screened by using double objective function, the risk threshold is optimized and determined, and the high risk period is identified and merged by using run theory. The core innovation of the application lies in that the risk population data on the annual scale is used to calibrate the risk model on the daily scale, thereby effectively solving the problem that the historical data granularity does not match the early warning demand. Through the establishment of the variable processing system under the physical constraint condition and the multi-model integrated assessment, the timeliness and accuracy of the risk assessment are significantly improved, and the transformation from passive disaster statistics to active risk early warning is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of earth science and disaster risk management, and particularly relates to a rural drought-induced drinking water difficulty risk dynamic assessment and early warning method. BACKGROUND

[0002] In rural areas, drought-induced drinking water difficulty problems seriously affect the basic life and production activities of residents.

[0003] At present, the drought-induced drinking water difficulty situation in rural areas is mainly grasped by relying on disaster statistics and reports of county and township grassroots, which is a typical passive response mechanism. This method has obvious lag, and the situation can be known only after the residents have drinking water difficulty, so it is difficult to achieve early warning and timely intervention. At the same time, the quality of reported data is affected by human factors, and there are problems such as inconsistency, untimeliness, etc., which makes it difficult to provide reliable basis for scientific decision-making.

[0004] Although there are standard precipitation index (SPI) and standard soil moisture index (SSI) technical means in the field of drought monitoring, these indexes are mainly used for monitoring and evaluation of meteorological drought and agricultural drought, and are not designed for rural drinking water difficulty risk. In fact, drought-induced drinking water difficulty in rural areas is a complex process involving natural conditions and social factors, and lacks direct correlation with general drought index.

[0005] More importantly, due to the lack of evaluation index and daily-scale monitoring system specifically for rural drinking water difficulty risk, it is difficult for relevant departments to grasp the risk development trend and spatial distribution characteristics in time. At the same time, without a scientific risk threshold determination method, it is difficult to accurately reflect the actual risk situation in different regions.

[0006] These problems lead to that drought-resistant water supply decision-making is often in a "fire-fighting" response state, making it difficult to achieve proactive early warning and precise intervention, affecting the efficiency of drought relief and the safety guarantee of rural residents' drinking water.

[0007] Therefore, it is of great significance to develop a rural drought-induced drinking water difficulty risk dynamic monitoring and early warning method to improve the monitoring and early warning capability of drinking water safety and the level of disaster prevention and reduction. SUMMARY

[0008] In view of the problems in the prior art, the present application aims to solve the passive response problem caused by relying on disaster reporting in the risk management of rural drought drinking difficulty, and provides a system solution integrating data fusion, model construction, risk period identification and determination. Through the establishment of the complete technical route from the acquisition of environmental variables and daily meteorological and hydrological data to the determination of physical constraint relationship, data dimension reduction and variable structure analysis, and then the construction of linear, nonlinear and machine learning risk assessment models, and finally the calibration of model parameters based on historical reporting data and the realization of fine identification of risk period, the passive reception of disaster reporting is changed to active monitoring and early warning, the timeliness and accuracy of risk identification are improved, scientific basis is provided for drought resistance water supply decision-making, and the safety of rural residents' drinking water is effectively guaranteed.

[0009] The purpose of the present application is achieved by the following technical solutions:

[0010] The present application provides a rural drought drinking difficulty risk dynamic assessment and early warning method, comprising the following steps:

[0011] Step 1: Multi-source data acquisition and preprocessing

[0012] Collect basic data, including environmental variable data and daily input variables; the environmental variable data includes population density, elevation, slope, historical average evapotranspiration and historical average precipitation; the daily input variables include maximum temperature, temperature anomaly, soil moisture content, consecutive high temperature days, average wind speed, reference crop evapotranspiration (ET0) and consecutive rainless days, and multi-time scale precipitation anomaly percentage;

[0013] After all the data are subjected to missing value interpolation, abnormal value identification and elimination, and standardization processing, a spatiotemporally consistent multi-source data set is formed, providing data support for subsequent physical constraint relationship determination and data dimension reduction and variable structure analysis.

[0014] Step 2: Physical constraint relationship determination

[0015] According to the different influences of different input variables on drought risk, the physical constraint relationship between each input variable and the rural drought drinking difficulty risk is clearly defined, and the influence direction of each variable on the risk is determined according to the principles of hydrology and meteorology and actual observation analysis. Specifically:

[0016] For factors including maximum temperature, temperature anomaly and consecutive rainless days, the increase of the values will lead to an increase in drought risk, marked as positive relationship (+);

[0017] For factors including soil moisture content, consecutive high temperature days and precipitation anomaly percentage, the increase of the values will help to alleviate the drought risk, marked as negative relationship (-);

[0018] For complex factors including average wind speed and reference crop evapotranspiration, due to the different impact mechanisms that may be exhibited under different geographical and climatic conditions, they are marked as uncertain relationships (?);

[0019] The direction and strength of their effects need to be clarified through subsequent parameter calibration and sensitivity analysis; The determination of the above physical constraint relationships provides scientific basis for data dimensionality reduction and variable structure analysis, and also provides constraint conditions for parameter setting during model construction, ensuring the physical reasonableness of the model.

[0020] Step 3: Data dimensionality reduction and variable structure analysis

[0021] After determining the physical constraint relationships, data processing is performed on the obtained multi-source data set;

[0022] Due to the existence of collinearity characteristics among the input variables, it may lead to model redundancy and reduced computational efficiency, therefore principal component analysis is used for dimensionality reduction: the correlation matrix is constructed with the standardized environmental variable data and daily input variables, the eigenvalues and eigenvectors are calculated, and the principal components are extracted in order of eigenvalue size;

[0023] The key formulas for principal component extraction are as follows:

[0024]

[0025] In the formula: PC i is the i-th principal component; a ij is the load coefficient of the j-th original variable in the i-th principal component; X j is the j-th standardized original variable; m is the total number of original variables.

[0026] Through this method, redundant information in the data can be effectively eliminated, and key feature variables are retained, which not only reduces the data processing amount, but also ensures the accuracy of the risk assessment model.

[0027] Step 4: Risk assessment model construction

[0028] Three types of models are constructed in parallel: linear model, nonlinear model and machine learning model, to comprehensively assess the risk of rural drought drinking difficulties; All three types of models use the idea of dimensionality reduction, further reducing the multiple principal components extracted in step 3 to obtain a one-dimensional sequence as the final risk value;

[0029] (1) Linear model

[0030] The linear model is constructed by linear combination of principal components, and its mathematical expression is:

[0031]

[0032] In the formula: Rlinear Risk value obtained by linear model; w i Weight coefficient of the ith principal component; PC i The ith principal component; n is the number of retained principal components;

[0033] (2) Nonlinear model

[0034] The nonlinear model selects a quadratic function form and introduces the interaction between principal components, and its expression is:

[0035]

[0036] In the formula: R nonlinear Risk value obtained by nonlinear model; a i The first-order term coefficient of the ith principal component; b ij The coefficient of the interaction term between the ith principal component and the jth principal component; PC i and PC j The sequence of the ith and jth principal components, respectively;

[0037] (3) Machine learning model

[0038] The machine learning model uses a clustering algorithm to cluster the sample points in the principal component space, and calculates the distance from each sample to the cluster center as a risk measure:

[0039] R ml = dist(PC sample , C k ) (4)

[0040] In the formula: R ml Risk value obtained by machine learning model; dist represents distance function; PC sample The coordinates of the sample point in the principal component space; C k The corresponding cluster center;

[0041] The three types of models reduce the principal component space to one-dimensional risk sequence through different methods, providing multi-dimensional risk quantification results for subsequent comprehensive evaluation.

[0042] Step 5: Model parameter calibration

[0043] Based on the training set data, the parameters and risk threshold of the three types of models constructed in step 4 are calibrated to determine the best risk determination standard. Through the establishment of a double-objective function evaluation system, the consistency of the model prediction results and the historical actual reporting data is comprehensively considered to ensure the accuracy and practicality of the evaluation results. Specifically, it includes:

[0044] S51: Model parameter initialization and optimization

[0045] The initial weight coefficient of the linear model is set in proportion to the variance contribution rate of each principal component, and is iteratively optimized by the least square method combined with the gradient descent algorithm; the non-linear model determines the first-order term coefficient and the interaction term coefficient through the stepwise regression method; the machine learning model determines the cluster center coordinates through K-means clustering, and standardizes the calculated risk value sequence and maps it to the [0, 1] interval;

[0046] S52: Risk distribution analysis and initial threshold setting

[0047] The risk value sequence obtained by the three types of models is analyzed for probability distribution, and an initial risk threshold is set based on historical data as the initial judgment boundary between risk and non-risk states.

[0048] S53: Double objective function construction and parameter optimization

[0049] A double objective function evaluation system is established, considering two optimization objectives: first, maximizing the Pearson correlation coefficient between the annual high-risk day sequence and the reported risk population number, to ensure the consistency of the change trend; second, minimizing the root mean square error between the annual high-risk day sequence and the reported risk population number, to ensure the closeness in numerical value.

[0050] A comprehensive evaluation index J = correlation coefficient / root mean square error is constructed, and the model parameters and risk threshold are iteratively adjusted to find the optimal parameter combination that makes J value optimal; through multiple rounds of optimization, the optimal model and its parameter settings are determined, providing a reliable basis for subsequent risk period identification.

[0051] S54: Model verification

[0052] The optimized model is applied to the test set data, the correlation between the high-risk day number and the actual reported risk population number is calculated, the comprehensive evaluation index is calculated, and the generalization ability and prediction stability of the model are verified to ensure that the model can adapt to different climate conditions and social and economic backgrounds.

[0053] Step 6: Risk period determination

[0054] Based on the run theory, by identifying and merging high-risk periods, accurate monitoring and early warning of the risk of drinking difficulty due to drought are achieved.

[0055] The beneficial effects of the present application compared with the prior art are:

[0056] 1. Fill in the technical blank, realize the transformation from passive response to active early warning: Break through the traditional disaster statistics mode: through multi-source data real-time acquisition and dynamic analysis, realize the early identification and early warning of rural drought drinking difficulty risk, change the passive mechanism of relying on county and township grassroots disaster reporting, solve the lag problem of "responding after disaster occurs" in the prior art. Build a whole-process early warning system: through the complete technical route of "data acquisition-physical constraint analysis-variable dimension reduction-multi-model construction and optimization-risk period identification", a systematic active early warning capability is formed, which can provide technical support for rural drinking water safety guarantee.

[0057] 2. Innovatively solve the matching problem of historical data granularity and early warning demand: Cross-scale data fusion technology: skillfully use the risk population data reported on the annual scale to calibrate the daily scale risk model, realize the organic combination of different time granularity data through double objective function optimization, break through the bottleneck of insufficient time resolution of historical data in traditional methods. Improve model applicability: This technology enables the model to inverse the daily scale risk evolution process based on limited annual scale statistical data, solves the model construction problem caused by lack of historical monitoring data in rural areas, especially suitable for data accumulation insufficient in rural scattered water supply areas.

[0058] 3. Multi-model integration and physical constraint system improve evaluation accuracy: Multi-dimensional risk quantification capability: linear model, nonlinear model and machine learning model are constructed in parallel, different algorithms are used to cross-verify the risk, and principal component analysis method is used to reduce the collinearity variables, which effectively avoids the limitations of single model, and makes the risk assessment result more comprehensive. Physical mechanism driven variable processing: Establish the physical constraint relationship between input variables and risk (such as positive / negative impact), provide scientific constraints for model construction, and ensure that the evaluation result conforms to the principles of hydrology and meteorology.

[0059] 4. Dynamic risk period identification improves the timeliness of early warning: Run theory fine analysis: Based on the run theory, high-risk periods are identified and merged, the daily risk assessment results are converted into continuous risk periods, and the risk duration is accurately located, providing a clear time window for drought resistance decision-making. Early warning: Through dynamic monitoring, the risk period is identified in time, so that relevant departments can deploy water supply guarantee measures in advance, effectively reducing the impact range and degree of drinking water difficulty events. BRIEF DESCRIPTION OF DRAWINGS

[0060] The application will be further described below in combination with the drawings and examples:

[0061] Figure 1 The figure is a risk assessment model framework diagram for drought drinking difficulty;

[0062] Figure 2 The figure is a risk value frequency distribution and threshold determination diagram;

[0063] Figure 3 Example chart for identifying and merging risk periods (April-May 2013). Detailed Implementation

[0064] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0065] Example 1

[0066] like Figure 1 The process shown in this embodiment, taking the risk assessment of rural drinking water difficulties due to drought in Yanting County, Sichuan Province as an example, provides a method for dynamic assessment and early warning of rural drinking water difficulties due to drought, including the following steps: Step 1, multi-source data acquisition and preprocessing

[0067] First, obtain two types of basic data, including environmental variable data and daily input variables.

[0068] (1) Obtaining environmental variable data

[0069] Population density data: Using LandScan global population density data from 2000 to 2019, after cropping and statistical analysis, the average population density within the administrative region of Yanting County was determined to be 220.5 people / square kilometer;

[0070] Elevation data: Obtained from the SRTM 90m resolution digital elevation model, the elevation range of Yanting County is 350-650 meters;

[0071] Slope data: Based on the above DEM data, the area with slopes of 6-15 degrees (inclusive) accounts for 66.09% of the county.

[0072] Historical average evapotranspiration: Calculated using the Penman-Monteith formula based on observational data from the Yanting County Meteorological Station and five surrounding meteorological stations from 1980 to 2019. The annual average reference crop evapotranspiration is 800-1200 mm.

[0073] Historical average precipitation: Based on historical observations from the same meteorological station network, the annual average precipitation in the region was calculated to be 907 mm.

[0074] (2) Obtaining variables by daily input

[0075] Meteorological elements: The maximum temperature, temperature anomaly, average wind speed, and ET0 data were extracted from the daily-scale observation records of the Yanting County Meteorological Station from 1990 to 2019.

[0076] Soil moisture content: Combined with MODIS satellite soil moisture products and data from nine local automatic soil moisture monitoring stations, the soil moisture content was obtained after bias correction.

[0077] Continuous high-temperature days: The number of consecutive days with daily maximum temperature ≥ 35°C was counted.

[0078] Continuous rainless days: The number of consecutive days with daily precipitation < 0.5mm was counted.

[0079] Multi-time scale precipitation anomaly percentage: The cumulative precipitation of 30, 60, 90, 120, 180, and 300 days was calculated, and the percentage value was obtained by comparing it with the historical average value of the same period.

[0080] (3) Data preprocessing method

[0081] Missing value processing: Linear interpolation was used to fill in missing values with a missing rate of less than 2% in meteorological station data. Remote sensing data missing was mainly caused by cloud cover, and a combination of time series interpolation and spatial interpolation was used to fill in the missing values.

[0082] Abnormal value identification and elimination: The 3σ principle was used to identify abnormal values, and historical extreme value records were used for judgment. Abnormal values accounted for less than 0.5% of the total data.

[0083] Data standardization: Z-score standardization method was used to process variables with different dimensions, converting them into standard normal distribution data with mean of 0 and standard deviation of 1.

[0084] Temporal and spatial consistency processing: All data were unified within the administrative boundary of Yanting County, with a spatial resolution of 1km x 1km and a temporal resolution of daily scale.

[0085] After missing value interpolation, abnormal value identification and elimination, and standardization processing, the above all data formed a spatiotemporally consistent multi-source data set.

[0086] Step 2: Physical constraint relationship determination

[0087] Based on the geographical and climatic characteristics of Yanting County and the historical data of rural people's drinking difficulties due to drought, the physical constraint relationship between each input variable and risk was defined in this embodiment. By analyzing the corresponding relationship between the historical data of rural people's drinking difficulties due to drought in Yanting County from 1990 to 2019 and each input variable, and combining the principles of regional hydrology and meteorology, the following physical constraint relationships were determined:

[0088] (1) Positive relationship (+) variable

[0089] The following variable value increase will lead to an increase in the risk of rural drought drinking difficulties, and is determined to be a positive relationship:

[0090] Maximum temperature: Historical data from Yanting County show that when the daily maximum temperature exceeds 32°C, the probability of rural drought drinking difficulty incidents significantly increases;

[0091] Temperature anomaly: When the temperature is more than 3°C higher than the same period in the previous year, the risk increases significantly;

[0092] Number of consecutive days without rain: This variable is significantly positively correlated with risk, especially when consecutive days without rain exceed 25 days, the risk increases sharply;

[0093] Number of consecutive high-temperature days: Analysis shows that for every additional day of consecutive high-temperature days, the risk increases by approximately 12%.

[0094] (2) Negative relationship (-) variable

[0095] The following variable value increase will help to mitigate the risk, and is determined to be a negative relationship:

[0096] Soil moisture content: When the surface 20cm soil moisture content is less than 18%, the risk increases significantly;

[0097] Precipitation anomaly percentage: The average precipitation anomaly at each time scale is negatively correlated with risk.

[0098] (3) Clarification of uncertain relationship (?) variable

[0099] Through historical data analysis and parameter calibration in Yanting County, the initial uncertain variable relationship is clarified:

[0100] Average wind speed: Under specific climatic conditions in Yanting County, wind speed is positively correlated with risk (+), mainly because increased wind speed exacerbates evaporation;

[0101] ET0 (reference crop evapotranspiration): ET0 in this region is positively correlated with risk (+).

[0102] Step 3: Data dimensionality reduction and variable structure analysis

[0103] Based on the multi-source data of Yanting County from 1990 to 2019 and the physical constraint relationships determined, this embodiment performs data dimensionality reduction and variable structure analysis on all 17 variables to reduce data dimensionality and extract key features. The specific implementation process is as follows:

[0104] (1) Correlation matrix construction

[0105] First, Pearson correlation coefficients were calculated for the 17 standardized variables, and a 17×17 correlation matrix was constructed. The analysis results showed significant collinearity among the variables. For example, the correlation coefficient between consecutive rainless days and soil moisture content reached -0.83, the correlation coefficient between temperature anomaly and maximum temperature was 0.76, and the correlation coefficients between the percentage of precipitation anomalies at each time scale ranged from 0.65 to 0.92. These high correlations confirmed the existence of redundant information among the variables, necessitating dimensionality reduction through principal component analysis.

[0106] (2) Eigenvalue calculation and principal component extraction

[0107] Based on the correlation matrix, eigenvalues ​​and eigenvectors were calculated. The results showed that the first four eigenvalues ​​were 5.24, 2.87, 1.53, and 1.42, with corresponding variance contribution rates of 40.3%, 22.1%, 11.8%, and 10.9%, respectively. The cumulative variance contribution rate was 85.1%, exceeding the preset 85% threshold. Therefore, the first four principal components were selected as the basic parameters for subsequent model construction.

[0108] (3) Principal component loading analysis

[0109] The formulas for calculating the loading coefficients of each principal component and the original variables are as follows:

[0110] |R-λI|=0 (1)

[0111] In the formula: R is the correlation matrix, λ is the eigenvalue, and I is the identity matrix.

[0112] Formula for calculating the variance contribution rate of each eigenvalue:

[0113]

[0114] In the formula: λ i Let be the i-th eigenvalue, p be the total number of variables (17), and j be the index or sequence number of each eigenvalue when summing all eigenvalues.

[0115] Formula for calculating cumulative variance contribution rate:

[0116]

[0117] In the formula: k is the number of principal components selected.

[0118] The calculation results show that the first four eigenvalues ​​are 5.24, 2.87, 1.53, and 1.42, with corresponding variance contribution rates of 40.3%, 22.1%, 11.8%, and 10.9%, respectively. The cumulative variance contribution rate is 85.1%, which has reached the preset threshold of 85%. Therefore, the first four principal components are selected as the basic parameters for subsequent model construction.

[0119] After the maximum variance rotation, the loading coefficients of each principal component and the original 17 variables are as follows:

[0120] Principal component 1 (PC1): Water supply index. Mainly reflects the availability of water, with soil moisture content (0.85), 90-day precipitation anomaly percentage (0.89), 120-day precipitation anomaly percentage (0.92), and 180-day precipitation anomaly percentage (0.87) as the main positive loading variables, and consecutive dry days (-0.82) as the main negative loading variable. This principal component explains 40.3% of the total variance, and a lower score indicates a lower water supply and a higher risk of drought.

[0121] Principal component 2 (PC2): Heat accumulation index. Mainly reflects the heat condition, with maximum temperature (0.83), temperature anomaly (0.86), and consecutive high temperature days (0.79) as the main positive loading variables. This principal component explains 22.1% of the total variance, and a higher score indicates a more severe heat accumulation and a higher risk of drought.

[0122] Principal component 3 (PC3): Water consumption index. Mainly reflects the intensity of evaporation, with average wind speed (0.75) and reference crop evapotranspiration ET0 (0.81) as the main positive loading variables. This principal component explains 11.8% of the total variance, and a higher score indicates faster water consumption and a higher risk of drought.

[0123] Principal component 4 (PC4): Time scale characteristic index. Mainly reflects the difference in precipitation characteristics at different time scales, with short-term precipitation anomaly percentage (30 days: 0.84, 60 days: 0.82) as the main positive loading variable and long-term precipitation anomaly percentage (300 days: -0.77) as the main negative loading variable. This principal component explains 10.9% of the total variance, reflecting the unevenness of the time distribution of precipitation, which is important for judging the trend of drought development.

[0124] (4) Principal component score calculation

[0125] According to formula (1), the principal component scores for each day are calculated using the standardized values of each variable and the loading coefficients. Specifically, for each day (7670 days) in Yanting County from 1990 to 2010, the scores of the four principal components are calculated, forming a 7670x4 principal component score matrix. This time period covers multiple typical drought events, and the principal component score sequence can fully reflect the development law of drought in Yanting County and the influence of hydro-meteorological conditions on drinking water risk. This matrix retains 85.1% of the information of the original 17 variables and significantly reduces the data dimension, providing a scientific basis for subsequent risk assessment model construction and parameter calibration.

[0126] Step 4: Risk assessment model construction

[0127] Based on the four principal components extracted from Yanting County, this embodiment constructs three different types of risk assessment models in parallel to achieve a comprehensive assessment of the risk of drinking water difficulties due to drought. The specific implementation process of the three models is as follows:

[0128] (1) Linear model construction

[0129] The linear model is constructed by using the weighted linear combination of principal components, and its mathematical expression is:

[0130]

[0131] In the formula: R linear is the risk value obtained by the linear model; w i is the weight coefficient of the i th principal component; PC i is the i th principal component.

[0132] (2) Nonlinear model construction

[0133] The nonlinear model introduces the interaction between principal components and is constructed in the form of a quadratic function, and its mathematical expression is:

[0134]

[0135] In the formula: R nonlinear is the risk value obtained by the nonlinear model; α i is the first-order term coefficient of the i th principal component; β ij is the coefficient of the interaction term between the i th principal component and the j th principal component, PC i and PC j are the i th and j th principal components, respectively.

[0136] (3) Machine learning model construction

[0137] The machine learning model is implemented using the K-means clustering algorithm. First, the four principal component data of the training set from 1990 to 2010 are analyzed by clustering to determine the optimal number of clusters as 2, representing the risk and non-risk states, respectively. The distance function uses the Euclidean distance to calculate the distance between the daily principal component vector and the risk cluster center. The smaller the distance value, the higher the risk. To make the risk value consistent with other models (the larger the value, the higher the risk), the reciprocal of the distance is used as the risk indicator:

[0138]

[0139] In the formula: R ml is the risk value obtained by the machine learning model; PC i is the i th principal component of the sample point; C 1iis the coordinate value of the i-th dimension of the risk cluster center.

[0140] Step 5: Model parameter calibration

[0141] Based on the training set data from 1990 to 2010, the parameters and risk thresholds of the three types of models are calibrated to determine the best risk determination standard.

[0142] S51: Model parameter initialization and optimization

[0143] (1) Linear model parameter calibration

[0144] The initial weight coefficients of the linear model are set in proportion to the variance contribution rate of each principal component, and are iteratively optimized by least squares method combined with gradient descent algorithm. After 500 iterations, the optimal weight coefficients are obtained: w1 = -0.45, w2 = 0.38, w3 = 0.32, w4 = 0.15. Among them, w1 is negative, which is consistent with the physical meaning of PC1 (water supply index); w2, w3, w4 are positive, which are consistent with the physical constraints.

[0145] (2) Nonlinear model parameter calibration

[0146] The nonlinear model determines the coefficients of the first-order terms and the interaction terms by stepwise regression method. After optimization calculation, the optimal parameter values are obtained:

[0147] First-order term coefficients: a1 = -0.42, a2 = 0.35, a3 = 0.30, a4 = 0.14

[0148] Main interaction term coefficients: β 12 = -0.18, β 13 = -0.15, β 23 = 0.12, β 14 = -0.08, β 24 = 0.06, and the absolute values of the remaining interaction term coefficients are all less than 0.05.

[0149] (3) Machine learning model parameter calibration

[0150] The cluster center coordinates determined by K-means clustering result are:

[0151] Cluster center C1 (at risk): [-1.25, 0.85, 0.65, 0.40];

[0152] Cluster center C2 (no risk): [0.95, -0.45, -0.55, -0.35];

[0153] The risk value sequence calculated by the machine learning model is Min-Max standardized, mapped to the [0, 1] interval, and the value closer to 1 indicates higher risk.

[0154] S52: Risk distribution analysis and initial threshold setting

[0155] (1) Probability distribution analysis of risk value sequences obtained by three types of models:

[0156] The risk value of the linear model shows an approximate normal distribution, with a mean of 0.05 and a standard deviation of 0.73.

[0157] The risk value of the nonlinear model shows a right-skewed distribution, with a mean of 0.08 and a standard deviation of 0.82.

[0158] The standardized risk value of the machine learning model shows a bimodal distribution, with a mean of 0.42 and a standard deviation of 0.18.

[0159] (2) Based on the historical data of drought-related drinking water difficulties in Yanting County from 1990 to 2010, set the initial risk threshold:

[0160] Linear model: r0_linear = 0.56 (corresponding to about 85% quantile);

[0161] Nonlinear model: r0_nonlinear = 0.58 (corresponding to about 85% quantile);

[0162] Machine learning model: r0_ml = 0.66 (corresponding to about 85% quantile).

[0163] S53: Construction of double-objective function and parameter optimization

[0164] (1) Based on the initial threshold, the risk value sequence from 1990 to 2011 is statistically analyzed year by year, and the number of high-risk days in each year is obtained:

[0165] Table 1: Simulation results of various models and actual reported risk population

[0166]

[0167] (2) Based on the above training set statistical results (1990-2010), calculate the correlation coefficient and root mean square error of the number of high-risk days and the actual reported risk population for each model:

[0168] The correlation coefficient of the linear model is 0.74, and the root mean square error is 0.70; the correlation coefficient of the nonlinear model is 0.88, and the root mean square error is 0.48; the correlation coefficient of the machine learning model is 0.84, and the root mean square error is 0.54. The comprehensive evaluation index J = correlation coefficient / root mean square error is constructed, and the calculation shows that the linear model is J = 1.06; the nonlinear model is J = 1.83; and the machine learning model is J = 1.565.

[0169] (3) Based on the nonlinear model (because its J value is the highest), parameter optimization adjustment is carried out:

[0170] Threshold optimization: Adjust the threshold r0 in the range of [0.55, 0.67], and find that when r0 = 0.61, the J value reaches the maximum 1.83 (as shown in Figure 2 ).

[0171] First-order coefficient optimization: Adjust the coefficients of α1 to α4 in the range of ±10%, and the optimized coefficients are α1 =-0.44, α2 = 0.37, α3 = 0.32, and α4 = 0.14.

[0172] Interaction term coefficient optimization: Adjust the main interaction term coefficient, and the optimized coefficient is β 12 =-0.20, β 13 =-0.17, and β 23 =0.14.

[0173] After 100 rounds of optimization, the comprehensive performance index J value of the nonlinear model is stable at 1.83, and it is determined as the optimal model.

[0174] S54: Model verification

[0175] The optimized nonlinear model (threshold r0 = 0.61) is applied to the risk assessment of the test set from 2011 to 2019, and the statistical results show that there is a significant correlation between the high-risk days and the actual reported risk population (correlation coefficient 0.91, root mean square error 0.33, and comprehensive evaluation index J value 2.76), indicating that the model has good generalization ability and prediction stability.

[0176] Step 6: Risk period determination

[0177] Based on the nonlinear model and its risk threshold optimized in step 5, this embodiment applies the run theory method to identify and merge the risk period of drought drinking difficulties in rural areas of Yanting County, and realizes fine risk monitoring (see Figure 3 ).

[0178] The specific implementation process is as follows:

[0179] (1) Initial risk day identification

[0180] First, the threshold value of 0.61 is applied to the optimized nonlinear model risk value sequence for binary classification judgment, and the initial risk state sequence is obtained (1 represents risk and 0 represents no risk). The drought process of Yanting County from April to May 2013 is analyzed, and the initial identification result is shown as the blue line "original value" in the figure. It can be observed that from April 8 to May 8, the risk state fluctuates several times, and there are several short-term interruptions.

[0181] (2) Risk period merging implementation

[0182] The merging rule of the run theory is applied to process the initial risk state sequence: the period of 3 consecutive days of risk value exceeding the threshold value of 0.61 is confirmed as the initial high-risk group; the interval between adjacent high-risk groups is analyzed, such as the interval of 1 day between April 13 and 15, which meets the merging condition of "interval ≤ 3 days"; according to this rule, the high-risk groups from April 8 to May 8 are iteratively merged into a complete risk period. The red line "corrected value" in the figure shows the result after merging. It can be clearly seen that the originally intermittent multiple high-risk periods are merged into a continuous complete risk period.

[0183] (3) Verification of merging termination condition

[0184] The boundaries of the merged risk period are confirmed: the risk values of the 7 days before the front end of the risk period (April 1-7) do not exceed the threshold value; the risk values of the 7 days after the back end of the risk period (May 9-15) also do not exceed the threshold value. The termination condition of "no risk in the 7 days before and after" is met, so it is confirmed that April 8 to May 8 is a complete risk period, with a total duration of 31 days. It is basically consistent with the reported risk period of drinking water in 2013.

[0185] (4) Actual impact verification

[0186] The risk period identified by the model is compared with the actual reporting data of Yanting County in 2013: the actual reporting shows that on April 10, 2013, the first report of local rural drinking water difficulty was reported, involving about 18,700 people; on April 30, the number of reports reached a peak of 2.15; on May 10, the drought situation was basically alleviated. The risk period identified by the model (April 8 to May 8) is basically consistent with the actual reporting situation.

[0187] Through the above risk period determination method, the daily risk assessment result is successfully converted into an actual operable risk period, providing a scientific basis for the monitoring and early warning of drought drinking water difficulty risk in Yanting County, significantly improving the timeliness and accuracy of risk management.

[0188] It should be pointed out finally that the above merely serves to illustrate the technical solutions of the present application but not to limit the same, and although the present application has been described in detail with reference to the preferred arrangement, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A rural drought drinking difficulty risk dynamic assessment and early warning method, characterized in that, The method comprises the following steps: Step 1: Multi-source data acquisition and preprocessing Collect basic data, including environmental variable data and daily input variables; the environmental variable data includes population density, elevation, slope, historical average evapotranspiration and historical average precipitation; the daily input variables include maximum temperature, temperature anomaly, soil moisture, consecutive high temperature days, average wind speed, reference crop evapotranspiration and consecutive rainless days, and precipitation anomaly percentage of multiple time scales; After all data are subjected to missing value interpolation, abnormal value identification and elimination, and standardization treatment, a multi-source data set consistent in time and space is formed; Step 2: Determination of physical constraint relationship According to the different effects of different input variables on drought risk, the physical constraint relationship between each input variable and the risk of drinking water shortage due to drought in rural areas is clearly defined, specifically: For factors including maximum temperature, temperature anomaly and consecutive rainless days, the increase in the value will lead to an increase in drought risk, marked as positive relationship (+); For factors including soil moisture, consecutive high temperature days and precipitation anomaly percentage, the increase in the value will help to alleviate the drought risk, marked as negative relationship (-); For complex factors including average wind speed and reference crop evapotranspiration, since they may exhibit different influence mechanisms under different geographical environments and climate conditions, they are marked as uncertain relationship (?); Step 3: Data dimension reduction and variable structure analysis After determining the physical constraint relationship, the multi-source data set obtained is processed; Dimension reduction is performed using principal component analysis: a correlation matrix is constructed from the standardized environmental variable data and daily input variables, the eigenvalues and eigenvectors are calculated, and the principal components are extracted in order of eigenvalue size; The key formula for principal component extraction is as follows: where: PC i is the i-th principal component; a ij is the loading coefficient of the j-th original variable in the i-th principal component; X j is the j-th original variable after standardization; and m is the total number of original variables. Step 4: Construction of risk assessment model Three types of models are constructed in parallel: linear model, nonlinear model and machine learning model; all three types of models use the dimension reduction idea to further reduce the multiple principal components extracted in step 3 to obtain a one-dimensional sequence as the final risk value; (1) Linear model The linear model is constructed using linear combination of principal components, and its mathematical expression is: wherein: R linear is the risk value from the linear model; w i is the weight coefficient of the i-th principal component; PC i is the i-th principal component; n is the number of retained principal components; (2) Nonlinear model The nonlinear model uses a quadratic function form and introduces the interaction between principal components, and its expression is: where: R nonlinear is the risk value from the nonlinear model; a i is the coefficient of the first order term of the ith principal component; β ij is the coefficient of the interaction term of the ith principal component and the jth principal component; PC i and PC j are the sequences of the ith and jth principal components, respectively; (3) Machine learning model The machine learning model uses a clustering algorithm to cluster the sample points in the principal component space, and calculates the distance from each sample to the cluster center as a risk measure: R ml = dist(PC sample ,C k ) (4) wherein: R ml is the risk value obtained by the machine learning model; dist denotes the distance function; PC sample is the coordinate of the sample point in the principal component space; C k is the corresponding cluster center; Step 5: Model parameter calibration Based on the training set data, the parameters and risk threshold values of the three types of models constructed in step 4 are calibrated to determine the best risk determination standard, specifically including: S51: Model parameter initialization and optimization The initial weight coefficients of the linear model are set in proportion to the variance contribution rate of each principal component, and are iteratively optimized by least squares method combined with gradient descent algorithm; the one-term coefficients and interaction term coefficients of the nonlinear model are determined by stepwise regression method; the cluster center coordinates of the machine learning model are determined by K-means clustering, and the risk value sequence calculated is standardized to map to the [0, 1] interval; S52: Risk distribution analysis and initial threshold setting The probability distribution of the risk value sequence obtained by the three types of models is analyzed, and the initial risk threshold is set based on historical data as the initial judgment boundary between risk and non-risk states; S53: Double objective function construction and parameter optimization A double objective function evaluation system is established, considering two optimization objectives: first, maximize the Pearson correlation coefficient between the annual high-risk day sequence and the reported risk population number, to ensure the consistency of the trend; second, minimize the root mean square error between the annual high-risk day sequence and the reported risk population number, to ensure the closeness in value; A comprehensive evaluation index J = correlation coefficient / root mean square error is constructed, and the model parameters and risk threshold are adjusted iteratively to find the optimal parameter combination that makes J value optimal; through multiple rounds of optimization, the optimal model and its parameter settings are determined; S54: Model verification The optimized model is applied to the test set data, the correlation between the high-risk day number and the actual reported risk population number is calculated, and the comprehensive evaluation index is calculated to verify the generalization ability and prediction stability of the model; Step 6: Risk period determination Based on the run theory, by identifying and merging high-risk periods, accurate monitoring and early warning of the risk of drinking water difficulties due to drought are achieved.

2. The rural drought risk dynamic assessment and early warning method according to claim 1, characterized in that, In step 1, the elevation and slope data are extracted from the SRTM 90m resolution digital elevation model, with an accuracy of not less than 90m; the population density uses LandScan global population density data; the historical average evapotranspiration and precipitation are calculated based on more than 30 years of observation records from meteorological stations; the daily input variables are obtained through automatic weather stations, rain gauge networks and remote sensing monitoring.

3. The rural drought risk dynamic assessment and early warning method according to claim 1, characterized in that, In step 3, when performing principal component analysis, start from principal component 1 and sequentially accumulate to principal component n until the cumulative variance contribution rate reaches 85%.

4. The rural drought risk dynamic assessment and early warning method according to claim 1, characterized in that, In step 5 S52, the initial risk threshold corresponds to the high quantile of the risk value distribution.

5. The rural drought risk dynamic assessment and early warning method according to claim 1, characterized in that, In step 6, the idea of the run theory specifically includes: Analyze the risk value sequence obtained in step 5, if the risk value exceeds the set threshold for 3 consecutive days, it is identified as a group of initial high-risk days; if the interval between two high-risk day groups is less than or equal to 3 days, the two high-risk day groups are merged into a longer high-risk period; The merging process is iterated until there are no high-risk days within 7 days before and after the start of the high-risk period.

Citation Information

Patent Citations

  • Typhoon disaster risk assessment and dynamic forecasting method

    CN116757305A

  • Dynamic risk control method and device, equipment and storage medium

    CN120163653A