A method for screening drought cause indicators
By constructing a comprehensive drought index and screening drought cause indicators using SHAP model, the problem of ignoring the lower surface factors and spatial characteristics in the existing technology is solved, and the accuracy of drought prediction and monitoring accuracy are improved.
Patent Information
- Application Number
- CN202411695284.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-11-25
AI Technical Summary
When obtaining the causal relationship between observation data and drought, the prior art ignores the impact of the lower surface factor on drought events, and cannot dig out the relationship between agricultural, hydrological drought conditions and observation data, and ignores spatial characteristics, making it difficult to reflect the spatial evolution characteristics of the drought process.
By obtaining multi-scale drought data on time and space, extracting drought characteristics, constructing a comprehensive drought index, analyzing the data set using the drought evaluation model and SHAP model, screening drought cause indicators, and considering the lower surface factors and spatiotemporal characteristics.
It improves the accuracy of drought prediction, can more effectively reflect the comprehensive situation of various drought types in agriculture, meteorology and hydrology, provides a more comprehensive and accurate basis for drought monitoring, and realizes quantitative cause analysis of the factors affecting drought events.
Smart Images

Figure CN119624176B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of screening drought cause indicators, and specifically, to a method for screening drought cause indicators. Background Art
[0002] Drought has the characteristics of long duration, serious disaster losses, wide influence range, etc., and is a natural disaster that seriously affects agricultural production and human livelihood. In recent years, under the background of global climate change, the development process of local drought is accelerating, such as the intensification of sudden drought, and these situations will seriously hinder the development of human society and threaten food security.
[0003] Using drought indicators to monitor and predict drought events is an effective method for preventing drought disasters. In the field of drought monitoring, currently, the causal relationship between each observation data and drought is usually analyzed to select the most suitable features for subsequent drought prediction, understand the physical process principle of comprehensive drought events, and improve the accuracy of drought prediction.
[0004] Regarding the causal relationship between observation data and drought, currently, most use easily obtainable meteorological drought monitoring indicators such as SPI, SPEI, and MCI, and methods such as correlation coefficient and model simulation are used to complete it, but this method has the following drawbacks:
[0005] 1) Usually, the influence of underlying surface factors on drought events is ignored, and the relationship between agricultural and hydrological drought conditions and observation data cannot be explored;
[0006] 2) To a certain extent, the temporal characteristics of the drought process are considered, but the spatial characteristics are ignored, and it is difficult to reflect the spatial evolution characteristics of the drought process. Summary of the Invention
[0007] In order to solve the problems that in obtaining the causal relationship between observation data and drought, the existing methods ignore the influence of underlying surface factors on drought events, cannot explore the relationship between agricultural and hydrological drought conditions and observation data, and ignore spatial characteristics, making it difficult to reflect the spatial evolution characteristics of the drought process, the present invention provides a method for screening drought cause indicators, and the method includes:
[0008] Obtain multi-scale drought data in terms of time and space, and extract the drought characteristics of the multi-scale drought data; based on the drought characteristics, construct a comprehensive drought index; obtain a data set based on the comprehensive drought index and the multi-scale drought data; analyze the data set based on a drought assessment model to obtain a loss value, input the loss value into a SHAP model to obtain SHAP values, and screen drought cause indicators based on the SHAP values to obtain a screening result.
[0009] This method incorporates the arid characteristics of the underlying surface factor, taking into account more comprehensive factors to improve the accuracy of drought prediction; a comprehensive drought index is constructed using multiple indicators, which can reflect the comprehensive drought index of various drought types in agriculture, meteorology, and hydrology. This index can integrate the characteristic information of different drought types, more effectively represent the comprehensive situation of drought, and provide a more comprehensive and accurate basis for drought monitoring; the drought assessment model is used to extract features, spatio-temporal features are extracted through the ConvLSTM module, and then the VGG module and the attention mechanism module are added, which can adaptively assign weights to the extracted spatio-temporal features, enhance the ability to extract feature information, improve the model's attention to the channels and internal features of spatio-temporal features, and enhance the prediction accuracy of the model for drought events; the SHAP model is constructed based on the loss value, and combined with the drought level to screen out drought features with higher contribution degrees, which can obtain more effective drought impact factors, and the SHAP model can calculate the marginal contribution to achieve a quantitative causal analysis of the impact factors of drought events. The accurately, efficiently and spatio-temporally featured drought index is screened out, which can provide effective assistance for the subsequent monitoring and prediction of drought events.
[0010] The underlying surface factor refers to the surface located at the bottom of the atmosphere that can interact with the atmosphere during the heat and moisture exchange processes. The underlying surface factors mainly include types such as soil, grassland, and water bodies, and these factors play an important role in climate formation.
[0011] Further, the specific steps for constructing the comprehensive drought index include: constructing a standardized anomaly index based on the drought characteristics, performing a dimensionless treatment on the standardized anomaly index to obtain the first data, and obtaining a correlation coefficient matrix based on the first data; obtaining the principal components based on the correlation coefficient matrix and the standardized anomaly index; obtaining the contribution rate of the principal components based on the correlation coefficient matrix, and obtaining the weight coefficients of the principal components based on the contribution rate; and obtaining the comprehensive drought index based on the principal components and the weight coefficients.
[0012] Constructing a comprehensive drought index that can reflect various drought types in agriculture, meteorology, and hydrology can uncover the relationships between agricultural and hydrological drought conditions and the observed data.
[0013] Further, in the order of invocation, the drought assessment model sequentially includes: the ConvLSTM module and the VGG module, and the VGG module includes the CBAM module.
[0014] Based on the ConvLSTM model that extracts spatio-temporal features, the VGG model and the attention mechanism module are added to improve the model's attention to the channels and internal features of spatio-temporal features.
[0015] Further, the specific steps for obtaining the loss value include: extracting spatio-temporal features from the dataset based on the ConvLSTM module, performing a convolution operation on the spatio-temporal features by the VGG module to obtain a first feature; the CBAM module obtaining an attention map based on the spatio-temporal features and the first feature, and obtaining the loss value based on the attention map.
[0016] Further, the specific steps for obtaining the loss value based on the attention map include: obtaining a regression loss based on the attention map and a first loss function, obtaining a classification loss based on the attention map and a second loss function, and obtaining the loss value based on the regression loss and the classification loss.
[0017] Construct an improved Smooth L1Loss-Cross Entropy Loss hybrid loss function, comprehensively consider the regression and drought classification characteristics of the model, and fuse spatio-temporal features in the calculation to better measure the performance of the model.
[0018] Further, the specific steps for obtaining the SHAP value include: presetting a drought level, weighting the loss value based on the drought level to obtain a weighted loss value; obtaining an output result based on the weighted loss value, obtaining an input condition based on the drought feature, and inputting the output result and the input condition into the SHAP model to obtain the SHAP value of the drought feature.
[0019] Construct a GradientExplainer interpreter based on the SHAP theory. For the benchmark model parameters of the interpreter, use the loss value that fuses multi-scale drought features as the explanation target to screen out drought features with higher contribution degrees, so as to solve the problem that it is difficult to calculate the marginal contribution for a convolutional model of multi-source spatio-temporal data.
[0020] Further, the specific steps for obtaining the screening result include: dividing the SHAP value based on the drought level to obtain a level SHAP value; obtaining the average value of the level SHAP value based on the time dimension and the space dimension to obtain an average SHAP value; screening drought cause indicators based on the SHAP value and the average SHAP value to obtain the screening result.
[0021] For each drought level, collect the corresponding SHAP values, and take the mean in the time dimension and the space dimension to obtain the comprehensive average contribution of each feature and the average contribution to a specific drought level. By visualizing the spatial distribution of the cumulative values of the features before the occurrence of drought, verify the spatial distribution of the drought event, and analyze the significance of different features for the occurrence of the drought event.
[0022] Further, the method further includes: obtaining the correlation degree between the comprehensive drought index and the meteorological drought index and the agricultural drought index, and obtaining a verification result based on the correlation degree.
[0023] Verify the correlation between the selected comprehensive drought index and historical drought events from the spatio-temporal characteristics, and analyze that the selected drought index can better monitor and predict drought events.
[0024] Further, the standardized anomaly index includes a standardized precipitation anomaly index, a standardized soil moisture anomaly index, and a standardized surface runoff anomaly index.
[0025] Further, the first loss function is the SmoothL1Loss loss function, and the second loss function is the cross-entropy loss function.
[0026] Comprehensively consider the regression and classification characteristics of the model, and integrate spatio-temporal features in the calculation to better measure the performance of the model. And the hybrid loss function can improve the robustness to outliers while maintaining the model accuracy.
[0027] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:
[0028] 1. Incorporate the drought characteristics of the underlying surface factor, with more comprehensive considerations, to improve the accuracy of drought prediction; construct a comprehensive drought index with multiple indicators, which can reflect the comprehensive drought index of multiple drought types in agriculture, meteorology, and hydrology. This index can integrate the characteristic information of different drought types, more effectively characterize the comprehensive situation of drought, and provide a more comprehensive and accurate basis for drought monitoring.
[0029] 2. Use the drought assessment model to extract features, extract spatio-temporal features through the ConvLSTM module, and then add the VGG module and the attention mechanism module, which can adaptively assign weights to the extracted spatio-temporal features, enhance the ability to extract feature information, improve the model's attention to the channels and internal features of spatio-temporal features, and enhance the prediction accuracy of the model for drought events.
[0030] 3. Construct the SHAP model with the loss value, and combine the drought level to screen for more contributing drought characteristics, which can obtain more effective drought impact factors. And the SHAP model can calculate the marginal contribution, realize the quantitative causal analysis of the impact factors of drought events, screen for an accurate, efficient and spatio-temporal characteristic drought index, and can provide effective help for the subsequent monitoring and prediction of drought events.
[0031] 4. Construct an improved Smooth L1Loss-Cross Entropy Loss hybrid loss function, comprehensively consider the regression and drought classification characteristics of the model, and integrate spatio-temporal features in the calculation to better measure the performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of the present invention, and do not limit the embodiments of the present invention;
[0033] Figure 1 is a schematic flow chart of a method for screening drought cause indicators in the present invention;
[0034] Figure 2 is a schematic structural diagram of a drought assessment model in the present invention; wherein, X1, X2…X t represent the input data of the current time series, C0, C1…C t represent the unit states of the previous moment, H0, H1…H t represent the hidden layer states of the previous moment. After inputting C0 and H0 corresponding to the X1 moment, zero vectors are used for initialization. ConvLSTM represents the ConvLSTM module, conv+relu represents the combination of the convolutional layer and the ReLu activation function, max pooling represents the max pooling, CBAM represents the attention mechanism, Sigmoid represents the Sigmoid activation function, Fully connectes+relu represents the combination of the fully connected layer and the ReLu activation function, and SDCI_pre represents the comprehensive drought index;
[0035] Figure 3 is a schematic diagram of calculating the feature importance of the drought disaster assessment model in a typical drought event in 2010 using the SHAP model in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0037] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0038] Embodiment 1
[0039] Reference Figures 1 - 3, this embodiment provides a method for screening drought cause indicators, and the method includes:
[0040] Obtain multi-scale drought data in terms of time and space, and extract the drought characteristics of the multi-scale drought data; for example, use China Global Land Reanalysis data (CRA), Gross Domestic Product data (GDP), meteorological station data, land use data, disaster data, elevation data, etc. CRA data, land use data, and elevation data all belong to raster data, and deep learning can be used for image convolution operations to improve the prediction accuracy of drought events.
[0041] In this embodiment, the multi-scale drought data can also be preprocessed, and the drought characteristics are extracted based on the preprocessed multi-scale drought data. The preprocessing can include format conversion, normalization, and mean value filling, etc. For example, convert the collected NC format data to tif format data, unify the spatial resolution of the data to 34 km, and perform normalization and mean value filling on the data.
[0042] In this embodiment, the drought characteristics can include soil heat flux, latent heat flux, instantaneous surface temperature, vegetation transpiration, total canopy water storage, instantaneous soil temperature, specific humidity at 2 m, wind speed at 10 m, and bare soil evaporation rate, etc.
[0043] The standardized anomaly index is a method for comparing the deviations between different data sets. It obtains a standardized anomaly value by dividing the distance of each data point from the mean value by the standard deviation. This value represents the degree of deviation of the data point from the mean value in the entire data set.
[0044] Using the PCA (Principal Component Analysis) method, based on the drought characteristics, construct a comprehensive drought index; the specific steps include: constructing a standardized anomaly index based on the drought characteristics, and the standardized anomaly index includes a standardized precipitation anomaly index, a standardized soil moisture anomaly index, and a standardized surface runoff anomaly index; its calculation method can be:
[0045]
[0046] Among them, StX m,n represents the standardized anomaly index, m and n respectively represent the nth day in the mth year, and X m,n represents the target data sought on the nth day of the mth year, represents the average value of the nth day in the set number of years, and σ n represents the standard deviation of the nth day in the inter-annual time series in the set number of years.
[0047] Nondimensionalization is an important processing concept in scientific research. Through an appropriate variable substitution, part or all of the units of an equation involving physical quantities are removed to simplify experiments or calculations. Nondimensionalization is an important step in the comprehensive evaluation process, used to eliminate the influence of dimensions so that indicators of different magnitudes can be compared and integrated.
[0048] Perform nondimensionalization on the standardized anomaly index to ensure that the same scale exists between each variable, and obtain the first data; its calculation method can be:
[0049]
[0050] where, x′ ij represents the j-th data value in the i-th element after standardization, x ij represents the j-th data value in the i-th element, σ i represents the standard deviation of the i-th element, represents the average of the i-th element, and both i and j represent serial numbers.
[0051] Obtain the correlation coefficient matrix based on the first data; calculate the correlation coefficient matrix between variables, and form a two-dimensional matrix with the standardized multi-dimensional data. Subsequently, calculate the covariance matrix based on this matrix, that is, the correlation coefficient matrix. Its calculation method can be:
[0052]
[0053] where, S ij represents the correlation coefficient matrix, represents the series mean of the i-th element after standardization, represents the series mean of the h-th element after standardization, x′ ij represents the j-th data value in the i-th element after standardization, x′ hj represents the j-th data value in the h-th element after standardization, m represents the number of samples, and both i and j represent serial numbers.
[0054] Obtain the principal components based on the correlation coefficient matrix and the standardized anomaly index; its calculation method can be:
[0055] D i = v1P r + v2P f + v3S m ; (4)
[0056] where, D i represents the principal component, P r represents the standardized precipitation anomaly index, P fIndicates the standardized surface runoff anomaly index, S m Indicates the standardized soil moisture anomaly index. v1, v2, and v3 respectively represent the eigenvectors of the correlation coefficient matrices of the precipitation index, surface runoff index, and soil moisture index, and i represents the serial number.
[0057] Obtain the contribution rate of the principal component based on the correlation coefficient matrix; its calculation method can be:
[0058] |S ij -λ i I| = 0; (5)
[0059]
[0060] Among them, S ij Represents the correlation coefficient matrix, I represents the identity matrix, λ i Represents the eigenvalue, g i Represents the contribution rate, p represents the number of eigenvalues, k represents the number of principal components, and i represents the serial number.
[0061] Obtain the weight coefficient of the principal component based on the contribution rate; its calculation method can be:
[0062]
[0063] Among them, w i Represents the weight coefficient, gi represents the contribution rate, p represents the number of eigenvalues, and i represents the serial number.
[0064] Obtain the comprehensive drought index based on the principal component and the weight coefficient. Its calculation method can be:
[0065]
[0066] Among them, Index represents the comprehensive drought index, w i Represents the weight coefficient, D i Represents the principal component, p represents the number of eigenvalues, and i represents the serial number. The lower the Index value, the more severe the drought, and the higher the value, the lower the severity of the drought.
[0067] Obtain a data set based on the comprehensive drought index and the multi-scale drought data;
[0068] Analyze the data set based on the drought assessment model to obtain a loss value. Among them, in the order of call, the drought assessment model successively includes: the ConvLSTM module and the VGG module, and the VGG module includes the CBAM module.
[0069] ConVLSTM is an architecture that combines Convolutional Neural Network (CNN) and Long Short-Term Memory Network (LSTM) for processing time-series data. Different from traditional LSTM, ConvLSTM applies convolutional operations at each time step, which helps capture spatial information in time-series data.
[0070] VGG is a deep Convolutional Neural Network (CNN) proposed by Karen Simonyan and Andrew Zisserman from the Visual Geomety Group of the University of Oxford in 2014. VGG is famous for the simplicity and consistency of its structure. It mainly consists of stacked small 3x3 convolutional kernels and 2x2 max pooling layers.
[0071] Convolutional Block Attention Module (CBAM) is a feed-forward convolutional neural network attention module. Given an intermediate feature map, CBMA sequentially infers attention maps along two separate dimensions (channel and space), and then multiplies the attention maps by the input feature map for adaptive feature refinement. Moreover, CBAM is a lightweight general module that can be seamlessly integrated into any CNN architecture.
[0072] ConvLSTM extracts spatial pattern information from images using convolutional operations, extracts spatio-temporal features of the input data, and converts them into tensors. At the same time, it analyzes time information in a similar way to traditional LSTM through convolutional gates. ConvLSTM can achieve this goal with fewer parameters and higher computational efficiency. Replace the fully connected layer in the ConvLSM model with VGG, and in the VGG model, introduce the CBAM module, add it after the convolutional layer in VGG, first extract channel attention features, then extract spatial attention features, and finally put the data with extracted features back into the convolutional layer of VGG for operation. Finally, VGG outputs the prediction result. Using the improved SmoothL1Loss-CrossEntropyLoss as the loss function during the training of the drought assessment model can solve the problem that it is difficult to calculate the marginal contribution for convolutional models of multi-source spatio-temporal data.
[0073] The specific steps to obtain the loss value include: extracting spatio-temporal features in the dataset based on the ConvLSTM module; using ConvLSTM to extract temporal and spatial features in the dataset. Its calculation method can be:
[0074]
[0075] where, f t represents the output result of the forget gate, x t represents the input at the current moment, h t-1represents the hidden layer state at the previous moment, C t-1 represents the cell state at the previous moment, W xf 、W hf and W cf all represent the weights of the forget gate, b f represents the bias term of the forget gate, σ represents the activation function, i t and respectively represent the two output results of the input gate, W xi 、W hi and W ci all represent the weights of the input gate, b i represents the bias term of the input gate, tanh represents the activation function, W xc and W hc all represent the weights of the cell state, b c represents the bias term of the cell state, C t represents the cell state at the current moment, ot represents the output result of the output gate, W xo 、W ho and W co all represent the weights of the output gate, b o represents the bias term of the output gate, h t represents the hidden layer state at the current moment, represents element-wise multiplication, also known as the Hadamard product, * represents convolution.
[0076] The VGG module performs a convolution operation on the spatio-temporal features to obtain the first feature; the VGG is used to perform a convolution operation on the spatio-temporal feature tensor extracted by the ConvLSTM again to extract their spatial features and enhance the model's ability to extract spatial features.
[0077] The CBAM module obtains an attention map based on the spatio-temporal features and the first feature, and obtains the loss value based on the attention map. The CBAM convolutional attention module is introduced into the VGG module to calculate the attention map sequentially from two different dimensions of channels and space, and multiply the obtained attention map by the input feature map to perform adaptive feature refinement. Through this step, the attention of the entire model to different channels can be improved, and the prediction accuracy of the model for drought events is enhanced.
[0078] Since the prediction result is both the regression value based on the comprehensive drought index and can be used to classify the drought level according to the value range where the regression value is located, the loss function of the drought assessment model is improved.
[0079] Among them, the specific steps for obtaining the loss value based on the attention map include:
[0080] The regression loss is obtained based on the attention map and the first loss function; for the regression loss, the SmoothL1Loss loss function with a penalty term coefficient is used for calculation. For the prediction results with a greater distance between the regression value range and the label, a larger penalty term coefficient is adopted to reflect the loss impact of the drought level division based on the regression value range on the prediction results. Its calculation method can be:
[0081] l(x,y) = L = {l1,…,l N} T ; (10)
[0082]
[0083] where l() represents the SmoothL1Loss loss function, L represents the loss vector, l N represents the loss value of a single sample, T represents the transpose, x represents the predicted value of the model, y represents the true label value, and SmoothL1Loss() represents the SmoothL1Loss loss function.
[0084] The classification loss is obtained based on the attention map and the second loss function; for the classification loss, the Cross Entropy Loss loss function with a penalty term coefficient is used for calculation. The regression value after drought level division is converted into a one-hot value as the score (Logits), and the penalty coefficient is weighted based on the distance between the regression value and the label, and the cross entropy is used to calculate the weighted result. Its calculation method can be:
[0085]
[0086] where H() represents the cross entropy loss function, y1 represents the one-hot encoded vector of the true label, x1 represents the probability distribution vector predicted by the model, C represents the number of categories, y i represents the one-hot encoded vector of the i-th label, x i represents the probability of the i-th category predicted by the model, and i represents the serial number.
[0087] Based on the regression loss and the classification loss, the loss value is obtained.
[0088] Among them, the first loss function is the SmoothL1Loss loss function, and the second loss function is the cross entropy loss function.
[0089] The Smooth L1 Loss is a quadratic function when |x - y| < 1, which makes it more sensitive to small errors and can better fit the data. When |x - y| ≥ 1, it is a linear function, which results in a relatively smaller penalty for large errors (outliers) and shows better robustness in the presence of outliers. When dealing with regression tasks, the Smooth L1 Loss treats all errors in a balanced manner and does not overly focus on a certain part of the data. This is particularly beneficial for small-class samples in imbalanced datasets and can prevent the model from overfitting to majority-class samples.
[0090] The logarithmic function in the cross-entropy loss function imposes a larger penalty on small probability values and a smaller penalty on large probability values. This means that when the model's predicted probability for a certain sample is very low (i.e., the prediction is incorrect), the loss will increase significantly, but it will not overly amplify the impact of outliers. The cross-entropy loss function can also handle imbalanced data by weighting the losses of different classes. For example, for minority-class samples, higher weights can be assigned to make the model pay more attention to the prediction of minority-class samples. Therefore, combining these two loss functions has unique advantages in dealing with outliers and imbalanced data, which is very important for handling complex drought data.
[0091] The Smooth L1 Loss is a commonly used loss function mainly for regression tasks. It is a combination of the L1 Loss and the L2 Loss, aiming to balance the advantages of both while overcoming their disadvantages.
[0092] By combining the losses of regression and classification tasks, this loss function not only reflects the model's performance in drought prediction but also the accuracy of the model in classifying different drought levels, thus providing a more comprehensive assessment of drought disasters.
[0093] Input the loss value into the SHAP model to obtain SHAP values. Since the drought characteristics in different regions may vary, by weighting the loss function that combines multiple features, the interpreter can better adapt to the drought characteristics of different regions and improve the generalization ability of the model. At the same time, because when using certain gradient calculation methods in the calculation process of SHAP, the output must be a scalar, using the loss value as the output of the interpretation model can solve the problem that it is difficult to calculate the marginal contribution of the convolutional model for multi-source spatio-temporal data.
[0094] Among them, the specific steps to obtain SHAP values include: presetting drought levels, weighting the loss value based on the drought levels to obtain weighted loss values;
[0095] The prediction results in different types of loss functions are weighted according to the drought level division. Among them, for the regression loss function, the dynamic weighting method is adopted, and the weight is dynamically adjusted according to the error size between the predicted value and the true value. The weighting rule can be:
[0096] R_Loss = (Loss1 * 0.1 + 1) · w; (13)
[0097] Among them, R_Loss represents, Loss1 represents the loss value output by the Smooth L1 Loss function, and w represents the fixed weight corresponding to the error distance between the predicted value and the true value after the drought level division.
[0098] For the classification loss function, the exponential function is used to adjust the weight, so that the weight of the category with a larger error increases faster. The weighting rule can be:
[0099] C Loss = Loss2 · (1 + e β·(d-1) ), β > 0; (14)
[0100] Among them, C_Loss represents, Loss2 represents the loss value output by the cross-entropy loss function, β represents the parameter controlling the weight growth rate, and d represents the drought level distance between the predicted value and the true value.
[0101] Through the weighting mechanism, the sensitivity to different drought levels is enhanced, ensuring that the model pays more attention to those drought levels with larger prediction errors, making the model more flexible and effective in dealing with drought data, which helps to improve the accuracy and practicality of the model in actual applications; dynamically adjusting the weight according to the performance of the model further optimizes the model performance and ensures that the model can be continuously improved during the training process; the drought characteristics in different regions may be different. By combining multiple characteristics and weighting them, the hybrid loss function can better adapt to the drought characteristics in different regions and improve the generalization ability of the model.
[0102] Based on the weighted loss value, the output result is obtained. Based on the drought characteristics, the input conditions are obtained. The output result and the input conditions are input into the SHAP model to obtain the SHAP value of the drought characteristics. The calculation method can be:
[0103]
[0104] Among them, φj() represents the SHAP value of feature j for the input x2, f represents the prediction function of the model, x2 represents the input sample, that is, the input feature vector of the model, S represents the feature subset that does not contain the feature, M represents the total number of features, N represents the set of all features, |S| represents the number of elements in the set S, f x (S) represents the predicted value of the model when only using the features in the set S, fx (S ∪ {j}) represents the predicted value of the model when only using the features in set S and feature j. The exclamation mark represents factorial, and j represents a feature.
[0105] SHAP (SHapley Additive exPlanations) is a game-theoretic method for interpreting the output of any machine learning model. It combines optimal credit assignment with local explanations using the classical game-theoretic Shapley value and its related extensions. The closer the SHAP value is to zero, the smaller the contribution of the feature to the prediction; while the SHAP value far from zero indicates a greater contribution of the feature to the prediction. To perform model interpretation in SHAP, an explainer needs to be created first. SHAP supports many types of explainers, such as deep, gradient, kernel, tree, and sampling explainers.
[0106] The SHAP model provides an intuitive way to understand how features affect model predictions and provides interpretability for the screening process of drought factors in this method.
[0107] During the visualization process of SHAP values, to further analyze the feature importance under different drought levels, the prediction results of the model are divided according to the defined drought intervals. For each drought level, the corresponding SHAP values are collected and averaged in the time dimension and the space dimension to obtain the comprehensive average contribution of each feature and the average contribution for a specific drought level.
[0108] Based on the SHAP values, drought cause indicators are screened to obtain the screening results. The specific steps include: based on the drought level, dividing the SHAP values to obtain the level SHAP values; for example, dividing the prediction results of the model into no drought, light drought, moderate drought, severe drought, and extreme drought, and thus dividing the SHAP values; based on the time dimension and the space dimension, obtaining the average value of the level SHAP values to obtain the mean SHAP values; based on the SHAP values and the mean SHAP values, screening the drought cause indicators to obtain the screening results.
[0109] Example Two
[0110] Based on Example One, in this example, the method further includes:
[0111] Obtaining the correlation degree between the comprehensive drought index and the meteorological drought index and the agricultural drought index, and obtaining the verification result based on the correlation degree.
[0112] If the SPI (Standardized Precipitation Index) representing meteorological drought is selected as the verification target, the relationship between the SDCI (Synthetic Drought Index) and the SPI is analyzed, and the correlation coefficient between the SDCI and the SPI within a certain number of years is calculated according to the Pearson correlation analysis; the Pearson correlation coefficient between the soil relative humidity index representing agricultural drought and the SDCI is calculated; the change trend of the total surface water resources is selected for analysis, and the change trend of the SDCI drought area within a certain number of years is statistically analyzed.
[0113] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0114] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for screening drought cause indicators, characterized in that: The method comprises: Acquire multi-scale drought data in time and space, and extract drought characteristics of the multi-scale drought data, wherein the drought characteristics include soil heat flux, latent heat flux, instantaneous value of surface temperature, vegetation transpiration, total canopy water storage, instantaneous value of soil temperature, 2-meter specific humidity, 2-meter air temperature, total evapotranspiration, latent evapotranspiration, instantaneous value of surface radiation temperature, surface albedo, sensible heat flux, surface air pressure, 10-meter wind speed, and bare soil evaporation rate; Based on the drought characteristics, construct a comprehensive drought index; based on the comprehensive drought index and the multi-scale drought data, obtain a data set; Analyze the data set based on the drought assessment model to obtain a loss value, input the loss value into the SHAP model to obtain a SHAP value, and screen drought cause indicators based on the SHAP value to obtain a screening result; The specific steps of constructing a comprehensive drought index include: constructing a standardized anomaly index based on the drought characteristics, performing dimensionless processing on the standardized anomaly index to obtain first data, and obtaining a correlation coefficient matrix based on the first data; obtaining a principal component based on the correlation coefficient matrix and the standardized anomaly index; obtaining a contribution rate of the principal component based on the correlation coefficient matrix, and obtaining a weight coefficient of the principal component based on the contribution rate; and obtaining the comprehensive drought index based on the principal component and the weight coefficient; According to the calling order, the drought assessment model includes: a ConvLSTM module and a VGG module, and the VGG module includes a CBAM module; The specific steps of obtaining the loss value include: extracting the spatiotemporal features in the data set based on the ConvLSTM module, the VGG module performing a convolution operation on the spatiotemporal features to obtain a first feature; the CBAM module obtains an attention map based on the spatiotemporal features and the first feature, and obtains the loss value based on the attention map; The specific steps of obtaining the loss value based on the attention map include: obtaining a regression loss based on the attention map and a first loss function, obtaining a classification loss based on the attention map and a second loss function, and obtaining the loss value based on the regression loss and the classification loss; The specific steps of obtaining the SHAP value include: presetting the drought level, weighting the loss value based on the drought level, and obtaining a weighted loss value; obtaining an output result based on the weighted loss value, obtaining an input condition based on the drought characteristic, inputting the output result and the input condition into the SHAP model, and obtaining the SHAP value of the drought characteristic; The specific steps of obtaining the screening results include: based on the drought level, dividing the SHAP value to obtain the level SHAP value; based on the time dimension and the space dimension, obtaining the average value of the level SHAP value to obtain the mean SHAP value; based on the SHAP value and the mean SHAP value, screening the drought cause index to obtain the screening result; The standardized anomaly index includes the standardized precipitation anomaly index, the standardized soil moisture anomaly index and the standardized surface runoff anomaly index.
2. A method for screening drought cause indicators according to claim 1, characterized in that: The method further comprises: The correlation between the comprehensive drought index and the meteorological drought index and the agricultural drought index is obtained, and a verification result is obtained based on the correlation.
3. A method for screening drought cause indicators according to claim 1, characterized in that: The first loss function is a SmoothL1Loss loss function, and the second loss function is a cross entropy loss function.
Citation Information
Patent Citations
City rainstorm waterlogging influence factor quantitative analysis method based on SHAP
CN115936490A
Composite high-temperature drought extreme event attribution method
CN116028767A
Drought monitoring method for constructing comprehensive drought index based on GRACE data
CN117993771A
Underground water hardness influence factor analysis method based on random forest and SHAP
CN119202951A