A composite disaster spatial variability identification method based on collaborative kriging interpolation
By using co-kriging interpolation, combined with semivariograms and co-semivariograms, the problem of insufficient data for risk assessment of complex disasters in coastal areas was solved, and the accurate identification and assessment of the spatial variability of complex disasters was achieved.
Patent Information
- Application Number
- CN202310437686.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-21
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-04-21
AI Technical Summary
Existing hydrological statistics cannot meet the needs of risk assessment for complex disasters in coastal areas. In particular, under the influence of global warming and human activities, the hydrological conditions and river geomorphological characteristics of coastal areas have changed significantly, making it impossible for existing methods to accurately assess the spatial variability of complex disasters.
Co-kriging interpolation was employed. By collecting relevant disaster data, correlation analysis and normality tests were performed. Highly correlated and significant main and secondary variables were selected. Multi-model fitting was performed using semivariograms and co-variograms. The model with the smallest RMSE value was selected for spatial interpolation. Combined with cross-validation, the spatial variability of compound disasters was identified.
It enables accurate spatial interpolation and variability identification of complex disasters in coastal areas, improves the accuracy and efficiency of disaster risk assessment, and can simulate the spatial variability of hydrological statistics.
Smart Images

Figure CN116541681B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the analysis technology of water disasters and their derivative geological disasters, and in particular to a method for identifying the spatial variability of composite disasters based on co-kriging interpolation. Background Technology
[0002] Due to the intensified global climate change caused by global warming, the natural climate conditions in coastal areas have undergone tremendous changes, leading to frequent extreme disasters. Extreme events resulting from the superposition of multiple catastrophic factors are known as compound disasters. Compound disasters are characterized by their wide-ranging impact and significant losses. Therefore, accurately, objectively, and efficiently assessing the disaster risks of compound disasters is of great guiding significance for effectively preventing extreme disasters in coastal areas.
[0003] However, under the influence of high-intensity human activities over the years, the hydrological conditions and river geomorphological features of coastal areas have changed significantly. Moreover, due to the limited river network density and the location of hydrological stations in coastal areas, the existing hydrological statistics are far from meeting the needs of risk assessment for complex disasters. Summary of the Invention
[0004] Purpose of the invention: The purpose of this invention is to provide a method for identifying the spatial variability of complex disasters based on co-kriging interpolation.
[0005] Technical solution: The present invention provides a method for identifying the spatial variability of composite disasters based on co-kriging interpolation, comprising the following steps:
[0006] S1. Collect various related disaster data within the study area, remove data with a large overall distribution distance from the various related disaster data, and create a related disaster dataset. The related disaster dataset includes various related disaster data and their latitude and longitude after removing data with a large overall distribution distance from the various related disaster data;
[0007] S2. Determine the main variable sequence of the composite disaster event based on the relevant disaster dataset, and pairwise combine the main variable sequence with other data sequences in the relevant disaster dataset. Use Spearman and / or Kendall correlation coefficients to perform correlation analysis on the pairwise combinations, and select the data sequences with high and significant correlation as the secondary variable sequence. Perform normality tests on the main variable sequence and the secondary variable sequence respectively. If the main variable sequence and the secondary variable sequence are non-normally distributed, convert the main variable sequence and the secondary variable sequence to normal distribution respectively.
[0008] S3. Using co-kriging interpolation, analyze and evaluate the normally distributed principal variable and secondary variable sequences obtained in step S2. First, based on the normally distributed principal variable or secondary variable sequences, calculate the corresponding semivariogram values of the principal and secondary variables using the semivariogram formula, and then calculate the corresponding co-variogram values of the principal and secondary variables using the co-variogram formula. Second, perform multi-model analysis on the obtained principal variable semivariogram value sequences, secondary variable semivariogram value sequences, and co-variogram value sequences. For model fitting, the root mean square error (RMSE) was used to evaluate the fitting results of each model. The model with the smallest RMSE value was selected as the optimal model, yielding the best semivariograms of the main variables, secondary variables, and covariograms. Subsequently, based on the best semivariograms of the main variables, secondary variables, and covariograms, co-kriging interpolation was used to spatially interpolate the main variable and secondary variable sequences to obtain the spatial distribution of the main variables within the study area. Finally, the RMSE was used to cross-validate the spatial interpolation of the main variable and secondary variable sequences.
[0009] S4. Based on the parameters of the optimal cooperative semivariogram obtained in step S3 and the spatial interpolation of the main variable sequence, spatial variability identification is performed on the composite disaster.
[0010] Furthermore, in step S2, the Spearman correlation coefficient γ s The calculation formula is:
[0011]
[0012] Where n is the length of the relevant disaster dataset used for correlation analysis; R i and S i Let be the rank sum of the main variable sequence and the rank of other data sequences formed by pairwise combinations of the main variable sequence, respectively. and These are the average of the main variable sequence and the average of other data sequences that are paired with the main variable sequence, respectively.
[0013] The formula for calculating the Kendall correlation coefficient τ is:
[0014]
[0015]
[0016] Where, x i and x j The data for the i-th and j-th main variable sequences are y, respectively. i and y jThese are the data from other data sequences that are paired with the main variable sequence, specifically the i-th and j-th data.
[0017] Furthermore, in step S3, for a single variable of either the main variable sequence or the secondary variable sequence, the formula for its semivariogram is:
[0018]
[0019] Where γ(h) is the calculated sequence of semivariogram values of the main or secondary variables, and χ² i For data points in the primary variable sequence or secondary variable sequence, χ² i +h is the data point χ in the primary or secondary variable sequence. i Data points that are a straight-line distance of h, where h is the distance between data points χ. i With data point χ i The linear distance between +h; N(h) is the logarithm of data points in the main variable sequence or secondary variable sequence that are a linear distance of h; Z(χ i ) and Z(χ i +h) represent the data points χ in the main variable sequence or the secondary variable sequence, respectively. i With data point χ i The value of the variable at +h;
[0020] The formula for the cooperative semivariogram formed by the main variable sequence and the secondary variable sequence is:
[0021]
[0022] Where, γ 12 (h) represents the calculated sequence of cooperative semivariogram values, Z1(x) i ) and Z1(x i +h) represent data points x in the main variable sequence. i With data point x i The variable value at +h; Z2(y′) j ) and Z2(y′ j +h) represent the data points y′ in the secondary variable sequence. j With data point y′ j The value of the variable at +h.
[0023] Furthermore, the fitting model in step S3 includes:
[0024] (1) Spherical model:
[0025]
[0026] Where γ1(h) is the semivariogram or covariogram fitted by the spherical model; C0 is the nugget value; C is the arch height, i.e., the partial sill value; (C0+C) is the sill value; a is the range; h is the data point χ in the main variable sequence or secondary variable sequence. i With data point χ i The straight-line distance between +h;
[0027] (2) Gaussian model:
[0028]
[0029] Wherein, γ2(h) is the semivariogram or co-variogram obtained by fitting the Gaussian model;
[0030] (3) Exponential Model:
[0031]
[0032] Wherein, γ3(h) is the semivariogram or co-variogram fitted by the exponential model;
[0033] (4) Power exponent model:
[0034] γ4(h)=Ah θ 0<θ<2
[0035] Where γ4(h) is the semivariogram or co-variogram fitted by the power exponent model; A is a constant; and θ is the power exponent.
[0036] Furthermore, the specific process of spatial interpolation of the main variable sequence and the secondary variable sequence based on the optimal semivariogram and the optimal cooperative semivariogram in step S3 includes:
[0037] The formula for calculating the point interpolation using the co-kriging method is as follows:
[0038]
[0039] Among them, Z * (a0) represents the predicted value of the main variable at the point to be estimated (a0); n1 and n2 are the sample lengths of the main variable sequence and the secondary variable sequence, respectively, Z1(x i ) data points x in the main variable sequence i The variable value at point Z2(y') j ) represents the data point y' in the secondary variable sequence. j The value of the variable at λ; 1i To calculate the interpolation, assign Z1(x) i The weighting coefficients, λ 2j Then assign Z2(y') j The weighting coefficients of ) and
[0040] Using the first-order stationarity condition and the assumption of the weighting coefficients, combined with the Lagrange function, we obtain the following matrix:
[0041]
[0042] Where C1 is the covariance of the main variable, C2 is the covariance of the secondary variable, and C... 12 Let μ1 and μ2 be the cross-covariance between the two variables, and μ1 and μ2 be the Lagrange coefficients.
[0043] By solving the weight coefficient λ 1i With λ 2j Then we obtain the predicted value Z of the main variable at the point to be estimated a0. * (a0) is used to calculate the predicted values at all points to be estimated in the study area, thus obtaining the spatial interpolation results of the main variables.
[0044] Furthermore, in step S3, the spatial interpolation results of the main variable and the secondary variable are evaluated using K-fold cross-validation. The specific method of K-fold cross-validation is as follows: the main variable sequence or the secondary variable sequence is divided into K parts, of which K-1 parts are used as training samples and the remaining 1 part is used as validation samples. This is repeated K times to obtain the predicted variable values after cross-validation. Then, the root mean square error (RMSE) is used to evaluate the residuals of the main variable and the secondary variable respectively. The smaller the RMSE value, the better.
[0045] The formula for calculating RMSE is:
[0046]
[0047] Among them, Z pred (χ i ) represents the data points χ² in the main variable sequence or secondary variable sequence after cross-validation. i The predictor variable value at point Z(χ) i ) data points χ in the main variable sequence or secondary variable sequence i The value of the variable at the specified position, and N is the length of the primary or secondary variable sequence.
[0048] Furthermore, step S4 specifically involves:
[0049] The parameters obtained from the co-variogram fitting include nugget value, partial sill value, and range. The nugget value C0 represents the nugget effect, indicating the spatial heterogeneity of the random component. The sill value C+C0 is the stable constant value that the semivariogram value will reach when the sample distance h increases to a certain value, representing the maximum degree of variation of the variable. The distance between sample points when the semivariogram value reaches the sill value is called the range a. Therefore, within the range, there is spatial correlation, and outside the range, there is no spatial correlation. The spatial correlation within the region is represented by C / (C+C0). The smaller the ratio, the stronger the spatial correlation.
[0050] The present invention provides a composite disaster spatial variability identification system based on co-kriging interpolation, comprising:
[0051] The data acquisition and processing module is used to collect various related disaster data within the study area, remove data with large overall distribution distances from the various related disaster data, and create a related disaster dataset. The related disaster dataset includes various related disaster data and their latitude and longitude after removing data with large overall distribution distances from the various related disaster data.
[0052] The correlation analysis module is used to determine the main variable sequence of compound disaster events based on the relevant disaster dataset, and to combine the main variable sequence with other data sequences in the relevant disaster dataset in pairs, and to perform correlation analysis on the pairwise combinations, selecting the data sequences with high and significant correlation as the secondary variable sequences;
[0053] The normality test module is used to perform normality tests on the main variable sequence and the secondary variable sequence respectively. If the main variable sequence and the secondary variable sequence are non-normally distributed, the main variable sequence and the secondary variable sequence will be converted into normal distribution respectively.
[0054] The co-kriging interpolation module is used to analyze and evaluate normally distributed principal and secondary variable sequences using co-kriging interpolation. First, based on the normally distributed principal or secondary variable sequences, the corresponding semivariogram values for the principal and secondary variables are calculated using the semivariogram formula, and the corresponding co-variogram values for the principal and secondary variables are calculated using the co-variogram formula. Second, the obtained principal, secondary, and co-variogram value sequences are then processed... Multiple model fitting was performed, and the root mean square error (RMSE) was used to evaluate the fitting results of each model. The model with the smallest RMSE value was selected as the optimal model, yielding the best semivariograms of the main variables, the second variables, and the covariograms. Subsequently, based on the best semivariograms of the main variables, the second variables, and the covariograms, co-kriging interpolation was used to spatially interpolate the main variable and second variable sequences to obtain the spatial distribution of the main variables within the study area. Finally, the RMSE was used to cross-validate the spatial interpolation of the main variable and second variable sequences.
[0055] The identification module is used to identify the spatial variability of complex disasters based on the parameters of the obtained optimal cooperative semivariogram and the spatial interpolation of the main variable sequence.
[0056] An apparatus of the present invention includes a memory and a processor, wherein:
[0057] Memory is used to store computer programs that can run on a processor;
[0058] A processor, used to execute, while running the computer program, the steps of a composite disaster spatial variability identification method based on co-kriging interpolation as described above.
[0059] The present invention provides a storage medium storing a computer program, which, when executed by at least one processor, implements the steps of a composite disaster spatial variability identification method based on co-kriging interpolation as described above.
[0060] Beneficial effects: Compared with the prior art, the significant technical effects of the present invention are as follows: By introducing the co-kriging method and combining it with the characteristic that compound disasters are caused by multiple factors, the main factors are analyzed using multivariate analysis. The spatial interpolation results have been verified and are more accurate. The obtained semivariogram parameters can characterize the degree of spatial variation. It can be used to interpolate and simulate hydrological statistics in coastal areas and to identify and analyze the spatial variability of compound disaster characteristics. Attached Figure Description
[0061] Figure 1This is a flowchart of the method of the present invention;
[0062] Figure 2 This is a map showing the distribution of hydrological stations in the study area;
[0063] Figure 3 It is a scatter plot of the semivariogram of rainfall and water level and the cooperative semivariogram;
[0064] Figure 4 This is a spatial distribution map of rainfall calculated using the co-kriging interpolation method. Detailed Implementation
[0065] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0066] Disasters resulting from the superposition of multiple catastrophic factors are called complex disasters, such as the complex floods caused by heavy rainfall coinciding with astronomical tides and typhoon storm surges in coastal areas. Extreme disaster events can also bring about a series of derivative disasters. The various hydrological characteristics within the time frame of a complex disaster are correlated; therefore, co-kriging can be used to spatially interpolate complex disasters and identify their spatial variability. This invention provides a method for identifying spatial variability of complex disasters based on co-kriging interpolation. First, relevant disaster data (such as rainfall intensity, coastal tide levels, related derivative disasters and their losses) within the study area are collected and organized. After selecting the main disaster events, the correlation between the main variable data and other event data is calculated. Two types of data with strong correlation are selected to represent the complex disaster events, and they are used as the main and secondary variables, respectively. Subsequently, after selecting the main and secondary variables, a normality test is performed. Co-kriging is used to fit the semivariograms and co-variograms of the main and secondary variables. Cross-validation is used to evaluate and analyze the spatial interpolation results. Finally, based on the parameters of the best-fit co-variogram and the spatial distribution characteristics of the main variables, the spatial variability of the complex disaster characteristics within the study area can be identified and analyzed.
[0067] like Figure 1 As shown, the present invention provides a method for identifying the spatial variability of composite disasters based on co-kriging interpolation, comprising the following steps:
[0068] S1. Collect various related disaster data within the study area, remove data with a large overall distribution distance from the various related disaster data, and create a related disaster dataset. The related disaster dataset includes various related disaster data and their latitude and longitude after removing data with a large overall distribution distance from the various related disaster data;
[0069] In this embodiment of the invention, various hydrological data from the Pearl River Delta river network region were collected and filtered. The following only lists the relevant disaster dataset created after removing data with relatively large overall distribution distances from various related disaster data. The relevant disaster dataset uses data from 78 hydrological stations within the Pearl River Delta river network region. Of these 78 stations, 35 are rainfall stations and 43 are water level stations. The relevant disaster data includes rainfall data from the 35 rainfall stations and water level data from the 43 water level stations. The relevant disaster dataset includes the latitude and longitude of the hydrological stations and the corresponding rainfall or water level data. The locations of each hydrological station are shown below. Figure 2 As shown.
[0070] S2. Select the main variable and secondary variable;
[0071] Based on the relevant disaster dataset, the main variable sequence to be studied is determined, and the main variable sequence is paired with other data sequences in the relevant disaster dataset. The Spearman and / or Kendall correlation coefficients are used to perform correlation analysis on the pairwise combinations. The pair with high and significant correlation is selected as the research sample, that is, the research sample includes the main variable sequence sample and the secondary variable sequence sample. The normality test is performed on the main variable sequence and the secondary variable sequence respectively. If the main variable sequence and the secondary variable sequence are not normally distributed, they are converted to normal distribution respectively.
[0072] In this embodiment of the invention, correlation analysis is performed on the collected data. Only the research samples with high and significant correlations are listed below as the main variable sequence samples and secondary variable sequence samples used in this analysis. This example selects rainfall and water level data as important factors causing complex floods, using rainfall as the main variable. Since rainfall and coastal water level are controlled to some extent by common large-scale weather patterns, water level can be chosen as a secondary variable to aid in the calculation. The following are the results of the correlation analysis using Spearman and Kendall correlation coefficients.
[0073] The formula for calculating the two types of correlation coefficient is as follows:
[0074] Spearman correlation coefficient:
[0075]
[0076] Where n is the length of the relevant disaster dataset used for correlation analysis; R i and S i Let be the rank sum of the main variable sequence and the rank of other data sequences formed by pairwise combinations of the main variable sequence, respectively. and These are the average of the main variable sequence and the average of other data sequences that are paired with the main variable sequence, respectively.
[0077] Kendall correlation coefficient:
[0078]
[0079]
[0080] Where, x i and x j The data for the i-th and j-th main variable sequences are y, respectively. i and y j These are the data from the i-th and j-th data sequences that are paired with the main variable sequence, respectively.
[0081] The significance of a correlation is determined by the probability value (p-value) that there is no correlation between the two variables. In this invention, when p < 0.05, it means that the data has passed the 95% significance test.
[0082] The correlation coefficients between the two are shown in Table 1.
[0083] Table 1. Correlation coefficients and p-values between rainfall and water level
[0084]
[0085] As shown in Table 1, there is a positive correlation between rainfall and water level, and the p-values are all less than 0.05, indicating a 95% significant correlation. Therefore, water level can be used as a secondary variable to help obtain the spatial distribution of rainfall.
[0086] Normality tests are performed on the main variable series and the secondary variable series respectively. Different test methods are selected according to the length of the selected main variable series / secondary variable series, such as the Shapiro-Wilk (SW) test and the Kolmogorov-Smirnov (KS) test. If the data series is non-normally distributed, it needs to be transformed into a normal distribution. Common transformation methods include logarithmic transformation, reciprocal transformation, and square root transformation. The transformed main / secondary variable series are then used for subsequent calculations.
[0087] In this embodiment of the invention, the QQ plot and the Shapiro-Wilk test are used to test the rainfall data series and the water level data series respectively. It is found that the rainfall data series needs to be converted to a normal distribution through square root, while the water level data series conforms to a normal distribution and does not need to be converted.
[0088] S3. Cokriging interpolation is used to analyze and evaluate the normally distributed main variable and secondary variable sequences obtained in step S2. First, multiple models are used to fit the main variable sequences, secondary variable sequences, and combinations thereof, resulting in various fitted semivariograms of the main variables, secondary variables, and their co-variograms. The root mean square error (RMSE) is used to evaluate the results of each fitted model, and the model with the smallest RMSE value is selected as the optimal model. The semivariogram and co-variogram fitted by the optimal model are the best semivariogram and best co-variogram, respectively. The parameters of the best co-variogram (gold value, sill value, range, and spatial correlation) are used to analyze the variability. Second, based on the best semivariogram and best co-variogram, cokriging interpolation is used to spatially interpolate the main variable and secondary variable sequences to obtain the spatial distribution of the main variables within the study area. Finally, the root mean square error is used to cross-validate the spatial interpolation of the main variable and secondary variable sequences.
[0089] Co-kriging can be used to analyze and evaluate primary and secondary variables within a study area that have undergone normal distribution transformation. This process consists of three parts: First, based on the normally distributed primary or secondary variable sequences, the corresponding semivariogram values for the primary and secondary variables are calculated using the semivariogram formula. Then, the co-variogram values for the primary and secondary variables are calculated using the co-variogram formula. Second, multiple models are fitted to the obtained primary, secondary, and co-variogram value sequences to obtain the fitted semivariograms of the primary and secondary variables. The covariance function of both variables was evaluated using the root mean square error (RMSE) to assess the fitting results of each model. The model with the smallest RMSE value was selected as the optimal model. The covariance function and covariance function fitted by the optimal model were identified as the best principal variable covariance function, the best secondary variable covariance function, and the best covariance function, respectively. Subsequently, based on the best principal variable covariance function, the best secondary variable covariance function, and the best covariance function, co-kriging interpolation was used to spatially interpolate the principal variable and secondary variable sequences to obtain the spatial distribution of the principal variable within the study area. Finally, the RMSE was used to cross-validate the spatial interpolation of the principal variable and secondary variable sequences.
[0090] S31, Best semivariogram fitting
[0091] The semivariograms and cooperative semivariograms of the main / secondary variables after conversion to a normal distribution are calculated. Scatter plots of these functions are then generated using R software, as shown below. Figure 3 As shown.
[0092] For a single variable, its semivariogram formula is:
[0093]
[0094] Where γ(h) is the calculated sequence of semivariogram values of the main or secondary variables, and χ² i For data points in the primary variable sequence or secondary variable sequence, χ² i +h is the data point χ in the primary or secondary variable sequence. i Data points that are a straight-line distance of h, where h is the distance between data points χ. i With data point χ i The linear distance between +h; N(h) is the logarithm of data points in the main variable sequence or secondary variable sequence that are a linear distance of h; Z(χ i ) and Z(χ i +h) represent the data points χ in the main variable sequence or the secondary variable sequence, respectively. i With data point χ i The value of the variable at +h;
[0095] The formula for the cooperative semivariogram formed by the main variable sequence and the secondary variable sequence is:
[0096]
[0097] Where, γ 12 (h) represents the calculated sequence of cooperative semivariogram values, Z1(x) i ) and Z1(x i +h) represent data points x in the main variable sequence. i With data point x i The variable value at +h; Z2(y′) j ) and Z2(y′ j +h) represent the data points y′ in the secondary variable sequence. j With data point y′ j The value of the variable at +h.
[0098] In fitting the semivariogram, a suitable continuous distribution function is selected to characterize the scatter plot based on its distribution and trend. In this example, four models are selected to fit the data:
[0099] (1) Spherical model
[0100]
[0101] Where γ1(h) is the semivariogram or covariogram fitted by the spherical model; C0 is the nugget value; C is the arch height, i.e., the partial sill value; (C0+C) is the sill value; a is the range; h is the data point χ in the main variable sequence or secondary variable sequence.i With data point χ i The straight-line distance between +h;
[0102] (2) Gaussian model
[0103]
[0104] Wherein, γ2(h) is the semivariogram or co-variogram obtained by fitting the Gaussian model;
[0105] (3) Exponential Model
[0106]
[0107] Wherein, γ3(h) is the semivariogram or co-variogram fitted by the exponential model;
[0108] (4) Power Exponent Model
[0109] γ4(h)=Ah θ 0<θ<2
[0110] Where γ4(h) is the semivariogram or co-variogram fitted by the power exponent model; A is a constant, and θ is the power exponent.
[0111] To select the best-fitting model from the semi-variogram models, the root mean square error (RMSE) is used to evaluate the fitting model results. The model with the smallest RMSE value is selected as the optimal model, and its model parameters are used to analyze the variation.
[0112] The formula for calculating RMSE is:
[0113]
[0114] Among them, Z pred (χ i ) represents the data points χ² in the main variable sequence or secondary variable sequence after cross-validation. i The predictor variable value at point Z(χ) i ) data points χ in the main variable sequence or secondary variable sequence i The value of the variable at the specified position, and N is the length of the primary or secondary variable sequence.
[0115] The root mean square errors of the four fitting models are shown in Table 2.
[0116] Table 2. Root Mean Square Error (RMSE) Fitting Accuracy of Different Semivariogram Models
[0117]
[0118]
[0119] As shown in Table 2, a Gaussian model should be selected to fit the semivariograms of the main / secondary variables and the covariograms.
[0120] S32, Spatial Interpolation Calculation
[0121] The spatial distribution of variables within the study area is analyzed and evaluated using the co-kriging method.
[0122] The formula for point interpolation calculation using the co-kriging method is as follows:
[0123]
[0124] Among them, Z * (a0) represents the predicted value of the main variable at the point to be estimated (a0); n1 and n2 are the sample lengths of the main variable sequence and the secondary variable sequence, respectively, Z1(x i ) data points x in the main variable sequence i The variable value at point Z2(y') j ) represents the data point y' in the secondary variable sequence. j The value of the variable at the location; λ 1i To calculate the interpolation, assign Z1(x) i The weighting coefficients, λ 2j Then assign Z2(y') j The weighting coefficients of ) and
[0125] Using the first-order stationarity condition and the assumption of the weighting coefficients, combined with the Lagrange function, we obtain the following matrix:
[0126]
[0127] Where C1 is the covariance of the main variable, C2 is the covariance of the secondary variable, and C... 12 Let μ1 and μ2 be the cross-covariance between the two variables, and μ1 and μ2 be the Lagrange coefficients.
[0128] By solving the weight coefficient λ 1i With λ 2j Then we obtain the predicted value Z of the main variable at the point to be estimated a0. * (a0). First, a 400×400 grid was established within the study area. Then, co-kriging interpolation was performed using R software with the optimal semivariogram and the optimal co-variogram to obtain the spatial distribution of interpolated rainfall. Figure 4 ).
[0129] Depend on Figure 4 It can be seen that the spatial distribution of rainfall generally shows a trend of gradually increasing from northeast to southwest, but in the northeastern part of Taishan City in the southwest direction, the rainfall shows a trend of gradually decreasing from the surrounding areas to the center.
[0130] S33, Cross-validation
[0131] For the spatial interpolation results of rainfall, K-fold cross-validation was used to evaluate the results. The specific method of K-fold cross-validation is as follows: the sample is divided into K parts, of which K-1 parts are used as training samples and the remaining part is used as validation samples. This is repeated K times to obtain the predicted rainfall after cross-validation. Then, the root mean square error (RMSE) is used to evaluate the residuals of rainfall and water level values respectively. The smaller the RMSE value, the better.
[0132] The formula for calculating RMSE is:
[0133]
[0134] Among them, Z pred (χ i The results are the predictions from the cross-validation. In this study, a 5-fold cross-validation method was used, and the results of the cross-validation between the main variables and secondary variables are shown in Table 3.
[0135] Table 3 Cross-validation results
[0136]
[0137] As shown in Table 3, the root mean square error between the predicted and observed values of the primary variable rainfall is smaller than that of the secondary variable water level, indicating that the co-kriging method can perform more accurate spatial interpolation of the primary variable.
[0138] S4. Identification of spatial variability in complex disasters;
[0139] Based on the parameters of the cooperative semivariogram and the interpolation results of the main variables obtained in step S3, spatial variability of composite disasters can be identified.
[0140] The parameters obtained from the co-variogram fitting have different spatial meanings. The nugget value C0 represents the nugget effect, indicating the spatial heterogeneity of the random component; the sill value C+C0 is the stable constant value that the semivariogram value will reach after the sample distance h increases to a certain value, representing the maximum degree of variation of the variable; the distance between sample points when the semivariogram value reaches the sill value is called the range a. Therefore, within the range, there is spatial correlation, and outside the range, there is no spatial correlation; the spatial correlation within the region is represented by C / (C+C0), and the smaller the ratio, the stronger the spatial correlation.
[0141] The spatial interpolation results of the main variables can be presented within the study area. By combining their spatial distribution characteristics with the parameters of the cooperative semivariogram, the spatial variability of disasters can be further identified.
[0142] In this embodiment of the invention, the optimal model is a Gaussian model, and the parameters of the cooperative semi-variogram function fitted by the Gaussian model are:
[0143] The sill value C = 1.82, the nugget value C0 = 1.93, and the range a = 65.84 km. Therefore, spatial variability analysis can be performed on the study area. The random portion within the area exhibits spatial heterogeneity, with an error of 1.93 at the smallest sampling scale; the maximum variability within the area is 3.75 km. 2 When the straight-line distance between two points exceeds 65.84 km, the spatial correlation of rainfall distribution gradually decreases; the spatial correlation C / (C+C0) = 0.49, indicating that the two variables under study have moderate spatial correlation. The spatial interpolation distribution of rainfall shows that spatial variation occurs in the northeastern part of Taishan City, with a decreasing trend in rainfall compared to nearby areas.
[0144] The present invention provides a composite disaster spatial variability identification system based on co-kriging interpolation, comprising:
[0145] The data acquisition and processing module is used to collect various related disaster data within the study area, remove data with large overall distribution distances from the various related disaster data, and create a related disaster dataset. The related disaster dataset includes various related disaster data and their latitude and longitude after removing data with large overall distribution distances from the various related disaster data.
[0146] The correlation analysis module is used to determine the main variable sequence of compound disaster events based on the relevant disaster dataset, and to combine the main variable sequence with other data sequences in the relevant disaster dataset in pairs, and to perform correlation analysis on the pairwise combinations, selecting the data sequences with high and significant correlation as the secondary variable sequences;
[0147] The normality test module is used to perform normality tests on the main variable sequence and the secondary variable sequence respectively. If the main variable sequence and the secondary variable sequence are non-normally distributed, the main variable sequence and the secondary variable sequence will be converted into normal distribution respectively.
[0148] The co-kriging interpolation module is used to analyze and evaluate normally distributed principal and secondary variable sequences using co-kriging interpolation. First, based on the normally distributed principal or secondary variable sequences, the corresponding semivariogram (SVA) values for the principal and secondary variables are calculated using the semivariogram formula, and the corresponding co-variogram (CVA) values for the principal and secondary variables are calculated using the co-variogram formula. Second, multiple model fitting is performed on the obtained SVA, CVA, and CVA values to obtain the optimal SVA, optimal CVA, and optimal CVA. Then, based on the optimal SVA, CVA, and CVA, spatial interpolation is used to interpolate the principal and secondary variable sequences to obtain the spatial distribution of the principal variables within the study area. Finally, root mean square error (RMSE) is used to cross-validate the spatial interpolation of the principal and secondary variable sequences.
[0149] The identification module is used to identify the spatial variability of complex disasters based on the parameters of the obtained optimal cooperative semivariogram and the spatial interpolation of the main variable sequence.
[0150] An apparatus of the present invention includes a memory and a processor, wherein:
[0151] Memory is used to store computer programs that can run on a processor;
[0152] The processor is configured to execute the steps of the composite disaster spatial variability identification method based on co-kriging interpolation as described above when running the computer program, and to achieve the technical effects described above.
[0153] The present invention provides a storage medium storing a computer program, which, when executed by at least one processor, implements the steps of a composite disaster spatial variability identification method based on co-kriging interpolation as described above, and achieves the technical effects described above.
Claims
1. A method for identifying the spatial variability of composite disasters based on co-kriging interpolation, characterized in that, Includes the following steps: S1. Collect various related disaster data within the study area, remove data with large overall distribution distances from the various related disaster data, and create a related disaster dataset. The related disaster dataset includes various related disaster data and their latitude and longitude after removing data with large overall distribution distances from the various related disaster data. Related disaster data includes: rainfall event intensity, coastal tide level, related derivative disasters and their losses. S2. Determine the main variable sequence of the composite disaster event based on the relevant disaster dataset, and pairwise combine the main variable sequence with other data sequences in the relevant disaster dataset. Use Spearman and / or Kendall correlation coefficients to perform correlation analysis on the pairwise combinations, and select the data sequences with high and significant correlation as the secondary variable sequence. Perform normality tests on the main variable sequence and the secondary variable sequence respectively. If the main variable sequence and the secondary variable sequence are non-normally distributed, convert the main variable sequence and the secondary variable sequence to normal distribution respectively. Spearman correlation coefficient The calculation formula is: ; Where n is the length of the relevant disaster dataset used for correlation analysis; and Let be the rank sum of the main variable sequence and the rank of other data sequences formed by pairwise combinations of the main variable sequence, respectively. and These are the average of the main variable sequence and the average of other data sequences that are paired with the main variable sequence, respectively. Kendall correlation coefficient The calculation formula is: ; , ; Where, x i and x j The data for the i-th and j-th main variable sequences are y, respectively. i and y j These are the data from other data sequences that are paired with the main variable sequence, specifically the i-th and j-th data sequences. S3. Using co-kriging interpolation, analyze and evaluate the normally distributed principal variable and secondary variable sequences obtained in step S2. First, based on the normally distributed principal variable or secondary variable sequences, calculate the corresponding semivariogram values of the principal and secondary variables using the semivariogram formula, and then calculate the corresponding co-variogram values of the principal and secondary variables using the co-variogram formula. Second, perform multi-model analysis on the obtained principal variable semivariogram value sequences, secondary variable semivariogram value sequences, and co-variogram value sequences. For model fitting, the root mean square error (RMSE) was used to evaluate the fitting results of each model. The model with the smallest RMSE value was selected as the optimal model, yielding the best semivariograms of the main variables, secondary variables, and covariograms. Subsequently, based on the best semivariograms of the main variables, secondary variables, and covariograms, co-kriging interpolation was used to spatially interpolate the main variable and secondary variable sequences to obtain the spatial distribution of the main variables within the study area. Finally, the RMSE was used to cross-validate the spatial interpolation of the main variable and secondary variable sequences. S4. Based on the parameters of the optimal cooperative semivariogram obtained in step S3 and the spatial interpolation of the main variable sequence, spatial variability identification is performed on the composite disaster.
2. The method for identifying spatial variability of composite disasters based on co-kriging interpolation according to claim 1, characterized in that, In step S3, for a single variable of either the main variable sequence or the secondary variable sequence, the semivariogram formula is: ; in, For the calculated sequence of semivariogram values of the main or secondary variables, Data points in the primary variable sequence or secondary variable sequence In the sequence of primary or secondary variables, the data points Data points that are a straight-line distance of h, For data points With data points The straight-line distance between them; The linear distance between the main variable sequence and the secondary variable sequence is The logarithm of the data points; and These are data points in the main variable sequence or the secondary variable sequence, respectively. With data points The value of the variable at that location; The formula for the cooperative semivariogram formed by the main variable sequence and the secondary variable sequence is: ; in, For the calculated sequence of cooperative semivariogram values, Z1(x) i ) and Z1(x i +h) represent the data points in the main variable sequence. With data points The value of the variable at that location; and These are the data points in the secondary variable sequence. With data points The variable value at that location.
3. The method for identifying spatial variability of composite disasters based on co-kriging interpolation according to claim 1, characterized in that, The fitting model in step S3 includes: (1) Spherical model: ; in, The semivariogram or co-variogram fitted to the spherical model; Value of a nugget; This refers to the arch height, i.e., the value of the eccentric abutment. This is the base value; For variable range; Data points in the primary variable sequence or secondary variable sequence With data points The straight-line distance between them; (2) Gaussian model: ; in, The semivariogram or cooperative semivariogram fitted to the Gaussian model; (3) Exponential model: ; in, The semivariogram or co-variogram fitted to the exponential model; (4) Power exponent model: ; in, The semivariogram or co-variogram fitted to the power-law model; It is a constant; It is a power exponent.
4. The method for identifying spatial variability of composite disasters based on co-kriging interpolation according to claim 1, characterized in that, Step S3, which involves spatial interpolation of the main variable sequence and the secondary variable sequence based on the optimal semivariogram and the optimal cooperative semivariogram, includes the following steps: The formula for calculating the point interpolation using the co-kriging method is as follows: ; Among them, Z * (a0) represents the predicted value of the main variable at the point to be estimated (a0); n1 and n2 are the sample lengths of the main variable sequence and the secondary variable sequence, respectively, Z1(x i Data points in the main variable sequence The variable value at point Z2(y') j ) represents the data point y' in the secondary variable sequence. j The value of the variable at that location; To calculate the interpolation, assign Z1(x) i The weighting coefficients of ) Then assign Z2(y') j The weighting coefficients of ) and ; Using the first-order stationarity condition and the assumption of the weighting coefficients, combined with the Lagrange function, we obtain the following matrix: ; in, The covariance of the main variable, The covariance of the sub-variable, Let the cross-covariance be the variance between the two variables. and For Lagrange coefficients; By solving the weight coefficients and Then we obtain the predicted value Z of the main variable at the point to be estimated a0. * (a0) is used to calculate the predicted values at all points to be estimated in the study area, thus obtaining the spatial interpolation results of the main variables.
5. The method for identifying spatial variability of composite disasters based on co-kriging interpolation according to claim 1, characterized in that, In step S3, the spatial interpolation results of the main variable and the secondary variable are evaluated using K-fold cross-validation. The specific method of K-fold cross-validation is as follows: the main variable sequence or the secondary variable sequence is divided into K parts, of which K-1 parts are used as training samples and the remaining 1 part is used as validation samples. This is repeated K times to obtain the predicted variable values after cross-validation. Then, the root mean square error (RMSE) is used to evaluate the residuals of the main variable and the secondary variable respectively. The smaller the RMSE value, the better. The formula for calculating RMSE is: ; in, For data points in the main variable sequence or secondary variable sequence after cross-validation The predictor variable values at that location, Data points in the primary variable sequence or secondary variable sequence The variable value at that location, The length of the primary variable sequence or the secondary variable sequence.
6. The method for identifying spatial variability of composite disasters based on co-kriging interpolation according to claim 1, characterized in that, Step S4 is as follows: The parameters obtained by fitting the co-variogram function include nugget value, partial sill value, and range. The size of the nugget effect represents the spatial heterogeneity of the random component; the sill value... When the sample point distance After increasing to a certain value, the semivariogram value will reach a stable constant value, representing the maximum degree of variation of the variable; and the distance between sample points when the semivariogram value reaches the sill value is called the range. Therefore, within the range of variation, spatial correlation exists; outside the range of variation, spatial correlation does not exist; spatial correlation within the region is used... This indicates that the smaller the ratio, the stronger the spatial correlation.
7. A composite disaster spatial variability identification system based on co-kriging interpolation, characterized in that, include: The data acquisition and processing module is used to collect various related disaster data within the study area, remove data with large overall distribution distances from the various related disaster data, and create a related disaster dataset. The related disaster dataset includes various related disaster data and their latitude and longitude after removing data with large overall distribution distances from the various related disaster data. The related disaster data includes: rainfall event intensity, coastal tide levels, related derivative disasters and their losses; The correlation analysis module is used to determine the main variable sequence of compound disaster events based on the relevant disaster dataset, and to combine the main variable sequence with other data sequences in the relevant disaster dataset in pairs. Spearman and / or Kendall correlation coefficients are used to perform correlation analysis on the pairwise combinations, and the data sequences with high and significant correlation are selected as secondary variable sequences. Spearman correlation coefficient The calculation formula is: ; Where n is the length of the relevant disaster dataset used for correlation analysis; and Let be the rank sum of the main variable sequence and the rank of other data sequences formed by pairwise combinations of the main variable sequence, respectively. and These are the average of the main variable sequence and the average of other data sequences that are paired with the main variable sequence, respectively. Kendall correlation coefficient The calculation formula is: ; , ; Where, x i and x j The data for the i-th and j-th main variable sequences are y, respectively. i and y j These are the data from other data sequences that are paired with the main variable sequence, specifically the i-th and j-th data sequences. The normality test module is used to perform normality tests on the main variable sequence and the secondary variable sequence respectively. If the main variable sequence and the secondary variable sequence are non-normally distributed, the main variable sequence and the secondary variable sequence will be converted into normal distribution respectively. The co-kriging interpolation module is used to analyze and evaluate normally distributed principal and secondary variable sequences using co-kriging interpolation. First, based on the normally distributed principal or secondary variable sequences, the corresponding semivariogram values for the principal and secondary variables are calculated using the semivariogram formula, and the corresponding co-variogram values for the principal and secondary variables are calculated using the co-variogram formula. Second, the obtained principal, secondary, and co-variogram value sequences are then processed... Multiple model fitting was performed, and the root mean square error (RMSE) was used to evaluate the fitting results of each model. The model with the smallest RMSE value was selected as the optimal model, yielding the best semivariograms of the main variables, the second variables, and the covariograms. Subsequently, based on the best semivariograms of the main variables, the second variables, and the covariograms, co-kriging interpolation was used to spatially interpolate the main variable and second variable sequences to obtain the spatial distribution of the main variables within the study area. Finally, the RMSE was used to cross-validate the spatial interpolation of the main variable and second variable sequences. The identification module is used to identify the spatial variability of complex disasters based on the parameters of the obtained optimal cooperative semivariogram and the spatial interpolation of the main variable sequence.
8. A device, characterized in that, Includes memory and processor, wherein: Memory is used to store computer programs that can run on a processor; A processor, configured to, while running the computer program, perform the steps of the method for identifying spatial variability of a composite hazard based on co-kriging interpolation as described in any one of claims 1-6.
9. A storage medium, characterized in that, The storage medium stores a computer program that, when executed by at least one processor, implements the steps of the composite disaster spatial variability identification method based on co-kriging interpolation as described in any one of claims 1-6.
Citation Information
Patent Citations
Hydrological water resource feature space variability identification method
CN108009562A
Rainstorm disaster weather risk assessment method based on Flood Area model, device and equipment
CN111915158A