Method for analyzing internal correlation among disaster-inducing factors of multiple types of disasters
By using the intrinsic correlation analysis method of multiple disaster-causing factors and the kernel principal component analysis and Bayesian network model, the problems of low data fusion and cross-domain data sharing delay in multi-hazard risk assessment are solved, and high accuracy and rapid response of multi-hazard risk early warning are achieved.
Patent Information
- Application Number
- CN202511561296.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-01-23
AI Technical Summary
Existing technologies struggle to effectively capture the complex nonlinear relationships between multiple hazard-causing factors, suffer from low data fusion, and experience delays in cross-domain data sharing, resulting in insufficient accuracy in multi-hazard risk assessment and a lack of universally applicable analytical frameworks.
The method of analyzing the intrinsic correlation of multiple disaster-causing factors is adopted, including data preprocessing, Bayesian network correlation model construction, spatiotemporal dynamic evolution analysis and risk warning threshold model. Through kernel principal component analysis, Bayesian network and spatiotemporal attention mechanism, the correlation and risk of disaster-causing factors are quantified.
It has improved the accuracy of multi-hazard risk early warning, shortened the analysis time, supported emergency response needs, and enhanced the region's disaster prevention and mitigation capabilities.
Smart Images

Figure CN121389781A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural disaster risk assessment technology, and in particular to a method for analyzing the intrinsic correlation between multiple disaster-causing factors. Background Technology
[0002] Current analysis of the correlation between multiple hazard-causing factors faces three major technical bottlenecks: First, linear models struggle to capture complex nonlinear relationships; for example, traditional multiple regression explains less than 45% of the chain effect of "rainstorm-landslide-flood." Second, physical models have numerous parameters and simplified coupling mechanisms; for instance, a hydrological-geological coupling model in a certain watershed, which ignored dynamic changes in soil moisture content, resulted in a debris flow early warning error of 32%. Third, data integration is low; meteorological, geological, and hydrological data are scattered across 12 departments, forming "information silos," with cross-domain data sharing delays exceeding 72 hours. Therefore, establishing an analytical method that integrates multi-source heterogeneous data, dynamically quantifies correlation strength, and adapts to complex scenarios is crucial for enhancing regional disaster risk resilience.
[0003] While existing patents attempt to address these issues, significant limitations remain: CN118917058A proposes a multi-disaster monitoring and early warning method, simulating flood evolution through hydrodynamic models, but focuses only on the physical processes of a single disaster type, without addressing the correlation analysis of disaster-causing factors; CN117196029A employs spatiotemporal data cube technology to mine correlation patterns of extreme climate events, but its static correlation rules cannot adapt to the dynamic evolution of disaster-causing factors; CN120494541A assesses forest fire risk based on a complex chain disaster evolution mechanism, targeting only a specific disaster chain (drought-high temperature-fire risk) and lacking a universal analytical framework. None of these technologies achieve quantitative analysis and multi-scale coupling of the intrinsic correlations of disaster-causing factors, making it difficult to support the needs of cross-disaster risk joint prevention and control.
[0004] To address the aforementioned problems, this invention proposes a method for analyzing the intrinsic correlation between multiple hazard-causing factors. This method comprises four steps: data preprocessing, Bayesian network correlation model construction, spatiotemporal dynamic evolution analysis, and risk warning threshold modeling. This enables deep coupling analysis of multiple hazard-causing factors, such as flash floods and debris flows. This invention solves the problem that traditional single-hazard analysis methods cannot capture the spatiotemporal linkage effects of hazard-causing factors, improves the accuracy of multi-hazard risk warnings, and provides scientific decision support for regional disaster prevention and mitigation. Summary of the Invention
[0005] To achieve deep coupling analysis of multiple disaster-causing factors such as flash floods and debris flows, and to solve the problem that traditional single-hazard analysis methods cannot capture the spatiotemporal linkage effect of disaster-causing factors, this invention provides a method for analyzing the intrinsic correlation between multiple disaster-causing factors, including four steps: S100 data preprocessing, S200 Bayesian network correlation model construction, S300 spatiotemporal dynamic evolution analysis, and S400 risk warning threshold model.
[0006] Furthermore, in step S100 data preprocessing, data on 17 disaster-causing factors such as topography and precipitation are collected, and dimensionality is reduced by kernel principal component analysis and dynamic weights are calculated using the entropy method.
[0007] As a preferred option, the 17 disaster-causing factors in step S100 data preprocessing include slope, slope length, slope aspect, elevation variation coefficient, hourly rainfall, maximum 3-day cumulative rainfall, soil clay content, lithology type, vegetation NDVI index, vegetation coverage, road density, land use type, peak ground acceleration, river network density, topographic relief, soil erodibility, and surface wind speed.
[0008] Preferably, in step S100 data preprocessing, the data sampling resolution is 30 meters × 30 meters, and the time span covers 20 years;
[0009] As a preferred embodiment, in step S100 data preprocessing, the 3σ criterion is used to remove extreme values, and bilinear interpolation is used to unify non-isochronous data to a 15-minute time interval.
[0010] Preferably, in step S100 data preprocessing, spatiotemporal kriging interpolation is used, the fill rate should be no less than 98%, and the root mean square error should be no more than 5%.
[0011] Preferably, in step S100 data preprocessing, the kernel principal component analysis dimensionality reduction uses a radial basis kernel function with γ = 0.02 to compress the 17-dimensional features into 8-dimensional principal components, with a cumulative contribution rate of not less than 92%; the entropy method dynamic weight calculation interval is 15 minutes, and the weight ω i The updated formula for (t) is:
[0012]
[0013] Among them, H i (t) is the information entropy of the i-th factor at time t, calculated using the following formula:
[0014] H i (t)=-∑p ij (t)·ln(p ij (t))
[0015] Where, p ij (t) represents the normalized probability of the j-th sample at time t.
[0016] Furthermore, in step S200, the Bayesian network association model is constructed to quantify the conditional probability association of flash flood-debris flow disaster-causing factors by building a Bayesian network model.
[0017] Furthermore, in step S200, the Bayesian network association model is constructed, and the Bayesian network model contains 12 hidden nodes and 5 observation nodes.
[0018] As a preferred option, the 12 hidden nodes of the Bayesian network model are: slope-precipitation-soil clay content composite factor, slope length-NDVI index composite factor, elevation variation coefficient-road density composite factor, soil saturation, slope stability, runoff intensity, geological hazard potential, and other principal components.
[0019] As a preferred option, the five observation nodes of the Bayesian network model are the probability of debris flow, the probability of flash flood, the probability of landslide, the probability of collapse, and the comprehensive disaster risk level.
[0020] Furthermore, in step S200, the Bayesian network association model is constructed by using the expectation-maximization algorithm to estimate the conditional probability table, with a conditional probability error of less than or equal to 3%; the model structure is determined by combining mutual information testing with expert knowledge correction.
[0021] Furthermore, the conditional probability table learns from historical disaster data over the past 20 years.
[0022] Furthermore, in step S300, the spatiotemporal dynamic evolution analysis introduces a spatiotemporal attention mechanism to optimize the multi-scale coupling weights.
[0023] Furthermore, in step S300, the spatiotemporal dynamic evolution analysis includes a temporal attention weight α. t Spatial attention weight β s The calculation method is as follows:
[0024]
[0025] Q t K t For time-based queries and key matrices, dimension d k Q is 16. s K s For spatial queries and key matrices, dimension d s It is 24;
[0026] The formula for calculating the multi-scale coupling weight W is as follows:
[0027]
[0028] in Represents the tensor product.
[0029] Furthermore, in step S400, the risk warning threshold model is used to establish a coupled risk warning threshold and output a correlation strength matrix.
[0030] As a preferred option, the XGBoost early warning model is selected for risk warning, which uses the coupling weight W as input to construct a multi-hazard risk index.
[0031] As a preferred option, the risk warning threshold model is constructed using an extreme gradient boosting tree, with the input being the coupled weight vector optimized in step S300 and the output being a risk index of 0-100.
[0032] Furthermore, the warning thresholds are divided into three levels: high risk (index ≥ 75), medium risk (50 ≤ index < 75), and low risk (index < 50);
[0033] Preferably, the threshold is determined by the maximum Oden index method.
[0034] In summary, this application includes the following beneficial technical effects:
[0035] Compared with the prior art, the present invention has the following significant advantages:
[0036] 1. Improved accuracy of correlation analysis: An improved Bayesian network model is adopted to improve the accuracy of calculating the correlation strength of disaster-causing factors and successfully identify the third-order correlation threshold of "precipitation-soil moisture content-landslide";
[0037] 2. By using kernel principal component analysis for dimensionality reduction and GPU parallel computing, the correlation analysis time of time series data is shortened, supporting emergency response needs;
[0038] 3. Improve the accuracy of joint early warning for multiple disasters and reduce potential economic losses.
[0039] This method provides a quantitative analysis tool for multi-hazard risk assessment, which can directly serve decision-making in areas such as land spatial planning and emergency resource allocation, and is of great significance for improving regional disaster prevention and mitigation capabilities. Attached Figure Description
[0040] Figure 1 This is a flowchart of a method for analyzing the intrinsic correlation between multiple disaster-causing factors provided by the present invention. Detailed Implementation
[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0042] First Embodiment
[0043] like Figure 1 As shown in the figure, this embodiment details the specific implementation of the method for analyzing the intrinsic correlation between multiple disaster-causing factors in the data preprocessing stage.
[0044] 1. Data collection and organization
[0045] Taking a typical flash flood-debris flow-prone watershed in Southwest China as an example, the system collected data on 17 disaster-causing factors across four categories: topography, precipitation, soil, vegetation, and geology. These factors included slope, slope length, aspect, coefficient of variation of elevation, hourly rainfall, maximum 3-day cumulative rainfall, soil clay content, lithology, vegetation NDVI index, vegetation cover, road density, land use type, peak ground acceleration, river network density, topographic relief, soil erodibility, and surface wind speed. All data were sampled at a uniform resolution of 30 meters × 30 meters, covering a time span of 20 years.
[0046] 2. Data Cleaning and Standardization
[0047] Extreme outliers were identified and removed using the 3σ criterion. Non-isochronous data were standardized to 15-minute time intervals using bilinear interpolation. Missing data were filled using spatiotemporal kriging interpolation, ensuring a filling rate of no less than 98% and a root mean square error of no more than 5%.
[0048] 3. Kernel Principal Component Analysis for Dimensionality Reduction
[0049] Kernel principal component analysis was performed using a radial basis kernel function with γ = 0.02, compressing the 17-dimensional original features into 8-dimensional principal components with a cumulative contribution rate of no less than 92%, effectively extracting the nonlinear principal component features from the disaster-causing factors.
[0050] 4. Entropy method for dynamic weight calculation
[0051] The information entropy of each factor is dynamically calculated at 15-minute intervals, and the weights are updated accordingly. The weight update formula is:
[0052]
[0053] Among them, H i (t) is the information entropy of the i-th factor at time t, calculated using the following formula:
[0054] H i (t)=-∑p ij (t)·ln(p ij (t))
[0055] The final output is the dimensionality-reduced principal component features and their dynamic weight vectors, providing high-quality input for subsequent association modeling.
[0056] Second Embodiment
[0057] like Figure 1 As shown in the figure, this embodiment details the specific implementation of the method in the Bayesian network association model construction stage.
[0058] 1. Network Structure Design
[0059] Based on mutual information testing and expert knowledge correction, the Bayesian network structure was determined. The network contains 12 hidden nodes and 5 observation nodes. The hidden nodes are: slope-precipitation-soil clay content composite factor, slope length-NDVI index composite factor, elevation variation coefficient-road density composite factor, soil saturation, slope stability, runoff intensity, geological hazard potential, and other principal components. The observation nodes are: debris flow probability, flash flood probability, landslide probability, collapse probability, and comprehensive hazard risk level.
[0060] 2. Learning Conditional Probability Tables
[0061] The expectation-maximization algorithm is used to learn the conditional probability table for each node from nearly 20 years of historical disaster data, ensuring that the conditional probability error is ≤3%. During the learning process, the preprocessed 8-dimensional principal component features and their dynamic weights are used as input to simulate the nonlinear dependencies between multiple factors.
[0062] 3. Model Validation and Application
[0063] The model's accuracy was evaluated using leave-one-out cross-validation to ensure its high reliability in quantifying the correlation intensity of flash flood-debris flow chain disasters. The model outputs the posterior probability of each observation node, providing a probabilistic basis for spatiotemporal dynamic analysis.
[0064] Third Embodiment
[0065] like Figure 1 As shown in the figure, this embodiment details the specific implementation of the method in the spatiotemporal dynamic evolution analysis stage.
[0066] 1. Introduction of Spatiotemporal Attention Mechanism: To capture the dynamic coupling effect of disaster-causing factors in time and space, a spatiotemporal attention mechanism is introduced. This mechanism includes a temporal attention weight α. t Spatial attention weight β s The calculation formula is as follows:
[0067]
[0068] Q t K t For time-based queries and key matrices, dimension d k Q is 16. s K s For spatial queries and key matrices, dimension d s It is 24.
[0069] 2. Multi-scale coupling weight calculation
[0070] By fusing the temporal and spatial attention weights through tensor product, we obtain the multi-scale coupled weight W:
[0071]
[0072] in This represents the tensor product. The weight matrix dynamically reflects the relative importance of disaster-causing factors in different time periods and regions, optimizing the input features of subsequent risk warning models.
[0073] Fourth embodiment
[0074] like Figure 1 As shown in the figure, this embodiment details the specific implementation of the method in the risk warning threshold model construction and output stage.
[0075] 1. Model Building
[0076] An extreme gradient boosting tree is used to construct a risk warning model. The input is the optimized multi-scale coupled weight vector W in the third embodiment, and the output is a comprehensive risk index in the range of 0-100.
[0077] 2. Threshold Division
[0078] Based on historical disaster samples, the risk warning threshold is determined using the maximum Yoden index method and divided into three levels: high risk (index ≥ 75), medium risk (50 ≤ index < 75), and low risk (index < 50).
[0079] 3. Output of correlation strength matrix
[0080] The model also outputs a correlation strength matrix between each disaster-causing factor and disaster type, which is presented in the form of a visual management platform to support regional disaster joint prevention and control decision-making.
Claims
1. A method for analyzing the intrinsic correlation among multiple disaster-causing factors, characterized in that, The method includes: S100 data preprocessing: Data on 17 disaster-causing factors, including topography and precipitation, were collected, and dimensionality was reduced by kernel principal component analysis and dynamic weights were calculated using the entropy method. S200 Bayesian Network Association Model Construction: Construct a Bayesian network model to quantify the conditional probability association of flash flood-debris flow disaster-causing factors; S300 Spatiotemporal Dynamic Evolution Analysis: Introducing a spatiotemporal attention mechanism to optimize multi-scale coupling weights; S400 Risk Warning Threshold Model: Establishes coupled risk warning thresholds and outputs a correlation strength matrix.
2. The method for analyzing the intrinsic correlation among multiple disaster-causing factors according to claim 1, characterized in that... The S100 data preprocessing includes 17 disaster-causing factors, such as slope, slope length, slope aspect, elevation variation coefficient, hourly rainfall, maximum 3-day cumulative precipitation, soil clay content, lithology, vegetation NDVI index, vegetation cover, road density, land use type, peak ground acceleration, river network density, topographic relief, soil erodibility, and surface wind speed. The data sampling resolution is 30 meters × 30 meters, and the time span covers 20 years.
3. The method for analyzing the intrinsic correlation among multiple disaster-causing factors according to claim 1, characterized in that... In the S100 data preprocessing, kernel principal component analysis (KPCA) was used for dimensionality reduction with a radial basis function (RBF) of γ = 0.02, compressing the 17-dimensional features into 8-dimensional principal components with a cumulative contribution rate of no less than 92%. The entropy method was used to calculate dynamic weights at 15-minute intervals, with weights ω... i The updated formula for (t) is: Among them, H i (t) represents the information entropy of the i-th factor at time t.
4. The method for analyzing the intrinsic correlation among multiple disaster-causing factors according to claim 1, characterized in that... The S200 Bayesian network model contains 12 hidden nodes and 5 observation nodes. The 12 hidden nodes are: slope-precipitation-soil clay content composite factor, slope length-NDVI index composite factor, elevation variation coefficient-road density composite factor, soil saturation, slope stability, runoff intensity, geological hazard potential, and other principal components. The 5 observation nodes are: debris flow probability, flash flood probability, landslide probability, collapse probability, and comprehensive disaster risk level. The conditional probability table is estimated using the expectation-maximization algorithm, with a conditional probability error of less than or equal to 3%. The model structure is determined by a combination of mutual information testing and expert knowledge correction.
5. The method for analyzing the intrinsic correlation among multiple disaster-causing factors according to claim 1, characterized in that... The spatiotemporal attention mechanism in S300 includes a temporal attention weight α. t Spatial attention weight β s The calculation method is as follows: Q t K t For time-based queries and key matrices, dimension d k Q is 16. s K s For spatial queries and key matrices, dimension d s It is 24; The formula for calculating the multi-scale coupling weight W is as follows: in Represents the tensor product.
6. The method for analyzing the intrinsic correlation among multiple disaster-causing factors according to claim 1, characterized in that... The risk warning threshold model in S400 is constructed using an extreme gradient boosting tree. The input is the coupled weight vector optimized in step S300, and the output is a risk index of 0-100. The warning thresholds are divided into three levels: high risk (index ≥ 75), medium risk (50 ≤ index < 75), and low risk (index < 50). The thresholds are determined by the maximum Yorden index method.
Citation Information
Patent Citations
Real-time interactive extreme climate disaster event association mining method
CN117196029A
Method for monitoring, predicting and early warning various disasters
CN118917058A
Forest fire risk assessment method based on composite chain disaster evolution mechanism
CN120494541A
Cited By
Geological disaster multi-scale environment control factor identification system based on machine learning
CN121834240A
Dynamic quantification method and device for seepage risk
CN121835522A