A method for realizing multi-dimensional classification and prediction of debris flow disaster risk
Patent Information
- Application Number
- CN202311414102.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-27
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2043-10-27
AI Technical Summary
[0007]针对目前泥石流灾害风险预测存在的局限问题,本发明将在泥石流灾害风险分类、泥石流长短期预测等方面,提出一种适用于复杂山区的泥石流灾害风险长短期分类预测方法
[0015] 1) Debris flow hazards in complex mountainous areas differ from those in other regions. For important locations such as railways, bridges, and highways, or densely populated towns, precision sensors can be installed to assist in debris flow risk prediction. Precision sensors can collect more accurate monitoring data, resulting in better debris flow risk prediction performance under the same prediction model. However, the complex geological conditions and topography of mountainous areas make them difficult for monitoring personnel to access; traditional debris flow hazard classification requires on-site exploration by professionals, which is unsuitable for the actual conditions in complex mountainous areas. This invention classifies potential debris flow hazard areas into valley-type and hillside-type debris flow hazard areas based on their watershed morphology, and then uses corresponding prediction models to predict debris flow risk in the monitoring areas. This invention proposes a watershed morphology classification method for debris flow hazard areas. The classification model performs binary classification on the DEM files of potential debris flow areas, dividing them into valley-type and hillside-type debris flow hazard areas according to their watershed morphology. This solves the problem of surveyors being unable to conduct on-site exploration and manual classification in complex mountainous areas.
Smart Images

Figure CN117390555B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of debris flow disaster risk prediction, and in particular to a prediction method for realizing multidimensional classification of debris flow disaster risk. Background Technology
[0002] Debris flows are fluids composed of a large amount of soil, rock fragments, water, and air, and they often occur in mountainous or hilly areas. Debris flows typically form on steep slopes or in gullies, and can reach speeds of tens of kilometers per hour. They possess powerful impact and destructive force, capable of destroying houses, roads, bridges, and farmland, causing casualties and property damage.
[0003] my country is a mountainous country, with mountains covering about one-third of its total area. Mountains are generally characterized by large elevation differences, steep slopes, and thin soil layers. If mountains, hills, and relatively rugged plateaus are collectively referred to as mountainous areas, then mountainous areas account for about two-thirds of my country's total area. Many mountainous areas are still traversed by roads or railways and inhabited by residents. During the rainy season, some mountainous areas frequently experience mudslides, causing significant loss of life and property. Therefore, accurate mudslide risk prediction is of paramount importance.
[0004] Debris flow disasters are classified into gully-type debris flows and hillside-type debris flows based on the morphology of the watershed where they occur. Currently, there is more research on risk prediction for gully-type debris flows, while research on hillside-type debris flow disaster prediction is relatively limited. In debris flow disaster risk prediction in densely populated mountainous areas, most methods utilize real-time monitoring parameter data acquired by multiple sensors, followed by data analysis and processing to achieve risk prediction. This type of method can meet the requirements for real-time risk prediction, but it requires a significant investment in monitoring equipment, resulting in high costs. For vast mountainous areas, the economic cost of widely adopting sensor-based monitoring methods is prohibitive.
[0005] In addition to real-time monitoring of debris flow disaster risks using sensor equipment, manual on-site exploration and visual methods based on expert experience can also be used. These methods are relatively low-cost, but the accuracy of risk prediction is limited by the experience of the surveyors, and the efficiency is too low.
[0006] Furthermore, in terms of debris flow disaster risk prediction models, traditional methods usually employ a single prediction model, which does not fully consider the differences and characteristics between short-term and long-term predictions. This makes it difficult to meet people's requirements for both short-term and long-term predictions of debris flow disaster risks. Therefore, the selection of debris flow disaster risk prediction models still needs continuous improvement. Summary of the Invention
[0007] To address the limitations of current debris flow disaster risk prediction methods, this invention proposes a long-term and short-term debris flow disaster risk classification and prediction method applicable to complex mountainous areas. The specific technical solution is as follows:
[0008] Step 1: Delineate the debris flow hazard monitoring area. Delineate the scope of the debris flow hazard monitoring area on the satellite map. Obtain the latitude and longitude of the hazard monitoring area in the four directions of east, west, south, and north using map tools. Use the latitude and longitude range to outline the scope of the debris flow hazard monitoring area.
[0009] Step 2: Download the digital elevation model (DEM) of the monitoring area. For the delineated debris flow hazard monitoring area, download the DEM data of the monitoring area from a professional website system.
[0010] Step 3: Classification of watershed morphology in debris flow hazard areas. After downloading the DEM data of the monitoring area, the data is cropped according to the actual size of the debris flow hazard monitoring area and saved as a new DEM file. The classification algorithm is used to classify the monitoring area to determine whether it is a valley-type debris flow hazard area or a hillside-type debris flow hazard area.
[0011] Step 4: Collect data on disaster impact factors in debris flow hazard areas. Obtain debris flow hazard point information from the daily geological disaster reports published on the website of the Ministry of Natural Resources, and convert the hazard point information into specific latitude and longitude coordinates through a map website. Obtain rainfall, soil moisture content, elevation, slope, stratum lithology, vegetation cover, and calculate surface deformation data for the corresponding areas from professional website systems.
[0012] Step 5: Learning a multidimensional classification prediction model for debris flow risk. Based on the acquired sample data of different types of debris flow disasters, a short-term debris flow disaster prediction model based on AdaBoost is learned using the AdaBoost algorithm combined with short-term influencing factors; a long-term debris flow disaster risk prediction model based on SVM is learned using Support Vector Machine (SVM) combined with long-term influencing factors. The prediction models are then validated using a dataset. Finally, long-term gully debris flow disaster risk prediction model based on SVM, short-term gully debris flow disaster risk prediction model based on AdaBoost, long-term hillside debris flow disaster risk prediction model based on SVM, and short-term hillside debris flow disaster risk prediction model based on AdaBoost are obtained from four dimensions.
[0013] Step 6: Multidimensional classification prediction and result presentation of debris flow risk. Based on the watershed classification results of the monitoring area, the corresponding prediction models are selected to process the monitoring data, realize the long-term and short-term prediction of debris flow disaster risk in the monitoring area, and visualize the prediction results.
[0014] The technical solutions of the embodiments of the present invention can bring at least the following beneficial effects:
[0015] 1) Debris flow hazards in complex mountainous areas differ from those in other regions. For important locations such as railways, bridges, and highways, or densely populated towns, precision sensors can be installed to assist in debris flow risk prediction. Precision sensors can collect more accurate monitoring data, resulting in better debris flow risk prediction performance under the same prediction model. However, the complex geological conditions and topography of mountainous areas make them difficult for monitoring personnel to access; traditional debris flow hazard classification requires on-site exploration by professionals, which is unsuitable for the actual conditions in complex mountainous areas. This invention classifies potential debris flow hazard areas into valley-type and hillside-type debris flow hazard areas based on their watershed morphology, and then uses corresponding prediction models to predict debris flow risk in the monitoring areas. This invention proposes a watershed morphology classification method for debris flow hazard areas. The classification model performs binary classification on the DEM files of potential debris flow areas, dividing them into valley-type and hillside-type debris flow hazard areas according to their watershed morphology. This solves the problem of surveyors being unable to conduct on-site exploration and manual classification in complex mountainous areas.
[0016] 2) Existing debris flow risk prediction methods do not consider the practical application needs of short-term and long-term predictions. Short-term and long-term predictions differ in various aspects, making the use of a single model unreasonable. This invention addresses the problem of short-term and long-term debris flow prediction by proposing separate long-term and short-term prediction models for gully-type and hillside-type debris flows, enabling debris flow predictions within three days and classification predictions within four to seven days. This invention also proposes a debris flow short-term and long-term classification prediction method combining Support Vector Machine (SVM) and AdaBoost algorithms. Based on the watershed morphology classification results, the debris flow is classified into short-term predictions within three days and long-term predictions within four to seven days according to the prediction time. The short-term prediction model based on AdaBoost is used to predict debris flow risk for 1 to 3 days, while the long-term prediction model based on SVM is used to predict debris flow risk for 4 to 7 days. This invention solves the problem of insufficient accuracy of traditional single models for short-term and long-term predictions. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings without creative effort.
[0018] Figure 1A flowchart of a method for multidimensional classification and prediction of debris flow disaster risk is provided in an embodiment of the present invention;
[0019] Figure 2 A flowchart of the DEM download process provided in this embodiment of the invention;
[0020] Figure 3 A schematic diagram of a gully-type debris flow basin;
[0021] Figure 4 A schematic diagram of a hillside debris flow basin;
[0022] Figure 5 The flowchart illustrates the learning process of the multidimensional classification prediction model for debris flow disasters provided in this embodiment of the invention. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] To ensure the accuracy of debris flow disaster prediction, this invention provides a method for multi-dimensional classification and prediction of debris flow disaster risk. For example... Figure 1 As shown, the method includes the following steps.
[0025] Step 1: Delineate the debris flow hazard monitoring area. Delineate the scope of the debris flow hazard monitoring area on a satellite map. Obtain the latitude and longitude range of the hazard monitoring area using map tools, and then use the latitude and longitude range to outline the scope of the debris flow hazard monitoring area.
[0026] Specifically, when obtaining the latitude and longitude range, the longitude range of the debris flow hazard monitoring area is obtained in the east-west direction using map tools, and the latitude range of the debris flow hazard monitoring area is obtained in the north-south direction.
[0027] Step 2: Download the digital elevation model (DEM) of the monitoring area. For the delineated debris flow hazard monitoring area, download the digital elevation model (DEM) data of the monitoring area from a professional website system.
[0028] A digital elevation model (DEM) realizes the digital simulation of ground topography (i.e., the digital representation of the surface morphology of the terrain). It is a physical ground model that represents the ground elevation using an ordered array of numerical values. This invention utilizes the acquired DEM of the monitoring area to perform watershed morphology classification processing within the monitoring area.
[0029] This invention acquires DEMs with 30m and 12.5m resolution. The 12.5m resolution DEM can be extracted from ALOS PALSAR satellite imagery. However, ALOS PALSAR satellite imagery coverage is incomplete, missing some areas. For areas without ALOS PALSAR satellite imagery, a 30m resolution DEM is used as a substitute. The DEM download process is as follows... Figure 2 As shown:
[0030] Step 21: Select the monitoring area to be studied from the NASA website EarthData (https: / / search.asf.alaska.edu);
[0031] Step 22: Determine whether there is ALOS PALSAR image data in the area. If the result is yes, proceed to step 1023; if the result is no, proceed to step 1025.
[0032] Step 23: Select images that cover the monitored area. Specifically, filter out ALOSPALSAR satellite images that completely contain the area. If the target area is too large, one satellite image cannot completely cover it. The large area needs to be divided into several smaller areas that can be completely covered by a single image. These smaller areas are then processed and used as the search benchmark to search for image data for that area.
[0033] Step 24: Download the Hi-Res Terrain Corrected file. Specifically, select the high-resolution terrain-corrected (Hi-Res Terrain Corrected) image of the area and download it to your local computer.
[0034] Step 25: Obtain the Geospatial Data Cloud DEM data. Specifically, select ASTERGDEM 30M (Advanced Spaceborne Thermal Emission and Reflection Radiometer Global Digital Elevation Model, with a global spatial resolution of 30 meters) resolution digital elevation data from the publicly available data on the Geospatial Data Cloud website (https: / / www.gscloud.cn / ) under the Computer Network Information Center of the Chinese Academy of Sciences.
[0035] Step 26: Input the latitude and longitude of the monitoring area;
[0036] Step 27: Download the ASTER GDEM 30M resolution digital elevation data for the corresponding area;
[0037] Step 28: Obtain the DEM file via step 1024 or step 1027.
[0038] Step 3: Classification of watershed morphology in debris flow hazard areas. After downloading the DEM data of the monitoring area, the data is cropped according to the actual size of the debris flow hazard monitoring area and saved as a new DEM file. The artificial intelligence support vector machine algorithm is used for classification to determine whether the monitoring area is a valley-type debris flow hazard area or a hillside-type debris flow hazard area.
[0039] For debris flow disasters, based on the watershed morphology of the disaster area, debris flow disasters can be classified into valley-type debris flow disasters and slope-type debris flow disasters. Therefore, when conducting debris flow disaster risk prediction for potential debris flow disaster areas, the disaster area can be divided into valley-type debris flow disaster risk areas and slope-type debris flow disaster risk areas according to the watershed morphology of the potential debris flow disaster area, as shown in the figure below. Figure 3 and Figure 4 As shown.
[0040] The downloaded DEM file typically covers a geographical area far larger than the debris flow hazard monitoring area. If this DEM is directly analyzed and processed, most of the processed area will not be the monitoring area, severely interfering with the identification results. Therefore, ArcGIS is needed to crop the original DEM file so that the cropped DEM file only contains the monitoring area.
[0041] Drag the downloaded original DEM file into ArcGIS, crop it to the actual size of the debris flow disaster monitoring area, and save it as a new DEM file.
[0042] Detailed steps include:
[0043] (1) Import the DEM file into ArcGIS;
[0044] (2) Open ArcGIS ArcToolbox, find “Data Management Tools”—“Raster”—“Raster Processing”—“Crop” tool;
[0045] (3) Select the current DEM file in “Input Raster”;
[0046] (4) According to the actual size of the debris flow disaster monitoring area, fill in the maximum and minimum values of latitude and longitude in sequence within the rectangular area;
[0047] (5) Click OK to generate the cropped DEM data and save it.
[0048] Historical debris flow disaster locations were obtained from the daily geological disaster reports published on the website of the Ministry of Natural Resources. These locations were then converted into specific latitude and longitude coordinates using a map website. Based on these coordinates, the corresponding DEM files were obtained and cropped to an appropriate size using the method described above. Debris flow disasters were classified using methods such as visual inspection, and labeled as either valley-type or hillside-type debris flow, which served as positive and negative samples for machine learning training.
[0049] For the already cropped and labeled DEM files, the plane curvature, profile curvature, slope, and elevation of the monitoring area are extracted from the DEM files as a dataset to train the support vector machine.
[0050] The training process of a support vector machine is as follows: input training set {(X) i ,y i )},i=1~N, where y i = +1 or -1, representing the positive or negative sample label, X i The main influencing factors include planar curvature, profile curvature, slope, and elevation; the objective function to be maximized is to be solved.
[0051]
[0052] Restrictions: y j = +1 or -1, representing positive or negative sample labels;
[0053] Where N represents the total number of training sets, K(X) i ,X j ) represents the kernel function, and C represents a penalty factor greater than 0;
[0054] In selecting the kernel function, a polynomial kernel function is chosen, namely...
[0055] K(X i ,X j )=(X i ·X j +c) d
[0056] Where c is a constant term used to adjust the nonlinearity of the kernel function, and d is the degree of the polynomial, which determines the complexity of the polynomial kernel function; C, c, and d need to be tuned during the learning process;
[0057] When the objective function is maximized, all values of α constitute the optimal solution α. * , where α=(α1,α2,...α N ), and Represents α iThe optimal solution after convergence;
[0058] The optimal solution α is calculated using the following algorithm. * The specific process for determining the value of is as follows:
[0059] Select two parameters α that need to be updated. i and α j Fix other α k k = 1 to N, k ≠ i, k ≠ j, due to constraints At this point, the constraints become...
[0060] α i y i +α j y j =c
[0061] 0≤α i ≤C
[0062] 0≤α j ≤C
[0063] in Therefore, we can conclude that... Using α i The expression replacing α j The objective problem can be transformed into an optimization problem with only one constraint: 0 ≤ α. i ≤C;
[0064] For an optimization problem with only one constraint, for α i Find the partial derivative, set it to 0, and then find the value of the variable α. inew According to α inew Find α jnew Iterate multiple times until convergence, calculating each... The optimal solution α is obtained * Construct the classification decision function f(x):
[0065]
[0066]
[0067] Where T is the transpose and sign(·) is the transition function:
[0068]
[0069] The labeled DEMs, categorized into gully-type and hillside-type debris flows, were used as positive and negative samples, respectively, and input into the training model for classification training. The trained classification model was then used as a watershed morphology classification model for debris flow hazard areas.
[0070] Finally, the cropped DEM of the debris flow hazard area is input into the watershed morphology classification model of the debris flow hazard area to obtain the watershed morphology classification results of the debris flow hazard area, namely, the valley-type debris flow hazard area or the hillside-type debris flow hazard area.
[0071] Step 4: Collect data on disaster-influencing factors in debris flow hazard areas. Obtain debris flow hazard point information from the daily geological disaster reports published on the website of the Ministry of Natural Resources, and convert the hazard point information into specific latitude and longitude coordinates through a map website. Obtain meteorological, regional topographic and geomorphological, rock strata lithology, and vegetation coverage data for the corresponding areas from professional website systems, and use Sentinel-1 data analysis to obtain surface deformation data for the corresponding areas.
[0072] The main influencing factors of debris flow disasters are selected as follows:
[0073] Many factors influence debris flow disasters, such as slope, elevation, lithology, geological structure, surface cover type, vegetation cover index, rainfall, soil moisture content, distance from water systems, distance from buildings, distance from roads, and surface deformation. Other scholars have already conducted thorough research on the factors influencing debris flow disasters, and their weight, from highest to lowest, is as follows: lithology > rainfall > geological structure > slope > vegetation cover index > elevation > soil moisture content > surface deformation > surface cover type > distance from water systems > distance from buildings > distance from roads. This invention selects lithology, rainfall, slope, vegetation cover, elevation, soil moisture content, and surface deformation as the main influencing factors for debris flow geological hazards.
[0074] The influencing factors selected in this invention differ between short-term and long-term forecasts. For short-term forecasts, short-term influencing factors include lithology, rainfall, slope, vegetation cover, elevation, soil moisture content, and cumulative rainfall. For long-term forecasts, long-term influencing factors include lithology, rainfall, slope, vegetation cover, elevation, soil moisture content, and surface deformation. The difference in the selection of influencing factors between short-term and long-term forecasts lies in the fact that short-term forecasts include cumulative rainfall, while long-term forecasts include surface deformation. The main reason for this is that cumulative rainfall has a significant impact on the occurrence of debris flow disasters within three days, but only actual rainfall data is available, not predicted data. Surface deformation, on the other hand, has a six-day data cycle, and the data obtained through interpolation is not actual data and contains a certain degree of error, making it unsuitable for sensitive short-term forecasts. Therefore, surface deformation is only used for long-term forecasts.
[0075] The data on factors influencing debris flow disasters are as follows:
[0076] Historical debris flow disaster locations were obtained from the daily geological disaster reports published on the website of the Ministry of Natural Resources. These locations were then converted into specific latitude and longitude coordinates using map websites. Based on these coordinates, meteorological, hydrological, and soil moisture data for each location were collected from various data websites and compiled into a positive sample dataset. Randomly selected areas that had never experienced debris flow disasters were used to collect the same data, serving as negative samples for the dataset. For an area that had experienced a debris flow disaster, data from two different time points were collected: data from a randomly selected day between 1 and 3 days before the disaster, and data from a randomly selected day between 4 and 7 days before the disaster. For areas that had never experienced a debris flow disaster, a random time point was also selected as the comparison time, and data from a randomly selected day between 1 and 3 days before the comparison time, and data from a randomly selected day between 4 and 7 days before the comparison time, were also collected.
[0077] The website of the Ministry of Natural Resources (https: / / www.mnr.gov.cn / ) publishes daily reports on geological disaster situations and risks. A Python script was written to collect geological disaster reports from the past ten years from this website and extract data such as the time, location, and type of disaster from the reports. Data on debris flow disasters over the years was then filtered out to obtain the location and watershed morphology of debris flow disasters. The website was then used to locate the disaster sites, converting the textual geological information into specific latitude and longitude coordinates.
[0078] Historical meteorological data was obtained from the website of the U.S. National Center for Environmental Information (NCIE), accessible at (https: / / www.ncei.noaa.gov / data / global-forecast-system / access / historical / analysis / ). The collected meteorological data includes rainfall, soil moisture content, and other data.
[0079] Regional topographic data can be obtained from the NASA website (https: / / urs.earthdata.nasa.gov / ) as raw ASTER GDEM 30M resolution elevation data. A data acquisition program was written to download the regional elevation data to the local machine. Then, a Python program was used to parse the data and perform preliminary analysis of the regional topography using slope and aspect algorithms to extract and calculate the regional elevation and slope data.
[0080] Regional stratigraphic and lithological data can be obtained from the website of the National Geological Archives (http: / / www.ngac.org.cn / DataSpecial / geomap.html). By locating the corresponding latitude and longitude points in the National Geological Map Data Thematic Service application, and based on the legend data, the geological and lithological data for that region can be manually collected.
[0081] Regional vegetation cover data was obtained from the European Centre for Medium-Range Weather Forecasts (https: / / www.ecmwf.int / en / forecasts / datasets / reanalysis-datasets / era5) website to represent the vegetation cover status. A Python program was used to call the website's data download API to collect regional leaf area index data.
[0082] The surface deformation data was obtained by downloading Sentinel-1 data from NASA's EarthData website and calculating the surface deformation using SBAS-InSAR technology. Because the shortest period for Sentinel-1 data is 6 days, the calculated surface deformation data is also in 6-day intervals, which cannot directly meet the requirements for predicting debris flow disasters and requires further data processing.
[0083] First, using the Lagrange interpolation method, the 6-day periodic surface deformation data were supplemented with 1-day periodic surface deformation data.
[0084] Lagrange polynomial
[0085] After completing the landform variables with a 1-day interval, support vector regression (SVR) is used to predict the landform variables, resulting in predicted values for the next 7 days.
[0086] Step 5: Learning a multidimensional classification prediction model for debris flow risk. Based on the acquired sample data of different types of debris flow disasters, a short-term debris flow disaster prediction model based on AdaBoost is learned using the AdaBoost algorithm combined with short-term influencing factors. A long-term debris flow disaster risk prediction model based on SVM is learned using Support Vector Machine (SVM) combined with long-term influencing factors. The prediction models are then validated using a dataset. Finally, long-term and short-term prediction models for gully-type debris flows and slope-type debris flows are obtained.
[0087] Rainfall is a crucial influencing factor for debris flow disasters. Current rainfall forecasts generally show higher accuracy for more recent data and lower accuracy for more distant data. To address this characteristic of rainfall forecasting, this invention employs the AdaBoost algorithm for both short-term and long-term predictions, and the Support Vector Machine (SVM) algorithm for predicting debris flow geological hazards.
[0088] Support Vector Machines (SVMs) and AdaBoost are both commonly used machine learning algorithms, each with its own advantages and disadvantages in different scenarios. SVM is a binary classification model that aims to find an optimal hyperplane to classify data into two classes. Its advantages include good performance in high-dimensional spaces, the ability to handle non-linear classification problems, and strong generalization ability. AdaBoost is an ensemble learning algorithm that builds multiple weak classifiers by progressively adjusting the weights of the dataset, ultimately combining them into a strong classifier. AdaBoost's advantages include improved classifier accuracy and robustness to noisy data. Compared to SVMs, AdaBoost is more sensitive and better suited for short-term forecasting with higher accuracy in rainfall prediction.
[0089] After selecting debris flow hazard areas, based on the watershed morphology classification results, these areas were divided into gully-type and hillside-type debris flow zones. Furthermore, based on predictions for three days and four to seven days, they were further categorized into short-term and long-term predictions. SVM risk prediction models were trained using debris flow influencing factor data, including surface deformation, for both gully-type and hillside-type debris flows. AdaBoost risk prediction models were trained using debris flow influencing factor data, including cumulative rainfall, for both gully-type and hillside-type debris flows. Ultimately, four models were developed: a long-term SVM-based gully-type debris flow hazard risk prediction model, a short-term AdaBoost-based gully-type debris flow hazard risk prediction model, a long-term SVM-based hillside-type debris flow hazard risk prediction model, and a short-term AdaBoost-based hillside-type debris flow hazard risk prediction model. These four models predict debris flow hazards from four dimensions. The flowchart is shown below. Figure 5 As shown.
[0090] The distinction between short-term and long-term forecasts is primarily based on the impact of cumulative rainfall on debris flow disasters. Previous studies have shown that cumulative rainfall within three days carries a higher weight in determining the impact of debris flow disasters. Therefore, this invention positions debris flow disaster risk prediction within three days as a short-term forecast, and debris flow disaster risk prediction within four to seven days as a long-term forecast.
[0091] AdaBoost, a representative algorithm in ensemble boosting, focuses on correcting errors made by weak classifiers. It uses the entire training set during the learning phase of each tree, iterates by changing the weights of the training samples, and then learns multiple classifiers. Finally, these "weak" classifiers are linearly combined to obtain the final classifier model. The training process is as follows:
[0092] (1) Given a training sample dataset: S={(x1z1),...,(x i zi ),...,(x N z N )}, where N is the total number of samples, x i The main influencing factors include stratigraphic lithology, rainfall, slope, vegetation cover, elevation, soil moisture content, and cumulative rainfall. i Let z represent the risk label of the training sample, where z i =1 or -1, representing risk and no risk respectively.
[0093] (2) The initial weight distribution of the training sample set is as follows:
[0094]
[0095] Where: D1(i) is the initial weight distribution of the training sample set;
[0096] W 1i Each training sample is initially assigned the same weight.
[0097] (3) Perform m iterations, using a decision tree as the base classifier, and adjust the weights D. m Learning from the training set, calculating the weak classifier G m (x).
[0098] (4) Calculate G m (x) Classification error rate e on the training set m
[0099]
[0100] In the formula: I(G m (x i The value of ≠z is 0 (correct classification) or 1 (incorrect classification).
[0101] (5) Calculate the weak classifier G m The coefficient a of (x) m
[0102]
[0103] (6) Update the weight distribution of the training data
[0104] (7) Construct the final classifier
[0105] According to the weight a of the weak classifier m Combine the various weak classifiers, i.e.
[0106]
[0107] By applying the sign function, a strong classifier is obtained as follows:
[0108]
[0109] The long-term training model is trained using the following method:
[0110] Input training set {(X i ,y i )},i=1~N, where y i = +1 or -1, representing the positive or negative sample label, X i The main influencing factors include stratigraphic lithology, rainfall, slope, vegetation cover, elevation, soil moisture content, and surface deformation; the objective function to be maximized is to be solved.
[0111]
[0112] Restrictions: y j = +1 or -1, representing positive or negative sample labels;
[0113] Where N represents the total number of training sets, K(X) i ,X j ) represents the kernel function, and C represents a penalty factor greater than 0;
[0114] In selecting the kernel function, the Gaussian kernel function is chosen, i.e.
[0115]
[0116] C and δ need to be tuned during the learning process;
[0117] When the objective function is maximized, all values of α constitute the optimal solution α. * , where α=(α1,α2,...α N ), and Represents α i The optimal solution after convergence;
[0118] The optimal solution α is calculated using the following algorithm. * The specific process for determining the value of is as follows:
[0119] Select two parameters α that need to be updated. i and α j Fix other α k k = 1 to N, k ≠ i, k ≠ j, due to constraints At this point, the constraints become...
[0120] α i y i +α j yj =c
[0121] 0≤α i ≤C
[0122] 0≤α j ≤C
[0123] in Therefore, we can conclude that... Using α i The expression replacing α j The objective problem can be transformed into an optimization problem with only one constraint: 0 ≤ α. i ≤C;
[0124] For an optimization problem with only one constraint, for α i Find the partial derivative, set the derivative to 0, and then find the value of the variable. according to Find Iterate multiple times until convergence, and calculate each The optimal solution α is obtained * Construct the classification decision function f(x):
[0125]
[0126]
[0127] Where T is the transpose and sign(·) is the transition function:
[0128]
[0129] The dataset was divided into training and testing sets in a 7:3 ratio. In the training set, each region contained two sets of data: data from one to three days prior to the target time (positive samples representing debris flow occurrence times, negative samples representing comparison times) and data from one to seven days prior to the target time. Based on the regional watershed morphology classification results, short-term prediction models for gully and hillside debris flows were trained using data from days 1 to 3 in the training set, respectively; and long-term prediction models for gully and hillside debris flows were trained using data from days 4 to 7 in the training set, respectively. The trained models were then validated using the testing set data.
[0130] The trained model will be used as a multidimensional classification and prediction model for debris flow disasters.
[0131] Step 6: Multidimensional classification prediction and result presentation of debris flow risk. Based on the watershed classification results of the monitoring area, the corresponding prediction models are selected to process the monitoring data, realize the long-term and short-term prediction of debris flow disaster risk in the monitoring area, and visualize the prediction results.
[0132] Data from debris flow hazard areas are input into a multidimensional classification and prediction model for debris flow disasters. The model first classifies the hazard areas into hillside debris flow or valley debris flow types based on watershed morphology. Then, a short-term prediction model is used to predict the debris flow risk for 1 to 3 days, and a long-term prediction model is used to predict the debris flow risk for 4 to 7 days. The final prediction results are then obtained.
[0133] Table 1. Results of Multidimensional Classification Prediction of Debris Flow Risk
[0134]
[0135] To verify the technical effectiveness of the method of this invention, we conducted comparative experiments on different prediction models. This experimental scheme used a total of 367 debris flow test samples, comparing watershed morphology classification and long-term / short-term classification models. First, we compared the debris flow disaster risk prediction results under the conditions of no classification and classified debris flow hazard areas. Under the same prediction algorithm, the watershed morphology classification prediction model proposed in this invention has an average accuracy 3.5% higher than the unclassified watershed morphology prediction model. Next, we compared the risk prediction results of the long-term debris flow risk prediction model, the short-term debris flow risk prediction model, and the long-term / short-term classification prediction model, respectively. Under the same test samples, comparing the long-term / short-term prediction model, the long-term / short-term prediction model proposed in this invention has an average accuracy 3.2% higher than the long-term prediction model in short-term prediction and an average accuracy 4.3% higher than the short-term prediction model in long-term prediction.
[0136] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A method for multidimensional classification and prediction of debris flow disaster risk, characterized in that, The method includes: Step 1: Delineate the debris flow hazard monitoring area. Delineate the scope of the debris flow hazard monitoring area on the satellite map. Obtain the latitude and longitude of the hazard monitoring area in the four directions of east, west, south, and north using map tools. Use the latitude and longitude range to outline the scope of the debris flow hazard monitoring area. Step 2: Download the digital elevation model of the monitoring area. For the delineated debris flow hazard monitoring area, download the digital elevation model data of the monitoring area from a professional website system. Step 3: Watershed morphology classification of debris flow hazard areas. The downloaded digital elevation model (DEM) data for the monitoring area is cropped to the actual size of the debris flow hazard monitoring area and saved as a new DEM file. The planar curvature, profile curvature, slope, and elevation of the monitoring area are extracted from the DEM file as a dataset. A support vector machine (SVM) algorithm is used for classification to determine whether the monitoring area is a valley-type or hillside-type debris flow hazard area. The classification training method is as follows: Input training set {(X i ,y i )},i=1~N, where y i = +1 or -1, representing the positive or negative sample label, X i The main influencing factors include planar curvature, profile curvature, slope, and elevation; the objective function to be maximized is to be solved. Restrictions: y j = +1 or -1, representing positive or negative sample labels; Where N represents the total number of training sets, K(X) i ,X j ) represents the kernel function, and C represents a penalty factor greater than 0; In selecting the kernel function, a polynomial kernel function is chosen, namely... K(X i ,X j )=(X i ·X j +c) d Where c is a constant term used to adjust the nonlinearity of the kernel function, and d is the degree of the polynomial, which determines the complexity of the polynomial kernel function; C, c, and d need to be tuned during the learning process; When the objective function is maximized, all values of α constitute the optimal solution α. * , where α=(α1,α2,...α N ), and Represents α i The optimal solution after convergence; The optimal solution α is calculated using the following algorithm. * The specific process for determining the value of is as follows: Select two parameters α that need to be updated. i and α j Fix other α k k = 1 to N, k ≠ i, k ≠ j, due to constraints At this point, the constraints become... a i y i +a j y j =c 0≤α i ≤C 0≤α j ≤C in Therefore, we can conclude that... Using α i The expression replacing α j The objective problem can be transformed into an optimization problem with only one constraint: 0 ≤ α. i ≤C; For an optimization problem with only one constraint, for α i Find the partial derivative, set the derivative to 0, and then find the value of the variable. according to Find Iterate multiple times until convergence, and calculate each The optimal solution α is obtained * Construct the classification decision function f(x): Where T is the transpose and sign(·) is the transition function: Step 4: Collect data on disaster influencing factors in debris flow hazard areas, obtain information on debris flow hazard points, and convert the hazard point information into specific latitude and longitude coordinates through a map website; obtain rainfall, soil moisture content, elevation, slope, stratum lithology, vegetation cover, and calculate surface deformation data for the corresponding areas from professional website systems. Step 5: Learning a multidimensional classification prediction model for debris flow risk. Based on the acquired sample data of different types of debris flow disasters, the AdaBoost algorithm is used to learn a short-term debris flow disaster prediction model based on AdaBoost in combination with short-term influencing factors; a support vector machine (SVM) is used to learn a long-term debris flow disaster risk prediction model based on SVM in combination with long-term influencing factors. The prediction models are then validated using a dataset. Finally, from four dimensions, a long-term gully-type debris flow disaster risk prediction model based on SVM, a short-term gully-type debris flow disaster risk prediction model based on AdaBoost, a long-term hillside-type debris flow disaster risk prediction model based on SVM, and a short-term hillside-type debris flow disaster risk prediction model based on AdaBoost are obtained. Step 6: Multidimensional classification prediction and result presentation of debris flow risk. Based on the watershed classification results of the monitoring area, the corresponding prediction models are selected to process the monitoring data, realize the long-term and short-term prediction of debris flow disaster risk in the monitoring area, and visualize the prediction results.
2. The method according to claim 1, characterized in that, The short-term influencing factors include lithology, rainfall, slope, vegetation cover, elevation, soil moisture content, and cumulative rainfall.
3. The method according to claim 1, characterized in that, The long-term influencing factors include stratigraphic lithology, rainfall, slope, vegetation cover, elevation, soil moisture content, and surface deformation.
4. The method according to claim 1, characterized in that, The short-term debris flow disaster prediction model is trained using the following method: (1) Given a training sample dataset: S={(x1z1),...,(x i z i ),...,(x N z N )}, where N is the total number of samples, x i The main influencing factors include stratigraphic lithology, rainfall, slope, vegetation cover, elevation, soil moisture content, and cumulative rainfall. i Let z represent the risk label of the training sample, where z i =1 or -1, representing risk and no risk respectively; (2) The initial weight distribution of the training sample set is as follows: Where: D1(i) is the initial weight distribution of the training sample set; W 1i Each training sample is initially assigned the same weight; (3) Perform m iterations, using a decision tree as the base classifier, and adjust the weights D. m Learning from the training set, calculating the weak classifier G m (x), G m The output value of (x) is {-1, 1}; (4) Calculate G m (x) Classification error rate e on the training set m In the formula: I(G m (x i )≠z i A value of 0 indicates correct classification, while a value of 1 indicates incorrect classification. (5) Calculate the weak classifier G m The coefficient a of (x) m (6) Update the weight distribution of the training data (7) Construct the final classifier According to the weight a of the weak classifier m Combine the various weak classifiers, i.e. By applying the sign function, a strong classifier is obtained as follows:
5. The method according to claim 1, characterized in that, The long-term debris flow disaster prediction model was trained using the following method: Input training set {(X i ,y i )},i=1~N, where y i = +1 or -1, representing the positive or negative sample label, X i The main influencing factors include stratigraphic lithology, rainfall, slope, vegetation cover, elevation, soil moisture content, and surface deformation; the objective function to be maximized is to be solved. Restrictions: y j = +1 or -1, representing positive or negative sample labels; Where N represents the total number of training sets, K(X) i ,X j ) represents the kernel function, and C represents a penalty factor greater than 0; In selecting the kernel function, the Gaussian kernel function is chosen, i.e. C and δ need to be tuned during the learning process; When the objective function is maximized, all values of α constitute the optimal solution α. * , where α=(α1,α2,...α N ), and Represents α i The optimal solution after convergence; The optimal solution α is calculated using the following algorithm. * The specific process for determining the value of is as follows: Select two parameters α that need to be updated. i and α j Fix other α k k = 1 to N, k ≠ i, k ≠ j, due to constraints At this point, the constraints become... a i y i +a j y j =c 0≤α i ≤C 0≤α j ≤C in Therefore, we can conclude that... Using α i The expression replacing α j The objective problem can be transformed into an optimization problem with only one constraint: 0 ≤ α. i ≤C; For an optimization problem with only one constraint, for α i Find the partial derivative, set the derivative to 0, and then find the value of the variable. according to Find Iterate multiple times until convergence, and calculate each The optimal solution α is obtained * Construct the classification decision function f(x): Where T is the transpose and sign(·) is the transition function:
6. The method according to claim 1, characterized in that, Short-term forecasts refer to mudslide disaster risk predictions within three days.
7. The method according to claim 1, characterized in that, Long-term forecast for debris flow disaster risk is four to seven days.