Debris flow susceptibility evaluation method based on different negative sample sampling strategies

By comparing different negative sample sampling strategies, combining the random forest model and multiple evaluation indicators, the optimal strategy is selected and a high-quality negative sample set is constructed, which solves the problem of poor generalization ability in the mudslide flow proneness evaluation model, and achieves higher prediction accuracy and generalization ability, supporting disaster management and prevention and control.

CN120429573APending Publication Date: 2025-08-05CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510508677.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

In the existing mudslide proneness evaluation model, the negative sample sampling strategy lacks systematicity, resulting in poor generalization capabilities of the model and is prone to misjudgment when facing new non-disaster scenarios.

Method used

By comparing different negative sample sampling strategies, combining the random forest model and a variety of evaluation indicators, the optimal negative sample sampling strategy is selected to improve the quality of negative samples, build a high-quality negative sample set, and improve the prediction accuracy and generalization ability of the model.

Benefits of technology

The accuracy and generalization ability of the mudslide flow susceptibility evaluation model are improved, and the mudslide flow susceptibility can be predicted more accurately, providing a scientific basis for disaster management and prevention and control strategies, and reducing potential risks and economic losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429573A_ABST
    Figure CN120429573A_ABST
Patent Text Reader

Abstract

The invention discloses a debris flow susceptibility evaluation method based on different negative sample sampling strategies. The debris flow susceptibility evaluation method comprises the steps of collecting and sorting basic data of a research area; based on the collected and sorted basic data, evaluation factors are selected, and correlation analysis is carried out on the evaluation factors; dividing a research area drainage basin unit based on an ArcGIS hydrological analysis technology, and selecting a drainage basin unit with a threshold value of 1k as an analysis object; selecting a negative sample sampling strategy; constructing models of different negative sample sampling strategies through a random forest; and comprehensively evaluating each model and selecting an optimal negative sample sampling strategy. By comparing different negative sample sampling strategies, the optimal sampling strategy can be selected, the high-quality negative sample set is further constructed, high-quality data support is provided for subsequent analysis, the precision and generalization ability of the susceptibility evaluation model can be effectively improved, and the debris flow susceptibility can be predicted more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an optimization method for different negative sample sampling strategies, specifically a debris flow susceptibility evaluation method based on different negative sample sampling strategies, belonging to the technical field of geological disaster prediction and risk assessment. Background Art

[0002] Debris flows are a typical sudden geological disaster, primarily occurring during heavy rain, glacier melt, or extreme weather conditions. Characterized by the high-velocity flow of a mixture of solid matter and water, they are extremely destructive. In recent years, with the intensification of global climate change and the frequent occurrence of extreme weather events, the incidence of debris flows has shown a significant shift. On the one hand, the increasing intensity and frequency of precipitation, particularly in mountainous areas, has led to rapid confluence of surface runoff, increasing the probability of debris flows. On the other hand, glacial melt has exacerbated the potential for debris flows in high-altitude areas. Furthermore, human activities, such as irrational land use, excessive logging, reclamation of steep slopes, and mineral resource development, have also increased the frequency and severity of debris flows, particularly in areas with low vegetation cover and steep terrain. These combined factors have made the prevention and control of debris flows increasingly severe, necessitating the urgent need to strengthen monitoring, early warning, and comprehensive management capabilities to ensure regional security and sustainable development.

[0003] In recent years, machine learning technology has been widely used in geological hazard risk assessment and prediction. In the prior art, 1) the geological disaster risk assessment model that integrates soil moisture, deformation and rainfall information disclosed in the patent with publication number CN119204644A, by introducing soil moisture and deformation information, establishes a soil moisture-deformation-rainfall fusion model, thereby effectively improving the accuracy of rainfall-induced geological disaster monitoring and early warning, optimizing the existing geological disaster risk assessment model, and providing technical support for disaster prevention and mitigation work; 2) a method for assessing debris flow activity in arid areas based on the identification of sedimentary fan sediment characteristics disclosed in the patent with publication number CN118379649A, which can automatically and repeatedly perform disaster analysis on a large area and a large range, reducing the analysis time and cost of inefficient methods such as manual surveys and on-site monitoring, and the results are stable, standardized, iteratively optimized and applicable to different environments and scenarios; 3) a method for assessing the susceptibility of debris flow in a watershed unit based on multi-model Stacking integration disclosed in the patent with publication number CN119312204A, which efficiently fuses multi-source data and creatively combines the mixed probability vector with the original Sample splicing enables the meta-model to simultaneously utilize the intrinsic characteristics of the original data and the predictive information of the base model during training, which can reduce the interference of irrelevant information and the dependence on the quality of the base model, and can more comprehensively capture the key information in the data, thereby enhancing the model's ability to interpret complex data and the prediction accuracy; 4) A typhoon-rainstorm-type debris flow susceptibility evaluation method based on an integrated learning algorithm, disclosed in publication number CN119337248A, selects influencing factors for the study area, establishes an evaluation index system, and obtains relevant data based on the index requirements; uses ArcGIS to extract factor data and preprocesses it; then uses the SMOTE algorithm to oversample the debris flow occurrence samples; then uses an integrated optimization algorithm based on the optimization model to optimize the CatBoost regression model parameters; and uses Sklearn to perform k-fold cross validation to evaluate the model stability and identify the best model; finally, the best model is used to draw a debris flow susceptibility map, and analyzes the impact of rainfall data in different time ranges on the accuracy of debris flow susceptibility prediction. Through the above techniques, we can see that the quality of the sample set has a crucial impact on the accuracy of the model. However, most current debris flow research focuses on the acquisition and analysis of positive samples during the data collection and processing process. For negative samples, they often simply select data from periods or areas where no disasters have occurred, lacking a systematic sampling strategy. This arbitrary sampling method leads to many problems with negative samples, resulting in poor model generalization ability. When faced with new non-disaster scenarios, they are easily misjudged as disaster scenarios, resulting in large deviations in the prediction results. Summary of the Invention

[0004] The purpose of the present invention is to provide a debris flow susceptibility evaluation method based on different negative sample sampling strategies in order to solve at least one of the above technical problems. By comparing different negative sample sampling strategies, combining the random forest model and using multiple evaluation indicators, the optimal negative sample sampling strategy is selected to improve the quality of negative samples, thereby improving the prediction accuracy and generalization ability of the geological disaster susceptibility model.

[0005] The present invention achieves the above-mentioned object through the following technical solutions: a debris flow susceptibility evaluation method based on different negative sample sampling strategies, the debris flow susceptibility evaluation method comprising the following steps:

[0006] S100, collect and organize basic data of the study area;

[0007] S200. Based on the collected and organized basic data, select evaluation factors and conduct correlation analysis on the evaluation factors;

[0008] S300, based on ArcGIS hydrological analysis technology, the watershed units of the study area were divided, and the watershed units with a threshold of 1k were selected as the analysis objects;

[0009] S400, selection of negative sample sampling strategy;

[0010] S500, building models with different negative sample sampling strategies through random forest;

[0011] S600: Comprehensively evaluate each model and select the optimal negative sample sampling strategy.

[0012] As a further solution of the present invention: in S100, the basic data includes but is not limited to existing regional geological data, field survey data, meteorological and hydrological data, remote sensing image data and other relevant documents.

[0013] As a further embodiment of the present invention: in S200, the selected evaluation factors include but are not limited to slope, aspect, curvature, elevation variation coefficient, water flow intensity index, distance to road, distance to river, normalized difference vegetation index (NDVI), soil type, land use and lithology;

[0014] The correlation analysis of the evaluation factors was conducted using the Pearson correlation coefficient. The calculation formula for the Pearson correlation coefficient is:

[0015]

[0016] Where: PCC is the Pearson correlation coefficient; x i ,y iThe value of r is between -1 and +1. If r>0, it indicates that the two variables are positively correlated; if r<0, it indicates that the two variables are negatively correlated. The larger the absolute value of r, the stronger the correlation. If r=0, it indicates that there is no linear correlation between the two. When r>0.7, the correlation feature is obvious and the correlation factor needs to be eliminated.

[0017] As a further solution of the present invention: in S300, the specific operation process of dividing the watershed units of the study area through ArcGIS hydrological analysis technology is: filling the basic DEM data, calculating the flow direction, calculating the cumulative water volume, and setting the flow threshold to 1k, generating the river network, generating the watershed and converting the raster into a surface.

[0018] As a further solution of the present invention: in S400, the selected negative sample sampling strategy includes but is not limited to random sampling, gentle slope sampling, frequency ratio sampling and semi-supervised learning sampling.

[0019] As a further solution of the present invention: in S500, an evaluation model for different negative sample sampling strategies is constructed based on random forest, including the following steps:

[0020] S501, determining negative sample sets of four sampling methods: random sampling, gentle slope sampling, frequency ratio sampling, and semi-supervised learning sampling;

[0021] S502, machine learning dataset construction;

[0022] S503. Build four negative sample sampling strategy evaluation models based on RF.

[0023] As a further solution of the present invention: in S501, random sampling is to randomly select negative sample units in the watershed unit; gentle slope sampling is to select negative sample units in the gentle slope area of the watershed unit; frequency ratio sampling is to select negative sample units in the extremely low and low susceptibility intervals based on the susceptibility evaluation of the watershed unit; semi-supervised learning sampling first randomly selects negative samples based on the watershed unit, and then uses the RF model to evaluate the susceptibility, and based on the RF model evaluation, further reselects, that is, reselects negative sample units in the extremely low and low susceptibility intervals; ordinary negative sample sets are obtained based on random sampling and gentle slope sampling, and high-quality negative sample sets are obtained based on frequency ratio sampling and semi-supervised learning.

[0024] As a further solution of the present invention: in S502, according to the number of positive sample sets, the same number of negative samples are selected for the four sampling strategies, that is, positive samples: negative samples = 1:1, thereby constructing a machine learning data set.

[0025] As a further solution of the present invention: In S503, four negative sample sampling strategy evaluation models are built based on random forest (RF), and the susceptibility of the four negative sample sampling strategies is evaluated using Python; the calculation formula of the random forest model is expressed as follows:

[0026]

[0027] Where: I(y t =ξ) is the indicator function, if y t =ξ, then the label is 1, otherwise it is 0; the number of nodes is represented by T; y t For decision tree.

[0028] As a further solution of the present invention: In S600, the four different negative sample sampling strategy evaluation models are comprehensively evaluated and the optimal selection is made through the area under the ROC curve, the mean sensitivity, and the standard deviation of sensitivity. The principles and formulas of the area under the ROC curve (AUC), the mean sensitivity (MS), and the standard deviation of sensitivity (SD) are as follows:

[0029] The ROC curve shows the performance of the model at different thresholds by plotting the relationship between the true positive rate (TPR) and the false positive rate (FPR). The calculation formula is as follows:

[0030]

[0031] Where: TP represents the number of true positives, FP represents the number of false positives, TN represents the number of true negatives, and FN represents the number of false negatives;

[0032] The mean sensitivity (MS) represents the average value of the model sensitivity, reflecting the overall average level of the model; the standard deviation (SD) of sensitivity measures the dispersion of sensitivity and reflects the consistency of model performance. The calculation formula is as follows:

[0033]

[0034] Where: x i represents the sensitivity of the i-th sample; n represents the total number of samples.

[0035] The beneficial effects of the present invention are:

[0036] (1) By comparing different negative sample sampling strategies, the present invention not only selects the optimal strategy to construct a high-quality negative sample set, but also improves the accuracy and generalization ability of the susceptibility evaluation model. At the same time, it will provide an optimization method for disaster management and prevention and control strategies, thereby effectively reducing potential risks and economic losses. By horizontally comparing multiple negative sample sampling strategies and selecting the optimal strategy, the quality of negative samples in the susceptibility model is improved, thereby improving the prediction accuracy and generalization ability of the model, thereby better serving the prediction and risk assessment of geological disasters.

[0037] (2) The present invention has significant advantages by comparing different negative sample sampling strategies: on the one hand, it can select the optimal sampling strategy and then construct a high-quality negative sample set to provide high-quality data support for subsequent analysis; on the other hand, it can effectively improve the accuracy and generalization ability of the susceptibility evaluation model, so that it can more accurately predict the susceptibility of debris flows and provide a scientific optimization method for disaster management and prevention and control strategies, thereby effectively reducing potential risks and reducing economic losses. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0039] Figure 2 This is the Pearson correlation coefficient heat map of the evaluation factors in Example 2;

[0040] Figure 3 This is the flow chart for extracting watershed units in Example 3;

[0041] Figure 4 These are negative sample images of different sampling strategies in Example 5;

[0042] Figure 5 This is a result diagram of the susceptibility evaluation model generated based on different negative sample data sets in Example 5;

[0043] Figure 6 1 is the ROC curve diagram of the four evaluation models in Example 6;

[0044] Figure 7 This is the probability distribution diagram of the sensitivity of the four evaluation models in Example 6. DETAILED DESCRIPTION

[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0046] Example 1, as Figure 1 As shown, this embodiment provides a debris flow susceptibility evaluation method based on different negative sample sampling strategies, and the debris flow susceptibility evaluation method specifically includes the following steps:

[0047] S100. Collect and organize basic data of the study area, including existing regional geological data, field survey data, meteorological and hydrological data, remote sensing image data and other relevant literature;

[0048] S200, based on the collected and organized basic data, selecting evaluation factors and performing correlation analysis on the evaluation factors;

[0049] S300, based on ArcGIS hydrological analysis technology, the watershed units of the study area were divided, and the watershed units with a threshold of 1k were selected as the analysis objects;

[0050] S400, selection of negative sample sampling strategy;

[0051] S500, building models with different negative sample sampling strategies through random forest;

[0052] S600: Comprehensively evaluate each model and select the optimal negative sample sampling strategy.

[0053] In addition to all the technical features of the first embodiment, the second embodiment also includes:

[0054] like Figure 2 As shown in Figure 2, in S200, the following 11 evaluation factors were selected through the collection and organization of basic data: slope, aspect, curvature, coefficient of elevation variation (ECV), stream intensity index (SPI), distance to roads, distance to rivers, normalized difference vegetation index (NDVI), soil type, land use, and lithology.

[0055] The Pearson correlation coefficient (PCC) was used to examine the correlation between the evaluation factors. Figure 2 The following is a heat map of the Pearson correlation coefficients among the 11 evaluation factors. Statistics show that the absolute values of the correlation coefficients among the 11 factors are all less than 0.7, indicating that these factors can be used in subsequent research.

[0056] The Pearson correlation coefficient formula is:

[0057]

[0058] Where: PCC is the Pearson correlation coefficient; x i ,y i The mean of .

[0059] Example 3, as Figure 3 As shown, in addition to all the technical features of the first embodiment, this embodiment also includes:

[0060] In S300, the watershed units of the study area are divided by ArcGIS hydrological analysis technology. The specific operation process is: fill the basic DEM data, calculate the flow direction, calculate the cumulative water volume, set the flow threshold to 1k, generate the river network, generate the watershed and convert the raster to surface. The specific extraction process is as follows: Figure 3 shown.

[0061] In addition to all the technical features of the first embodiment, the fourth embodiment further includes: in S400, the negative sample sampling strategies selected include random sampling, gentle slope sampling, frequency ratio sampling (FR), and semi-supervised learning sampling. The principles of the four sampling methods are described as follows:

[0062] Random sampling is the simplest and most fundamental sampling method. It randomly selects non-disaster data points from a sample set, with each sample having equal probability of being selected. This method is simple, intuitive, easy to implement, and requires no prior knowledge. However, in an unevenly distributed dataset, this sampling method can result in samples from important categories or regions being selected with low probability, potentially selecting non-disaster samples in potential disaster areas.

[0063] Gentle slope sampling is a strategy specifically designed for terrain-related research. Its core concept is to sample negative samples from areas with gentler slopes within the study area's watershed units. In practice, debris flows often occur in areas with high slopes. Therefore, sampling negative samples from these areas is beneficial for obtaining high-quality negative samples. Compared to random sampling, this method significantly reduces the potential for blindness. However, it is undeniable that this method may overlook the representation of non-gentle slope areas, resulting in insufficient model generalization.

[0064] Frequency ratio model sampling is a sampling strategy that combines the results of geohazard susceptibility zoning with efficient negative sample screening. Its core is to reselect negative samples from the extremely low and low susceptibility intervals in the initial susceptibility zoning. This eliminates low-quality negative samples that may contain noise or mislabeling, thereby generating a high-quality negative sample set.

[0065] Semi-supervised learning sampling is a sampling method that combines unlabeled data with a small amount of labeled data. Its core concept is to improve understanding of the overall data distribution by leveraging unlabeled data. This method fully exploits the information in unlabeled data and is significantly effective in improving negative datasets. The specific steps are: 1) training a preliminary model using a small amount of existing labeled data; 2) using this model to predict the unlabeled data and generate pseudo-labels; 3) selecting samples from the pseudo-labeled data with high uncertainty or potentially rich information for further labeling or sampling.

[0066] Example 5: In addition to all the technical features of Example 1, this example also includes:

[0067] like Figure 4 and Figure 5 As shown, in S500, an evaluation model for different negative sample sampling strategies is constructed based on random forest, including the following steps:

[0068] S501, determine the negative sample sets of the four sampling methods. Random sampling is to randomly select negative sample units in the watershed unit ( Figure 4 a); Slope sampling is to select negative sample units in the gentle slope area of the watershed unit ( Figure 4 b); Frequency ratio sampling is based on the evaluation of the susceptibility of watershed units, and negative sample units are selected in the extremely low and low susceptibility intervals ( Figure 4 c); Semi-supervised learning sampling first randomly selects negative samples based on the watershed unit, then uses the RF model to evaluate the susceptibility. Based on the RF evaluation, further reselection is performed, that is, reselecting negative sample units in the extremely low and low susceptibility ranges. Based on random sampling and gentle slope sampling, a common negative sample set is obtained, and a high-quality negative sample set is obtained through frequency ratio sampling (FR) and semi-supervised learning ( Figure 4 d).

[0069] S502: Machine learning dataset construction: Based on the number of positive sample sets, the same number of negative samples is selected for the four sampling strategies, i.e., positive sample: negative sample = 1:1, thereby constructing a machine learning dataset.

[0070] S503. Build four negative sample sampling strategy evaluation models based on RF. Four negative sample sampling strategy evaluation models are built based on random forest (RF).

[0071] The calculation formula of the random forest model is as follows:

[0072]

[0073] Where: I(y t =ξ) is the indicator function, if y t =ξ, then the label is 1, otherwise it is 0; the number of nodes is represented by T; y t For decision tree.

[0074] The susceptibility of four negative sample sampling strategies was evaluated using Python. In order to facilitate the comparison of the prediction results of different negative sample sampling strategies, the natural breakpoint method was used to divide the susceptibility levels into extremely low susceptibility area, low susceptibility area, medium susceptibility area, high susceptibility area and extremely high susceptibility area ( Figure 5). According to the results of the susceptibility level classification, the susceptibility level distribution of the four negative sample sampling strategies is roughly the same, but in the gentle slope sampling, the susceptibility value in the central and eastern parts of the study area is generally low, while the susceptibility of the other three types is extremely high. Analysis shows that this may be caused by the sampling method. However, in other areas of the study area, among the four sampling strategies, the extremely high and high susceptibility areas are mostly distributed in the central and northern regions, while the susceptibility levels in the southern region are all low. However, in semi-supervised sampling, the distribution range of extremely high susceptibility areas is wider ( Figure 5 d).

[0075] In order to further compare the distribution of disasters in each susceptibility range under different sampling strategies, the disaster statistics of the four sampling strategies are shown in Table 1. It can be seen from the table that for random sampling, the total area occupied by the extremely high and high susceptibility areas is 58.15 km 2 , accounting for 45.22% of the study area. However, the disaster area in the extremely high and high-prone areas accounts for 91.85%. For the gentle slope sampling, the debris flow disaster area in the extremely high and high-prone areas is 0.887km 2 , accounting for about 83.13% of the debris flow area, while the extremely high and high prone areas account for about 42.07% of the study area. For frequency ratio sampling, the extremely high and high prone areas account for an area of 64.24 km 2 , accounting for approximately 49.96% of the study area, and the identified debris flow area accounts for approximately 95.03% of the total area. For semi-supervised sampling, although 98.24% of debris flow disasters occurred in extremely high and high susceptibility areas, the area of these areas also accounted for 66.14%. From the perspective of frequency ratio, the frequency ratio values of various sampling methods increase with increasing susceptibility level. In particular, the frequency ratio values are exceptionally high in extremely high and high susceptibility areas, indicating a good segmentation effect. For the four sampling methods, the frequency ratio values for extremely high susceptibility areas are ranked from highest to lowest: frequency ratio sampling, random sampling, gentle slope sampling, and semi-supervised sampling. For semi-supervised sampling, the frequency ratio values for the high susceptibility area are only 0.480, and the frequency ratio for the extremely high susceptibility area is only 1.552, indicating that this sampling strategy does not provide a good segmentation effect. In summary, the FR negative sampling method achieves the best segmentation effect.

[0076] Table 1 Statistics of different negative sample sampling strategies

[0077]

[0078] Example 6: Figure 6 and Figure 7 As shown, in addition to all the technical features of the first embodiment, this embodiment also includes:

[0079] The four different negative sample sampling strategy evaluation models are comprehensively evaluated by the area under the ROC curve (AUC), mean sensitivity (MS), and standard deviation of sensitivity (SD). The principles and formulas of the area under the ROC curve (AUC), mean sensitivity (MS), and standard deviation of sensitivity (SD) are as follows:

[0080] The ROC curve is a tool used to evaluate the performance of a binary classification model. It shows the performance of the model at different thresholds by plotting the relationship between the true positive rate (TPR) and the false positive rate (FPR). The calculation formula is as follows:

[0081]

[0082] Where: TP represents the number of true positives, FP represents the number of false positives, TN represents the number of true negatives, and FN represents the number of false negatives;

[0083] The mean sensitivity (MS) represents the average of the model's sensitivity, reflecting the overall average level of the model. The standard deviation (SD) of sensitivity measures the dispersion of sensitivity and reflects the consistency of model performance. The calculation formula is as follows:

[0084]

[0085] Where: x i represents the sensitivity of the i-th sample; n represents the total number of samples.

[0086] The ROC curves and AUC values of the four susceptibility models are as follows Figure 6 As shown in the figure, random sampling, a relatively simple and untargeted sampling method, performs poorly, with an AUC value of only 0.798. This means that the model under this strategy has a large error in distinguishing positive and negative samples, making it prone to misclassification and limiting overall prediction accuracy. In contrast, the AUC value for gentle slope sampling increases to 0.833. This improvement clearly demonstrates the advantage of targeting gentle slopes for negative sample screening. Compared to random sampling, this strategy can more effectively capture representative features of negative samples, thereby reducing misclassifications during the model discrimination process and significantly improving model accuracy. However, two more advanced and sophisticated sampling methods, frequency ratio sampling and semi-supervised sampling, demonstrate even superior performance, driving models with far higher accuracy than the previous two sampling methods. FR sampling achieves an AUC value of 0.855, and semi-supervised sampling achieves an AUC value of 0.862. This shows that based on the basic susceptibility results, negative sample sampling in the extremely low and low susceptibility intervals, whether based on FR sampling or semi-supervised sampling, can further improve the model accuracy.

[0087] The sensitivity probability distribution is as follows Figure 7As shown in the figure, in random sampling ( Figure 7 a), the probability of susceptibility is more abundant in the extremely low and extremely high areas, while in other intervals it shows a trend of first decreasing and then increasing. Figure 7 b), the number of distributions in the middle area (0.3-0.7) is relatively small, and the probability values of susceptibility are mostly distributed in the range of <0.2 and >0.8. Figure 7 c), except for the distribution of extreme susceptibility probability values, the distribution of other values is roughly the same. In semi-supervised sampling ( Figure 7 d) The distribution is relatively extreme, mainly distributed at the two ends of 0-1, and the largest number is close to 1, which means that the overall susceptibility value of the study area is high, which is extremely unreasonable.

[0088] For the mean sensitivity (MS) of various sampling methods, semi-supervised sampling > gentle slope sampling > random sampling > FR sampling; for the standard deviation (SD) of sensitivity, FR sampling < random sampling < gentle slope sampling < semi-supervised sampling. This suggests that FR sampling can reflect a large number of debris flow hazards using a small number of units with high susceptibility indices.

[0089] In summary, comparing all sampling methods, the evaluation model generated by the frequency ratio and semi-supervised learning negative sampling strategies outperforms random sampling and gentle slope sampling. However, while the semi-supervised sampling strategy achieves the highest accuracy, its performance is not as good as expected based on the susceptibility statistics, the predicted probability distribution of debris flow susceptibility, and the mean sensitivity (MS) and standard deviation (SD) of the sensitivity. Therefore, in this example, frequency ratio sampling is selected as the optimal sampling strategy.

[0090] Working principle: First, for the study area, all kinds of basic data are comprehensively collected and systematically organized; then, evaluation factors related to debris flow susceptibility are selected, and factor correlation analysis is carried out to lay a solid foundation for subsequent research; then, the watershed units of the study area are reasonably divided to provide a clear spatial range for subsequent negative sample sampling and model construction; in the negative sample sampling link, different negative sample sampling strategies are selected to fully explore the advantages of different strategies; based on the random forest algorithm, corresponding models are constructed for each negative sample sampling strategy; finally, a variety of evaluation indicators are used to conduct a comprehensive evaluation of each model to screen out the optimal negative sample sampling strategy.

[0091] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

[0092] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. A debris flow susceptibility assessment method based on different negative sample sampling strategies, characterized by: The debris flow susceptibility assessment method comprises the following steps: S100, collect and organize basic data of the study area; S200, based on the collected and organized basic data, selecting evaluation factors and performing correlation analysis on the evaluation factors; S300, based on ArcGIS hydrological analysis technology, the watershed units of the study area were divided, and the watershed units with a threshold of 1k were selected as the analysis objects; S400, selection of negative sample sampling strategy; S500, building models with different negative sample sampling strategies through random forest; S600: Comprehensively evaluate each model and select the optimal negative sample sampling strategy.

2. The debris flow susceptibility assessment method according to claim 1, characterized in that: In S100, the basic data includes but is not limited to existing regional geological data, field survey data, meteorological and hydrological data, remote sensing image data and other relevant documents.

3. The debris flow susceptibility assessment method according to claim 1, characterized in that: In S200, the evaluation factors include but are not limited to slope, aspect, curvature, elevation variation coefficient, flow intensity index, distance to road, distance to river, normalized difference vegetation index, soil type, land use and lithology; The correlation analysis of the evaluation factors was conducted using the Pearson correlation coefficient. The calculation formula for the Pearson correlation coefficient is: Where: PCC is the Pearson correlation coefficient; x i ,y i The value of r is between -1 and +1. If r>0, it indicates that the two variables are positively correlated; if r<0, it indicates that the two variables are negatively correlated. The larger the absolute value of r, the stronger the correlation. If r=0, it indicates that there is no linear correlation between the two. When r>0.7, the correlation feature is obvious and the correlation factor needs to be eliminated.

4. The debris flow susceptibility assessment method according to claim 1, characterized in that: In the above S300, the specific operation process of dividing the watershed units of the study area based on ArcGIS hydrological analysis technology is: filling the basic DEM data, calculating the flow direction, calculating the cumulative water volume, setting the flow threshold to 1k, generating the river network, generating the watershed and converting the raster into a surface.

5. The debris flow susceptibility assessment method according to claim 1, characterized in that: In S400, the negative sample sampling strategy includes but is not limited to random sampling, gentle slope sampling, frequency ratio sampling and semi-supervised learning sampling.

6. The debris flow susceptibility assessment method according to claim 1, characterized in that: In S500, constructing models of different negative sample sampling strategies by random forest includes the following steps: S501, determining negative sample sets of four sampling methods: random sampling, gentle slope sampling, frequency ratio sampling, and semi-supervised learning sampling; S502, machine learning dataset construction; S503. Build four negative sample sampling strategy evaluation models based on RF.

7. The debris flow susceptibility assessment method based on different negative sample sampling strategies according to claim 6, characterized in that: In S501, the random sampling is to randomly select negative sample units in the watershed unit; the gentle slope sampling is to select negative sample units in the gentle slope area of the watershed unit; the frequency ratio sampling is to select negative sample units in the extremely low and low susceptibility intervals based on the susceptibility evaluation of the watershed unit; the semi-supervised learning sampling first randomly selects negative samples based on the watershed unit; Then, the RF model is used to evaluate the susceptibility. Based on the RF model evaluation, further reselection is performed, that is, negative sample units are reselected in the extremely low and low susceptibility intervals. Ordinary negative sample sets are obtained based on random sampling and gentle slope sampling, while high-quality negative sample sets are obtained through frequency ratio sampling and semi-supervised learning.

8. The debris flow susceptibility assessment method according to claim 7, characterized in that: In S502, according to the number of positive sample sets, the same number of negative samples are selected for the four sampling strategies, that is, positive samples: negative samples = 1:1, thereby constructing a machine learning data set.

9. The debris flow susceptibility assessment method according to claim 8, characterized in that: In S503, four negative sample sampling strategy evaluation models are built based on random forest, and the susceptibility of the four negative sample sampling strategies is evaluated using Python; the calculation formula of the random forest model is as follows: Where: I(y t =ξ) is the indicator function, if y t =ξ, then the label is 1, otherwise it is 0; the number of nodes is represented by T; y t For decision tree.

10. The debris flow susceptibility assessment method according to claim 9, characterized in that: In S600, four different negative sample sampling strategy evaluation models are comprehensively evaluated and the optimal selection is made through the area under the ROC curve, the average sensitivity, and the standard deviation of the sensitivity; The ROC curve shows the performance of the model at different thresholds by plotting the relationship between the true positive rate and the false positive rate. The calculation formula is as follows: Where: TP represents the number of true positives, FP represents the number of false positives, TN represents the number of true negatives, and FN represents the number of false negatives; The average sensitivity value represents the average value of the model sensitivity, reflecting the average level of the model as a whole. The standard deviation of sensitivity measures the degree of dispersion of sensitivity and reflects the consistency of model performance. The calculation formula is as follows: Where: x i represents the sensitivity of the i-th sample; n represents the total number of samples.

Citation Information

Patent Citations

  • Method for evaluating debris flow activity in arid region based on accumulation fan deposition feature recognition

    CN118379649A

  • Geological disaster risk assessment model fusing soil humidity, deformation and rainfall information

    CN119204644A

  • Basin unit debris flow susceptibility evaluation method based on multi-model Stacking integration

    CN119312204A

  • Typhoon rainstorm debris flow susceptibility evaluation method based on ensemble learning algorithm

    CN119337248A