Geological disaster susceptibility evaluation method and system based on sample enhancement

By using sample expansion, credibility assessment, and screening techniques, combined with the XGBoost model, the problems of insufficient sample size and difficulty in quantifying credibility in geological hazard susceptibility assessment were solved, achieving high-precision and high-reliability hazard risk assessment.

CN120931077APending Publication Date: 2025-11-11SUN YAT SEN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511029950.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing methods for assessing geological hazard susceptibility suffer from insufficient sample size and difficulty in quantifying reliability, resulting in insufficient applicability and accuracy of machine learning models in new spatiotemporal contexts.

Method used

A new expanded sample set is generated through the sample expansion stage. The contribution of the influence factor is introduced as a weight coefficient to calculate the weighted cosine similarity, construct a credibility evaluation index, screen high-confidence samples, form an optimized training set, and train it using the XGBoost model.

Benefits of technology

It significantly improved the prediction accuracy and reliability of the geological hazard susceptibility assessment model, and enhanced the model's applicability and generalization ability in different regions and time periods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931077A_ABST
    Figure CN120931077A_ABST
Patent Text Reader

Abstract

The invention discloses a geological disaster susceptibility evaluation method and system based on sample enhancement, and the method comprises the steps: collecting original disaster sample data, carrying out the spatial expansion of a positive sample through the analysis of the spatial features of a disaster influence range, and generating a new enhanced sample set; introducing an influence factor contribution degree as a weight coefficient, and respectively calculating weighted cosine similarity between the expanded sample and the original positive sample; constructing a credibility evaluation index based on a weighted cosine similarity result; performing sample screening according to a preset credibility threshold value; using the screened sample set to train an XGBoost model; and applying the trained model to a target research area. The system comprises a sample enhancement module, a model training module and a model application module. According to the invention, the accuracy and reliability of geological disaster susceptibility evaluation are effectively improved. The method can be widely applied to the field of disaster prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of disaster prediction, and in particular to a method and system for assessing the susceptibility of geological disasters based on sample augmentation. Background Technology

[0003] In recent years, with the development of Geographic Information Systems (GIS) and big data technologies, machine learning methods (such as logistic regression, support vector machines, random forests, and XGBoost) have been widely applied to geological hazard susceptibility modeling due to their ability to automatically learn nonlinear relationships. However, the performance of machine learning models is highly dependent on the quality of training samples, and existing training samples often have the following shortcomings: First, the number of samples is insufficient. Real disaster events are scarce and spatially limited, making it difficult for the sample size to cover all possible combinations of conditions. Therefore, the model is prone to overfitting to a few cases or misjudging in rare scenarios. To increase the number of samples, traditional approaches often simply expand the spatiotemporal range, but due to the evolution of geological environment, climate, and human activities over time and space, the training samples exhibit spatiotemporal heterogeneity, reducing the model's applicability to new spatiotemporal contexts. Second, the reliability of samples is difficult to quantify. Existing methods typically treat historical disaster points as "completely reliable" positive samples, while randomly selecting areas far from disaster points as negative samples, ignoring the uncertainty that may exist in the sample labels themselves. Although some studies have attempted to introduce reliability evaluation, they mostly rely on expert experience and lack objective quantitative indicators. In summary, existing geological hazard susceptibility assessments lack effective sample augmentation and quality control methods, making it difficult to improve the accuracy and generalization ability of machine learning models with limited samples. Summary of the Invention

[0004] In view of this, in order to solve the technical problem of insufficient data samples in existing geological hazard susceptibility assessment methods, which leads to low accuracy of the assessment, the present invention proposes a geological hazard susceptibility assessment method based on sample enhancement, the method comprising the following steps:

[0005] In the sample expansion stage, the original disaster sample data is first collected. By analyzing the spatial characteristics of the disaster's impact range, the positive samples are spatially expanded to generate a new enhanced sample set.

[0006] In the similarity weighting calculation stage, the contribution of the influence factor is introduced as a weight coefficient, and the weighted cosine similarity between the expanded sample and the original positive sample is calculated to quantify their spatial feature correlation.

[0007] In the credibility assessment phase, credibility evaluation indicators are constructed based on the weighted cosine similarity results to quantitatively assess the reliability of the expanded sample and the original sample data.

[0008] In the sample selection and optimization stage, based on the preset confidence threshold, high-confidence augmented samples and original samples are selected to form the optimized final training sample set.

[0009] During the model training phase, the XGBoost model is trained using the selected sample set.

[0010] In the prediction and application phase, the trained model is applied to the target study area to output a spatialized disaster susceptibility probability distribution map, thereby achieving regional disaster risk classification and assessment.

[0011] Secondly, this invention also proposes a geological hazard susceptibility assessment system based on sample enhancement, the system comprising:

[0012] The sample enhancement module is used to handle the sample expansion stage, the similarity weighting calculation stage, the credibility assessment stage, and the sample screening and optimization stage.

[0013] The model training module is used to handle the model training phase.

[0014] The model application module is used to handle the prediction application phase.

[0015] Based on the above scheme, this invention provides a geological hazard susceptibility assessment method and system based on sample augmentation. By combining the hazard impact range and prototype theory to expand the sample, the number of positive samples is effectively increased; the importance of influencing factors is quantified based on deterministic coefficients; a weighted environmental similarity and credibility measurement mechanism is introduced to achieve quantitative assessment of sample label credibility; on this basis, a multi-threshold screening strategy is established to significantly optimize the quality of the training sample set; and combined with the XGBoost model, high-precision hazard susceptibility prediction is achieved. In summary, this invention significantly improves the prediction accuracy and reliability of the geological hazard susceptibility assessment model through innovative sample augmentation and screening techniques, providing quantitative and operable technical support for geological hazard risk management. Attached Figure Description

[0016] Figure 1 This is a flowchart of the steps of a geological hazard susceptibility assessment method based on sample enhancement according to the present invention;

[0017] Figure 2 This is a schematic diagram of sample expansion based on the impact range of geological disasters in a specific embodiment of the present invention;

[0018] Figure 3 This is a structural block diagram of a geological hazard susceptibility assessment system based on sample enhancement according to the present invention. Detailed Implementation

[0019] In addition to the lack of sample size mentioned in the background technology, existing geological hazard susceptibility assessment methods also include physical mechanism methods and statistical empirical methods, which also have technical problems. Physical models can describe the occurrence mechanism of disasters, but require complete physical and hydrological data of soil and rock, and are not suitable for large-scale rapid assessment. Statistical models (such as the analytic hierarchy process, AHP) are relatively simple and effective, but rely on expert experience and have a certain degree of subjectivity.

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] It should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this application can be combined with each other.

[0022] It should be understood that the terms "system," "apparatus," "unit," and / or "module" used in this application are a method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.

[0023] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "a," and / or "the" are not specifically singular and may include the plural. Generally, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements. An element defined by the phrase "comprising an..." does not exclude the presence of other identical elements in the process, method, product, or apparatus that includes the element.

[0024] In the description of the embodiments of this application, "a plurality of" refers to two or more. The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0025] Furthermore, flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, the steps can be processed in reverse order or simultaneously. Additionally, other operations can be added to these processes, or one or more steps can be removed from them.

[0026] Reference Figure 1 This is a flowchart illustrating an optional example of the geological hazard susceptibility assessment method based on sample enhancement proposed in this invention. The method can be applied to computer equipment, and the assessment method proposed in this embodiment may include, but is not limited to, the following steps:

[0027] Step S1: Obtain sample data, using disaster locations as positive sample prototypes and non-disaster locations as negative samples, and expand the positive samples based on the disaster impact range to obtain expanded samples;

[0028] Step S2: Using the contribution of the impact factor as a weighting coefficient, calculate the weighted cosine similarity between the expanded sample and the original positive sample;

[0029] Step S3: Define the credibility of positive samples and negative samples, and calculate the credibility of augmented samples and sample data based on weighted cosine similarity;

[0030] Step S4: Based on the threshold and the confidence level calculated in step S3, the expanded samples and sample data are filtered to obtain the final sample set;

[0031] Step S5: Train the XGBoost model based on the final sample set from step S4;

[0032] Step S6: Based on the disaster probability generation model trained in step S5, process the study area to generate the corresponding disaster susceptibility probability.

[0033] In some feasible embodiments, before step S1, the following steps are also included:

[0034] Influencing factor selection: Topographic, meteorological, human activity, and geological data of the study area were collected, including DEM, elevation, slope, aspect, geological units, land use, distance from rivers / roads, and annual rainfall. The degree of collinearity of the influencing factors was detected using the Pearson correlation coefficient and VIF method, and redundant factors were eliminated. The calculation formulas are shown in Equations 1-2. Fifteen representative influencing factors (e.g., slope, aspect, soil moisture content) were selected for final modeling. Each factor was processed in a uniform raster format and used as input data for subsequent models.

[0035]

[0036] Among them, Xi ,Y i Let X and Y represent the i-th observations of factors X and Y, respectively. Let X and Y represent the average values ​​of factors X and Y, respectively. r is the Pearson correlation coefficient, ranging from -1 to 1, where 1 indicates a perfect positive correlation, -1 indicates a perfect negative correlation, and 0 indicates no correlation.

[0037]

[0038] Among them, R j 2 denoted as the coefficient of determination for a multiple linear regression model with the j-th factor as the dependent variable and all other factors as independent variables.

[0039] In some feasible embodiments, step S1 specifically includes:

[0040] Based on authoritative geological hazard datasets, typical landslides, collapses, ground subsidence, and subsidence points that occurred in the most recent year (e.g., 2024) were extracted as positive sample prototypes. For each prototype point, surrounding grids were extracted and expanded according to its hazard type: for slope-type hazard points such as landslides / collapses, all grid points within the same slope unit were obtained as expanded positive samples; for ground subsidence / subsidence points, grids within a certain range were extracted according to the influence radius of the Chinese geological hazard classification. This resulted in a large-scale positive sample expansion set. After completing the construction of the positive sample library, the remaining non-hazard points in the study area were used as candidate negative sample libraries. The specific expansion diagram is shown below. Figure 2 As shown.

[0041] In some feasible embodiments, the calculation process of the impact factor contribution in step S2 specifically includes:

[0042] To further clarify the contribution of influencing factors to disaster occurrence, a method for calculating the contribution of influencing factors based on the coefficient of determination (CF) is proposed. The coefficient of determination (CF) is used to analyze the sensitivity between various factors influencing the occurrence of disaster events. Its value ranges from -1 to 1, with positive values ​​indicating a positive promotion of disaster occurrence and negative values ​​indicating a negative inhibition of disaster occurrence. The larger the absolute value, the more significant the influence. The calculation method for the coefficient of determination is shown in Formula 3:

[0043]

[0044] Among them, PP a PP represents the ratio of the number of disaster points present in a unit within influence factor a to the area of ​​that unit; s This is a prior probability index for the occurrence of a disaster event across the entire study area, representing a quantitative estimate of the overall susceptibility of the disaster in the study area. PP sIt is calculated by the ratio of the number of disaster points in the study area to the area of ​​the study area.

[0045] Among them, PP a PP represents the ratio of the number of disaster points present in a unit within influence factor a to the area of ​​that unit; s This is a prior probability index for the occurrence of a disaster event across the entire study area, representing a quantitative estimate of the overall susceptibility of the disaster in the study area. PP s It is calculated by the ratio of the number of disaster points in the study area to the area of ​​the study area.

[0046] Based on the CF value of each partition for each impact factor, and combined with the partition weight calculated from the sample size ratio of each partition, the contribution of the impact factor is comprehensively calculated. The calculation formula is shown in Equation 4-5:

[0047]

[0048] Where F represents the contribution of the impact factor; ω i Indicates partition weight; CF i denoted by , indicating the partition certainty coefficient; n is the number of partitions for the impact factor. The screened impact factors are divided into discrete and continuous factors. Discrete factors are partitioned according to the total number of categories, while continuous factors are divided into 10 partitions using the natural discontinuity method. N i This represents the number of disaster points within the partition, where N is the total number of disaster points.

[0049] In some feasible embodiments, the calculation process of the weighted cosine similarity in step S2 specifically includes:

[0050] Based on the influence factor weights, the weighted cosine similarity method is used to evaluate the similarity between the expanded samples and the original positive samples. The formula for calculating the weighted cosine similarity is as follows:

[0051]

[0052] Where, sim(S) k ,P j F represents the similarity between sample B and the j-th prototype positive sample A. i This represents the contribution of the i-th influence factor. A i B i ... k ) represents the weighted cosine similarity of sample B.

[0053] In some feasible embodiments, step S3 specifically includes:

[0054] The weighted cosine similarity in step S2 is directly defined as the positive sample credibility R. P R P A higher value indicates that the point is closer to the environmental conditions of the positive sample, and the more likely a disaster is to occur; correspondingly, the confidence level R of the negative sample is defined. N The calculation formula is shown in Equation 8-9.

[0055] R P =sim(S k )#(6)

[0056] R N =1-sim(S) k )#(7)

[0057] Among them, R P With R N These represent the confidence levels for positive and negative samples, respectively, with values ​​ranging from 0 to 1.

[0058] In some feasible embodiments, step S4 specifically includes:

[0059] R is set based on experimental experience and sample distribution. P and R N The filtering interval, for example, taking R P The threshold is between 0.6 and 0.8, R N The threshold is between 0.45 and 0.6. Only those satisfying R are retained. P Positive samples ≥ threshold and R N Negative samples ≥ a threshold are used as the training set; simultaneously, spatial neighborhood tests are used to remove isolated sample points with discrete distributions, ensuring the spatial continuity of the sample set. Through this process, a high-purity and highly representative final sample set is constructed.

[0060] In some feasible embodiments, step S5 specifically includes:

[0061] The final sample set after screening was divided into a training set (70%) and a test set (30%) using stratified sampling. The XGBoost model was trained using Python's Scikit-learn library, and five-fold cross-validation was used for parameter tuning to optimize model performance. After model training, the accuracy, precision, recall, F1 score, and AUC were evaluated using the test set. In this invention, the XGBoost model outperformed other comparative models on multiple metrics, demonstrating the best overall recognition ability. The specific accuracy differences between different models are shown in Table 1.

[0062] Table 1. Differences in accuracy metrics among different models

[0063] Model accuracy Accuracy Recall rate F1 AUC Logistic Regression 0.859 0.918 0.788 0.848 0.917 Support Vector Machine 0.856 0.949 0.752 0.839 0.913 Random Forest 0.859 0.866 0.848 0.857 0.952 XGBoost 0.863 0.931 0.785 0.852 0.952

[0064] In some feasible embodiments, step S6 specifically includes:

[0065] The trained XGBoost model was used to generate the disaster susceptibility probability for each grid cell in the study area, and a susceptibility distribution map was output in GIS software. The prediction results under different thresholds were compared with historical disaster data to validate the model, analyzing its generalization ability and timeliness. The validation results can be used to guide differentiated disaster prevention and mitigation strategies.

[0066] Based on the above method embodiments, the present invention realizes a complete process from sample enhancement and credibility assessment to high-precision model construction, effectively improving the accuracy and reliability of geological disaster susceptibility assessment.

[0067] Based on the above method, this invention significantly improves upon existing methods in terms of model accuracy, AUC value, validation rate, and generalization ability. Experimental results show that the accuracy and AUC of the machine learning model are greatly improved after adopting the method of this invention. Table 2 shows the changes in model accuracy and AUC value under different positive sample confidence levels. Table 3 shows the changes in model accuracy and AUC value under different negative sample confidence levels.

[0068] Table 2. Changes in model accuracy and AUC value under different positive sample confidence levels.

[0069]

[0070] Table 3. Changes in model accuracy and AUC value under different negative sample confidence levels.

[0071]

[0072]

[0073] Further analysis shows that when the RP threshold is increased from 0 to 0.6, the model accuracy and AUC for each disaster reach their relative peak (e.g., accuracy reaches 0.886 and AUC reaches 0.958 for landslide disaster). Beyond this threshold, performance improvement plateaus or slightly decreases due to a sharp reduction in sample size, reflecting an optimal balance between reliability and quantity. Furthermore, experiments comparing different negative sample reliability thresholds reveal that increasing RN helps stabilize and improve model performance, and the model exhibits higher tolerance for fluctuations in the number of negative samples under high-quality positive sample conditions. Overall, the method of this invention effectively retains the information needed for training while eliminating noisy samples, resulting in a significant improvement in model accuracy and AUC compared to the unenhanced method. This invention also demonstrates superior generalization ability. The training set constructed through sample reliability screening makes the model more focused on robust environmental conditions, making it more applicable to disaster judgments across different times and regions.

[0074] like Figure 3 As shown, a geological hazard susceptibility assessment system based on sample augmentation includes:

[0075] The sample enhancement module is used to perform steps S1-S4;

[0076] The model training module is used to execute step S5;

[0077] The model application module is used to execute step S6.

[0078] The content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0079] A geological hazard susceptibility assessment device based on sample augmentation:

[0080] At least one processor;

[0081] At least one memory for storing at least one program;

[0082] When the at least one program is executed by the at least one processor, the at least one processor implements a sample-enhanced geological hazard susceptibility assessment method as described above.

[0083] The content of the above method embodiments is applicable to the device embodiments. The specific functions implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0084] A storage medium storing processor-executable instructions, which, when executed by a processor, are used to implement a sample-enhanced geological hazard susceptibility assessment method as described above.

[0085] The content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0086] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A method for assessing geological hazard susceptibility based on sample augmentation, characterized in that, Includes the following steps: Obtain sample data and expand the positive sample based on the scope of disaster impact to obtain the expanded sample; The weighted cosine similarity between the expanded sample and the original positive sample is calculated using the contribution of the impact factor as a weighting coefficient. The credibility of the augmented sample and the sample data is calculated based on the weighted cosine similarity. The expanded samples and the sample data are filtered according to the credibility level to obtain the final sample set; The machine learning model is trained based on the final sample set to obtain a trained disaster probability generation model; The disaster probability generation model, once trained, is used to process the study area and generate corresponding disaster susceptibility probabilities.

2. The geological hazard susceptibility assessment method based on sample augmentation according to claim 1, characterized in that, The step of acquiring sample data and expanding the positive sample based on the disaster impact range to obtain the expanded sample specifically includes: Obtain sample data, with disaster-affected areas as positive samples and non-disaster-affected areas as negative samples; For disaster points on a slope, all grid points located within the same slope unit are used as augmented samples; For disaster points on the ground, grid points within a preset range are extracted as expanded samples based on the radius of influence.

3. The geological hazard susceptibility assessment method based on sample augmentation according to claim 2, characterized in that, The formula for calculating the contribution of the impact factor is as follows: Where F represents the contribution of the impact factor; ω i Indicates partition weight; CF i Represents the coefficient of determination for the partition; n is the number of partitions for the impact factor; N i PP represents the number of disaster points within the partition; N is the total number of disaster points; a PP represents the ratio of the number of disaster points present in a unit within influence factor a to the area of ​​that unit; s This is the prior probability index of a disaster event occurring in the study area.

4. The geological hazard susceptibility assessment method based on sample augmentation according to claim 2, characterized in that, The formula for calculating the weighted cosine similarity is as follows: Where, sim(S) k ,P j F represents the similarity between sample B and the j-th prototype positive sample A. i A represents the contribution of the i-th impact factor. i B i Let sim(S) represent the i-th influence factor values ​​of the original positive sample A and sample B, respectively, and M represent the number of prototype positive samples of different geological hazard types. k ) represents the weighted cosine similarity of sample B.

5. The geological hazard susceptibility assessment method based on sample enhancement according to claim 1, characterized in that, The step of filtering the expanded samples and the sample data according to the credibility to obtain the final sample set specifically includes: Set the threshold according to the sample distribution; The expanded samples and the sample data are filtered according to the threshold and the confidence level, and isolated sample points with discrete distribution are removed by spatial neighborhood test to obtain the final sample set.

6. The geological hazard susceptibility assessment method based on sample enhancement according to claim 5, characterized in that, The step of training the machine learning model based on the final sample set to obtain the trained disaster probability generation model specifically includes: The final sample set is divided into a training set and a test set according to a preset ratio; The XGBoost model was trained based on the training set, and the parameters were adjusted using the five-fold cross-validation method to obtain the trained disaster probability generation model. The trained disaster probability generation model is evaluated based on the test set.

7. The geological hazard susceptibility assessment method based on sample enhancement according to claim 5, characterized in that, The step of processing the tested area based on the trained disaster probability generation model to generate the corresponding disaster susceptibility probability specifically includes: The trained disaster probability generation model is applied to the study area to generate the disaster susceptibility probability of each grid cell in the study area and output a susceptibility distribution map. The probability of disaster susceptibility under different thresholds was compared and verified with actual data.

8. A geological hazard susceptibility assessment system based on sample augmentation, characterized in that, include: The sample enhancement module is used to acquire sample data and expand positive samples based on the scope of disaster impact to obtain expanded samples; The weighted cosine similarity between the expanded sample and the original positive sample is calculated using the contribution of the impact factor as a weighting coefficient; the credibility of the expanded sample and the sample data is calculated based on the weighted cosine similarity; the expanded sample and the sample data are then filtered according to the credibility to obtain the final sample set. The model training module trains the machine learning model based on the final sample set to obtain a trained disaster probability generation model. The model application module processes the tested area based on the trained disaster probability generation model to generate the corresponding disaster susceptibility probability.

9. A geological hazard susceptibility assessment device based on sample augmentation, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the sample-enhanced geological hazard susceptibility assessment method as described in any one of claims 1-7.

Citation Information

Cited By

  • Sample dynamic supplement method, system and equipment based on environment covariable and medium

    CN121505392A