An interpretable landslide susceptibility assessment method and medium
By constructing mechanism-geography-physics-mathematical mapping and FP-Growth algorithms, the negative samples and mining correlation rules are systematically screened, and the subjectivity of factor selection and sample selection randomness in landslide prone modeling is solved, which achieves high accuracy and interpretability of landslide disaster prone assessment, and improves the model's prediction ability in large areas.
Patent Information
- Application Number
- CN202510765622.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing landslide prone modeling methods have insufficient subjectivity of factor selection, sample selection randomness and evaluation results, which leads to unstable model prediction performance and low reliability, making it difficult to achieve high-precision and interpretable landslide disaster assessment in large areas.
By constructing mechanism-geography-physics-mathematical mapping, negative samples were screened using the frequency ratio method, combining the FP-Growth algorithm to mine the correlation rules, an interpretable landslide disaster susceptibility assessment model was established, and a comprehensive frequency ratio and similarity calculation was used to guide the interpretation of the model results.
It improves the accuracy and interpretability of landslide disaster susceptibility assessment, reduces the subjectivity of factor selection and the randomness of sample selection, provides scientific, reasonable and easy-to-understand landslide susceptibility assessment results, and improves the generalization ability and early warning accuracy of the model in complex scenarios.
Smart Images

Figure CN120316618B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of spatiotemporal data mining, and in particular to an interpretable landslide hazard susceptibility assessment method and medium. Background Art
[0002] Landslides are a serious global geological hazard characterized by high risk and widespread distribution. Landslide occurrence is influenced by multiple factors, including geology, watershed conditions, land use, and construction activities. Their triggering factors are often related to rainfall or groundwater level fluctuations. These fluctuations reduce soil cohesion and slope stability, thereby suppressing the safety factor. Landslide susceptibility modeling is a crucial foundation for disaster management. Its core purpose is to analyze historical landslide data and related environmental factors to uncover implicit patterns between landslide instances and environmental factors, thereby accurately predicting the likelihood of future landslides. Therefore, conducting landslide susceptibility research can predict their spatial distribution and is of great significance to land use planning, ecological and environmental protection, and the formulation of disaster prevention and mitigation policies.
[0003] Landslides are actually the result of the interaction of geographical factors, such as geological structure, soil composition, and vegetation cover, which collectively influence their occurrence. The importance of these factors varies across landslide types and locations. Modeling requires factor analysis of susceptibility, including factor validity, collinearity, and importance. However, the relationship between landslide triggering factors and landslides is extremely complex, making landslide susceptibility modeling challenging. By establishing numerical relationships between the locations of past landslides and conditional factors, it is possible to predict areas where future landslides are likely to occur. Accurate landslide susceptibility modeling and prediction can provide scientific guidance for disaster prevention and mitigation, providing time and conditions for early action, thereby preventing risks from occurring and preventing landslide hazards from becoming human disasters, thereby significantly reducing the severity of disasters.
[0004] Existing landslide susceptibility modeling methods are primarily divided into two categories: deterministic and non-deterministic. Deterministic methods are primarily based on physical mechanics principles, using slope limit equilibrium models to assess slope stability and subsequently predict landslide susceptibility. These methods effectively explore the relationship between landslide hazards and conditioning factors, but they face challenges such as data collection difficulties, complex spatial variations in conditioning factors, and model reliability affected by rock unit complexity. They are generally suitable for small-scale landslide susceptibility assessments in areas with relatively uniform geological and geomorphological conditions and shallow landslides. Non-deterministic methods are further divided into knowledge-driven qualitative methods and mathematics-driven quantitative methods. The former are based on subjective judgment and qualitative analysis based on expert experience, such as the analytic hierarchy process (AHP) and the fuzzy comprehensive evaluation method. These methods rely on experts to assign and rank weights to landslide-influencing factors, but the assessment results lack objectivity and reproducibility, and their reliance on expert knowledge leads to uncertainty in weight assignment. The latter, on the other hand, primarily include conventional mathematical and statistical models and machine learning models. Conventional mathematical statistical models, such as the information method, coefficient of determination method, frequency ratio method, and weight of evidence method, are easily affected by the size and classification of raster pixel values and are difficult to reflect the nonlinear relationship between landslides and their underlying environmental factors. Machine learning models, such as random forests, logistic regression, support vector machines, and convolutional neural networks, possess powerful adaptive learning capabilities and can better capture the nonlinear relationship between factors and landslides. Despite the "black box" operation problem, their susceptibility modeling process is relatively simple and efficient.
[0005] In practical applications, selecting an appropriate modeling approach requires comprehensive consideration of factors such as data quality, study area characteristics, and specific needs. With the development of Earth observation sensors and data mining technologies, data-driven methods, such as machine learning, have become the primary means of landslide susceptibility modeling. However, existing sampling and factor selection strategies for landslide susceptibility modeling each have their own characteristics and limitations. Random sampling is the most commonly used method for obtaining negative samples, but it is generally only applicable to small areas and can lead to low modeling accuracy over large areas. To improve sampling accuracy and quality, various improved methods have been proposed, such as exploratory analysis of prior data, buffer-controlled sampling, and distance- and density-based metrics. Commonly used factor selection strategies include correlation testing, multicollinearity testing, and factor interaction analysis. Landslides result from the combined effects of multiple factors, and interactions between factors can increase or decrease landslide risk. Therefore, the selection of moderating factors is crucial for improving the accuracy of susceptibility assessments. However, most publicly available solutions rely solely on empirically selected factors, neglecting to identify universally applicable combinations of conditional factors.
[0006] In summary, existing landslide susceptibility modeling methods, especially mainstream machine learning methods, still have the following three limitations or deficiencies: First, the subjectivity of factor selection: the factor selection process is greatly influenced by human factors and lacks unified standards, which may lead to unsatisfactory factor combinations and affect the predictive performance and reliability of the model; Second, the randomness of sample selection: there is uncertainty in the selection of non-landslide samples, and random selection or constrained selection may lead to unstable model performance. In addition, methods that over-rely on the characteristic distribution of positive samples are prone to overfitting; Third, the interpretability of evaluation results: machine learning models are often regarded as "black boxes", and it is difficult to understand their decision-making basis and internal mechanisms, which limits the scientific interpretation and practical application value of model results. Summary of the Invention
[0007] The purpose of the present invention is to provide a large-area, interpretable landslide susceptibility assessment method to address the shortcomings of the above-mentioned background technology, so as to improve the generalization ability and warning accuracy of the model in complex scenarios and provide reliable technical support for geological disaster risk prevention and control.
[0008] To achieve the above, the present invention provides an interpretable landslide susceptibility assessment method, comprising the following steps:
[0009] S1. Based on the landslide hazard system theory, a mechanism-geography-physics-mathematical mapping is constructed to systematically establish a landslide hazard environmental factor system;
[0010] S2, through the frequency ratio method, according to the comprehensive frequency ratio frequency distribution, extracts the negative sample search threshold, infers the number and location of negative sample screening, and completes the negative sample selection;
[0011] The frequency ratio is the ratio of the percentage of disaster grids in a certain environmental factor classification interval to the percentage of grids in the total number of disaster grids in the study area; the comprehensive frequency ratio is the mean of the frequency ratios between factors after maximum and minimum normalization based on the spatial distribution of the frequency ratios of various environmental factors;
[0012] S3, using association rule mining method, extracts landslide rule knowledge subgraph and non-landslide rule knowledge subgraph according to positive sample and negative sample data respectively, and establishes an interpretable model of landslide susceptibility based on the similarity calculation of rule knowledge subgraph to realize interpretable landslide susceptibility assessment.
[0013] Furthermore, S1 specifically includes the following sub-steps:
[0014] S11, data preparation:
[0015] Collect landslide catalog data and landslide disaster-prone environment data in the target area, construct a mechanism-geography-physics-mathematical mapping of the disaster-prone environment from a multi-level perspective, and model various factors involved in the disaster-prone environment;
[0016] S12: perform spatial alignment on the collected environmental factor data to unify the coordinate system and grid resolution of all data; perform maximum and minimum value normalization on continuous factors, classify and encode discrete factors and normalize them to eliminate dimensional differences.
[0017] Furthermore, S2 specifically includes the following sub-steps:
[0018] S21, calculate the frequency ratio, the calculation formula is:
[0019] ;
[0020] Where FR is the frequency ratio, is the number of grids with geological hazards in the classification interval for a certain environmental factor, F is the total number of all geological hazard grids in the interval, is the number of grids within the classification interval of a certain environmental factor, is the total number of grids in the study area;
[0021] S22, calculate the comprehensive frequency ratio, the calculation formula is:
[0022] ;
[0023] Where m represents the number of integrated environmental factors, represents the frequency ratio distribution of the first grid factor after normalization, represents the spatial distribution of the comprehensive frequency ratio of the entire study area;
[0024] S23, using statistical methods to analyze the numerical distribution of the comprehensive frequency ratio, drawing a frequency distribution histogram based on the calculated comprehensive frequency ratio, and setting the peak value of the comprehensive frequency ratio as the negative sample search threshold q; according to the statistical properties of the frequency ratio, when When , it indicates that the corresponding spatial area is more likely to not have landslides;
[0025] Statistical comprehensive frequency ratio and The corresponding spatial range area is used as the sampling ratio of positive samples to negative samples, and the sampling quantity of negative samples is determined by using this ratio;
[0026] For each disaster point, a buffer zone with a preset range is built based on sampling experience, and further This discriminant condition randomly filters negative samples and saves their geographic coordinates and sample labels.
[0027] Furthermore, S2 further includes the following sub-steps:
[0028] S24, the variance inflation factor is used to perform statistical analysis on all frequency ratio environmental factors. The variance inflation factor is used to measure the degree to which a certain independent variable g is linearly explained by other independent variables. The larger the variance inflation factor, the more serious the collinearity. The variance inflation factor VIF is expressed as:
[0029] ;
[0030] in, It indicates the goodness of fit obtained by regressing a certain independent variable g on all other independent variables.
[0031] Furthermore, S3 includes the following sub-steps:
[0032] S31, through the FP-Growth algorithm, treats landslide disasters as events, and various environmental factors as different components of the event. The equally spaced classification of each factor is regarded as independent items. The frequent occurrence patterns between different items are extracted to obtain the statistical rules implied by the occurrence or non-occurrence of landslide disasters, reflecting the occurrence law of landslide disasters.
[0033] S32, for the position to be inferred, determining the feature interval and the interval combination relationship, and calculating the similarity between the matched rule and the rule set representing the prior statistical knowledge;
[0034] S33, construct an interpretation template for the susceptibility discrimination results, comprehensively consider attribute similarity and structural similarity, and complete the landslide susceptibility assessment.
[0035] Furthermore, S31 specifically includes the following sub-steps:
[0036] S311, traverse the data set, count the support count of each item; delete items whose support is lower than the preset threshold t , and sort the remaining items in descending order of support to form an item header table; each item header table contains the name of the item, the support count, and a pointer to the first node of the corresponding item in the FP-Tree;
[0037] S312, traverse each transaction in the data set and process each item in the order of the items in the item header table; starting from the root node, if the current item exists in the child node of the current node, increase the support count of the child node; otherwise, create a new child node and update the linked list of the item in the item header table;
[0038] S313, for each item in the item header table, starting from the end node of its linked list, recursively traverse the linked list to generate a conditional pattern base with the node as the suffix path, the conditional pattern base includes other items in the path except the current item and the corresponding support counts;
[0039] S314, for each item in the item header table, combine it with the conditional pattern base to form a new frequent item set; if the conditional pattern base is not empty, recursively call the FP-Tree construction and mining process with the conditional pattern base as input until no further mining is possible;
[0040] S315, calculate support and confidence based on frequent item sets and generate association rules; support represents item sets ( X , Y ) appears in the transaction database, and the confidence level indicates the frequency of occurrence of X The transaction also contains Y The probability of setting the rule support threshold E 1 or confidence threshold E 2. Screen effective rules for disaster occurrence or non-occurrence.
[0041] Furthermore, in S32, for the point O to be estimated, its characteristic interval is ,in Represents the frequency ratio belonging interval of the mth environmental factor; the rule set R contains multiple association rules, each association rule The form is ,in represents the kth rule item in the jth association rule;
[0042] Similarity The calculation formula is:
[0043] ;
[0044] Where M is the number of rules in the rule set R, is the number of terms in the jth rule, Represents the structural similarity weight, expressed as rule support or confidence, Indicates the structural similarity, expressed as the ratio of the length of the j-th rule to the length of all items in the point to be inferred. is the attribute similarity between the inferred point O and the jth rule, and They are:
[0045] ;
[0046] ;
[0047] Depending on whether the rule is in the feature factor set, the judgment value is 1 or 0, and u represents the number of significant feature items of the point O to be inferred.
[0048] Furthermore, in S33, the feature items in the successfully matched valid association rules are regarded as nodes, and edge links are established between the feature items in the rules to form graph structure 1; graph structure 2 is constructed based on all valid association rules; by combining the frequency ratio feature values of the environmental factors of the points to be inferred, the similarity between graph structure 1 and graph structure 2 is calculated to intuitively explain the susceptibility judgment results.
[0049] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the computer program implements the above-mentioned interpretable landslide susceptibility assessment method.
[0050] The above solution of the present invention has the following beneficial effects:
[0051] The present invention provides an interpretable landslide susceptibility assessment method and medium. By collecting landslide catalog data and landslide disaster-prone environment data in the target area, the method constructs a mechanism-geography-physics-mathematical mapping of the disaster-prone environment from a multi-level perspective, systematically models various factors involved in the disaster-prone environment, and avoids the uncertainty of susceptibility modeling caused by subjective selection of environmental factors. The present invention can comprehensively consider the influence of various factors when extracting negative samples, and use the numerical distribution of the comprehensive frequency ratio to extract the negative sample screening threshold, thereby guiding the determination of the number and location of negative sample sampling, and avoiding the low accuracy of susceptibility modeling caused by random selection of negative samples. By introducing the FP-Growth algorithm, the paper considers landslide hazards as events, and various environmental factors as distinct components of these events. The equally spaced gradations (or classifications) of each component can be considered independent items. By extracting the frequent occurrence patterns between these items, the paper then derives the statistical rules underlying the occurrence (or non-occurrence) of landslide hazards. These rules can, to a certain extent, reflect the patterns of disaster occurrence and thus guide the modeling of susceptibility. The paper further constructs an interpretation template for susceptibility modeling results. By comprehensively considering attribute similarity and structural similarity, the results are made more scientific, reasonable, and easy to understand. Therefore, the paper has significant translational and practical application value in the field of landslide susceptibility assessment.
[0052] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a flow chart of the overall steps of the present invention;
[0054] Figure 2 It is a detailed flow chart of the present invention;
[0055] Figure 3 The environmental factor mechanism-geographical-physical-mathematical characteristic mapping diagram provided by the present invention;
[0056] Figure 4 It is the comprehensive frequency ratio spatial distribution and negative sample sampling spatial distribution diagram provided by the present invention;
[0057] Figure 5 It is an effective association rule graph structure provided by the present invention;
[0058] Figure 6 It is the association rule graph structure of the successful matching of the inferred point provided by the present invention;
[0059] Figure 7 It is a spatial distribution map of the landslide hazard susceptibility assessment results that can be interpreted in the present invention. DETAILED DESCRIPTION
[0060] The following describes the embodiments of the present disclosure through specific concrete examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.
[0061] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement an apparatus and / or practice a method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this apparatus and / or practice this method.
[0062] It should also be noted that the diagrams provided in the following embodiments are merely schematic illustrations of the basic concepts of the present disclosure. The diagrams only show components relevant to the present disclosure and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the configuration, quantity, and proportion of each component may be varied at will, and the component layout may be more complex. Furthermore, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will appreciate that the described aspects may be practiced without these specific details.
[0063] like Figure 1 、 Figure 2 As shown, an embodiment of the present invention provides an interpretable landslide susceptibility assessment method, comprising the following steps:
[0064] S1. Based on the landslide disaster system theory, a mechanism-geography-physics-mathematical mapping is constructed to systematically establish a landslide disaster environmental factor system.
[0065] In this embodiment, this step specifically includes the following sub-steps:
[0066] S11, data preparation:
[0067] A regional disaster system is a system of abnormal changes in the Earth's surface layer, consisting of a disaster-prone environment, hazard-causing factors, and hazard-bearing bodies. The disaster-prone environment is a crucial component of the disaster system, providing the conditions and context for the formation and development of hazard-causing factors. The disaster-prone environment is unstable, and its stability is influenced by multiple factors, such as geological structure, topography, soil type, and vegetation cover. These factors interact to determine the stability of the disaster-prone environment and the likelihood of disasters. Within the regional disaster system, the instability of the disaster-prone environment, the dangerousness of the hazard-causing factors, and the vulnerability of the hazard-bearing bodies together constitute the functional framework of the disaster system, influencing the occurrence, development, and extent of disaster losses. Therefore, when preparing data, this embodiment first collects landslide catalog data and landslide disaster-prone environment data for the target area. Then, from a multi-layered perspective, a mechanistic-geographical-physical-mathematical mapping of the disaster-prone environment is constructed to systematically model the various factors involved in the disaster-prone environment. Specifically, these factors include but are not limited to topography (slope, elevation, curvature), geological structure (lithology, distance to fault zones), meteorology and hydrology (rainfall intensity, soil moisture content), and socio-economics (land use, road density).
[0068] S12, data preprocessing:
[0069] The collected data on various environmental factors (influencing factors) are spatially aligned, that is, the coordinate system (such as WGS84) and grid resolution (such as 30m×30m) of all data are unified to ensure spatial consistency; the maximum and minimum values of continuous factors (such as slope) are normalized, and discrete factors (such as lithology categories) are classified, coded and normalized to eliminate dimensional differences.
[0070] S2, through the frequency ratio method, extracts the negative sample search threshold according to the comprehensive frequency ratio frequency distribution, infers the number and position of negative sample screening, and realizes the optimization of negative samples.
[0071] In this embodiment, this step specifically includes the following sub-steps:
[0072] S21, frequency ratio calculation:
[0073] Frequency Ratio, FR ) can be summarized as the ratio of the percentage of disaster grids within a certain factor classification interval to the percentage of all disaster grids to the percentage of grids within this classification interval to the total number of grids in the study area. The calculation formula is:
[0074] ;
[0075] in, is the number of grids with geological hazards in the classification interval for a certain environmental factor, F is the total number of all geological hazard grids in the interval, is the number of grids within the classification interval of a certain environmental factor, is the total number of grids in the study area.
[0076] It should be noted that FR indicates the degree of influence of each classification interval of environmental factors on the occurrence of geological disasters. FR>1 indicates that the classification interval of the environmental factor has a strong influence on the occurrence of geological disasters, and FR≤1 indicates that the classification interval of the environmental factor has little influence on the occurrence of disasters.
[0077] As described above, the frequency ratio method first requires the classification or grading of geohazard environmental factors. To improve operability and ease understanding, this example uses the equal interval method to divide each environmental factor into n levels (n is set to a natural number between 1 and 101). According to the FR calculation formula, each grid position for each environmental factor in the entire study area will obtain a frequency ratio value.
[0078] It should be noted that in the actual modeling process, the classification or grading of geological hazard environmental factors can also be more appropriately carried out by adopting the natural breakpoint method, empirical grading method, parameter optimization method, etc.
[0079] S22, comprehensive frequency ratio calculation:
[0080] In order to comprehensively consider the influence of various environmental factors when extracting negative samples, the concept of comprehensive frequency ratio is proposed in this embodiment. The numerical distribution of the comprehensive frequency ratio is used to extract the negative sample screening threshold, thereby guiding the determination of the number and location of negative sample sampling.
[0081] Specifically, the comprehensive frequency ratio is the mean of the frequency ratios among factors after normalization based on the spatial distribution of the frequency ratios of various environmental factors. The comprehensive frequency ratio can be formally expressed as:
[0082] ;
[0083] Where m represents the number of integrated environmental factors, represents the frequency ratio distribution of the first grid factor after normalization, Represents the spatial distribution of the comprehensive frequency ratio of the entire study area.
[0084] S23, negative sample screening:
[0085] It should be noted that the sampling ratio of positive and negative samples is particularly important for sample distribution balance. In order to determine the appropriate number of negative samples, this embodiment first uses a statistical method to analyze the numerical distribution of the comprehensive frequency ratio, draws a frequency distribution histogram based on the comprehensive frequency ratio calculated in S22, and sets the peak value of the comprehensive frequency ratio as the negative sample search threshold q. According to the statistical properties of the frequency ratio, when , it indicates that the corresponding spatial area is more likely to not experience landslides.
[0086] Then, the statistical comprehensive frequency ratio and The corresponding spatial range area is used as the sampling ratio of positive samples to negative samples, and the sampling number of negative samples is determined using this ratio.
[0087] Then, a 3km buffer zone was constructed for each disaster point based on sampling experience. On this basis, a further 3km buffer zone was constructed outside the buffer zone based on This discriminant condition randomly filters negative samples and saves their geographic coordinates and sample labels.
[0088] S24, feature multicollinearity analysis:
[0089] After negative sample screening in S23, the feature composition of different sample points has also been optimized. The feature value is no longer the normalized attribute value obtained in S12, but the statistically significant frequency ratio calculated value obtained in S21, which improves the representativeness of the feature. In order to determine the factors that really affect the occurrence of landslides among different environmental factors, this embodiment further uses the variance inflation factor (VIF) to perform statistical analysis on all frequency ratio factors. The variance inflation factor is used to measure the degree to which a certain independent variable g is linearly explained by other independent variables. The larger the VIF value, the more serious the collinearity. VIF can be expressed in the form of:
[0090] ;
[0091] in, It represents the goodness of fit obtained by regressing a specific independent variable g on all other independent variables. It should be noted that when the VIF is greater than or equal to 10, it is considered to be severely collinear and the independent variable needs to be deleted.
[0092] S3, using association rule mining method, extracts landslide rule knowledge subgraph and non-landslide rule knowledge subgraph according to positive sample and negative sample data respectively, and establishes an interpretable model of landslide susceptibility based on the similarity calculation of rule knowledge subgraph to realize interpretable landslide susceptibility assessment.
[0093] In this embodiment, this step specifically includes the following sub-steps:
[0094] S31, Association rule mining method:
[0095] The FP-Growth (Frequent Pattern Growth) algorithm is a highly efficient algorithm for association rule mining. Its core concept is to compactly store frequent itemsets in a database in a tree format by constructing a frequent pattern tree (FP-Tree). This avoids the frequent steps of generating and verifying candidate itemsets in the traditional Apriori algorithm, significantly improving mining efficiency. The FP-Tree is a special prefix tree structure, where each node represents an item and records the number of times it appears in the transaction database. The tree construction process effectively compresses data while preserving the associations between itemsets, enabling efficient subsequent frequent itemset mining based on the tree structure.
[0096] In this embodiment, the FP-Growth algorithm is used to treat landslide disasters as events, and various environmental factors can be regarded as different components of the event. The equally spaced classification (or classification) of each element can be regarded as independent items. By extracting the frequent occurrence patterns between different items, the statistical rules implied in the occurrence (or non-occurrence) of landslide disasters are extracted. These rules can reflect the laws of disaster occurrence to a certain extent, thereby guiding the interpretable modeling of susceptibility.
[0097] Specifically, the association rule mining method includes the following sub-steps:
[0098] S311, build item header table: traverse the data set, count the support count of each item; delete items whose support is lower than the preset threshold t , and sort the remaining items in descending order of support to form an item header table; each item header table contains the name of the item, the support count, and a pointer to the first node of the item in the FP-Tree.
[0099] S312, constructing an FP-Tree: traverse each transaction in the data set and process each item in the order of the items in the item header table; starting from the root node, if the current item exists in the child node of the current node, increase the support count of the child node; otherwise, create a new child node and update the linked list of the item in the item header table.
[0100] S313, Constructing a Conditional Pattern Base: For each item in the item header table, recursively traverse the linked list starting from the end node of its linked list to generate a conditional pattern base with the node as the suffix path. The conditional pattern base includes all items in the path except the current item and their corresponding support counts.
[0101] S314, recursively mine the FP-Tree: for each item in the item header table, combine it with the conditional pattern base to form a new frequent item set; if the conditional pattern base is not empty, use the conditional pattern base as input and recursively call the FP-Tree construction and mining process until mining can no longer continue.
[0102] S315, association rule generation: After obtaining the frequent item sets, association rules are generated by calculating indicators such as support and confidence. Among them, support represents the frequency of the item set (X, Y) in the transaction database, and confidence measures the reliability of the rule, that is, the probability that Y is also included in the transaction containing X. For example, for rule A→B, its support is , the confidence level is , that is, the probability that a transaction containing A also contains B. By setting the rule support threshold E1 or the confidence threshold E2, effective disaster occurrence (or non-occurrence) rules can be further screened.
[0103] S32, similarity calculation:
[0104] Through S31, several rules with certain support and confidence are obtained. Suppose a rule in the rule set is expressed as , where F, G, and H represent environmental factors, and 1, 2, and 3 represent the grading (or classification) numbers of these environmental factors using the equal-interval method. Therefore, the different rules mined can form a knowledge graph structure that represents the prior statistical knowledge underlying the occurrence (or non-occurrence) of disasters. For the location to be inferred, the corresponding feature intervals and interval combinations can be identified based on the feature values calculated in S23. The similarity between the matched rules and the rule set representing the prior statistical knowledge can then be calculated.
[0105] In a specific embodiment, it is assumed that there is a point O to be estimated, and its characteristic interval is ,in Indicates the frequency ratio of the mth environmental factor belonging interval. The rule set R contains multiple association rules, each association rule The form is ,in Indicates the kth rule item in the jth association rule. In order to ensure that the association rules can effectively reflect the laws of disaster occurrence, the process of obtaining the rule set R needs to first filter out the intersection of the positive sample rule set and the negative sample rule set from the positive sample rule set.
[0106] Therefore, the similarity The calculation method can be expressed in the form of:
[0107] ;
[0108] Where M is the number of rules in the rule set R, is the number of terms in the jth rule, Represents the structural similarity weight, which can be expressed by rule support or confidence. Indicates the structural similarity, expressed as the ratio of the length of the j-th rule to the length of all items in the point to be inferred. is the attribute similarity between the point to be inferred O and the jth rule. and It can be further formulated as follows:
[0109] ;
[0110] ;
[0111] therefore, Depending on whether the rule is in the feature factor set, the judgment value is 1 or 0, and u represents the number of significant feature items of the point O to be inferred.
[0112] S33, Explanation of susceptibility determination results:
[0113] In this embodiment, an explanation template for the susceptibility discrimination result is constructed, and by comprehensively considering the attribute similarity and the structure similarity, the susceptibility discrimination result is made more scientific, reasonable and easy to understand.
[0114] In a specific embodiment, the interpretation template is expressed as follows:
[0115] In this susceptibility modeling, we first extracted [number of positive sample association rules] significant and reliable landslide occurrence rules and [number of negative sample association rules] significant and reliable landslide non-occurrence rules from the positive and negative sample sets. By calculating the intersection of these rules, we filtered out [number of intersection association rules] ambiguous rules. Ultimately, we obtained [number of positive sample association rules - number of intersection association rules] valid association rules. Specifically, they include:
[0116] {Rule 1}
[0117] {Rule 2}
[0118] …
[0119] {rule }
[0120] These association rules are pre-mined from positive samples based on the FP-Growth algorithm, and the fuzzy rules that appear simultaneously in positive and negative samples are filtered out. These association rules reflect, to a certain extent, the intrinsic connections and combination patterns between the characteristic items of different environmental factors.
[0121] For the point to be inferred [specific point coordinates], extract the matching association rules based on the frequency ratio distribution of its environmental factors. The association rules successfully matched include:
[0122] {Rule 1}
[0123] {Rule 2}
[0124] …
[0125] {rule }
[0126] The feature items in these successfully matched valid association rules are treated as nodes, and edge links are established between the feature items in the rules to form Graph Structure 1. Simultaneously, Graph Structure 2 is constructed based on all valid association rules. By combining the frequency ratio feature values of the environmental factors at the inferred points, the similarity between Graph Structure 1 and Graph Structure 2 (attribute similarity and structural similarity) is calculated, which allows for an intuitive interpretation of the susceptibility judgment results.
[0127] For example, the probability of occurrence is divided into five levels according to the equal interval classification method: very low probability, low probability, medium probability, high probability, and very high probability: 0.2, 0.4, 0.6, 0.8, 1. If the similarity between graph structure 1 and graph structure 2 is high ( ), indicating that the environmental factor feature item of the inferred point has a high degree of consistency with the susceptibility feature pattern in the positive sample, thus supporting that the point has a high susceptibility. On the contrary, if the similarity is low ( ), it indicates that its characteristic pattern is quite different from the known prone situations and the proneness is relatively low.
[0128] As can be seen above, this interpretation method based on graph structural similarity not only considers the characteristic terms of each environmental factor itself but also fully reflects the combined relationships and interactions between them. By comprehensively considering attribute similarity and structural similarity, the susceptibility judgment results are more scientific, reasonable, and easy to understand, significantly improving the susceptibility assessment of landslide hazards.
[0129] The following uses a specific case to further illustrate the effects of the present invention. Considering that Yunnan Province is located in the sloping area where the Qinghai-Tibet Plateau transitions to the Yunnan-Guizhou Plateau, it has active geological structures, severe terrain dissection, and concentrated and intense rainfall. This example selects Yunnan Province as the research area. The following will use a landslide disaster example to specifically illustrate the specific implementation steps of this solution for landslide susceptibility assessment:
[0130] S1, constructing the mechanism-geography-physics-mathematical feature mapping, e.g. Figure 3 As shown in the figure, based on environmental factors such as topography, geology, meteorology, hydrology, and socioeconomics, an index system was constructed, including but not limited to slope, aspect, NDVI, human density, low-grade road density, low-grade road distance, profile curvature, land use type, soil type, terrain relief, engineering geological map, plan curvature, average annual rainfall, fault density, fault distance, water system distance, elevation, high-grade road density, and high-grade road distance. These environmental factor characteristics were extracted from multiple perspectives: mechanistic, geographical, physical, and mathematical. The data were sourced from the Internet Geographic Information Public Platform.
[0131] Preprocess each indicator variable. Spatially align the collected environmental factor data, unifying the coordinate system (e.g., WGS84) and grid resolution (e.g., 30m × 30m) for all data to ensure spatial consistency. Furthermore, normalize the maximum and minimum values of continuous factors (e.g., slope), and classify and normalize discrete factors (e.g., lithology) to eliminate dimensional differences.
[0132] S2, negative sample optimization. Before sampling negative samples, it is necessary to first determine the sampling area and sampling ratio of negative samples. According to the comprehensive frequency ratio calculation method, the comprehensive frequency ratio distribution of the study area is obtained as follows: Figure 4 At the same time, the peak value of the frequency distribution histogram of the comprehensive frequency ratio of the study area was further extracted. The results showed that the comprehensive frequency ratio of the grid in the whole area peaked in the range of 0.55~0.6. Therefore, the negative sample screening threshold was set to 0.55, and the number of negative sample samples was Figure 3 The area ratio between those with a medium comprehensive frequency ratio less than 0.55 and those greater than 0.55 is multiplied by the number of landslide hazard points, that is, 1 / 0.56×1397=2495 negative samples, of which 1397 is the number of landslide hazard points in the study area in the past ten years.
[0133] S3, Explainable Modeling of Susceptibility. First, each environmental factor is regarded as a different component of the event. The equally spaced classification (or classification) of each factor can be regarded as independent items. By extracting the frequent occurrence patterns between different items, the statistical rules implied in the occurrence (or non-occurrence) of landslide disasters are extracted. Among them, when running the FP-Growth algorithm, this case sets the support threshold t of the feature item to 0.5 and the support threshold T of the item set to 0.75. The threshold indicates that items with a frequency of less than 0.5 in positive sample (or negative sample) events will be eliminated, and then, item sets with a frequency of less than 0.75 in positive sample (or negative sample) events will not be included as association rules. The partial item list and item set list obtained during the implementation process are shown in Table 1:
[0134] Table 1 Some items and item sets (the letters of the items represent a certain environmental factor, and the numbers represent the classification (or classification) numbers of the environmental factors obtained by the equal interval method)
[0135]
[0136] By filtering the intersection of the positive sample item set list and the negative sample item set list, a total of 23 false rules were filtered out, and finally 110 valid association rules related to landslide occurrence were obtained. The graph structure of the valid association rules is as follows: Figure 5 As shown in Figure 2. Taking a landslide event outside the sample data set as an example, the inferred point successfully matches 102 valid association rules, and the graph structure is as follows: Figure 6 shown.
[0137] Based on the interpretation template in the previous article, the susceptibility of this landslide event is explained as follows:
[0138] In the results of this susceptibility modeling, we first started with the positive and negative sample sets and extracted
[133] significant and reliable landslide occurrence rules and
[45] significant and reliable landslide non-occurrence rules. After calculating the intersection of the rules, a total of
[23] rules with ambiguous interpretations were filtered out. Finally,
[110] valid association rules were obtained. Specifically, they include:
[0139] {D4, B3}
[0140] {D4, R0}
[0141] {D4, K4}
[0142] {D4, O1}
[0143] {D4, B3, R0}
[0144] …
[0145] {B3, R0, D4, O1, K4, Q2}
[0146] For the inferred point [105.011, 27.480519], the matching association rules are extracted based on the frequency ratio distribution of its environmental factors. The successfully matched association rules include:
[0147] {D4, B3}
[0148] {D4, R0}
[0149] {D4, K4}
[0150] {D4, O1}
[0151] {D4, B3, R0}
[0152] …
[0153] {O1, K4, B3, R0, D4}
[0154] The susceptibility is divided into five levels according to the equal interval classification method: extremely low susceptibility, low susceptibility, medium susceptibility, high susceptibility, and extremely high susceptibility: 0.2, 0.4, 0.6, 0.8, 1. According to the similarity calculation formula, the similarity between graph structure 1 and graph structure 2 is 0.92 (greater than the high susceptibility threshold of 0.6), indicating that the environmental factor characteristic items of the inferred point have a high degree of consistency with the susceptibility characteristic pattern in the positive sample, thus supporting that the point has a high susceptibility. Finally, the spatial distribution map of the interpretable landslide susceptibility assessment results is shown in Figure 1. Figure 7 shown.
[0155] Based on the same inventive concept, this embodiment further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the aforementioned interpretable landslide susceptibility assessment method is implemented.
[0156] Computer-readable media include, but are not limited to, any type of disk (including floppy disks, hard disks, optical disks, CD-ROMs, and magneto-optical disks), ROMs, RAMs, EPROMs (Erasable Programmable Read-Only Memory), EEPROMs, flash memories, magnetic cards, or optical cards. In other words, computer-readable media include any medium that can store or transmit information in a form that can be read by a device.
[0157] The computer-readable storage medium and the like provided in this embodiment have the same inventive concept and the same beneficial effects as the aforementioned method, and are not described in detail here.
[0158] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0159] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. An interpretable landslide susceptibility assessment method, characterized in that: The steps include: S1. Based on the landslide hazard system theory, a mechanism-geography-physics-mathematics mapping is constructed to systematically establish a landslide hazard environmental factor system; S2, through the frequency ratio method, extracts the negative sample search threshold according to the comprehensive frequency ratio frequency distribution, infers the number and location of negative sample screening, and completes the negative sample optimization; The frequency ratio is the ratio of the percentage of disaster grids in a certain environmental factor classification interval to the percentage of grids in the total number of disaster grids in the study area; the comprehensive frequency ratio is the mean of the frequency ratios between factors after maximum and minimum normalization based on the spatial distribution of the frequency ratios of various environmental factors; S3, using association rule mining methods, extracts landslide rule knowledge subgraphs and non-landslide rule knowledge subgraphs based on positive sample and negative sample data, and establishes an interpretable landslide susceptibility model based on the similarity calculation of rule knowledge subgraphs to achieve interpretable landslide susceptibility assessment; S3 includes the following sub-steps: S31, through the FP-Growth algorithm, treats landslide disasters as events, and various environmental factors as different components of the event. The equally spaced classification of each factor is regarded as independent items. The frequent occurrence patterns between different items are extracted to obtain the statistical rules implied by the occurrence or non-occurrence of landslide disasters, reflecting the occurrence law of landslide disasters. S32, for the position to be inferred, determining the feature interval and the interval combination relationship, and calculating the similarity between the matched rule and the rule set representing the prior statistical knowledge; S33, construct an interpretation template for the susceptibility discrimination results, comprehensively consider attribute similarity and structural similarity, and complete the landslide susceptibility assessment.
2. The interpretable landslide susceptibility assessment method according to claim 1, characterized in that: S1 specifically includes the following sub-steps: S11, data preparation: Collect landslide catalog data and landslide disaster-prone environment data in the target area, construct a mechanism-geography-physics-mathematical mapping of the disaster-prone environment from a multi-level perspective, and model various factors involved in the disaster-prone environment; S12, spatially align the collected environmental factor data to unify the coordinate system and grid resolution of all data; The continuous factors are normalized to their maximum and minimum values, and the discrete factors are classified, coded and normalized to eliminate dimensional differences.
3. The interpretable landslide susceptibility assessment method according to claim 2, characterized in that: S2 specifically includes the following sub-steps: S21, calculate the frequency ratio, the calculation formula is: ; in, is the frequency ratio, is the number of grids with geological hazards in the classification interval for a certain environmental factor, F is the total number of all geological hazard grids in the interval, is the number of grids within the classification interval of a certain environmental factor, is the total number of grids in the study area; S22, calculate the comprehensive frequency ratio, the calculation formula is: ; Where m represents the number of integrated environmental factors, represents the frequency ratio distribution of the first grid factor after normalization, represents the spatial distribution of the comprehensive frequency ratio of the entire study area; S23, using statistical methods to analyze the numerical distribution of the comprehensive frequency ratio, drawing a frequency distribution histogram based on the calculated comprehensive frequency ratio, and setting the peak value of the comprehensive frequency ratio as the negative sample search threshold q; according to the statistical properties of the frequency ratio, when When , it indicates that the corresponding spatial area is more likely to not have landslides; Statistical comprehensive frequency ratio and The corresponding spatial range area is used as the sampling ratio of positive samples to negative samples, and the sampling quantity of negative samples is determined by using this ratio; For each disaster point, a buffer zone with a preset range is built based on sampling experience, and further This discriminant condition randomly filters negative samples and saves their geographic coordinates and sample labels.
4. The interpretable landslide susceptibility assessment method according to claim 3, characterized in that: S2 also includes the following sub-steps: S24, the variance inflation factor is used to perform statistical analysis on all frequency ratio environmental factors. The variance inflation factor is used to measure the degree to which a certain independent variable g is linearly explained by other independent variables. The larger the variance inflation factor, the more serious the collinearity. The variance inflation factor VIF is expressed as: ; in, It indicates the goodness of fit obtained by regressing a certain independent variable g on all other independent variables.
5. The interpretable landslide susceptibility assessment method according to claim 3, characterized in that: S31 specifically includes the following sub-steps: S311, traverse the data set and count the support count of each item; delete items whose support is lower than a preset threshold t, and sort the remaining items in descending order of support to form an item header table; each item header table contains the name of the item, the support count, and a pointer to the first node of the corresponding item in the FP-Tree; S312, traverse each transaction in the data set and process each item in the order of the items in the item header table; starting from the root node, if the current item exists in the child node of the current node, increase the support count of the child node; Otherwise, create a new child node and update the linked list of the item in the header table; S313, for each item in the item header table, starting from the end node of its linked list, recursively traverse the linked list to generate a conditional pattern base with the node as the suffix path, the conditional pattern base includes other items in the path except the current item and the corresponding support counts; S314, for each item in the item header table, combine it with the conditional pattern base to form a new frequent item set; if the conditional pattern base is not empty, recursively call the FP-Tree construction and mining process with the conditional pattern base as input until no further mining is possible; S315. Calculate support and confidence based on frequent item sets to generate association rules. Support represents the frequency of item sets (X, Y) in the transaction database, and confidence represents the probability that transactions containing X also contain Y. Set the rule support threshold E1 or confidence threshold E2 to screen for effective disaster occurrence or non-occurrence rules.
6. The interpretable landslide susceptibility assessment method according to claim 5, characterized in that: In S32, for the point O to be estimated, Its characteristic interval is ,in Represents the frequency ratio belonging interval of the mth environmental factor; the rule set R contains multiple association rules, each association rule The form is ,in represents the kth rule item in the jth association rule; Similarity The calculation formula is: ; Where M is the number of rules in the rule set R, is the number of terms in the jth rule, Represents the structural similarity weight, expressed as rule support or confidence, Indicates the structural similarity, expressed as the ratio of the length of the j-th rule to the length of all items in the point to be inferred. is the attribute similarity between the inferred point O and the jth rule, and They are: ; ; Depending on whether the rule is in the feature factor set, the judgment value is 1 or 0, and u represents the number of significant feature items of the point O to be inferred.
7. The interpretable landslide susceptibility assessment method according to claim 6, characterized in that: In S33, the feature items in the successfully matched valid association rules are regarded as nodes, and edge links are established between the feature items in the rules to form graph structure 1; graph structure 2 is constructed based on all valid association rules; by combining the frequency ratio feature values of the environmental factors of the points to be inferred, the similarity between graph structure 1 and graph structure 2 is calculated to intuitively explain the susceptibility judgment results.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements an interpretable landslide susceptibility assessment method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Landslide risk prediction method and system based on knowledge graph construction
CN111639878A
Slope unit landslide susceptibility evaluation method based on optimized negative sample selection
CN118332434A