Explanatable landslide disaster susceptibility evaluation method and medium
The method addresses the limitations of existing slide hazard modeling by using frequency ratios and FP-Growth to create an interpretable model, improving prediction accuracy and reliability in large regions.
Patent Information
- Application Number
- CN202510765622.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing slide hazard modeling methods, particularly machine learning approaches, suffer from subjective factor selection, random sample choice, and lack of explainability, leading to inaccurate and unreliable predictions, especially in large regions.
A method involving frequency ratio-based sample selection and FP-Growth algorithm to construct an interpretable model by analyzing environmental factors, using frequency ratios and association rules to determine slide hazard likelihood.
Enhances model generalization and prediction accuracy in complex scenarios by providing an interpretable model that leverages factor interactions and reduces uncertainty in slide hazard assessment.
Smart Images

Figure CN120316618A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of spatio-temporal data mining, and particularly relates to an interpretable landslide hazard susceptibility assessment method and medium. Background Art
[0002] Landslides are serious global geological disasters with high risk and wide distribution. The occurrence of landslides is affected by multiple factors such as geology, watershed, land use, and construction activities, and its triggering factors are mostly related to rainfall or changes in groundwater levels. Rainfall or changes in groundwater levels reduce the soil cohesion and the slope stability, thus reducing the safety factor. Landslide susceptibility modeling is an important basis for disaster management. Its core purpose is to discover the implicit patterns between landslide instances and environmental factors by analyzing historical landslide data and related environmental factors, and then accurately predict the likelihood of future landslides. Therefore, conducting research on landslide susceptibility, predicting its spatial distribution, is of great significance for land use planning, ecological environment protection, and the formulation of disaster prevention and mitigation policies.
[0003] In fact, landslides are the result of the interaction of geographical factors, such as geological structure, soil composition, and vegetation cover, etc., which jointly affect their occurrence, and the importance of different factors varies among different landslide types and occurrence locations. In the specific modeling process, factor analysis of susceptibility needs to be carried out, including aspects such as the effectiveness, collinearity, and importance of factors. However, the relationship between landslide triggering factors and landslides is extremely complex, which poses many challenges to landslide susceptibility modeling. By establishing a numerical relationship between the locations of past landslides and conditional factors, the areas where future landslides may occur can be predicted. Accurate landslide susceptibility modeling and prediction can provide scientific guidance for disaster prevention and mitigation, provide time and conditions for humans to take early actions, thus preventing the occurrence of risks, preventing landslide hazards from evolving into human disasters, and significantly reducing the degree of disaster.
[0004] The existing landslide susceptibility modeling methods are mainly divided into two categories: deterministic methods and non-deterministic methods. Deterministic methods are mainly based on physical and mechanical principles, and evaluate slope stability through slope limit equilibrium models, and then predict landslide susceptibility. Such methods can better explore the relationship between landslide disasters and conditioning factors, but face problems such as difficult data collection, complex spatial variation of conditioning factors, and the reliability of the model being affected by the complexity of rock units. They are usually applicable to small-scale landslide susceptibility assessments in areas with relatively uniform geological and geomorphic conditions and shallow landslides. Non-deterministic methods are further divided into knowledge-driven qualitative methods and mathematics-driven quantitative methods. The former is based on subjective judgments and qualitative analyses of expert experience, such as the analytic hierarchy process, fuzzy comprehensive evaluation method, etc. Such methods rely on experts to assign weights and rank the influencing factors of landslides, but the evaluation results lack objectivity and repeatability, and the dependence on expert knowledge leads to uncertainty in weight assignment. The latter mainly includes conventional mathematical statistics models and machine learning models. Conventional mathematical statistics models such as the information value method, coefficient of determination method, frequency ratio method, weight of evidence method, etc. are easily affected by the size and classification of grid cell values and are difficult to reflect the non-linear relationship between landslides and their underlying environmental factors. Machine learning models such as random forest, logistic regression, support vector machine, convolutional neural network, etc. have powerful adaptive learning capabilities and can better capture the non-linear relationship characteristics between factors and landslides. Although there are problems with "black box" operations, their susceptibility modeling process is relatively simple and efficient.
[0005] In practical applications, selecting an appropriate modeling method requires comprehensive consideration of factors such as data quality, characteristics of the study area, and specific requirements. With the development of Earth observation sensors and data mining technologies, data-driven methods represented by machine learning have become the main means of susceptibility modeling. However, the sampling strategies and factor selection strategies in existing landslide susceptibility modeling methods have their own characteristics and limitations. In terms of sampling strategies, random sampling is the most commonly used method for obtaining negative samples, but it is usually only applicable to small areas and is prone to low modeling accuracy in large areas. In order to improve sampling accuracy and quality, various improved methods have been proposed at present, such as exploratory analysis of prior data, buffer zone control sampling, distance- and density-based measurement methods, etc. In terms of factor selection strategies, common methods include correlation tests, multicollinearity tests, and factor interaction analyses. Landslides are the result of the combined action of multiple factors, and the interaction between factors may increase or decrease the risk of landslides. Therefore, the principle of selecting conditioning factors is the key to improving the accuracy of susceptibility assessment. However, most of the relevant public solutions only select a series of factors based on experience and ignore finding a combination of conditioning factors with universality.
[0006] In summary, the existing landslide susceptibility modeling methods, especially the mainstream machine learning methods, still have the following three limitations or deficiencies: First, subjectivity in factor selection: The factor selection process is greatly affected by human factors and lacks a unified standard, which may lead to an unsatisfactory factor combination and affect the prediction performance and reliability of the model; Second, randomness in sample selection: There is uncertainty in the selection of non-landslide samples. Random selection or constrained selection may lead to unstable model performance, and methods that overly rely on the feature distribution of positive samples are prone to overfitting; Third, interpretability of evaluation results: Machine learning models are usually regarded as "black boxes", and it is difficult to understand their decision-making basis and internal mechanisms, which limits the scientific interpretation and practical application value of the model results. Summary of the Invention
[0007] The objective of the present invention is: aiming at the deficiencies existing in the above-mentioned background technology, to provide an interpretable landslide disaster susceptibility assessment method for large areas, so as to improve the generalization ability and early warning accuracy of the model in complex scenarios and provide reliable technical support for geological disaster risk prevention and control.
[0008] In order to achieve the above, the present invention provides an interpretable landslide disaster susceptibility assessment method, including the following steps: S1, according to the landslide disaster system theory, construct a mechanism-geography-physics-mathematics mapping, and systematically establish a landslide disaster environmental factor system; S2, through the frequency ratio method, extract the negative sample search threshold according to the frequency distribution of the comprehensive frequency ratio, infer the number and location of negative sample screening, and complete the optimization of negative samples; The frequency ratio is the ratio of the number of disaster grids in a classification interval of a certain environmental factor to the percentage of all disaster grids and the percentage of the number of grids in this classification interval to the total number of grids in the study area; the comprehensive frequency ratio is the mean of the frequency ratios between factors after maximum-minimum normalization based on the spatial distribution of the frequency ratios of various environmental factors. S3, adopt the association rule mining method, extract the landslide rule knowledge subgraph and non-landslide rule knowledge subgraph according to the positive sample and negative sample data respectively, and establish an interpretable model for landslide disaster susceptibility based on the similarity calculation of the rule knowledge subgraph, so as to realize the interpretable assessment of landslide disaster susceptibility.
[0009] Further, S1 specifically includes the following sub-steps: S11, data preparation: S11, collect the landslide catalog data and landslide disaster-forming environment data of the target area, construct a mechanism-geography-physics-mathematics mapping of the disaster-forming environment from a multi-level perspective, and model various factors involved in the disaster-forming environment; S12. Align the collected environmental factor data spatially to unify the coordinate system and grid resolution of all data; perform maximum-minimum normalization on continuous factors and classification coding and normalization on discrete factors to eliminate the dimensional difference.
[0010] Further, S2 specifically includes the following sub-steps: S21. Calculate the frequency ratio, and the calculation formula is: ; where FR is the frequency ratio, is the number of grid cells with geological disasters in the classification interval of a certain environmental factor, F is the total number of all geological disaster grid cells in the interval, is the number of grid cells of a certain environmental factor in the classification interval, is the total number of grid cells in the study area; S22. Calculate the comprehensive frequency ratio, and the calculation formula is: ; where m represents the number of integrated environmental factors, represents the frequency ratio distribution of the first normalized grid factor, represents the spatial distribution of the comprehensive frequency ratio of the entire study area; S23. Analyze the numerical distribution of the comprehensive frequency ratio by statistical methods, draw a frequency distribution histogram based on the calculated comprehensive frequency ratio, and set the peak value of the comprehensive frequency ratio as the negative sample search threshold q; according to the statistical properties of the frequency ratio, when , it indicates that the corresponding spatial area is more inclined not to have landslides; Statistically analyze the comprehensive frequency ratio and the corresponding spatial range area, take its ratio as the sampling ratio of positive and negative samples, and use this ratio to determine the sampling quantity of negative samples; For each disaster point, construct a buffer zone with a preset range according to sampling experience, and further randomly screen negative samples outside the buffer zone according to this discriminant condition, and save its geographical coordinates and sample labels.
[0011] Further, S2 also includes the following sub-steps: S24. Use the variance inflation factor to statistically analyze all frequency ratio environmental factors. The variance inflation factor is used to measure the degree to which an independent variable g is linearly explained by other independent variables. The larger the variance inflation factor, the more serious the collinearity. The variance inflation factor VIF is expressed as: ; where, represents the goodness of fit obtained by regressing a certain independent variable g on all other independent variables.
[0012] Furthermore, S2 includes the following sub-steps: S31: Regarding the landslide disaster as an event through the FP-Growth algorithm, regarding various environmental factors as different constituent elements in the event, regarding the equal-interval grading of each element as mutually independent items, extracting the frequently occurring patterns between different items, obtaining the statistical rules implied by the occurrence or non-occurrence of the landslide disaster, and reflecting the occurrence law of the landslide disaster; S32: For the position to be inferred, judging the characteristic interval and the interval combination relationship, and calculating the similarity between the matched rule and the rule set representing the prior statistical knowledge; S33: Constructing an interpretation template for the susceptibility discrimination result, comprehensively considering the attribute similarity and the structure similarity, and completing the evaluation of the susceptibility of the landslide disaster.
[0013] Furthermore, S31 specifically includes the following sub-steps: S311: Traversing the data set, counting the support count of each item; deleting the items with support lower than the preset threshold t and sorting the remaining items in descending order of support to form an item header table; each item header table contains the name of the item, the support count, and a pointer to the first node of the corresponding item in the FP-Tree; S312: Traversing each transaction in the data set and processing each item in turn according to the order of the items in the item header table; starting from the root node, if the current item exists in the child nodes of the current node, increasing the support count of the child node; otherwise, creating a new child node and updating the linked list of the item in the item header table; S313: For each item in the item header table, starting from the last node of its linked list, recursively traversing the linked list to generate a conditional pattern base with the node as the suffix path, and the conditional pattern base contains the other items in the path except the current item and the corresponding support count; S314: For each item in the item header table, combining it with the conditional pattern base to form a new frequent item set; if the conditional pattern base is not empty, using the conditional pattern base as the input and recursively calling the construction and mining process of the FP-Tree until no further mining can be carried out; S315: Calculating the support and confidence based on the frequent item set to generate association rules; the support represents the frequency of the item set ( X , Y ) appearing in the transaction database, and the confidence represents the probability that the transaction containing X also contains Y . Setting a rule support threshold E 1 or a confidence threshold E 2 to screen out the effective rules for the occurrence or non-occurrence of the disaster.
[0014] Further, for the point O to be speculated in S32, its characteristic interval is , where represents the frequency ratio attribution interval of the m-th environmental factor; the rule set R contains multiple association rules, and each association rule is in the form of , where represents the k-th rule item in the j-th association rule; Similarity The calculation formula is: ; Among them, M is the number of rules in the rule set R, is the number of items in the j-th rule, represents the structure similarity weight, expressed by rule support or confidence, represents the structure similarity, expressed by the length ratio of all items in the j-th rule to the point to be speculated, is the attribute similarity between the point O to be speculated and the j-th rule, And are respectively: ; ; According to whether the rule is in the characteristic factor set, the value is judged to be 1 or 0, and u represents the number of significant characteristic items of the point O to be speculated.
[0015] Further, in S33, the characteristic items in the successfully matched valid association rules are regarded as nodes, and the edge links between the characteristic items in the rule are established to form a graph structure one; a graph structure two is constructed according to all valid association rules; by combining the characteristic values of the environmental factor frequency ratios of the point to be speculated, the similarity between the graph structure one and the graph structure two is calculated to intuitively explain the susceptibility discrimination result.
[0016] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements an interpretable landslide disaster susceptibility assessment method as described above.
[0017] The above solution of the present invention has the following beneficial effects: The interpretable landslide hazard susceptibility assessment method and medium provided by the present invention collect landslide catalog data and landslide disaster-forming environment data of the target area, construct a mechanism-geography-physics-mathematics mapping of the disaster-forming environment from a multi-level perspective, and systematically model various factors involved in the disaster-forming environment, avoiding the uncertainty of susceptibility modeling caused by subjective selection of environmental factors; when extracting negative samples, the present invention can comprehensively consider the influence of various factors, use the numerical distribution of the comprehensive frequency ratio to extract the negative sample screening threshold, and then guide the determination of the negative sample sampling quantity and location, avoiding the low accuracy of susceptibility modeling caused by random selection of negative samples; by introducing the FP-Growth algorithm, the present invention regards landslide disasters as events, and various environmental factors can be regarded as different components in the events. The equal-interval grading (or classification) of each factor can be regarded as independent items. By extracting the frequently occurring patterns between different items, the statistical rules contained in the occurrence (or non-occurrence) of landslide disasters are obtained. To a certain extent, these rules can reflect the disaster occurrence law, so as to guide the modeling of susceptibility degree; the present invention further constructs an interpretation template for the susceptibility modeling result. By comprehensively considering the attribute similarity and structure similarity, the susceptibility modeling result is made more scientific, reasonable and easy to understand. Therefore, the present invention has significant transformation value and practical application value in the field of landslide hazard susceptibility assessment.
[0018] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is the overall step flow chart of the present invention; Figure 2 is the detailed process block diagram of the present invention; Figure 3 is the mechanism-geography-physics-mathematics feature mapping diagram of the environmental factors provided by the present invention; Figure 4 is the spatial distribution diagram of the comprehensive frequency ratio and the negative sample sampling space provided by the present invention; Figure 5 is the effective association rule graph structure provided by the present invention; Figure 6 is the association rule graph structure successfully matched by the point to be inferred provided by the present invention; Figure 7 is the spatial distribution diagram of the interpretable landslide hazard susceptibility assessment result in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0020] The following describes the embodiments of the present disclosure through specific examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.
[0021] It should be noted that the following describes various aspects of embodiments within the scope of the appended claims. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.
[0022] It should also be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present disclosure. The diagrams only show the components related to the present disclosure, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex. Additionally, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0023] As Figure 1 、 Figure 2 shown, an embodiment of the present invention provides an interpretable landslide hazard susceptibility assessment method, including the following steps: S1. According to the landslide hazard system theory, construct a mechanism - geography - physics - mathematics mapping, and systematically establish a landslide hazard environmental factor system.
[0024] In this embodiment, this step specifically includes the following sub - steps: S11. Data preparation: The regional disaster system is a system of earth surface anomalies jointly composed of the disaster-forming environment, disaster-causing factors, and disaster-bearing bodies. Among them, the disaster-forming environment is an important part of the disaster system, which provides conditions and background for the formation and development of disaster-causing factors. The disaster-forming environment is unstable, and its stability is affected by various factors, such as geological structure, topography, soil type, vegetation cover, etc. These factors interact with each other and jointly determine the stability degree of the disaster-forming environment and the possibility of disasters occurring. In the regional disaster system, the instability of the disaster-forming environment, the danger of disaster-causing factors, and the vulnerability of disaster-bearing bodies together constitute the functional system of the disaster system, affecting the occurrence, development, and loss degree of disasters. Therefore, when preparing data, in this embodiment, first, the landslide catalog data and landslide disaster-forming environment data of the target area are collected, and then a mechanism-geography-physics-mathematics mapping of the disaster-forming environment is constructed from a multi-level perspective to systematically model various factors involved in the disaster-forming environment. Specifically, these factors include, but are not limited to, topography (slope, elevation, curvature), geological structure (lithology, distance from the fault zone), meteorological hydrology (rainfall intensity, soil moisture content), and social economy (land use, road density), etc.
[0025] S12, Data preprocessing: Spatially align the data of each environmental factor (influence factor) collected, that is, unify the coordinate system (such as WGS84) and raster resolution (such as 30m×30m) of all data to ensure spatial consistency; perform maximum-minimum normalization processing on continuous factors (such as slope), and perform classification coding and normalization processing on discrete factors (such as lithology category) to eliminate the dimension difference.
[0026] S2, Through the frequency ratio method, extract the negative sample search threshold according to the frequency distribution of the comprehensive frequency ratio, infer the number and location of negative sample screening, and realize the optimization of negative samples.
[0027] In this embodiment, this step specifically includes the following sub-steps: S21, Frequency ratio calculation: The frequency ratio (Frequency Ratio, FR ) can be summarized as the ratio of the number of disaster grids in a classification interval of a certain factor to the percentage of all disaster grids to the percentage of the number of grids in this classification interval to the total number of grids in the study area. The calculation formula is: ; Among them, is the number of grids where geological disasters occur in a classification interval of a certain environmental factor, F is the total number of all geological disaster grids in the interval, is the number of grids of a certain environmental factor in the classification interval, is the total number of grids in the study area.
[0028] It should be noted that FR indicates the degree of influence of each classification interval of environmental factors on the occurrence of geological disasters. FR>1 indicates that the environmental factor classification interval has a strong influence on the occurrence of geological disasters, and FR≤1 indicates that the environmental factor classification interval has little influence on the occurrence of disasters.
[0029] As described above, the frequency ratio method first needs to classify or grade the environmental factors of geological disasters. To improve operability and facilitate understanding, in this embodiment, the equal interval method is selected to divide each environmental factor into n levels (in the present invention, n is set as a natural number between 1 and 101). According to the calculation formula of FR, each grid position of each environmental factor in the entire study area will obtain a frequency ratio value.
[0030] It should be noted that in the actual modeling process, the classification or grading of environmental factors of geological disasters can also be carried out by using natural breakpoint method, empirical grading method, parameter optimization method, etc. for more appropriate classification or grading.
[0031] S22, Comprehensive frequency ratio calculation: In order to comprehensively consider the influence of various environmental factors when extracting negative samples, in this embodiment, the concept of comprehensive frequency ratio is proposed, and the numerical distribution of the comprehensive frequency ratio is used to extract the negative sample screening threshold, so as to guide the determination of the number and location of negative sample sampling.
[0032] Specifically, the comprehensive frequency ratio is the mean of the frequency ratios between factors after maximum-minimum normalization based on the spatial distribution of the frequency ratios of various environmental factors. The comprehensive frequency ratio can be expressed formulaically as: ; Among them, m represents the number of environmental factors integrated, represents the frequency ratio distribution of the first grid factor after normalization, represents the spatial distribution of the comprehensive frequency ratio of the entire study area.
[0033] S23, Negative sample screening: It should be noted that the sampling ratio of positive and negative samples is particularly important for the balance of sample distribution. To determine the appropriate number of negative samples, in this embodiment, the statistical method is first used to analyze the numerical distribution of the comprehensive frequency ratio, and the frequency distribution histogram is drawn according to the comprehensive frequency ratio calculated in S22, and the peak value of the comprehensive frequency ratio is set as the negative sample search threshold q. According to the statistical properties of the frequency ratio, when it indicates that the corresponding spatial area is more inclined not to have landslides.
[0034] Then, count the comprehensive frequency ratio and For the corresponding spatial range area, use its ratio as the sampling ratio of positive and negative samples, and determine the sampling quantity of negative samples using this ratio.
[0035] Then, for each disaster point, construct a 3-km buffer based on sampling experience. On this basis, further randomly screen negative samples outside the buffer according to this discriminant condition, and save their geographical coordinates and sample labels.
[0036] S24, Feature Multicollinearity Analysis: After the negative sample screening in S23, the feature composition of different sample points has also been optimized. The feature values are no longer the normalized attribute values obtained in S12, but the frequency ratio calculation values with statistical significance obtained in S21, which improves the representativeness of the features. To determine the factors that truly have an impact on "landslide occurrence" among different environmental factors, in this embodiment, the Variance Inflation Factor (VIF) is further used to statistically analyze all frequency ratio factors. The variance inflation factor is used to measure the degree to which an independent variable g is linearly explained by other independent variables. The larger the VIF value, the more severe the collinearity. VIF can be formulated as: ; where represents the goodness of fit obtained by regressing a certain independent variable g on all other independent variables. It should be noted that when VIF is greater than or equal to 10, it is considered severe collinearity, and this independent variable needs to be deleted.
[0037] S3. Adopt the association rule mining method to extract the landslide rule knowledge subgraph and non-landslide rule knowledge subgraph from the positive and negative sample data respectively, and establish an interpretable model for landslide disaster susceptibility based on the similarity calculation of the rule knowledge subgraph, so as to realize the interpretable evaluation of landslide disaster susceptibility.
[0038] In this embodiment, this step specifically includes the following sub-steps: S31, Association Rule Mining Method: The FP-Growth (Frequent Pattern Growth) algorithm is an efficient algorithm for association rule mining. Its core idea is to construct a frequent pattern tree (FP-Tree), which compactly stores the frequent item sets in the database in the form of a tree, thus avoiding the steps of frequently generating and checking candidate item sets in the traditional Apriori algorithm and significantly improving the mining efficiency. Among them, the FP-Tree is a special prefix tree structure, where each node represents an item and records the number of times the item appears in the transaction database. The tree construction process can effectively compress the data and retain the association information between item sets, enabling subsequent frequent item set mining to be efficiently carried out on the tree structure.
[0039] In this embodiment, through the FP-Growth algorithm, the landslide disaster is regarded as an event, and each environmental factor can be regarded as different constituent elements in the event. The equal-interval grading (or classification) of each element can be regarded as mutually independent items. By extracting the frequent occurrence patterns between different items, the statistical rules contained in the occurrence (or non-occurrence) of the landslide disaster can be further extracted. To a certain extent, these rules can reflect the disaster occurrence law, thereby guiding the interpretable modeling of susceptibility.
[0040] Specifically, the association rule mining method includes the following sub-steps: S311, construct the item header table: traverse the data set, count the support count of each item; delete the items with support lower than the preset threshold t and sort the remaining items in descending order of support to form the item header table; each item header table contains the name of the item, the support count, and a pointer to the first node of the item in the FP-Tree.
[0041] S312, construct the FP-Tree: traverse each transaction in the data set, and process each item in turn according to the order of items in the item header table; starting from the root node, if the current item exists in the children nodes of the current node, increase the support count of the child node; otherwise, create a new child node and update the linked list of the item in the item header table.
[0042] S313, construct the conditional pattern base: for each item in the item header table, starting from the last node of its linked list, recursively traverse the linked list to generate the conditional pattern base with the node as the suffix path. Among them, the conditional pattern base contains the other items in the path except the current item and the corresponding support count.
[0043] S314, recursively mine the FP-Tree: for each item in the item header table, combine it with the conditional pattern base to form a new frequent item set; if the conditional pattern base is not empty, use the conditional pattern base as the input and recursively call the construction and mining process of the FP-Tree until mining can no longer continue.
[0044] S315, Association rule generation: After obtaining the frequent item sets, association rules are generated by calculating metrics such as support and confidence. Among them, support represents the frequency of the item set (X, Y) appearing in the transaction database, and confidence measures the reliability of the rule, that is, the probability of Y being included in the transactions that contain X. For example, for the rule A → B, its support is , and the confidence is , that is, the probability of B being included in the transactions that contain A. By setting the rule support threshold E1 or the confidence threshold E2, effective disaster occurrence (or non-occurrence) rules are further screened.
[0045] S32, Similarity calculation: After obtaining several rules with a certain degree of support and confidence through S31, assume that a certain rule in the rule set is expressed as , where F, G, and H respectively represent a certain environmental factor, and 1, 2, and 3 respectively represent the grading (or classification) numbers obtained by the equal interval method for the environmental factors. Therefore, the different rules mined can form a knowledge graph structure, which is used to represent the prior statistical knowledge reflected by the disaster occurrence (or non-occurrence). For the position to be inferred, the corresponding feature interval and the interval combination relationship can be discriminated according to the eigenvalue calculated by S23, and then the similarity between the matched rule and the rule set representing the prior statistical knowledge can be calculated.
[0046] In a specific embodiment, assume there is a point O to be speculated, and its feature interval is , where represents the frequency ratio attribution interval of the m-th environmental factor. The rule set R contains multiple association rules, and each association rule is in the form of , where represents the k-th rule item in the j-th association rule. To ensure that the association rules can effectively reflect the disaster occurrence law, the acquisition process of the rule set R needs to first filter out the intersection of the positive sample rule set and the negative sample rule set from the positive sample rule set.
[0047] Therefore, the similarity calculation method can be formulated as: ; where M is the number of rules in the rule set R, is the number of items in the j-th rule, represents the structure similarity weight, which can be expressed by the rule support or confidence, represents the structure similarity, which is expressed by the length ratio of all items in the j-th rule to the point to be speculated, is the attribute similarity between the point O to be speculated and the j-th rule. and can be further formulated as follows: ; ; Therefore, According to whether the rule is within the set of characteristic factors, the value is judged to be 1 or 0, and u represents the number of significant characteristic items of the point O to be inferred.
[0048] S33, Explanation of the susceptibility discrimination result: In this embodiment, an explanation template for the susceptibility discrimination result is constructed. By comprehensively considering the attribute similarity and the structure similarity, the susceptibility discrimination result is made more scientific, reasonable and easy to understand.
[0049] In a specific embodiment, the explanation template is expressed as follows: In the result of this susceptibility modeling, first start from the positive and negative sample sets, extract [number of positive sample association rules] significant and reliable landslide occurrence rules and [number of negative sample association rules] significant and reliable landslide non-occurrence rules. After calculating the intersection of the rules, a total of [number of intersection association rules] rules with ambiguous explanations are filtered out. Finally, [number of positive sample association rules - number of intersection association rules] effective association rules are obtained. Specifically include: {Rule 1} {Rule 2} … {Rule } These association rules are pre-mined from the positive samples based on the FP-Growth algorithm, and the ambiguous rules that appear in both the positive and negative samples are filtered out. These association rules reflect to a certain extent the internal connections and combination patterns between different environmental factor characteristic items.
[0050] For the point to be inferred [specific point coordinates], according to the frequency ratio distribution of its environmental factors, extract the association rules that match it. The successfully matched association rules specifically include: {Rule 1} {Rule 2} … {Rule }
[0051] Regard the characteristic items in these successfully matched effective association rules as nodes, and establish edge links between the characteristic items in the rules to form Graph Structure 1. At the same time, construct Graph Structure 2 according to all the effective association rules. By combining the characteristic values of the environmental factor frequency ratios of the point to be inferred, calculate the similarity (attribute similarity and structure similarity) between Graph Structure 1 and Graph Structure 2, and the susceptibility discrimination result can be intuitively explained.
[0052] For example, the susceptibility is classified into five levels of extremely low susceptibility, low susceptibility, medium susceptibility, high susceptibility, and extremely high susceptibility according to the equal-interval classification method: 0.2, 0.4, 0.6, 0.8, 1. If the similarity between Graph Structure 1 and Graph Structure 2 is high ( ), it indicates that the environmental factor feature items of the point to be inferred have a high degree of coincidence with the susceptibility feature pattern in the positive samples, thus supporting a high degree of susceptibility at this point. On the contrary, if the similarity is low ( ), it indicates that its feature pattern is quite different from the known susceptibility situation, and the degree of susceptibility is relatively low.
[0053] It can be seen from the above that this interpretation method based on graph structure similarity not only considers the feature items of each environmental factor itself, but also fully reflects the combination relationship and interaction among them. By comprehensively considering the attribute similarity and structure similarity, the susceptibility discrimination result is made more scientific, reasonable and easy to understand, thus significantly improving the susceptibility assessment of landslide disasters.
[0054] The effects of the present invention are further described below through specific cases. Considering that Yunnan Province is located in the slope zone from the Qinghai-Tibet Plateau to the Yunnan-Guizhou Plateau, with active geological structures, severe terrain cutting, concentrated and intense rainfall. This example selects Yunnan Province as the research area, and the following will specifically describe the specific implementation steps of the present solution for landslide disaster susceptibility assessment in combination with landslide disaster examples: S1, construct the mechanism-geography-physics-mathematical feature mapping, as Figure 3 shown. According to environmental factors such as landform, geological structure, meteorological hydrology, and social economy, construct an index system including but not limited to 19 kinds of susceptibility factor data such as slope, aspect, ndvi, human density, low-grade road density, low-grade road distance, profile curvature, land use type, soil type, terrain undulation degree, engineering geological map, plane curvature, annual average rainfall, fault density, fault distance, water system distance, elevation, high-grade road density, and high-grade road distance. The above data extracts the environmental factor features from multiple levels of mechanism, geography, physics, and mathematics, and the data comes from the Internet Geographic Information Public Platform.
[0055] Preprocess each index variable. Perform spatial alignment on the collected environmental factor data, that is, unify the coordinate system (such as WGS84) and raster resolution (such as 30m×30m) of all data to ensure spatial consistency. In addition, perform maximum-minimum normalization processing on continuous factors (such as slope), and perform classification coding and normalization processing on discrete factors (such as lithology category) to eliminate the dimension difference.
[0056] S2, Negative sample optimization. Before negative sample sampling, it is first necessary to determine the sampling area and sampling ratio of negative samples. According to the comprehensive frequency ratio calculation method, the comprehensive frequency ratio distribution of the study area is obtained as shown in Figure 4 . At the same time, the peak of the frequency distribution histogram of the comprehensive frequency ratio of the study area is further extracted. The results show that the comprehensive frequency ratio of the grid in the whole area peaks in the interval of 0.55 - 0.6. Therefore, the negative sample screening threshold is set to 0.55, and the number of negative samples is Figure 3 the area ratio between the comprehensive frequency ratio less than 0.55 and greater than 0.55 in the above multiplied by the number of landslide disaster points, that is, 1 / 0.56×1397 = 2495 negative samples, where 1397 is the number of landslide disaster points in the study area in the past ten years.
[0057] S3, Interpretability - enabled susceptibility modeling. First, each environmental factor is regarded as a different component element in the event, and the equal - interval grading (or classification) of each element can be regarded as independent items. By extracting the frequently occurring patterns between different items, the statistical rules implied by the occurrence (or non - occurrence) of landslides are further extracted. Among them, when running the FP - Growth algorithm, in this case, the support threshold t of the feature item is set to 0.5, and the support threshold T of the item set is set to 0.75. The threshold indicates that items with a frequency lower than 0.5 in the positive sample (or negative sample) event will be eliminated. Furthermore, item sets with a frequency lower than 0.75 in the positive sample (or negative sample) event will not be listed as association rules. Some of the item lists and item set lists obtained in the implementation process are shown in Table 1: Table 1 Some items and item sets (the letters of the items represent a certain environmental factor, and the numbers represent the grading (or classification) numbers of the environmental factors obtained by the equal - interval method)
[0058] By performing intersection filtering on the positive sample item set list and the negative sample item set list, a total of 23 pseudo - rules are filtered out, and finally 110 effective association rules related to landslide occurrence are obtained. The graph structure of the effective association rules formed is as shown in Figure 5 . Taking a certain landslide event outside the sample data set as an example, 102 effective association rules are successfully matched at the point to be inferred, and the graph structure formed is as shown in Figure 6 .
[0059] Based on the interpretation template in the previous text, the susceptibility of this landslide event is explained as follows: In the results of this susceptibility modeling, first starting from the positive and negative sample sets,
[133] significant and reliable landslide occurrence rules and
[45] significant and reliable landslide non - occurrence rules are extracted. After calculating the intersection of the rules, a total of
[23] rules with ambiguous interpretations are filtered out. Finally,
[110] effective association rules are obtained. Specifically including: {D4, B3} {D4, R0} {D4, K4} {D4, O1} {D4, B3, R0} … {B3, R0, D4, O1, K4, Q2} For the point to be inferred [105.011, 27.480519], according to the frequency ratio distribution of its environmental factors, the associated rules that match it are extracted. The specifically matched associated rules include: {D4, B3} {D4, R0} {D4, K4} {D4, O1} {D4, B3, R0} … {O1, K4, B3, R0, D4} The probability of occurrence is divided into five levels according to the equal-interval grading method: extremely low probability of occurrence, low probability of occurrence, medium probability of occurrence, high probability of occurrence, and extremely high probability of occurrence: 0.2, 0.4, 0.6, 0.8, 1. According to the similarity calculation formula, the similarity between the first graph structure and the second graph structure is 0.92, (greater than the high probability of occurrence threshold of 0.6), indicating that the environmental factor feature items at the point to be inferred have a high degree of coincidence with the probability of occurrence feature patterns in the positive samples, thus supporting that this point has a high probability of occurrence. Finally, the spatial distribution map of the interpretable landslide hazard probability of occurrence assessment results is as Figure 7 shown.
[0060] Based on the same inventive concept, this embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the aforementioned interpretable landslide hazard probability of occurrence assessment method.
[0061] The computer-readable medium includes but is not limited to any type of disk (including floppy disks, hard disks, optical disks, CD-ROMs, and magneto-optical disks), ROM, RAM, EPROM (Erasable Programmable Read-Only Memory), EEPROM, flash memory, magnetic cards or optical cards. That is to say, the computer-readable medium includes any medium that stores or transmits information in a form that can be read by a device.
[0062] The computer-readable storage medium provided in this embodiment, etc., has the same inventive concept and the same beneficial effects as the aforementioned method, and will not be elaborated here.
[0063] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0064] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it cannot be understood as a limitation to the scope of the application for this reason. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. An interpretable landslide hazard susceptibility assessment method, characterized in that It includes the following steps: S1. According to the landslide disaster system theory, construct a mechanism-geography-physics-mathematics mapping, and systematically establish a landslide disaster environmental factor system; S2. Through the frequency ratio method, extract the negative sample search threshold according to the frequency distribution of the comprehensive frequency ratio, infer the number and location of negative sample screening, and complete the optimization of negative samples; The frequency ratio is the ratio of the number of disaster grids in a classification interval of a certain environmental factor to the percentage of all disaster grids, and the percentage of the number of grids in this classification interval to the total number of grids in the study area; the comprehensive frequency ratio is the mean value of the frequency ratios between factors after maximum-minimum normalization based on the spatial distribution of the frequency ratios of each environmental factor; S3. Adopt the association rule mining method, extract the landslide rule knowledge subgraph and non-landslide rule knowledge subgraph according to the positive sample and negative sample data respectively, establish an interpretable model for landslide disaster susceptibility based on the calculation of the similarity of the rule knowledge subgraph, and realize the interpretable assessment of landslide disaster susceptibility.
2. The interpretable landslide disaster susceptibility assessment method according to claim 1, characterized in that, S1 specifically includes the following sub-steps: S11. Data preparation: S11. Collect the landslide catalog data and landslide disaster-causing environment data of the target area, construct a mechanism-geography-physics-mathematics mapping of the disaster-causing environment from a multi-level perspective, and model various factors involved in the disaster-causing environment; S12. Align the spatial data of each environmental factor collected, and unify the coordinate system and grid resolution of all data; Perform maximum-minimum normalization processing on continuous factors, and perform classification coding and normalization processing on discrete factors to eliminate the dimension difference.
3. An interpretable landslide disaster susceptibility assessment method according to claim 2, characterized in that, S2 specifically includes the following sub-steps: S21. Calculate the frequency ratio, and the calculation formula is: ; Wherein, is the frequency ratio, is the number of grid cells where geological disasters occur within the classification interval for a certain environmental factor, F is the total number of grid cells of all geological disasters within the interval, is the number of grid cells of a certain environmental factor within the classification interval, is the total number of grid cells in the study area; S22. Calculate the comprehensive frequency ratio, and the calculation formula is: ; where m represents the number of integrated environmental factors, represents the frequency ratio distribution of the first raster factor after normalization, represents the comprehensive frequency ratio spatial distribution of the entire study area; S23. Using a statistical method to analyze the numerical distribution of the comprehensive frequency ratio, drawing a frequency distribution histogram based on the calculated comprehensive frequency ratio, and setting the peak value of the comprehensive frequency ratio as the negative sample search threshold q; according to the statistical properties of the frequency ratio, when it indicates that the corresponding spatial region is more inclined not to have landslides. Statistical comprehensive frequency ratio and The area of the spatial range corresponding thereto is used as the sampling ratio of the positive sample and the negative sample, and the sampling number of the negative sample is determined by using this ratio; For each disaster point, a buffer zone with a preset range is constructed according to sampling experience. Further outside the buffer zone, according to this discrimination condition, negative samples are randomly selected, and their geographical coordinates and sample labels are saved.
4. An interpretable landslide disaster susceptibility assessment method according to claim 3, characterized in that, S2 also includes the following sub-steps: S24. Use the variance inflation factor to perform statistical analysis on all frequency ratio environmental factors. The variance inflation factor is used to measure the degree to which an independent variable g is linearly explained by other independent variables. The larger the variance inflation factor, the more serious the collinearity. The variance inflation factor VIF is expressed as: ; Among them, represents the goodness of fit obtained by regressing all other independent variables with a certain independent variable g.
5. An interpretable landslide disaster susceptibility assessment method according to claim 3, characterized in that, S2 includes the following sub-steps: S31. Through the FP-Growth algorithm, regard the landslide disaster as an event, regard each environmental factor as different components in the event, and regard the equal-interval classification of each element as independent items, extract the frequently occurring patterns between different items, obtain the statistical rules contained in the occurrence or non-occurrence of landslide disasters, and reflect the occurrence law of landslide disasters; S32. For the position to be inferred, judge the characteristic interval and the interval combination relationship, and calculate the similarity between the matched rule and the rule set representing the prior statistical knowledge; S33. Construct an interpretation template for the susceptibility discrimination result, comprehensively consider the attribute similarity and the structure similarity, and complete the assessment of landslide disaster susceptibility.
6. The interpretable landslide disaster susceptibility assessment method according to claim 5, characterized in that, S31 specifically includes the following sub-steps: S311. Traverse the data set, count the support count of each item; delete the items with support lower than the preset threshold t, and sort the remaining items in descending order of support to form an item header table; each item header table contains the name of the item, the support count, and a pointer to the first node of the corresponding item in the FP-Tree; S312. Traverse each transaction in the dataset and process each item in sequence according to the order of items in the item header table; starting from the root node, if the current item exists in the child nodes of the current node, increment the support count of that child node; otherwise, create a new child node and update the linked list of this item in the item header table; S313. For each item in the item header table, starting from the last node of its linked list, recursively traverse the linked list to generate a conditional pattern base with the node as the suffix path. The conditional pattern base contains other items in the path except the current item and their corresponding support counts; S314. For each item in the item header table, combine it with the conditional pattern base to form a new frequent item set; if the conditional pattern base is not empty, use the conditional pattern base as the input and recursively call the construction and mining process of the FP-Tree until mining can no longer continue; S315. Calculate the support and confidence based on the frequent item sets to generate association rules; the support represents the frequency of the item set (X, Y) appearing in the transaction database, and the confidence represents the probability that Y is contained in the transactions that contain X. Set the rule support threshold E1 or the confidence threshold E2 to filter out effective disaster occurrence or non-occurrence rules.
7. An interpretable landslide disaster susceptibility assessment method according to claim 6, characterized in that, In S32, for the point O to be inferred, Its characteristic interval is , where represents the frequency ratio attribution interval of the m-th environmental factor; the rule set R contains multiple association rules, and each association rule is in the form of , where represents the k-th rule item in the j-th association rule; Similarity The calculation formula is as follows: ; where M is the number of rules in the rule set R, is the number of items in the j-th rule, represents the structural similarity weight, expressed by rule support or confidence, represents the structural similarity, expressed by the length ratio of all items in the j-th rule to the items at the point to be inferred, is the attribute similarity between the point O to be inferred and the j-th rule, and are respectively: ; ; Judge whether the value is 1 or 0 according to whether the rule is within the set of characteristic factors. u represents the number of significant characteristic items of the point O to be inferred.
8. An interpretable landslide disaster susceptibility assessment method according to claim 7, characterized in that In S33, regard the feature items in the successfully matched effective association rules as nodes, and establish edge links between the feature items in the rules to form Graph Structure 1; construct Graph Structure 2 according to all effective association rules; by combining the frequency ratio eigenvalue of the environmental factors of the point to be inferred, calculate the similarity between Graph Structure 1 and Graph Structure 2 to intuitively explain the susceptibility discrimination result.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements an interpretable landslide disaster susceptibility assessment method according to any one of claims 1-8.
Citation Information
Patent Citations
Landslide risk prediction method and system based on knowledge graph construction
CN111639878A
Landslide susceptibility evaluation method and system
CN114091274A
Slope unit landslide susceptibility evaluation method based on optimized negative sample selection
CN118332434A
Landslide susceptibility evaluation method based on negative sample optimization selection strategy
CN118410534A
System and method for identifying sets of attributes in a database, system and method for constructing a tree structure for a database and computer program elements
WO2007055655A1