Mountain disaster susceptibility prediction method and device considering compound terrain distance

By constructing a composite terrain distance and optimizing the parameters of a geographically weighted random forest model, the problem of terrain differences not being considered in traditional methods is solved, and accurate and stable prediction of mountain disaster susceptibility assessment is achieved.

CN121881130BActive Publication Date: 2026-05-22YUNNAN NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YUNNAN NORMAL UNIV
Filing Date
2026-03-17
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Traditional methods for assessing the susceptibility of mountain disasters are insufficient to accurately reflect the true risk pattern of complex terrain, and geographically weighted random forest models do not fully consider terrain differences, resulting in insufficient prediction accuracy and stability.

Method used

A composite terrain distance model is constructed, which integrates spatial distance and terrain differences. The weights of terrain factors are calculated using the entropy weight method. The geographical weighted weights and random forest model parameters are optimized by combining the Bayesian optimization method, and a geographically weighted random forest model is constructed to achieve adaptive optimization of key parameters.

Benefits of technology

It improves the accuracy and stability of mountain disaster susceptibility assessment, accurately characterizes the spatial heterogeneity of complex mountainous areas, and enhances the accuracy and consistency of prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121881130B_ABST
    Figure CN121881130B_ABST
Patent Text Reader

Abstract

The application discloses a mountain disaster susceptibility prediction method and device considering composite terrain distance, and relates to the field of natural disaster risk assessment, wherein the method comprises the following steps: obtaining mountain disaster samples and multi-source environmental factors in a study area; selecting terrain factors from the multi-source environmental factors, and calculating the weight of each terrain factor based on an entropy weight method; introducing a weighted terrain difference distance on the basis of a Euclidean space distance, and fusing and constructing a composite terrain distance; sorting and screening the mountain disaster sample data according to the composite terrain distance, determining a local reference sample set and a geographical weighted weight corresponding to each local reference sample; adopting a Bayesian optimization method to jointly optimize the bandwidth parameter of the geographical weighted weight, the weight parameter of the composite terrain distance and the structure parameter of a random forest model, and obtaining an optimal parameter combination; and under the constraint of the optimal parameter combination, predicting the study area point by point. The application can improve the precision of mountain disaster susceptibility prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural disaster risk assessment technology, and in particular to a method and apparatus for predicting the susceptibility of mountain disasters that takes into account the distance of complex terrain. Background Technology

[0002] Mountain hazard susceptibility assessment is the process of quantitatively characterizing the probability of occurrence and spatial distribution characteristics of hazards such as landslides, collapses, and debris flows. Due to the large topographic relief and complex landforms in mountainous areas, the hazard formation process exhibits significant spatial non-stationarity and local heterogeneity. Traditional susceptibility assessment methods based on global statistical assumptions are difficult to accurately reflect the true hazard risk pattern in mountainous areas with complex topography.

[0003] In recent years, machine learning methods such as random forests have been widely used in assessing the susceptibility of mountain disasters due to their strong nonlinear fitting ability and good adaptability to multiple factors. However, traditional random forest models usually ignore the spatial correlation between samples, assuming that samples are independent, making it difficult to characterize the spatial differences in disaster occurrence mechanisms. To address this, the geographically weighted random forest method, by introducing spatial distance weights, enhances the model's ability to characterize spatial heterogeneity to some extent.

[0004] However, existing geographically weighted random forest methods generally use Euclidean spatial distance to measure the proximity relationship between samples, without fully considering the impact of differences in terrain elements such as elevation, slope, and aspect on disaster incubation. In complex mountainous areas, this can easily lead to samples with large differences in terrain environment but spatial proximity being incorrectly included in local modeling, weakening the physical rationality and prediction accuracy of the model. Furthermore, geographically weighted bandwidth and model parameters often rely on empirical settings or single-objective optimization, making it difficult to achieve adaptive adjustment under different terrain backgrounds, which further limits its application effect.

[0005] Therefore, in the practice of disaster risk assessment in complex mountainous areas, how to break through the limitations of traditional models that rely solely on distance measurement, achieve synergistic consideration of spatial location and terrain features, and improve model adaptability through parameter optimization strategies remains a core technical problem that urgently needs to be solved in the field of mountain disaster susceptibility assessment. Summary of the Invention

[0006] This invention provides a method for predicting the susceptibility of mountain disasters that takes into account complex terrain distances. This method constructs a composite terrain distance that integrates spatial distance and terrain differences, and adaptively optimizes key parameters within a geographically weighted random forest framework. This accurately characterizes the spatial heterogeneity of disaster susceptibility in complex mountainous areas, improving the accuracy and stability of susceptibility assessment results. The method includes:

[0007] Acquire mountain hazard sample data and multi-source environmental factor data within the study area;

[0008] Topographic factors are screened from the multi-source environmental factors, and the weights of each topographic factor are calculated based on the entropy weight method to quantify the contribution of different topographic factors to topographic differences. Based on Euclidean spatial distance, a weighted topographic difference distance is introduced based on the weights of each topographic factor, and a composite topographic distance is constructed to simultaneously characterize the spatial proximity and topographic similarity between samples.

[0009] Based on the composite terrain distance, the mountain disaster sample data around the predicted location are sorted and filtered to determine the local reference sample set, and the geographical weighting weight corresponding to each local reference sample is calculated based on the kernel function.

[0010] The Bayesian optimization method is used to jointly optimize the bandwidth parameter of the geographic weighting, the weight parameter of the composite terrain distance, and the structural parameters of the random forest model to obtain the optimal parameter combination.

[0011] Under the constraints of the optimal parameter combination, a geographically weighted random forest model is constructed based on local reference samples and their weights to make point-by-point predictions for the study area, and the probability value of mountain disaster susceptibility is obtained as the prediction result of mountain disaster susceptibility.

[0012] This invention also provides a mountain disaster susceptibility prediction device that takes into account complex terrain distances. This device constructs a composite terrain distance that integrates spatial distance and terrain differences, and adaptively optimizes key parameters within a geographically weighted random forest framework. This accurately characterizes the spatial heterogeneity of disaster susceptibility in complex mountainous areas, improving the accuracy and stability of susceptibility assessment results. The device includes:

[0013] The acquisition unit is used to acquire mountain disaster sample data and multi-source environmental factor data within the study area;

[0014] The composite terrain distance construction unit is used to screen terrain factors from the multi-source environmental factors, calculate the weight of each terrain factor based on the entropy weight method to quantify the contribution of different terrain factors to terrain differences; based on Euclidean spatial distance, a weighted terrain difference distance is introduced based on the weight of each terrain factor, and the composite terrain distance is constructed to simultaneously characterize the spatial proximity and terrain similarity between samples.

[0015] The local reference sample screening and weight calculation unit is used to sort and screen the mountain disaster sample data around the predicted location based on the composite terrain distance, determine the local reference sample set, and calculate the geographical weighting weight corresponding to each local reference sample based on the kernel function.

[0016] The joint optimization unit is used to jointly optimize the bandwidth parameter of the geographic weighting, the weight parameter of the composite terrain distance, and the structural parameters of the random forest model using the Bayesian optimization method to obtain the optimal parameter combination.

[0017] The prediction unit is used to construct a geographically weighted random forest model based on local reference samples and their weights under the constraints of the optimal parameter combination, and to make point-by-point predictions for the study area to obtain the probability value of mountain disaster susceptibility.

[0018] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned method for predicting the susceptibility of mountain disasters taking into account the distance of complex terrain.

[0019] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for predicting the susceptibility of mountain disasters taking into account the distance of complex terrain.

[0020] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-mentioned method for predicting the susceptibility of mountain disasters taking into account the distance of complex terrain.

[0021] In this embodiment of the invention, the beneficial technical effect of the mountain disaster susceptibility prediction scheme that takes into account composite terrain distance is achieved through: acquiring mountain disaster sample data and multi-source environmental factor data within the study area; screening terrain factors from the multi-source environmental factors, calculating the weight of each terrain factor based on the entropy weight method to quantify the contribution of different terrain factors to terrain differences; introducing weighted terrain difference distance based on the weight of each terrain factor, and constructing a composite terrain distance based on Euclidean spatial distance to simultaneously characterize the spatial proximity and terrain similarity between samples; sorting and screening mountain disaster sample data around the location to be predicted according to the composite terrain distance to determine the local reference sample set, and calculating the local reference sample set based on the kernel function. The geographically weighted weights corresponding to the local reference samples are determined. A Bayesian optimization method is used to jointly optimize the bandwidth parameters of the geographically weighted weights, the weight parameters of the composite terrain distance, and the structural parameters of the random forest model to obtain the optimal parameter combination. Under the constraint of this optimal parameter combination, a geographically weighted random forest model is constructed based on the local reference samples and their weights to perform point-by-point predictions for the study area. The resulting probability value of mountain disaster susceptibility is used as the prediction result. This approach enables the construction of a composite terrain distance that integrates spatial distance and terrain differences, and achieves adaptive optimization of key parameters within the geographically weighted random forest framework. This accurately characterizes the spatial heterogeneity of disaster susceptibility in complex mountainous areas, improving the accuracy and stability of susceptibility prediction results. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0023] Figure 1 This is a flowchart illustrating the mountain disaster susceptibility prediction method that takes into account the distance of complex terrain in an embodiment of the present invention.

[0024] Figure 2 This is a schematic diagram illustrating sample selection based on Euclidean distance in an embodiment of the present invention;

[0025] Figure 3 This is a schematic diagram illustrating sample selection based on composite terrain distance in an embodiment of the present invention;

[0026] Figure 4 This is a line graph showing the optimization of the α-AUC parameters in an embodiment of the present invention;

[0027] Figure 5 This is a diagram showing the bandwidth adaptive optimization results in an embodiment of the present invention;

[0028] Figure 6 This is a spatial distribution map of the predicted probability of mountain disaster susceptibility based on a composite terrain distance-weighted random forest model in an embodiment of the present invention.

[0029] Figure 7 This is a comparison of the prediction accuracy of different susceptibility assessment models in the embodiments of the present invention;

[0030] Figure 8 This is a schematic diagram illustrating the overlay verification of disaster point and susceptibility zoning results in an embodiment of the present invention;

[0031] Figure 9 This is a frequency ratio statistical analysis chart based on the susceptibility partitioning results in an embodiment of the present invention;

[0032] Figure 10 This is a schematic diagram of the mountain disaster susceptibility prediction device that takes into account the distance of complex terrain in an embodiment of the present invention. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0034] The acquisition, storage, use, and processing of data in this application comply with relevant laws and regulations.

[0035] This invention proposes a mountain disaster susceptibility prediction scheme that considers composite terrain distances. This method specifically improves the Geographically Weighted Random Forest (GWRF) model, constructing a composite terrain distance that integrates spatial location and terrain features. This allows spatial proximity and terrain similarity to jointly participate in local reference sample search and neighborhood weight calculation. A Bayesian optimization algorithm is used to jointly optimize the composite distance weights, geographic weighted bandwidth, and random forest structure parameters to determine the optimal parameter combination. Based on this, a Composite Terrain-distance Geographically Weighted Random Forest (CTD-GWRF) model is constructed and applied to mountain disaster susceptibility assessment to improve the accuracy, stability, and practicality of the assessment results under complex terrain conditions. The following is a detailed description of this mountain disaster susceptibility prediction scheme that considers composite terrain distances.

[0036] Figure 1 This is a flowchart illustrating the mountain disaster susceptibility prediction method considering composite terrain distances in an embodiment of the present invention. Figure 1 As shown, the specific method of the present invention includes the following steps:

[0037] Step 101: Obtain mountain hazard sample data and multi-source environmental factor data within the study area;

[0038] Step 103: Screen topographic factors from the multi-source environmental factors, calculate the weight of each topographic factor based on the entropy weight method to quantify the contribution of different topographic factors to topographic differences; based on the Euclidean spatial distance, introduce a weighted topographic difference distance based on the weight of each topographic factor, and fuse them to construct a composite topographic distance to simultaneously characterize the spatial proximity and topographic similarity between samples.

[0039] Step 104: Based on the composite terrain distance, sort and filter the mountain disaster sample data around the location to be predicted, determine the local reference sample set, and calculate the geographical weighting weight corresponding to each local reference sample based on the kernel function.

[0040] Step 105: Using the Bayesian optimization method, the bandwidth parameter of the geographic weighting, the weight parameter of the composite terrain distance, and the structural parameters of the random forest model are jointly optimized to obtain the optimal parameter combination;

[0041] Step 106: Under the constraints of the optimal parameter combination, a geographically weighted random forest model is constructed based on local reference samples and their weights to make point-by-point predictions for the study area, and the probability value of mountain disaster susceptibility is obtained as the prediction result of mountain disaster susceptibility.

[0042] The method for predicting mountain hazard susceptibility considering composite terrain distances provided in this invention involves: acquiring mountain hazard sample data and multi-source environmental factor data within the study area; filtering terrain factors from the multi-source environmental factors, calculating the weights of each terrain factor based on the entropy weight method to quantify the contribution of different terrain factors to terrain differences; introducing weighted terrain difference distances based on the weights of each terrain factor, and constructing composite terrain distances to simultaneously characterize the spatial proximity and terrain similarity between samples; sorting and filtering mountain hazard sample data around the location to be predicted based on the composite terrain distances to determine a local reference sample set, and calculating the geographical weights corresponding to each local reference sample based on a kernel function; and employing Bayesian optimization... This invention employs a method to jointly optimize the bandwidth parameters of the geographic weighted weights, the weight parameters of the composite terrain distance, and the structural parameters of the random forest model to obtain the optimal parameter combination. Under the constraint of this optimal parameter combination, a geographic weighted random forest model is constructed based on local reference samples and their weights to perform point-by-point predictions for the study area, obtaining the probability value of mountain disaster susceptibility as the prediction result. The beneficial technical effects of the mountain disaster susceptibility prediction method considering composite terrain distance provided in this embodiment are: it can achieve the construction of a composite terrain distance that integrates spatial distance and terrain differences, and realize adaptive optimization of key parameters within the geographic weighted random forest framework, accurately characterizing the spatial heterogeneity of disaster susceptibility in complex mountainous areas, and improving the accuracy and stability of susceptibility prediction results. The following is a detailed introduction to this mountain disaster susceptibility prediction method considering composite terrain distance.

[0043] In step 101 above, sample data of mountain hazards and multi-source environmental factors within the study area are acquired. To ensure the effectiveness of subsequent model training and susceptibility assessment, this step integrates high-quality data from multiple channels to construct a basic dataset covering the core influencing factors of disaster occurrence, specifically including the following sub-steps:

[0044] 1.1: Data Collection and Screening of Mountain Hazard Samples. Core information such as spatial coordinates (latitude and longitude or Cartesian coordinates), hazard type, and occurrence time of typical mountain hazards (landslides, collapses, debris flows, etc.) in the study area were collected through authoritative channels including geological hazard risk survey databases, professional landslide monitoring platforms, and regional geological survey results. During the screening process, invalid records with ambiguous spatial locations or missing attribute information were removed, retaining a sufficient number of spatially evenly distributed hazard point samples. This ensured that the samples covered different high-risk landform units such as deep canyons, steep slopes, fault zones, and gully confluence areas, meeting the model training requirements for sample representativeness.

[0045] 1.2: Acquisition of Multi-Source Environmental Factor Data: Based on the causal mechanisms of mountain disasters, core environmental factors covering topography, geology, hydrology, climate, vegetation, and human activities were selected to comprehensively cover the key elements for disaster formation and occurrence. Data sources included DEMs, regional geological maps, meteorological observation data, satellite remote sensing imagery, and basic geographic information data, from which corresponding types of factors were extracted. All data were uniformly collected using the CGCS2000_Gauss_KrugerCM_93E coordinate system to ensure spatial reference consistency.

[0046] 1.3: Data Format Standardization. Raw data from different sources and in different formats were uniformly converted into raster data format. The raster resolution was reasonably set according to the study area and evaluation accuracy requirements to ensure that the data accurately reflects local topographic and environmental differences. Simultaneously, edge matching and spatial alignment were performed on all raster data to eliminate data splicing gaps and coordinate offsets, providing complete spatial data for subsequent preprocessing and model building.

[0047] Following step 101 and preceding step 103, the following step 102 is included: preprocessing of multi-source environmental factors and construction of the sample dataset. This step eliminates data noise and interference factors through data cleaning, standardization, redundancy removal, and sample equalization to construct a high-quality sample dataset that meets the requirements for model training. Specifically, it includes the following sub-steps:

[0048] 2.1: Normalization of Continuous Factors. For continuous environmental factors such as elevation, slope, rainfall, and vegetation cover, a minimum-maximum normalization method is used for dimensionless processing to eliminate the influence of differences in the dimensions of different factors on model training. The calculation formula is as follows:

[0049] ;

[0050] in, For the original values ​​of the factor, and These are the minimum and maximum values ​​of the factor within the study area, respectively. The normalized value is taken in the range of [0,1], ensuring that continuous factors of different magnitudes can be directly used in the model calculation.

[0051] 2.2: Discrete Factor Coding Process. Discrete factors such as land cover type, lithology type, and soil type are coded hierarchically based on their classification system and hazard correlation, using integer sequences (1-n, where n is the total number of factor categories) to represent different categories. During the coding process, consistency within the same category and distinctiveness between different categories are ensured, enabling the discrete factors to be effectively identified and utilized by the model.

[0052] 2.3: Collinearity Test and Factor Selection. The variance inflation factor (VIF) is used to quantify the collinearity strength among environmental factors, avoiding model parameter estimation instability and decreased generalization ability due to information redundancy between factors. A reasonable VIF threshold is set, and redundant factors with VIF values ​​exceeding the threshold are eliminated, retaining core factors that have a significant impact on disaster occurrence and are independent of each other for global modeling. Simultaneously, to meet the requirements for constructing composite terrain distances, factors directly related to terrain features (such as aspect, curvature, and elevation) are selected from the core factors, and collinearity tests and selections are performed separately to ensure that the combination of terrain factors can comprehensively and without redundancy characterize terrain differences.

[0053] As can be seen from the above, in one embodiment, the above-mentioned method for predicting the susceptibility of mountain disasters taking into account the distance of complex terrain may further include: normalizing the multi-source environmental factor data and performing multicollinearity test processing.

[0054] 2.4: Balanced Sample Set Construction and Partitioning. A reverse selection method using Euclidean distance buffers is employed. Non-hazardous samples are randomly selected from stable areas far from known disaster points, maintaining a balance between the number of non-hazardous and hazardous samples to effectively mitigate the adverse effects of class imbalance on model training. The hazard and non-hazardous samples are combined to form a complete sample set, which is then randomly divided into a training set and a validation set according to a preset ratio (e.g., 8:2). The training set is used for model parameter optimization and model construction, while the validation set is used for model performance evaluation. During the partitioning process, consistency in spatial distribution and environmental characteristics between the training and validation sets is ensured to avoid distortion of model evaluation results due to sample partitioning bias.

[0055] As can be seen from the above, in one embodiment, the above-mentioned method for predicting the susceptibility of mountain disasters taking into account the distance of complex terrain further includes: normalizing the multi-source environmental factor data and performing multicollinearity test processing.

[0056] In step 103 above, the composite terrain distance is constructed. This step breaks through the limitation of traditional geographic weighted models that rely solely on Euclidean distance, and integrates spatial proximity and terrain similarity to construct a composite terrain distance that can accurately characterize the correlation features of complex mountainous environment samples. Specifically, it includes the following sub-steps:

[0057] 3.1: Topographic Factor Weight Quantification. Topographic-related factors (such as elevation, slope, aspect, topographic curvature, topographic humidity index, and topographic roughness) are extracted from the screened core factors. The entropy weight method is used to calculate the weight of each topographic factor, quantifying the contribution of different topographic factors to topographic differences. Assume there are m topographic factors and n sample points. After normalization, the value of the i-th sample under the f-th factor is... x if The specific calculation process is as follows:

[0058] 3.1.1: Calculate the weight of the f-th terrain factor:

[0059] ;

[0060] Let f be the weight of the f-th terrain factor. Let be the value of the i-th sample point under the f-th terrain factor, and n be the total number of sample points;

[0061] 3.1.2: Calculate the information entropy of the f-th terrain factor:

[0062] ;

[0063] in, Let f be the information entropy of the f-th terrain factor. k =1 / ln( n ) is the entropy coefficient, if =0 then define x =0, to avoid logarithmic calculation errors.

[0064] 3.1.3: Calculate the weight of the f-th terrain factor:

[0065] ;

[0066] denoted as the weight of the f-th topographic factor, a larger weight value indicates a more significant impact of the topographic factor on topographic differences. m The total number of terrain factors is represented by this method, which objectively quantifies the weights of terrain factors and avoids bias caused by subjective weighting.

[0067] As can be seen from the above, in one embodiment, calculating the weights of various terrain factors based on the entropy weight method includes:

[0068] Calculate the weight of the f-th terrain factor:

[0069] ;

[0070] in, Let f be the weight of the f-th terrain factor. For the first i The sample point at the th th f The values ​​under each terrain factor, where n is the total number of sample points;

[0071] Calculate the information entropy of the f-th terrain factor based on its weight:

[0072] ;

[0073] in, Let f be the information entropy of the f-th terrain factor. k =1 / ln( n ) is the entropy coefficient, if =0, then define x =0, to avoid logarithmic calculation errors;

[0074] Based on the information entropy of the f-th terrain factor, calculate the entropy weight of the f-th terrain factor as the weight:

[0075] ;

[0076] denoted as the weight of the f-th topographic factor, a larger weight value indicates a more significant impact of the topographic factor on topographic differences. m This represents the total number of terrain factors.

[0077] 3.2: The implementation of this invention may also include preprocessing the slope aspect data in the terrain-related factors. Considering that slope aspect is a circular variable, and there is a numerical abrupt change at 0° / 360° but the terrain features are continuous, the slope aspect is converted into a two-dimensional orthogonal component for calculation, as shown in the following formula:

[0078] ;

[0079] ;

[0080] Where aspect is the original slope aspect value (unit: degrees). and The transformed orthogonal components all have values ​​in the range of [-1, 1], ensuring the continuity of slope aspect in the annular space and avoiding distortion in terrain distance calculations caused by abrupt changes in angle.

[0081] 3.3: Composite Terrain Distance Fusion Calculation. A composite terrain distance is constructed by fusing Euclidean spatial distance and terrain difference distance, simultaneously representing the spatial proximity and terrain similarity between samples. The calculation formula is as follows:

[0082] ;

[0083] Where α is the Euclidean spatial distance weight and β is the terrain difference distance weight, satisfying α+β=1, and the values ​​of both are determined through subsequent Bayesian optimization. The Euclidean distance is locally normalized to eliminate the impact of differences in spatial scale across different regions. The weighted difference distance based on terrain factor weights is calculated using the following formula:

[0084] ;

[0085] In the formula, The entropy weight of the f-th terrain factor, , , where are the standardized values ​​of sample i and sample j on the f-th terrain factor, respectively, and m is the total number of terrain factors involved in the calculation. This fusion distance achieves the dual constraint of "spatial proximity and terrain similarity," serving as the basis for subsequent local sample selection.

[0086] As described above, in one embodiment, based on the Euclidean spatial distance, a weighted terrain difference distance is introduced based on the weights of various terrain factors to fuse and construct a composite terrain distance, including: fusing and constructing the composite terrain distance according to the following formula:

[0087] ;

[0088] in, D For composite terrain distance, α is the Euclidean spatial distance weight, and β is the terrain difference distance weight, satisfying α+β=1. The values ​​of both are determined through subsequent Bayesian optimization. The Euclidean distance is locally normalized to eliminate the impact of differences in spatial scale across different regions; This is a weighted terrain difference distance based on terrain factor weights. The calculation formula is:

[0089] ;

[0090] In the formula, The entropy weight of the f-th terrain factor, , are the standardized values ​​of samples i and j on the f-th terrain factor, respectively, and m is the total number of terrain factors involved in the calculation.

[0091] In step 104 above, local reference samples are screened and weights are calculated. Based on the composite terrain distance, local reference samples that match the terrain features of the location to be evaluated are screened, and the sample weights are calculated based on the distance decay law to highlight the contribution of nearby similar samples to the prediction results. This specifically includes the following sub-steps:

[0092] 4.1 Intelligent Filtering of Local Reference Samples. A ball tree spatial search structure is used to perform fast neighborhood retrieval of the location to be evaluated, efficiently filtering candidate reference samples. A geographically weighted bandwidth parameter is set, and when the composite terrain distance between a candidate sample and the location to be evaluated is less than a preset composite terrain distance threshold, it is included in the local reference sample set. The value of the bandwidth parameter is determined through subsequent Bayesian optimization to ensure that the number of selected reference samples meets the training requirements of the local model while avoiding the weakening of spatial non-stationarity due to excessive samples. If the number of selected reference samples for a location to be evaluated is less than the preset threshold, the prediction result of the global random forest model is used as a fallback to avoid model prediction instability caused by data sparsity.

[0093] As described above, in one embodiment, based on the composite terrain distance, the mountain hazard sample data around the predicted location are sorted and filtered to determine a local reference sample set, including:

[0094] A ball-tree spatial search structure is used to perform fast neighborhood retrieval of the location to be predicted and to select candidate reference samples.

[0095] Set the bandwidth parameter for geographic weighting. When the distance between the candidate reference sample and the composite terrain of the location to be predicted is less than the preset composite terrain distance threshold, the candidate reference sample will be included in the local reference sample set.

[0096] 4.2: Reference Sample Weight Calculation. In one embodiment, the geographic weight of each local reference sample is calculated based on a kernel function, including: constructing a distance decay weight model based on the kernel function, and calculating the weight of each local reference sample, as shown in the following formula:

[0097] ;

[0098] in, w Here, d represents the weight of the local reference sample, d is the composite terrain distance between the location to be evaluated and the reference sample, and h is the optimal bandwidth parameter. This kernel function follows a distance decay rule: samples that are closer and have more similar terrain have greater weights, while samples that are farther apart have smaller weights. This effectively highlights the dominant role of locally similar samples in model prediction. The calculated sample weights are normalized to ensure that the sum of all reference sample weights is 1, avoiding weight stacking bias that could affect model training performance.

[0099] 4.3: Validation of sample screening effect. Figure 2 This is a schematic diagram illustrating sample selection based on Euclidean distance in an embodiment of the present invention. Figure 3This is a schematic diagram of sample selection based on composite terrain distance in an embodiment of the present invention. It can be seen that samples selected based on traditional Euclidean distance only satisfy spatial proximity, but some samples differ significantly in terrain features such as elevation and slope. In contrast, samples selected based on composite terrain distance maintain both spatial proximity and high consistency in terrain features, effectively avoiding the problem of mismatch between neighborhood samples and terrain environment in traditional methods, and providing a more reasonable sample basis for local model construction.

[0100] In step 105 above, the key parameters of the model are jointly optimized. A Bayesian optimization algorithm based on the tree-structured Parzen estimator (TPE) is used to jointly optimize the key parameters of the geographically weighted random forest model, achieving adaptive optimization of parameter combinations and avoiding the subjectivity and limitations of empirical parameter settings. Specifically, this includes the following sub-steps:

[0101] 5.1: Determine the parameters to be optimized and the search range. Based on the model structure and the characteristics of mountain hazard assessment, key parameters that significantly affect model performance are selected for joint optimization, mainly including:

[0102] Geographically weighted bandwidth h: The core parameter for controlling the size of local reference samples. The search range is reasonably set according to the sample density and spatial scale of the study area.

[0103] Composite terrain distance weight α: A parameter that adjusts the contribution ratio of Euclidean spatial distance and terrain difference distance, with a search range of [0,1] and β=1-α;

[0104] Random forest structure parameters include the number of decision trees (affecting the ensemble learning effect of the model), the maximum tree depth (controlling the complexity of the decision tree), and the minimum number of leaf node samples (avoiding model overfitting). The search range of each parameter should be reasonably set according to the model training requirements and computational efficiency.

[0105] 5.2: Bayesian Optimization Implementation Process. The objective function is to maximize the average AUC (area under the receiver operating characteristic curve) of cross-validation, balancing model classification accuracy and stability. Several sets of random parameter combinations are initialized. A geographically weighted random forest model is trained based on each set of parameters, and the validation set AUC is calculated. A probabilistic mapping relationship between parameters and model performance is constructed based on the TPE algorithm. The next optimal parameter combination is intelligently selected through iterative search for model training. The number of iterations is set according to the parameter space complexity and optimization accuracy requirements.

[0106] 5.3: Determining the Optimal Parameter Combination. After iteration, the parameter combination that maximizes the objective function value is selected as the optimal parameters for the model, including the optimal bandwidth h, the optimal composite distance weights α and β, and the optimal random forest structure parameters. Figure 4The line graph for α-AUC parameter optimization in this embodiment of the invention illustrates the impact of different α values ​​on model performance and the process of determining the optimal α. Figure 5 This diagram illustrates the bandwidth adaptive optimization results in this embodiment of the invention, showing the relationship between bandwidth changes and the model's AUC value, providing a basis for selecting the optimal bandwidth. Through joint parameter optimization, a dynamic balance between model complexity and classification accuracy is achieved, improving the model's adaptability in complex mountainous environments.

[0107] As described above, in one embodiment, under the constraint of the optimal parameter combination, a geographically weighted random forest model is constructed based on local reference samples and their weights to perform point-by-point predictions for the study area, obtaining the probability value of mountain disaster susceptibility, including:

[0108] For each location to be evaluated within the study area, a local geographic weighted random forest model is independently trained based on local reference samples selected from composite terrain distances and calculated weights. This enables the model to adaptively capture the patterns of disaster factors at different spatial locations, effectively characterizing the spatial nonstationarity of the mountain disaster formation process.

[0109] Parallel computing technology is employed, with the number of parallel threads matched to the number of CPU cores; a model caching mechanism is enabled to cache and store the trained local geographic weighted random forest model to avoid repeated training.

[0110] The raster data of the study area is divided into several prediction units according to a preset block size. For each prediction unit, all the core environmental factor attribute values ​​corresponding to it are extracted as model inputs. The local geographic weighted random forest model corresponding to that location is called to make predictions and output the probability value of mountain disaster susceptibility for that prediction unit.

[0111] By stitching together the probability values ​​of mountain disaster susceptibility from all prediction units, a spatial distribution map of the probability of mountain disaster susceptibility across the entire study area is generated.

[0112] In step 106 above, a geographically weighted random forest model is constructed and point-by-point predictions are made. Under the constraint of optimal parameter combination, a geographically weighted random forest model that integrates spatial nonstationarity and nonlinear learning capabilities is constructed to predict the point-by-point susceptibility probability of the study area. This specifically includes the following sub-steps:

[0113] 6.1: Model Framework Construction. Based on the random forest framework, a localized modeling strategy is constructed using the geographically weighted approach to achieve accurate predictions with a "one model per location" approach. For each location to be evaluated within the study area, a local random forest model is independently trained based on local reference samples selected from composite terrain distances and calculated sample weights. This enables the model to adaptively capture the patterns of hazard factors at different spatial locations, effectively characterizing the spatial non-stationarity of the mountain hazard formation process.

[0114] 6.2: Model Training Configuration Optimization. To improve model training and prediction efficiency, parallel computing technology is adopted, with the number of parallel threads and CPU cores matched. A model caching mechanism is enabled to cache and store the trained local models, avoiding repeated training, which is especially suitable for point-by-point prediction over large research areas. Simultaneously, a global random forest model is set as a fallback mechanism, automatically invoked when the number of local reference samples is insufficient, ensuring the continuity and completeness of prediction results.

[0115] 6.3: Point-by-point prediction execution in the study area. The raster data of the study area is divided into several prediction unit blocks according to a preset block size. A block prediction strategy is adopted to improve computational efficiency. For each prediction unit, all core environmental factor attribute values ​​corresponding to it are extracted as model inputs. The local geographic weighted random forest model corresponding to that location is called to make predictions, and the probability of mountain disaster occurrence (range 0-1), i.e., the susceptibility probability value, is output.

[0116] 6.4: Integrated Output of Prediction Results. The susceptibility probability values ​​of all prediction units are spliced ​​and integrated to generate a spatial distribution map of the susceptibility probability of mountain disasters across the entire study area. Figure 6 This is a spatial distribution map of the predicted probability of mountain disaster susceptibility based on a composite terrain distance geographic weighted random forest model in this embodiment of the invention. It clearly shows the spatial gradient change of susceptibility probability in the study area, providing basic data for subsequent susceptibility classification.

[0117] This embodiment of the invention may further include step 107, generating susceptibility probability grading and zoning results. A scientifically sound grading method is used to classify the susceptibility probability values ​​into levels, generating spatial zoning results for mountain disaster susceptibility with clear risk gradients. Specifically, this includes the following sub-steps:

[0118] 7.1: Selection of Grading Method. The equidistant grading method is adopted to classify the susceptibility probability values ​​into levels. This method is based on the susceptibility probability value range [0,1], dividing the level intervals at equal intervals to ensure the objectivity and repeatability of the grading standard. Compared with other grading methods, the equidistant grading method has clear logic, is easy to calculate, and can intuitively present the gradient change of susceptibility probability, facilitating the comparison of disaster risks between different regions, while also meeting the requirement for simplicity in level classification in disaster risk management.

[0119] 7.2: Classification Standards. Based on the actual application scenarios and probability value distribution characteristics of mountain disaster risk management, the susceptibility probability values ​​are divided into five levels according to the principle of equal intervals: extremely low susceptibility area (probability ∈ [0, 0.2)), low susceptibility area (probability ∈ [0.2, 0.4)), medium susceptibility area (probability ∈ [0.4, 0.6)), high susceptibility area (probability ∈ [0.6, 0.8)), and extremely high susceptibility area (probability ∈ [0.8, 1.0)).

[0120] 7.3: Zoning Results Output and Visualization. Based on the aforementioned equidistant grading standard, the spatial distribution map of the susceptibility probability of the entire study area is converted into a zoning map of mountain hazard susceptibility levels. A differentiated color system is used to identify areas of each level (e.g., dark green for extremely low susceptibility areas, light green for low susceptibility areas, yellow for medium susceptibility areas, light red for high susceptibility areas, and dark red for extremely high susceptibility areas), and legends clearly define the spatial scope and risk meaning of each level. Simultaneously, detailed zoning statistics are output, including the area, area percentage, spatial distribution range, main geomorphic units, and environmental characteristics of each susceptibility level, forming a complete zoning result.

[0121] As can be seen from the above, in one embodiment, the above-mentioned method for predicting the susceptibility of mountain disasters taking into account the distance of composite terrain may further include: classifying the probability values ​​of mountain disaster susceptibility according to a preset threshold, and generating spatial zoning results of mountain disaster susceptibility.

[0122] In one embodiment, the probability values ​​of mountain disaster susceptibility are classified according to a preset threshold to generate spatial zoning results for mountain disaster susceptibility, including:

[0123] Based on the actual application scenarios and probability value distribution characteristics of mountain disaster risk management, the susceptibility probability value is divided into five levels according to the principle of equal intervals using the equal interval grading method.

[0124] Based on the five levels, the spatial distribution map of the probability of mountain disasters in the entire study area is converted into a zoning map of the susceptibility levels of mountain disasters.

[0125] This embodiment of the invention may further include step 108: model accuracy verification and result analysis. Through multi-index, multi-dimensional verification analysis, the reliability of model performance and susceptibility partitioning results is comprehensively evaluated, highlighting the technical advantages of the method of this invention. Specifically, this includes the following sub-steps:

[0126] 8.1: Multi-indicator comparison of model accuracy. A hierarchical cross-validation strategy was employed to conduct parallel comparative validation of the composite terrain distance geographic weighted random forest (CTD-GWRF) model constructed in this invention with traditional random forest (RF) and geographic weighted random forest (GWRF) models. A comprehensive evaluation system was constructed using six core indicators: AUC (Area Under the Receiver Operating Characteristic), Accuracy, Precision, Recall, F1 score, and Kappa coefficient, comprehensively covering the model's classification ability, disaster identification completeness, and result consistency. Figure 7 This is a graph showing the comparison of prediction accuracy of different susceptibility evaluation models in the embodiments of the present invention.

[0127] 8.2: Verification of Disaster Point Overlay. Historical disaster point data for the study area were spatially overlaid with the susceptibility zoning results map. The number, distribution density, and proportion of disaster points within each susceptibility zone were statistically analyzed to verify the degree of agreement between the zoning results and the actual disaster distribution. Figure 8 This is a schematic diagram illustrating the overlay verification of disaster point and susceptibility zoning results in an embodiment of the present invention.

[0128] 8.3: Frequency Ratio Statistical Analysis. The frequency ratio (FR) quantitative index is used to further verify the scientific validity and rationality of the susceptibility zoning. The frequency ratio is defined as the ratio of the proportion of disaster points in a certain susceptibility level area to the proportion of the area area. The larger the FR value, the higher the susceptibility of the area to disasters, and vice versa. Figure 9 This is a frequency ratio statistical analysis chart based on the susceptibility partitioning results in an embodiment of the present invention.

[0129] As can be seen from the above, in one embodiment, the above-mentioned method for predicting the susceptibility of mountain disasters taking into account the distance of complex terrain may further include: comprehensively evaluating the reliability of the geographically weighted random forest model and the prediction results of mountain disaster susceptibility through multi-indicator and multi-dimensional verification analysis.

[0130] Through the complete implementation process described above, the beneficial technical effects of the mountain disaster susceptibility prediction scheme considering complex terrain distances provided by this invention are as follows: Relying on the information entropy method to objectively quantify the weights of terrain factors, and combining this with the complex terrain distance to construct a geographically weighted model, it can effectively solve the problem of mismatch between neighborhood samples and terrain environment in traditional methods; simultaneously, through joint optimization of Bayesian parameters, it further improves the stability and adaptability of the model in complex mountainous scenarios. This invention uses an equal-interval grading method to complete the classification of susceptibility probabilities, combined with multi-dimensional accuracy verification, to fully ensure the reliability and practicality of the zoning results. The final zoning results retain the gradient characteristics of susceptibility probabilities and have strong practical operability, providing technical support for mountain disaster risk assessment and disaster prevention and mitigation planning in complex mountainous environments.

[0131] This invention also provides a mountain disaster susceptibility prediction device that takes into account the distance of complex terrain, as described in the following embodiments. Since the principle by which this device solves the problem is similar to the mountain disaster susceptibility prediction method that takes into account the distance of complex terrain, the implementation of this device can refer to the implementation of the mountain disaster susceptibility prediction method that takes into account the distance of complex terrain, and will not be repeated here.

[0132] Figure 10 This is a schematic diagram of the mountain disaster susceptibility prediction device that takes into account the distance of complex terrain in an embodiment of the present invention, as shown below. Figure 10 As shown, the device includes:

[0133] Acquisition Unit 01 is used to acquire mountain disaster sample data and multi-source environmental factor data within the study area;

[0134] The composite terrain distance construction unit 03 is used to screen terrain factors from the multi-source environmental factors, calculate the weight of each terrain factor based on the entropy weight method to quantify the contribution of different terrain factors to terrain differences; based on the Euclidean spatial distance, a weighted terrain difference distance is introduced based on the weight of each terrain factor, and a composite terrain distance is constructed to simultaneously characterize the spatial proximity and terrain similarity between samples.

[0135] The local reference sample screening and weight calculation unit 04 is used to sort and screen the mountain disaster sample data around the predicted location based on the composite terrain distance, determine the local reference sample set, and calculate the geographical weighting weight corresponding to each local reference sample based on the kernel function.

[0136] Joint optimization unit 05 is used to jointly optimize the bandwidth parameter of the geographic weighting, the weight parameter of the composite terrain distance, and the structural parameters of the random forest model using the Bayesian optimization method to obtain the optimal parameter combination;

[0137] Prediction unit 06 is used to construct a geographically weighted random forest model based on local reference samples and their weights under the constraints of the optimal parameter combination, and to make point-by-point predictions for the study area to obtain the probability value of mountain disaster susceptibility.

[0138] In one embodiment, the joint optimization unit is specifically used for:

[0139] The bandwidth parameter of the geographic weighted weight, the weight parameter of the composite terrain distance, and the structure parameter of the random forest model are used as the parameters to be optimized, and the search range of each parameter to be optimized is determined.

[0140] With the objective function of maximizing the average AUC value of cross-validation, several sets of random parameter combinations to be optimized are initialized. A geographically weighted random forest model is trained based on each set of parameters, and the AUC value of the validation set is calculated. The probability mapping relationship between parameters and model performance is constructed based on the TPE algorithm. The next optimal parameter combination is intelligently selected for model training through iterative search. The number of iterations is set according to the parameter space complexity and optimization accuracy requirements.

[0141] After the iteration is completed, the parameter combination that maximizes the objective function value is selected as the optimal parameter combination.

[0142] In one embodiment, the prediction unit is specifically used for:

[0143] For each location to be evaluated within the study area, a local geographic weighted random forest model is independently trained based on local reference samples selected from composite terrain distances and calculated weights. This enables the model to adaptively capture the patterns of disaster factors at different spatial locations, effectively characterizing the spatial nonstationarity of the mountain disaster formation process.

[0144] Parallel computing technology is employed, with the number of parallel threads matched to the number of CPU cores; a model caching mechanism is enabled to cache and store the trained local geographic weighted random forest model to avoid repeated training.

[0145] The raster data of the study area is divided into several prediction units according to a preset block size. For each prediction unit, all the core environmental factor attribute values ​​corresponding to it are extracted as model inputs. The local geographic weighted random forest model corresponding to that location is called to make predictions and output the probability value of mountain disaster susceptibility for that prediction unit.

[0146] By stitching together the probability values ​​of mountain disaster susceptibility from all prediction units, a spatial distribution map of the probability of mountain disaster susceptibility across the entire study area is generated.

[0147] In one embodiment, the local reference sample screening and weight calculation unit is specifically used for:

[0148] A ball-tree spatial search structure is used to perform fast neighborhood retrieval of the location to be predicted and to select candidate reference samples.

[0149] Set the bandwidth parameter for geographic weighting. When the distance between the candidate reference sample and the composite terrain of the location to be predicted is less than the preset composite terrain distance, the candidate reference sample is included in the local reference sample set.

[0150] In one embodiment, the geographic weighting weight corresponding to each local reference sample is calculated based on a kernel function, including:

[0151] A distance decay weighting model is constructed based on the kernel function to calculate the weight of each local reference sample, as shown in the following formula:

[0152] ;

[0153] in, w d represents the weight of the local reference sample, d represents the composite terrain distance between the location to be predicted and the reference sample, h represents the optimal bandwidth parameter, and the kernel function indicates that the closer the sample is and the more similar the terrain is, the greater the weight is, and the farther the sample is, the smaller the weight is.

[0154] In one embodiment, calculating the weights of each topographic factor based on the entropy weight method includes:

[0155] Calculate the weight of the f-th terrain factor:

[0156] ;

[0157] in, Let f be the weight of the f-th terrain factor. For the first i The sample point at the th th f The values ​​under each terrain factor, where n is the total number of sample points;

[0158] Calculate the information entropy of the f-th terrain factor based on its weight:

[0159] ;

[0160] in, Let f be the information entropy of the f-th terrain factor. k =1 / ln( n ) is the entropy coefficient, if =0, then define x =0, to avoid logarithmic calculation errors;

[0161] Based on the information entropy of the f-th terrain factor, calculate the entropy weight of the f-th terrain factor as the weight:

[0162] ;

[0163] The weight of the f-th terrain factor is... m The total number of terrain factors involved in the calculation; the larger the weight value, the more significant the impact of the terrain factor on terrain differences.

[0164] In one embodiment, based on the Euclidean spatial distance, a weighted terrain difference distance is introduced according to the weights of various terrain factors, and a composite terrain distance is constructed by fusing the distances, including: fusing the composite terrain distance according to the following formula:

[0165] ;

[0166] in, D For composite terrain distance, α is the Euclidean spatial distance weight, and β is the terrain difference distance weight, satisfying α+β=1. The values ​​of both are determined through subsequent Bayesian optimization. The distance is the Euclidean distance after local normalization to eliminate the impact of differences in spatial scale across different regions; This is a weighted terrain difference distance based on terrain factor weights. The calculation formula is:

[0167]

[0168] In the formula, The entropy weight of the f-th terrain factor, , are the standardized values ​​of samples i and j on the f-th terrain factor, respectively, and m is the total number of terrain factors involved in the calculation.

[0169] In one embodiment, the above-mentioned mountain disaster susceptibility prediction device that takes into account the distance of composite terrain may further include: a generation unit, used to classify the probability value of mountain disaster susceptibility according to a preset threshold, and generate a spatial zoning result of mountain disaster susceptibility.

[0170] In one embodiment, the generating unit is specifically used for:

[0171] Based on the actual application scenarios and probability value distribution characteristics of mountain disaster risk management, the susceptibility probability value is divided into five levels according to the principle of equal intervals using the equal interval grading method.

[0172] Based on the five levels, the spatial distribution map of the probability of mountain disasters in the entire study area is converted into a zoning map of the susceptibility levels of mountain disasters.

[0173] In one embodiment, the above-mentioned method for predicting the susceptibility of mountain disasters taking into account the distance of complex terrain may further include: a processing unit for normalizing and performing multicollinearity test processing on the multi-source environmental factor data.

[0174] In one embodiment, the above-mentioned method for predicting mountain hazard susceptibility considering complex terrain distance may further include: a validation unit, used to comprehensively evaluate the reliability of the geographically weighted random forest model and the mountain hazard susceptibility prediction results through multi-indicator and multi-dimensional validation analysis.

[0175] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned method for predicting the susceptibility of mountain disasters taking into account the distance of complex terrain.

[0176] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for predicting the susceptibility of mountain disasters taking into account the distance of complex terrain.

[0177] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-mentioned method for predicting the susceptibility of mountain disasters taking into account the distance of complex terrain.

[0178] In this embodiment of the invention, a mountain disaster susceptibility prediction scheme considering composite terrain distances involves: acquiring mountain disaster sample data and multi-source environmental factor data within the study area; selecting terrain factors from the multi-source environmental factors, calculating the weights of each terrain factor based on the entropy weight method to quantify the contribution of different terrain factors to terrain differences; introducing weighted terrain difference distances based on the weights of each terrain factor, and constructing composite terrain distances based on Euclidean spatial distances to simultaneously characterize the spatial proximity and terrain similarity between samples; and sorting and selecting mountain disaster sample data around the location to be predicted based on the composite terrain distances to determine a local reference sample set, and calculating the local reference distance for each sample set based on a kernel function. The samples are assigned geographical weights. A Bayesian optimization method is used to jointly optimize the bandwidth parameter of the geographical weights, the weight parameter of the composite terrain distance, and the structural parameters of the random forest model to obtain the optimal parameter combination. Under the constraint of this optimal parameter combination, a geographically weighted random forest model is constructed based on local reference samples and their weights to perform point-by-point predictions for the study area. The resulting probability value of mountain disaster susceptibility is used as the prediction result. This method can achieve the construction of a composite terrain distance that integrates spatial distance and terrain differences, and adaptive optimization of key parameters within the geographically weighted random forest framework. This accurately characterizes the spatial heterogeneity of disaster susceptibility in complex mountainous areas, improving the accuracy and stability of susceptibility prediction results.

[0179] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0180] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0181] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0182] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0183] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for predicting the susceptibility of mountain disasters considering the distance of complex terrain, characterized in that, include: Acquire mountain hazard sample data and multi-source environmental factor data within the study area; Topographic factors are screened from the multi-source environmental factors, and the weights of each topographic factor are calculated based on the entropy weight method to quantify the contribution of different topographic factors to topographic differences. Based on Euclidean spatial distance, a weighted terrain difference distance is introduced according to the weights of various terrain factors, and a composite terrain distance is constructed by fusing them together. This composite terrain distance is constructed according to the following formula: ;in, D For distances in complex terrain, α The weights are Euclidean distances. β As the weight of the terrain difference distance, satisfying α + β =1, and the values ​​of both are determined through subsequent Bayesian optimization; The Euclidean distance is locally normalized to eliminate the impact of differences in spatial scale across different regions; This is a weighted terrain difference distance based on terrain factor weights. The calculation formula is: In the formula, For the first f Entropy weight of each terrain factor, , Samples i With sample j In the f Values ​​under each terrain factor m The total number of terrain factors is used to simultaneously characterize the spatial proximity and terrain similarity between samples; Based on the composite terrain distance, the mountain hazard sample data around the predicted location are sorted and filtered to determine the local reference sample set. Then, the geographic weighting of each local reference sample is calculated based on a kernel function. This includes: constructing a distance decay weighting model based on the kernel function and calculating the weight of each local reference sample, as shown in the following formula: ;in, w For the weights of the local reference samples, d The composite terrain distance between the location to be predicted and the reference sample. h As the optimal bandwidth parameter, the kernel function indicates that samples that are closer together and have more similar terrain have a higher weight, while samples that are farther apart have a lower weight. The Bayesian optimization method is used to jointly optimize the bandwidth parameter of the geographic weighting, the weight parameter of the composite terrain distance, and the structural parameters of the random forest model to obtain the optimal parameter combination. Under the constraints of the optimal parameter combination, a geographically weighted random forest model is constructed based on local reference samples and their weights to perform point-by-point predictions for the study area, obtaining the probability value of mountain disaster susceptibility as the prediction result. This includes: for each location to be evaluated within the study area, independently training a local geographically weighted random forest model based on local reference samples selected from composite terrain distances and calculated weights. This enables the model to adaptively capture the influence patterns of disaster factors at different spatial locations, effectively characterizing the spatial non-stationarity of the mountain disaster incubation process; and employing parallel computing technology, setting the number of parallel threads and CPU... Core number adaptation; enable model caching mechanism to cache and store the trained local geographic weighted random forest model to avoid repeated training; divide the raster data of the study area into several prediction units according to the preset block size, extract all the corresponding core environmental factor attribute values ​​of each prediction unit as model input, call the local geographic weighted random forest model corresponding to the location to be evaluated to make predictions, and output the probability value of mountain disaster susceptibility of the prediction unit; stitch together and integrate the probability values ​​of mountain disaster susceptibility of all prediction units to generate a spatial distribution map of the probability of mountain disaster susceptibility of the entire study area.

2. The method as described in claim 1, characterized in that, A Bayesian optimization method is used to jointly optimize the bandwidth parameter of the geographic weighting, the weight parameter of the composite terrain distance, and the structural parameters of the random forest model to obtain the optimal parameter combination, including: The bandwidth parameter of the geographic weighted weight, the weight parameter of the composite terrain distance, and the structure parameter of the random forest model are used as the parameters to be optimized, and the search range of each parameter to be optimized is determined. With the objective function of maximizing the average AUC value of cross-validation, several sets of random parameter combinations to be optimized are initialized. A geographically weighted random forest model is trained based on each set of parameters, and the AUC value of the validation set is calculated. The probabilistic mapping relationship between parameters and model performance is constructed based on the TPE algorithm. The next optimal parameter combination is intelligently selected for model training through iterative search. The number of iterations is set according to the parameter space complexity and optimization accuracy requirements. After the iteration is completed, the parameter combination that maximizes the objective function value is selected as the optimal parameter combination.

3. The method as described in claim 1, characterized in that, Based on the composite terrain distance, the mountain hazard sample data around the predicted location are sorted and filtered to determine the local reference sample set, including: A ball-tree spatial search structure is used to perform fast neighborhood retrieval of the location to be predicted and to select candidate reference samples. Set the bandwidth parameter for geographic weighting. When the distance between the candidate reference sample and the composite terrain of the location to be predicted is less than the preset composite terrain distance threshold, the candidate reference sample is included in the local reference sample set.

4. The method as described in claim 1, characterized in that, The weights of various topographic factors are calculated based on the entropy weight method, including: Calculate the first f The proportion of each topographic factor: ; in, Let f be the weight of the f-th terrain factor. For the first i The sample point at the th th f Values ​​under each terrain factor n The total number of sample points; According to the f The weight of each topographic factor is calculated. f Information entropy of each terrain factor: ; in, For the first f Information entropy of a terrain factor k =1 / ln( n ) is the entropy coefficient, if =0, then define x =0, to avoid logarithmic calculation errors; According to the f The information entropy of the i-th terrain factor is calculated. f The entropy weights of the terrain factors are used as the weights: ; For the first f The weight of each topographic factor is assigned; a larger weight value indicates a more significant impact of that topographic factor on topographic differences. m This represents the total number of terrain factors.

5. The method as described in claim 1, characterized in that, Also includes: Based on preset thresholds, the probability values ​​of mountain disaster susceptibility are classified into levels, generating spatial zoning results for mountain disaster susceptibility.

6. The method as described in claim 5, characterized in that, Based on preset thresholds, the probability values ​​of mountain disaster susceptibility are classified into categories, generating spatial zoning results for mountain disaster susceptibility, including: Based on the actual application scenarios and probability value distribution characteristics of mountain disaster risk management, the susceptibility probability value is divided into five levels according to the principle of equal intervals using the equal interval grading method. Based on the five levels, the spatial distribution map of the probability of mountain disasters in the entire study area is converted into a zoning map of the susceptibility levels of mountain disasters.

7. The method as described in claim 1, characterized in that, Also includes: The multi-source environmental factor data were normalized and subjected to multicollinearity testing.

8. The method as described in claim 1, characterized in that, Also includes: Through multi-indicator and multi-dimensional validation analysis, the reliability of the geographically weighted random forest model and the prediction results of mountain disaster susceptibility are comprehensively evaluated.

9. A mountain disaster susceptibility prediction device that takes into account the distance of complex terrain, characterized in that, include: The acquisition unit is used to acquire mountain disaster sample data and multi-source environmental factor data within the study area; A composite terrain distance construction unit is used to screen terrain factors from the multi-source environmental factors and calculate the weight of each terrain factor based on the entropy weight method to quantify the contribution of different terrain factors to terrain differences. Based on Euclidean spatial distance, a weighted terrain difference distance is introduced according to the weights of various terrain factors, and a composite terrain distance is constructed by fusing them together. This composite terrain distance is constructed according to the following formula: ;in, D For distances in complex terrain, α The weights are Euclidean distances. β As the weight of the terrain difference distance, satisfying α + β =1, and the values ​​of both are determined through subsequent Bayesian optimization; The Euclidean distance is locally normalized to eliminate the impact of differences in spatial scale across different regions; This is a weighted terrain difference distance based on terrain factor weights. The calculation formula is: In the formula, For the first f Entropy weight of each terrain factor, , Samples i With sample j In the f Values ​​under each terrain factor m The total number of terrain factors is used to simultaneously characterize the spatial proximity and terrain similarity between samples; The local reference sample screening and weight calculation unit is used to sort and screen the mountain disaster sample data around the predicted location based on the composite terrain distance, determine the local reference sample set, and calculate the geographic weight of each local reference sample based on the kernel function. This includes: constructing a distance decay weight model based on the kernel function and calculating the weight of each local reference sample, as shown in the following formula: ;in, w For the weights of the local reference samples, d The composite terrain distance between the location to be predicted and the reference sample. h As the optimal bandwidth parameter, the kernel function indicates that samples that are closer together and have more similar terrain have a higher weight, while samples that are farther apart have a lower weight. The joint optimization unit is used to jointly optimize the bandwidth parameter of the geographic weighting, the weight parameter of the composite terrain distance, and the structural parameters of the random forest model using the Bayesian optimization method to obtain the optimal parameter combination. The prediction unit, under the constraints of the optimal parameter combination, constructs a geographically weighted random forest model based on local reference samples and their weights, performs point-by-point predictions for the study area, and obtains the probability value of mountain disaster susceptibility as the prediction result. This includes: independently training a local geographically weighted random forest model for each location to be evaluated within the study area, based on local reference samples selected from composite terrain distances and calculated weights. This enables the model to adaptively capture the influence patterns of disaster factors at different spatial locations, effectively characterizing the spatial non-stationarity of the mountain disaster incubation process; and employing parallel computing technology, setting the number of parallel threads. The system is adapted to the number of CPU cores; a model caching mechanism is enabled to cache and store the trained local geographic weighted random forest model to avoid repeated training; the raster data of the study area is divided into several prediction units according to a preset block size; for each prediction unit, all the corresponding core environmental factor attribute values ​​are extracted as model inputs, and the local geographic weighted random forest model corresponding to the location to be evaluated is called to make predictions, outputting the probability value of mountain disaster susceptibility for the prediction unit; the probability values ​​of mountain disaster susceptibility for all prediction units are stitched together and integrated to generate a spatial distribution map of the probability of mountain disaster susceptibility for the entire study area.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 8.

12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Geological disaster prediction method, device and equipment

    CN111144651A

  • Geographic weighting and random forest coupled surface temperature downscaling method

    CN117035066A