Landslide hazard risk assessment method, equipment, medium and product
By constructing a disaster-pregnancy factor system and artificial neural network, combined with prediction density and weighted cross-entropy loss function, the subjectivity and randomness of negative sample generation methods were solved, and high accuracy and rationality of landslide disaster risk assessment were achieved.
Patent Information
- Application Number
- CN202510919023.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-04
AI Technical Summary
In the existing technology, the geological hazard risk assessment method has the problems of subjectivity and randomness in the negative sample generation method, which leads to deviations in the model prediction results, and the single classification solution method cannot reasonably assess the probability distribution of geological hazard risks in geographic space.
A disaster-predisposing factor system was constructed, and an artificial neural network was used to build a landslide disaster risk assessment model. Negative samples were obtained through uniform sampling, and the prediction density and weighted cross entropy loss function were introduced for training. The weights of positive and negative samples were adjusted to improve the objectivity and accuracy of the model.
It improves the objectivity and accuracy of landslide disaster risk assessment, can achieve higher prediction accuracy at a lower prediction density, and identify risk areas more accurately.
Smart Images

Figure CN120410236B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of geological disaster risk assessment, and in particular to a landslide disaster risk assessment method, equipment, medium and product. Background Art
[0002] Currently, a large number of mature methods for geohazard risk assessment have been developed, primarily categorized as subjective judgment methods, binary regression methods, and single-class solution methods. The most typical subjective judgment method is the analytic hierarchy process (AHP), which uses expert scoring to determine the weights of risk factors that contribute to geohazards. However, this method is highly dependent on experts and is highly subjective. The key idea behind binary regression is to identify historical disaster sites within a region as positive samples, setting the probability of disaster occurrence to 1. Simultaneously, a certain number of sites with very low disaster probability are identified as negative samples, setting the probability of disaster occurrence to 0. By applying bidirectional constraints on these positive and negative samples, a relatively optimal solution for the risk assessment model is obtained. This process is difficult to objectively determine which locations in a region have a low probability of disaster occurrence, so a random generation method is often used to generate a certain number of negative samples. However, these randomly generated static negative samples may not truly represent locations with low regional geohazard risk, leading to significant bias in the model's predictions. To address the challenges of binary regression, researchers have proposed single-class solution methods to address the difficulty of obtaining negative samples. However, since the existing single-classification solution method can only consider the multidimensional spatial distribution of disaster-prone factors from a statistical numerical perspective, it is unable to reasonably evaluate the probability distribution of geological hazard risks in geographic space, and its evaluation results are difficult to further apply. Summary of the Invention
[0003] The purpose of this application is to provide a landslide hazard risk assessment method, equipment, medium and product to solve the subjectivity and randomness problems of traditional negative sample generation methods and the problem of insufficient spatial distribution assessment of traditional single-class solution methods, thereby improving the objectivity, rationality and accuracy of landslide hazard risk assessment.
[0004] To achieve the above objectives, this application provides the following solutions.
[0005] In a first aspect, the present application provides a landslide hazard risk assessment method, comprising:
[0006] Constructing a disaster-pregnancy factor system for landslide hazard risk assessment, wherein the disaster-pregnancy factor system includes multiple disaster-pregnancy factors; the multiple disaster-pregnancy factors include: elevation data, landform type, geological structure factors, engineering geological factors, earthquake parameters, average annual precipitation, water system, surface cover data, land use type and road distribution data;
[0007] The study area is uniformly sampled according to the preset sampling interval to obtain a uniform point set in the study area;
[0008] For the disaster points in the study area, the amount of information about the disaster points is collected based on the disaster-pregnancy factor system to form a positive sample set;
[0009] For uniform points in the study area, the information of uniform points is collected based on the disaster-prone factor system to form a negative sample set;
[0010] Constructing a landslide disaster risk assessment model based on artificial neural network;
[0011] The landslide disaster risk assessment model is trained using positive sample sets and negative sample sets, and the probability of landslide disaster occurrence for each sample is output;
[0012] The prediction accuracy and prediction density of the landslide hazard risk assessment model are calculated based on the probability of landslide hazard occurrence of each sample;
[0013] The weighted cross entropy loss function is calculated based on the predicted density and the error is back-propagated until the training is completed to obtain a trained landslide hazard risk assessment model;
[0014] The trained landslide hazard risk assessment model is used to conduct landslide hazard risk assessment on the area to be assessed.
[0015] Optionally, the construction of a disaster-predisposing factor system involved in landslide disaster risk assessment specifically includes:
[0016] Constructing a disaster-predisposing factor system for landslide risk assessment ;in, For the A pregnancy disaster factor, is the total number of pregnancy disaster factors; ;
[0017] For each fertility factor , using the preset classification criteria to divide it into kind.
[0018] Optionally, for the disaster points in the study area, the amount of information of the disaster points is collected based on the disaster-predisposing factor system to form a positive sample set, specifically including:
[0019] Statistics of each disaster-prone factor in the study area Belong to The number of disaster points in the sample class and disaster area ; ;
[0020] Using the formula Calculate pregnancy disaster factor No. Class information ;in is the total number of disaster points in the study area; is the total area of all disaster points in the study area;
[0021] All disaster points and corresponding information Constitute the positive sample set.
[0022] Optionally, the constructing of a landslide disaster risk assessment model based on an artificial neural network specifically includes:
[0023] With a size The two-dimensional tensor of is used as the input layer of the landslide hazard risk assessment model. Each neuron in the input layer receives the information of a sample. ;in, Indicates the number of input samples;
[0024] The two-dimensional tensor of the input layer passes through the first fully connected layer fc1 and the relu activation layer, outputting a size of The hidden layer of
[0025] Then, the neurons in the hidden layer pass through the second fully connected layer fc2, with an output size of The tensor of The tensor of each sample includes the predicted probability of the sample belonging to the positive sample and the negative sample and ;in, ;
[0026] Finally, through the softmax function, the formula Calculate the probability of landslide disaster for each sample .
[0027] Optionally, the calculation of the prediction accuracy and prediction density of the landslide disaster risk assessment model based on the landslide disaster occurrence probability of each sample specifically includes:
[0028] Will The positive samples are regarded as the disaster samples predicted correctly, and the number of disaster samples predicted correctly is counted. ;
[0029] Using the formula Calculating the prediction accuracy of landslide hazard risk assessment models ;in is the total number of positive samples participating in the evaluation;
[0030] statistics The number of negative samples , and use the formula Calculating the predicted density of landslide hazard risk assessment models ;in Indicates the total number of negative samples involved in the evaluation.
[0031] Optionally, the weighted cross entropy loss function is: ;in, and is a sign function, when the sample When they belong to negative samples and positive samples respectively, the value is 1, otherwise the value is 0; and The samples The predicted probability of belonging to negative samples and positive samples; is the calculated loss value.
[0032] Optionally, performing landslide hazard risk assessment on the area to be assessed using the trained landslide hazard risk assessment model specifically includes:
[0033] Use the trained landslide hazard risk assessment model to output the probability of landslide disasters in the area to be assessed ;
[0034] according to The value of will divide the area to be assessed into very low risk area, low risk area, medium risk area, high risk area or very high risk area; among them, The value of [0-0.2) is the extremely low risk area, [0.2-0.4) is the low risk area, [0.4-0.6) is the medium risk area, [0.6-0.8) is the high risk area, and [0.8-1] is the extremely high risk area.
[0035] In a second aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the landslide hazard risk assessment method.
[0036] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the landslide hazard risk assessment method when executed by a processor.
[0037] In a fourth aspect, the present application provides a computer program product, comprising a computer program, which implements the landslide hazard risk assessment method when executed by a processor.
[0038] According to the specific embodiments provided in this application, this application discloses the following technical effects.
[0039] The present application provides a landslide hazard risk assessment method, equipment, medium, and product. Ten disaster-predisposing factors, including elevation data, landform types, geological structural factors, engineering geological factors, earthquake parameters, average annual precipitation, water systems, surface cover data, land use types, and road distribution data, are selected as the most relevant factors for landslide hazards to construct a disaster-predisposing factor system for landslide hazard risk assessment. On this basis, a landslide hazard risk assessment model is constructed based on an artificial neural network, and prediction density is introduced as a negative evaluation factor for the model. During model training, uniform points are used as negative samples. By calculating the prediction density and weighted cross-entropy loss function in real time and performing error backpropagation, the positive and negative sample feedback weights of the artificial neural network can be set based on the prediction density. The present application method uses uniform points as negative samples, overcoming the error problem of random negative samples and achieving higher prediction accuracy at a lower prediction density. Compared to traditional single-class solution methods, the positive and negative sample sets constructed by the present application incorporate the spatial distribution of disasters and have higher landing area accuracy. Therefore, the present application method improves the objectivity, rationality, and accuracy of landslide hazard risk assessment as a whole. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0041] Figure 1 A flow chart of a landslide hazard risk assessment method for this application;
[0042] Figure 2 Schematic diagram of uniform point sampling;
[0043] Figure 3 This is a schematic diagram of the artificial neural network structure of the landslide hazard risk assessment model;
[0044] Figure 4 Schematic diagram of the model training process based on prediction density;
[0045] Figure 5 This is a schematic diagram of the distribution of landslide disaster points in City A in the embodiment of this application;
[0046] Figure 6 This is a schematic diagram of the random negative sample distribution of City A in the embodiment of this application;
[0047] Figure 7 For different models / algorithms Curves and precision points Schematic diagram of the comparison;
[0048] Figure 8 This is a schematic diagram comparing the spatial location accuracy of the proposed method and the logistic regression model;
[0049] Figure 9 This application method is compared with the single-class SVM algorithm Schematic diagram of curve comparison;
[0050] Figure 10 This is a schematic diagram of the risk area division results of City A obtained using this application method. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0052] Constructing a geological hazard risk assessment model and discovering disaster-prone areas are of great significance for landslide disaster early warning. Traditional methods use randomly generated negative samples for model training. Due to the randomness and subjectivity of negative sample selection, the accuracy of the trained model is limited. Traditional single-class solution methods are unable to reasonably evaluate the probability distribution of geological hazard risks in geographic space. In this regard, this application proposes a landslide hazard risk assessment method, equipment, medium and product, which aims to solve the subjectivity and randomness of traditional negative sample generation methods and the problem of insufficient spatial distribution assessment of traditional single-class solution methods, so as to improve the objectivity, rationality and accuracy of landslide hazard risk assessment.
[0053] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0054] In an exemplary embodiment, Figure 1 As shown, the present application provides a landslide disaster risk assessment method, including the following steps 1 to 9.
[0055] Step 1: Construct a disaster-pregnancy factor system involved in landslide disaster risk assessment, wherein the disaster-pregnancy factor system includes multiple disaster-pregnancy factors.
[0056] Disaster-pregnant factors refer to various dynamic factors generated by a disaster-pregnant environment, involving various aspects such as topography, geological factors, meteorology and hydrology, land cover, land use, and human activities. This application selects 10 disaster-pregnant factors, including elevation (Digital Elevation Model, DEM) data, landform types, geological structural factors, engineering geological factors, earthquake parameters, average annual precipitation, water systems, surface cover data, land use types, and road distribution data, to construct a disaster-pregnant factor system for landslide hazard risk assessment. Suppose the disaster-pregnant factor system for landslide hazard risk assessment is:
[0057] (1);
[0058] in, For the A pregnancy disaster factor, is the total number of pregnancy disaster factors; In the exemplary embodiment of the present application, .
[0059] For each pregnancy factor , using certain classification standards to divide it into Among them, for disaster-prone factors with continuous values, such as average annual precipitation, classification is generally carried out according to certain industry standards. For example, areas with an average annual precipitation greater than or equal to 800 mm are classified as humid areas, areas with an average annual precipitation between 400 mm and 800 mm are classified as semi-humid areas, and areas with an average annual precipitation less than 400 mm are classified as arid and semi-arid areas. Alternatively, classification breakpoints can be automatically generated using the equal spacing or natural breakpoint method, such as for DEM data. For discrete disaster-prone factors, statistics and classification are carried out according to their own classification status. For example, landform types can be divided into categories such as plains, hills and mountains; geological structural factors can be divided into categories such as folds, faults, joints, tectonic lines, and fault zones; engineering geological factors can be divided into categories such as stratigraphic lithology and surface geological effects; earthquake parameters can be divided into categories such as earthquake intensity, earthquake magnitude, earthquake probability, and earthquake recurrence period; water systems can be divided into categories such as rivers, lakes, swamps, and glaciers; surface cover data can be divided into land vegetation types / vegetation distribution categories such as forests, grasslands, farmlands, urban land, water bodies, and bare land; land use types can be divided into categories such as agricultural land, construction land, and unused land; and road distribution data can be divided into traffic road distribution categories such as highways and first- to fourth-class highways.
[0060] Step 2: Uniformly sample the study area according to the preset sampling interval to obtain a uniform point set in the study area.
[0061] Define a sampling interval within the study area , uniform sampling is performed in the horizontal and vertical directions to obtain a set of sample points uniformly distributed in the study area, which is called the uniform point set in the study area. Figure 2 In the embodiment shown, the polygonal boundary is the boundary of the study area, and the points inside the boundary are uniform points composed of equally spaced sampling. In an exemplary embodiment, the sampling spacing can be Set to 0.02°, where ° represents latitude and longitude.
[0062] Step 3: For the disaster points in the study area, collect the information of the disaster points based on the disaster-prone factor system to form a positive sample set.
[0063] When calculating the information content of disaster-prone factors, for the classified disaster-prone factors, firstly, the spatial superposition method is used to count the information of each disaster-prone factor in the study area. Belong to The number of disaster points in the sample class and disaster area , and then calculate the information volume of the disaster point according to the following formula (2):
[0064] (2);
[0065] in, The disaster-prone factor calculated for the disaster point No. Class information, . is the total number of disaster points in the study area; The total area of all disaster points in the study area and the amount of information corresponding to all disaster points in the study area Constitute the positive sample set.
[0066] Step 4: For uniform points in the study area, the amount of information of the uniform points is collected based on the disaster-prone factor system to form a negative sample set.
[0067] The calculation method of the information volume of uniform points is similar to that of disaster points. First, according to the spatial superposition method, each disaster-prone factor in the study area is statistically analyzed. Belong to The uniform number of points in the class sample and uniform point area , and then calculate the information content of the uniform point according to the following formula (3):
[0068] (3);
[0069] in, is the total number of uniform points in the study area; is the total area of the study area. is the disaster factor calculated for the uniform point No. For the convenience of the following description, the amount of information will not be combined with the disaster factor calculated for the disaster point. No. The information of each class is formally distinguished. All uniform points in the study area and their corresponding information Constitute the negative sample set.
[0070] Step 5: Construct a landslide disaster risk assessment model based on artificial neural network.
[0071] This application builds a landslide disaster risk assessment model based on the information of disaster-prone factors and the fully connected layer of an artificial neural network, hereinafter referred to as the model. The artificial neural network structure of the landslide disaster risk assessment model is as follows: Figure 3 As shown, its input layer is a two-dimensional tensor with a size of ,in represents the number of input samples, is the total number of hazard factors. Each neuron in the input layer receives the amount of information of a sample (including positive and negative samples). .
[0072] See also Figure 3 , the two-dimensional tensor of the input layer passes through the first fully connected layer fc1 and the relu activation layer, and outputs a size of Then, the neurons in the hidden layer pass through the second fully connected layer fc2, and the output size is Here, we solve the output probability according to the idea of classification problem. That is, for each input sample, it is divided into two categories, namely positive samples and negative samples. Positive samples refer to samples where disasters have occurred, and negative samples refer to samples where disasters will not occur. samples, and the tensor output by the second fully connected layer fc2 includes the predicted probability that the sample belongs to a positive sample and a negative sample and ; .
[0073] Finally, the sample probability of each sample is calculated through the softmax function The softmax function is defined as follows:
[0074] (4).
[0075] Take the final sample probability The output layer of the landslide disaster risk assessment model outputs the probability of landslide disaster occurrence for each sample. , the size is .
[0076] Step 6: Use the positive sample set and the negative sample set to train the landslide disaster risk assessment model and output the landslide disaster probability of each sample.
[0077] In the two-classification model, the information of each disaster factor of the positive sample (disaster point) and the negative sample (uniform point) is The training dataset serves as the training data set, and the relative weights of each hazard-prone factor are determined through training. The actual information content of the hazard-prone factors in the positive and negative samples in the training dataset determines the contribution of each hazard-prone factor to the trained model. Generally speaking, factors that are more sensitive to landslide hazards are given greater weights. This means that the quality of the training dataset is crucial to model accuracy. However, traditional negative samples are randomly generated, and the range of hazard-prone factors in these negative samples deviates significantly from the ideal state. This can cause the model to incorrectly identify the dominant hazard-prone factor, resulting in errors in the relative weights of the various hazard-prone factors in the model. Therefore, addressing the error in negative samples is key to improving the accuracy of binary classification models.
[0078] Here, we first define the prediction accuracy of the model and predicted density When the model is applied to the disaster site, The positive samples are regarded as the disaster samples predicted correctly, and the number of disaster samples predicted correctly is counted. The proportion of landslide disaster probability greater than 0.5 in the statistical disaster points is called prediction accuracy. , expressed as:
[0079] (5);
[0080] in is the total number of positive samples participating in the evaluation, is the number of disaster samples predicted correctly.
[0081] Similarly, when the model is applied to the sampling uniform points in the entire study area, the probability of landslide disaster occurrence at each uniform point is obtained. , the proportion of uniform points with a statistical probability value greater than 0.5 is used as the prediction density of the model , the formula is:
[0082] (6);
[0083] in, For statistics The number of negative samples, Indicates the total number of negative samples involved in the evaluation.
[0084] Prediction accuracy and predicted density These are two key indicators for evaluating the rationality of landslide risk assessment models, and neither of them can be missing. If a model marks most locations in the study area as having a high probability of landslide disasters, If the value is greater than 0.5, the model will likely achieve a high prediction accuracy. However, in practical applications, such a model will lose its application value due to a high false alarm rate. In other words, a good landslide hazard risk assessment model needs to be used at a lower prediction density. Get higher prediction accuracy Therefore, a reasonable model needs to consider not only the prediction accuracy , also consider the predicted density Prediction accuracy For neural networks, the prediction accuracy of the model can be continuously improved by backpropagating the error between the output probability and the real sample. How to reflect it in the neural network is the key problem that this application needs to solve. If the predicted density is calculated directly by counting the number of pixels, it will be difficult to train and calculate within a reasonable time due to the large number of pixels. Therefore, this application proposes the uniform sampling method in step 2 to approximate the predicted density. Obviously, applying the model to each uniform sampling point and counting the probability of landslide disasters will The negative sample ratio greater than 0.5 can be approximated as the predicted density .
[0085] In the landslide hazard risk assessment model, if the model only substitutes positive samples (actual disaster points), the training results of the model will be biased towards the positive direction as a whole. In extreme cases, it may even lead to the occurrence of all locations in the study area. are all greater than 0.5. Therefore, negative samples are essential for model training. However, the traditional method of generating random negative samples is highly subjective and random, which will limit the accuracy of the model. To solve this problem, this application uses uniform points as negative feedback for the entire region. That is, each uniform point is used as a negative sample, and the information content of their disaster-prone factors is obtained and substituted into the model to predict the probability of landslide disasters and calculate the predicted density.
[0086] Specifically, this application collects information about disaster points to form a positive sample set PS; collects information about uniform points to form a negative sample set NS. 80% of the positive sample set PS is randomly selected as the training data set, and the remaining 20% is used as the validation data set. During model training, uniform points are used as negative samples of the model and substituted into the model for training. Since the number of uniform points is greater than the number of disaster points, in order to avoid the problem of negative bias of the model caused by the imbalance of positive and negative samples, the predicted density Introduced into the loss function evaluation of the model.
[0087] Based on the predicted density The weighted cross entropy loss function is as follows:
[0088] (7);
[0089] in, and is a sign function, when the sample When they belong to negative samples and positive samples respectively, the value is 1, otherwise the value is 0. Including positive samples and negative samples, specifically, when the sample When it belongs to negative sample (0), The value is 1, otherwise the value is 0; when the sample When it belongs to the positive sample (1), The value is 1, otherwise the value is 0. and The samples The predicted probability of belonging to negative samples and positive samples, that is, the corresponding model and . is the calculated loss value.
[0090] From formula (7), we can see that the training weights of negative samples (uniform points) and positive samples (disaster points) are and . In this way, when the model is trained, if the model is trained with a positive bias, the predicted density Increase, the negative samples (uniform points) at this time will be set to a larger weight; when the training results are negatively offset, the predicted density The positive samples (uniform points) will be assigned a larger weight. The weighted cross entropy loss function of formula (7) will guide the model to improve the prediction accuracy while reducing the prediction density until a delicate balance between prediction density and prediction accuracy is achieved.
[0091] Step 7: Calculate the prediction accuracy and prediction density of the landslide hazard risk assessment model based on the landslide hazard occurrence probability of each sample.
[0092] This application is based on the predicted density The model training process is as follows Figure 4 As shown. The model is constructed with rounds as units using a weighted cross entropy loss function. Suppose the initial model of the i-th round of training is m i-1 , the trained model is m i In the i-th round of training, the probability of landslide disaster occurrence at each uniform point in the negative sample set NS is calculated as , statistical prediction density .Will Substitute into formula (7) to calculate the loss value of the i-th round of training , and perform error back propagation to obtain the model m i .
[0093] Step 8: Calculate the weighted cross entropy loss function based on the predicted density and perform error back propagation until the training is completed to obtain a trained landslide hazard risk assessment model.
[0094] During the training process, by minimizing the weighted cross entropy loss function, the model parameters can be adjusted to make the model's predicted value as close to the actual value as possible. The smaller it is, the closer the model's predicted results are to the actual results, and the better the model performance is.
[0095] The artificial neural network in this application is trained using the error back propagation algorithm. When the model prediction result is inconsistent with the actual result, the error is back propagated to update the weight parameters of the fully connected layers fc1 and fc2 in the model. When the stopping condition of the model training is reached, the model m at this time is i The trained landslide hazard risk assessment model is packaged for subsequent prediction.
[0096] Step 9: Use the trained landslide hazard risk assessment model to conduct landslide hazard risk assessment on the area to be assessed.
[0097] When landslide disaster risk assessment is needed for the area to be assessed, the information of historical disaster points is first collected based on the disaster factor system and input into the neurons of the model input layer. Then, the trained landslide disaster risk assessment model is used to output the probability of landslide disaster in the area to be assessed. This application divides the range of landslide disaster probability into 5 levels according to the equal division method. The larger the value, the greater the landslide disaster risk. Among them, [0-0.2) is the extremely low risk area, [0.2-0.4) is the low risk area, [0.4-0.6) is the medium risk area, [0.6-0.8) is the high risk area, and [0.8-1] is the extremely high risk area. The range into which the value of falls can be used to determine whether the area to be assessed is an extremely low-risk area, a low-risk area, a medium-risk area, a high-risk area, or an extremely high-risk area, and landslide disaster warnings can be issued for medium-, high-, and extremely high-risk areas.
[0098] In order to verify the effectiveness of the landslide hazard risk assessment method proposed in this application, City A was selected as the research area and landslide hazards were selected as the research object to verify the rationality of the method. City A has a undulating terrain with high mountains and deep valleys, and the terrain is strongly cut. In addition, the region has a subtropical monsoon mountain climate with an average annual precipitation of more than 1200mm. Under the dual effects of precipitation and geological structure, landslide geological disasters frequently occur in City A. A total of 545 landslide hazard points were collected in City A, and the specific distribution is as follows: Figure 5 shown.
[0099] To facilitate model construction and accuracy comparison, the study area uses the same disaster-pregnancy factor system, including 10 disaster-pregnancy factors: elevation data, landform type, geological structure factors, engineering geological factors, earthquake parameters, average annual precipitation, water system, surface cover data, land use type, and road distribution data. The deterministic coefficient method is used for heterogeneous data assimilation.
[0100] Among the 545 disaster points, 80% (436) disaster points were randomly selected as positive samples for training, and the remaining 109 disaster points were used for testing. Taking the circumscribed rectangle of City A as the sampling range, uniform sampling was performed at a spacing of 0.02°, and sampling points outside City A were deducted, resulting in a total of 3410 uniform points as negative samples. Based on the disaster-pregnancy factor system, the amount of information collected from disaster points and uniform points was used to form positive and negative sample sets, respectively. The positive and negative sample sets were substituted into the model for training, and the prediction accuracy of the test disaster points during the training process was recorded. and the model's predicted density , as shown in Table 1 below.
[0101] Table 1 Model training rounds and corresponding prediction accuracy and prediction density
[0102]
[0103] As can be seen from Table 1, as the number of training rounds i increases, the predicted density Gradually decreases and stabilizes. As the predicted density The reduction of prediction accuracy This shows that the prediction accuracy With predicted density There is a positive correlation. When the method of this application is used to predict the probability of landslide disasters in the same area, the predicted density The smaller, The smaller the area where the value is greater than 0.5, the lower the probability that the disaster point falls within the warning area.
[0104] The proposed method is compared with the traditional binary regression method and the single classification solution method. In order to compare the advantages and disadvantages of the models relatively objectively, the same training disaster points (436) and test disaster points (109) are sampled to construct a positive sample set; when comparing the models, the prediction accuracy of the test data set is mainly examined. and predicted density .
[0105] We selected models / algorithms such as logistic regression, support vector machine, random forest, and XGBoost (eXtreme Gradient Boosting) for accuracy comparison. Traditional binary regression methods use a method of randomly generating negative samples. When generating negative samples, it is required that each random point is at least 5 km away from the nearest disaster point, and the number of negative samples must be consistent with the number of disaster points in each study area. The distribution of random negative samples (random points) is as follows: Figure 6 shown.
[0106] Different models can produce different landslide disaster probability distribution maps, and their prediction density and prediction accuracy are also different. In order to objectively compare the accuracy of different models / algorithms, the method of this application is used. The curve is used as a reference, and other algorithms adjust parameters to obtain the optimal Combine the models and use them as a precision point Draw on On a plane, such as Figure 7 As shown. Figure 7 As can be seen from the figure, for the validation dataset, the accuracy points of traditional binary regression methods such as logistic regression, support vector machine, random forest, and XGBoost all fall within the range of the method in this application. This shows that at the same predicted density Under this condition, the method of this application can obtain higher prediction accuracy , that is, the risk warning area identified by the method of this application is more accurate. Similarly, with the same prediction accuracy The predicted density of the method of this application is That is, the method of this application solves the subjectivity and randomness problems of static negative samples and can achieve higher prediction accuracy with a smaller risk marking area.
[0107] In order to further verify the spatial accuracy of the proposed method, it is compared with the logistic regression model. To make the assessment process more reasonable, when the proposed method is used for model training, the training is stopped and the model is saved when the predicted density is the same as that of the logistic regression model. The model is then used to generate the landslide hazard risk assessment result map for City A, as shown in the figure below. Figure 8 shown. Figure 8 In the figure, a, b, c, and d are local areas of the logistic regression model, and e, f, g, and h are local areas of the method of this application. The rationality of the model / algorithm is compared by comparing the landing accuracy of the local areas.
[0108] from Figure 8The results show that, compared with the logistic regression model area a and the method of this application area e, the disaster point A falls in the medium risk area in the logistic regression model, but falls in the extremely high risk area in the method of this application; compared with area b and area f, the disaster point B falls in the medium risk area in area b and falls in the high risk area in area f; compared with area c and area g, the disaster point C falls in the medium risk area in area c and falls in the high risk area in area g; compared with area d and area h, the disaster point D falls in the medium risk area in area d and falls in the extremely high risk area in area h. It can be seen that the risk areas marked by the method of this application (probability of landslide disasters) >0.5), more disaster points can fall within it, thus having higher prediction accuracy.
[0109] To further verify the effectiveness of this method, we compare it with the single-class SVM algorithm. The single-class SVM algorithm can output different prediction densities and prediction accuracies by controlling the probability of outliers. Curve and the method of this application The curves are superimposed, and the result is as follows Figure 9 As shown. Figure 9 As can be seen from the figure, the single-class SVM algorithm Curve in this application method The lower portion of the curve indicates that, at the same prediction density, the proposed method achieves higher prediction accuracy, meaning that the risk warning areas delineated by the proposed method are more precise. Similarly, at the same prediction accuracy, the proposed method achieves a lower prediction density. This demonstrates that the proposed method performs better than the single-class SVM algorithm.
[0110] The probability of landslide disaster in City A is obtained The risk level is divided into five levels according to the equal division method. The larger the value, the more likely landslide disasters are to occur. [0-0.2) is an extremely low risk area, [0.2-0.4) is a low risk area, [0.4-0.6) is a medium risk area, [0.6-0.8) is a high risk area, and [0.8-1] is an extremely high risk area. Figure 10 This is the risk zone division result of City A obtained by this application method. Figure 10 The results show that most of the risk areas are concentrated in the north, south and east of City A, and the probability of landslide disasters is higher in areas near rivers and roads; in the remaining areas, the larger the DEM and annual precipitation, the greater the probability of landslide disasters, and the larger the earthquake peak motion value, the greater the probability of landslide disasters; and areas with greater terrain undulations are more prone to landslide disasters than other areas; this is consistent with the actual occurrence pattern of landslide disasters in City A.
[0111] The present application proposes a landslide disaster risk assessment method, which constructs a landslide disaster risk assessment model based on an artificial neural network and introduces prediction density as a negative evaluation factor of the model. In model training, uniform points are used as negative samples for model training, and the prediction density is calculated in real time through the model in training, and the positive and negative sample feedback weights of the neural network are set according to the prediction density. Through the landslide disaster risk assessment experiment in City A, it is proved that the method of the present application overcomes the error problem of random negative samples and can obtain higher prediction accuracy at a lower prediction density. Compared with the traditional single-classification solution method, the method of the present application incorporates the spatial distribution factor of the disaster and has a higher landing area accuracy.
[0112] The research results of this application provide a new idea and method for landslide disaster risk assessment and early warning. The predicted density is estimated by uniform points, and the positive and negative sample feedback weights of the neural network are set accordingly. This solves the subjectivity and randomness problems of the traditional negative sample generation method, as well as the problem of insufficient spatial distribution assessment of the traditional single-class solution method. It is a new attempt in the field of landslide disaster risk assessment and has important theoretical and practical value.
[0113] In an exemplary embodiment, the present application also provides a computer device, which can be a server or a terminal. The computer device includes a processor, a memory, an input / output interface, and a communication interface. The processor, the memory, and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the landslide hazard risk assessment method is implemented.
[0114] In an exemplary embodiment, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the landslide hazard risk assessment method when executed by a processor.
[0115] In an exemplary embodiment, the present application further provides a computer program product, including a computer program, which implements the landslide hazard risk assessment method when executed by a processor.
[0116] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by hardware associated with computer program instructions. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory or other media in the various embodiments provided herein may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, and the like. Volatile memory may include random access memory (RAM) or external cache memory, and the like. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0117] By studying the topographic and geomorphological characteristics and historical disaster characteristics of areas with more serious landslide disasters, this application collected topographic, geological and other relevant data in the study area, selected DEM data, geomorphological types, geological structural factors, engineering geological factors, earthquake parameters, average annual precipitation, water system, surface cover data, land use type and road distribution data, a total of 10 influencing factors to participate in the landslide disaster risk assessment, and calculated the information content of each disaster-prone factor to establish a landslide disaster risk assessment model. 80% of the positive sample set PS was randomly selected as the training data set, and the remaining 20% was used as the validation data set. The landslide disaster risk assessment model was constructed based on an artificial neural network, and the prediction density was introduced as a negative evaluation factor of the model. In the model training, uniform points were used as negative samples for model training. The prediction density was calculated in real time through the model in training, and the positive and negative sample feedback weights of the neural network were set according to the prediction density. Through the landslide disaster assessment experiment in City A, it was found that the method of this application overcame the error problem of random negative samples and could obtain higher prediction accuracy at a lower prediction density. Compared with the traditional single-class solution method, this application incorporates the spatial distribution factor of the disaster and has higher landing area accuracy.
[0118] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0119] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0120] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A landslide disaster risk assessment method, characterized in that: include: Constructing a disaster-pregnancy factor system for landslide disaster risk assessment, wherein the disaster-pregnancy factor system includes multiple disaster-pregnancy factors; Multiple disaster-predisposing factors include: elevation data, landform types, geological structural factors, engineering geological factors, earthquake parameters, average annual precipitation, water system, surface cover data, land use type and road distribution data; The hazard-prone factor system involved in landslide hazard risk assessment is assumed to be: ;in, For the A pregnancy disaster factor, is the total number of pregnancy disaster factors; ; For each fertility factor , using the preset classification criteria to divide it into Classification; For the disaster-prone factors with continuous values, that is, the average annual precipitation, classification is carried out according to certain industry standards, and the areas with average annual precipitation greater than or equal to 800 mm are divided into humid areas, the areas with average annual precipitation between 400 mm and 800 mm are divided into semi-humid areas, and the areas with average annual precipitation less than 400 mm are divided into arid and semi-arid areas; For DEM data, classification breakpoints are automatically generated according to the equal interval or natural breakpoint method; For discrete disaster-prone factors, statistics and classification are carried out according to their own classification status, and landform types are divided into plains, hills and mountains; For geological structure factors, The data are divided into folds, faults, joints, tectonic lines and fault zones; engineering geological factors are divided into stratum lithology and surface geological action; earthquake parameters are divided into earthquake intensity, earthquake magnitude, earthquake probability and earthquake recurrence period; water systems are divided into rivers, lakes, swamps and glaciers; surface cover data are divided into land vegetation types / vegetation distribution categories of forests, grasslands, farmlands, urban land, water bodies and bare land; land use types are divided into agricultural land, construction land and unused land; road distribution data are divided into traffic road distribution categories of high-speed highways and first- to fourth-class highways; The study area is uniformly sampled according to the preset sampling interval to obtain a uniform point set in the study area; For the disaster points in the study area, the amount of information about the disaster points is collected based on the disaster-pregnancy factor system to form a positive sample set; For uniform points in the study area, the information of uniform points is collected based on the disaster-prone factor system to form a negative sample set; Constructing a landslide disaster risk assessment model based on artificial neural network; Based on the information of disaster-prone factors, a landslide disaster risk assessment model is constructed based on the fully connected layer of the artificial neural network; the input layer of the landslide disaster risk assessment model is a two-dimensional tensor with a size of ,in represents the number of input samples, is the total number of disaster factors; each neuron in the input layer receives the amount of information of one sample ; Samples include positive samples and negative samples; Pregnancy disaster factor No. Class information, ; The two-dimensional tensor of the input layer passes through the first fully connected layer fc1 and the relu activation layer, outputting a size of Then, the neurons of the hidden layer pass through the second fully connected layer fc2, and the output size is Tensor of samples, and the tensor output by the second fully connected layer fc2 includes the predicted probability that the sample belongs to a positive sample and a negative sample and ; ; Finally, the sample probability of each sample is calculated through the softmax function ;The softmax function is defined as follows: ; Take the final sample probability As the probability of landslide disaster occurrence predicted by the model; the output layer of the landslide disaster risk assessment model outputs the probability of landslide disaster occurrence of each sample , the size is ; The landslide disaster risk assessment model is trained using positive sample sets and negative sample sets, and the probability of landslide disaster occurrence for each sample is output; The prediction accuracy and prediction density of the landslide hazard risk assessment model are calculated based on the probability of landslide hazard occurrence of each sample; The weighted cross entropy loss function is calculated based on the predicted density and the error is back-propagated until the training is completed to obtain a trained landslide hazard risk assessment model; Based on the predicted density The weighted cross entropy loss function is as follows: ;in, and is a sign function, when the sample When they belong to negative samples and positive samples respectively, the value is 1, otherwise the value is 0; the sample here Including positive samples and negative samples, when the sample When it is a negative sample, The value is 1, otherwise the value is 0; when the sample When it is a positive sample, The value is 1, otherwise the value is 0; and The samples The predicted probability of belonging to negative samples and positive samples, that is, the corresponding model and ; is the calculated loss value; The trained landslide hazard risk assessment model is used to conduct landslide hazard risk assessment on the area to be assessed.
2. The landslide disaster risk assessment method according to claim 1, characterized in that: For the disaster points in the study area, the amount of information about the disaster points is collected based on the disaster-pregnancy factor system to form a positive sample set, specifically including: Statistics of each disaster-prone factor in the study area Belong to The number of disaster points in the sample class and disaster area ; ; Using the formula Calculate pregnancy disaster factor No. Class information ;in is the total number of disaster points in the study area; is the total area of all disaster points in the study area; All disaster points and corresponding information Constitute the positive sample set.
3. The landslide disaster risk assessment method according to claim 2, characterized in that: The prediction accuracy and prediction density of the landslide disaster risk assessment model are calculated based on the probability of landslide disaster occurrence of each sample, specifically including: Will The positive samples are regarded as the disaster samples predicted correctly, and the number of disaster samples predicted correctly is counted. ; Using the formula Calculating the prediction accuracy of landslide hazard risk assessment models ;in is the total number of positive samples participating in the evaluation; statistics The number of negative samples , and use the formula Calculating the predicted density of landslide hazard risk assessment models ;in Indicates the total number of negative samples involved in the evaluation.
4. The landslide disaster risk assessment method according to claim 3, characterized in that: The landslide hazard risk assessment model trained to perform landslide hazard risk assessment on the area to be assessed specifically includes: Use the trained landslide hazard risk assessment model to output the probability of landslide disasters in the area to be assessed ; according to The value of will divide the area to be assessed into very low risk area, low risk area, medium risk area, high risk area or very high risk area; among them, The value of [0-0.2) is the extremely low risk area, [0.2-0.4) is the low risk area, [0.4-0.6) is the medium risk area, [0.6-0.8) is the high risk area, and [0.8-1] is the extremely high risk area.
5. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the landslide hazard risk assessment method according to any one of claims 1 to 4.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the landslide hazard risk assessment method according to any one of claims 1 to 4 is implemented.
7. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the landslide hazard risk assessment method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Landslide risk assessment method and system
CN120125038A