Landslide surge height prediction method and system based on sparrow-optimized random forest model
By building a landslide surge physical model and using the sparrow-optimized random forest model, the problem of ignoring trajectory changes and morphological changes in existing technologies was solved, and accurate prediction of landslide surge height was achieved, thereby improving the accuracy and stability of the prediction.
Patent Information
- Application Number
- CN202411401030.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-10-09
AI Technical Summary
The existing landslide surge height prediction method only considers the starting and ending positions of the landslide body, ignores the trajectory changes and the morphological changes of the landslide body, can only calculate a certain moment, cannot reflect the entire process of the landslide, and assumes too many conditions, making it difficult to obtain accurate results.
The sparrow optimized random forest model was adopted to build a landslide surge physical model. Experimental groups with different landslide volumes and water depths were set up to obtain the experimental data set, which was then scaled proportionally according to the Froude similarity criterion. The sparrow optimized random forest model was built using Python language and the sklearn and mealpy libraries. Parameters were set and trained. A performance monitoring layer and a feature sharing network were introduced. New samples were generated by combining the adversarial generative network. Batch normalization and random dropout strategies were implemented to predict the maximum height of landslide surge.
It improves the accuracy and stability of landslide surge height prediction, can comprehensively consider the relationship between multiple factors, reduce research complexity, and achieve accurate prediction of the maximum height of landslide surge.
Smart Images

Figure CN119416303B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data-driven and landslide surge disaster prevention and control, and in particular to a landslide surge height prediction method and system based on a sparrow-optimized random forest model. Background Art
[0002] Landslide surge research has long been a key focus of research on the landslide surge hazard chain. For the safety of major engineering projects, studying the mechanisms and patterns of landslide surges helps assess geological hazard risks, implement effective preventive measures, and ensure project safety. In mountainous river channels, landslides can cause severe landslide surges, posing a significant threat to reservoir banks, hydraulic structures, and other structures. Therefore, the impact of landslide surges cannot be ignored. Accurately predicting the maximum height of landslide surges, and thus enabling targeted prevention and control, is of great scientific and engineering significance. In recent years, an increasing number of studies have applied machine learning methods to landslide research. Compared to existing research methods, these methods not only consider multiple factors influencing landslides but also improve the accuracy of research results. Machine learning methods can rapidly process large amounts of data in a short period of time and comprehensively consider the relationships between various characteristic factors, reducing research complexity.
[0003] There are several main existing methods for predicting landslide surge height, each with its own limitations: the energy method only considers the starting and ending positions of the landslide during calculation, ignoring changes in trajectory and landslide morphology; the Pan Jiazheng method considers the influence of landslide morphology and water, but only calculates at a specific moment and cannot reflect the entire landslide process; the American Society of Civil Engineers recommended method considers the influence of landslide shape, but its excessive assumptions make accurate results difficult; the Scheidegger method establishes a relationship between friction coefficient and landslide volume, but is only applicable to debris materials. The method of predicting maximum landslide surge height using a sparrow-optimized random forest model not only considers multiple factors affecting landslides but also comprehensively considers the relationships between these factors, reducing research complexity and improving the accuracy of the results.
[0004] Prior art one, Chinese patent, application number 202311133111.5 discloses a method for predicting the probability of exceeding the maximum height of a landslide-surge wave, including: constructing candidate models in dimensionless form, collecting test data of the model in each experimental design, establishing each experimental database, performing Bayesian processing on the candidate models according to each experimental database, generating Bayesian factors for each candidate model, determining the candidate model with the largest Bayesian factor as the maximum height prediction model, updating the maximum height prediction model through Bayesian, determining the parameter probability distribution, and predicting the probability of exceeding the maximum height of the landslide-surge wave in the target reservoir area according to the parameter probability distribution. Although each candidate model is processed by Bayesian, the maximum height prediction model is determined, the parameter probability distribution is determined, and the uncertainty of the experimental design is quantified, thereby predicting the probability of exceeding the maximum height of the landslide-surge wave in the target reservoir area. However, only the starting position and the ending position of the landslide body are considered, and the trajectory changes and the morphological changes of the landslide body are ignored, resulting in reduced prediction accuracy;
[0005] Prior art two, Chinese patent application number 201910437509.5, discloses a model test device for studying the energy dissipation effect of landslide surge waves and wave height prediction. The device includes a landslide model, a sliding device, a test flume, a bank roughness and slope shape device, and a wave height measurement device. The sliding device and the bank roughness and slope shape device are arranged on the wall surface of the test flume, the landslide model is placed on the sliding device, and the wave height measurement device is installed above or on the side of the test flume via a movable frame. A test method for the model test device for studying the energy dissipation effect of landslide surge waves and wave height prediction is also disclosed. The test device can simulate the surge changes caused by the sliding process of different landslide bodies, and the influence of bank slope shape and bank roughness on the surge propagation process. Although it is possible to analyze and predict the propagation law and attenuation trend of landslide surge waves caused by different bank slope shapes and bank roughness, and establish a prediction model for the wave height distribution along the landslide surge path, it can only be calculated at a specific moment and cannot reflect the entire landslide process.
[0006] Prior art three, Chinese patent, application number 202410014296.6 discloses a method for judging the type of surge waves induced by landslides in water areas and predicting the wave height, including establishing a dynamic model of slope instability, obtaining the landslide velocity u(t) and landslide displacement s(t) process, calculating the dimensionless parameters Fr and Sz, distinguishing the surge wave type, and calling the oscillation wave / single wave formula to calculate the wave height m of the corresponding wave type. Although the classification and discrimination of landslide surge waves and the proposal of wave height prediction methods for different surge waves provide reliable technical support for the determination of relevant disaster prevention and mitigation measures and the formulation of emergency plans, there are too many assumptions, and it is difficult to obtain accurate results.
[0007] Currently, existing technologies 1, 2, and 3 only consider the starting and ending positions of the landslide during calculation, ignoring changes in trajectory and morphology. This allows calculations to be performed at a specific moment, failing to reflect the entire landslide process. These assumptions make obtaining accurate results difficult. Therefore, the present invention provides a method and system for predicting landslide surge height based on a sparrow-optimized random forest model. This method, which uses a sparrow-optimized random forest model to predict maximum landslide surge height, not only considers multiple factors influencing landslides but also comprehensively considers the relationships between these characteristic factors, reducing research complexity and improving the accuracy of research results. Summary of the Invention
[0008] The main purpose of the present invention is to provide a method and system for predicting landslide surge height based on the sparrow optimized random forest model, so as to solve the problem that the existing technology only considers the starting and ending positions of the landslide body, ignores the trajectory changes and the morphological changes of the landslide body, can only calculate a certain moment, cannot reflect the entire process of the landslide, assumes too many conditions, and is difficult to obtain accurate results.
[0009] To achieve the above object, the present invention provides the following technical solutions:
[0010] A landslide surge height prediction method based on a sparrow-optimized random forest model, comprising:
[0011] For the target landslide area, a corresponding landslide surge physical model was built, and test groups with different landslide volumes and water depths were set up for testing to obtain the corresponding test data set, which was then scaled according to the Froude similarity criterion.
[0012] The experimental data were preprocessed to obtain the corresponding data set, which was then divided into a training set and a test set in a ratio of 7:3, i.e. the test set data accounted for 30% of the total data;
[0013] Using Python language, a sparrow optimized random forest model was built based on sklearn and mealpy libraries. The input values of various parameters were set and the training set was used to train the sparrow optimized random forest model. The labeled dataset and the trained sparrow optimized random forest model were used to predict the maximum height of the landslide surge.
[0014] As a further improvement of the present invention, the process of performing proportional scaling according to the Froude similarity criterion includes:
[0015] Calculate the Froude number of the target landslide area and the model; introduce the characteristic flow frequency F c and the flow resistance factor D:
[0016]
[0017] Matching of the Froude number of the model and the target landslide area: The Froude number of the test model is equal to that of the target landslide area, that is:
[0018]
[0019] Among them, V m and V p are the flow velocities in the model and target areas respectively; L m and L p are the characteristic lengths of the model and target regions, respectively;
[0020] Determine the ensemble scaling ratio λ of the model and the target landslide area;
[0021] According to the Froude number formula, the scaling factor is defined as:
[0022]
[0023] Update the flow rate of the model to meet the proportional scaling condition and introduce an adjustment coefficient C v Consider the differences between actual flows:
[0024]
[0025] Adjustment coefficient C v Adjusted based on previous trial data and similarity;
[0026] The test flow rate, pressure, gravitational acceleration and other related physical quantities are adjusted proportionally;
[0027] The influence of the density λ′ and dynamic viscosity of the introduced fluid on the pressure is:
[0028]
[0029] Among them C p is the pressure adjustment coefficient, which is calibrated by test data;
[0030] After adjusting all relevant physical quantities, experiments are conducted on the model, and data are collected and recorded accordingly; the experimental data are scaled to the target landslide area in each experiment:
[0031] Prediction result = C data Actual test results
[0032] Among them C data is the conversion factor used to fit the model to the experimental data.
[0033] As a further improvement of the present invention, the preprocessed data set is divided into a training set and a test set, including:
[0034] Represent the dataset as a set containing multiple samples, each of which consists of features and corresponding target variables; preprocess the data to generate a feature matrix and target variables; calculate a weight for each sample, which reflects the importance of the sample;
[0035] The cumulative distribution function of the samples is calculated based on the weights, and the cumulative distribution function is used for weighted random sampling. Samples are randomly selected in a 7:3 ratio as the training set, and the remaining samples are used as the test set;
[0036] The obtained training set and test set meet the required ratio in terms of sample number.
[0037] As a further improvement of the present invention, the process of predicting the maximum height of a landslide surge includes:
[0038] A performance monitoring layer is introduced into the Sparrow Optimization Algorithm to track the performance of the slope surge height prediction model on the validation set in real time during the optimization process. The number of global searches is dynamically increased during the search process. Whenever the Sparrow Optimization Algorithm detects a local optimal bottleneck, it enters a global reinforcement phase to increase the mutation rate and crossover rate.
[0039] A feature sharing network is established between individual sparrows, allowing them to share and receive feature information through a weighted adjacency graph. Each sparrow receives excellent features found by its neighbors through a graph propagation mechanism, converging on a set of features that are effective for predicting swell height. After completing their respective tasks, sparrows use a feature evaluation and voting algorithm to decide whether to adopt a feature, increasing the frequency of use of the optimal feature and achieving efficient information sharing and integration within the group.
[0040] In the feature integration stage, different sparrow individuals are allowed to introduce heterogeneous features from their own fields to form diversified feature clusters. They are grouped by clustering the effects of different feature combinations, and the optimal feature combination is selected for training through similarity metrics.
[0041] During the operation of the random forest model, the decision tree monitors the performance of split nodes through weighted hysteresis tracking, and determines in real time whether nodes under certain conditions need to be reconstructed. Each node in the decision tree has a lifespan, and a new tree structure is dynamically created. When constructing each tree, a feature subset is selected based on feature importance, and a voting mechanism is used to weight the important features of each tree. During feature analysis, new features are introduced and integrated through sequential dependency selection.
[0042] Generate new samples for the training dataset using a generative adversarial network, and construct different versions of the training set by combining the generated samples with the actual samples; set the parameters of the standard distribution selected in the generator pre-training phase; and implement batch normalization and random dropout strategies in the training of the Sparrow Optimized Random Forest model.
[0043] The data of each experiment, including landslide volume and water depth, are obtained to form an input vector. The corresponding maximum surge height is the label. The input vector and label are combined to form a labeled dataset. The labeled dataset is represented by a one-dimensional array, in which each element corresponds to the maximum surge height of each set of feature data. The maximum height of landslide surge is predicted using the labeled dataset and the trained sparrow-optimized random forest model.
[0044] As a further improvement of the present invention, a process for achieving efficient integration of information sharing in a group includes:
[0045] Suppose there are multiple sparrow individuals, each sparrow individual is connected to its neighbor set, an adjacency matrix is defined, and the feature vector of the sparrow individual is also defined. After one propagation, the individual's features are obtained by weighted average of the features of its neighbors;
[0046] There are N′ sparrow individuals, each sparrow individual a and its neighbor set N a ′ connection. We define the adjacency matrix in:
[0047] A ab = {1, if sparrow b is a neighbor of sparrow a 0, otherwise
[0048] Define the eigenvector of sparrow individual a as Where M′ is the feature dimension. After one propagation, the feature of individual a can be obtained by weighted average of the features of its neighbors:
[0049] f a″ =σ(∑b∈N′ a w ab f b )
[0050] Among them, w ab is the influence weight of sparrow b on sparrow a, σ(·) is the activation function used to introduce nonlinearity;
[0051] For the newly generated features, each sparrow calculates its prediction ability and the prediction error of the sparrow for the new features. All sparrows vote based on the error evaluation results to select the features.
[0052] Among them, for the newly generated feature f a″ , each sparrow calculates its prediction ability, let ε a is the prediction error of sparrow a for the new feature, and the formula is as follows:
[0053]
[0054] Where K is the number of validation set samples, is the output of predicting sample k using the new feature, y k is the actual maximum height of the swell;
[0055] All sparrows vote based on the error evaluation results to select features. After the features pass the evaluation, the number of votes for each feature is V. c , then the probability P of a feature c being adopted b for:
[0056]
[0057] The sparrow adopts the probability P b The features that are greater than the set threshold θ form a new feature set for subsequent training.
[0058] As a further improvement of the present invention, the process of selecting the optimal feature combination for training by similarity measurement includes:
[0059] Generate various combinations from all available features, each feature combination represents a specific selection of a set of features, using a binary method to indicate whether each feature is selected; use a predefined random forest to train each feature combination and evaluate the performance on the validation set, assigning a performance score to each feature combination based on the evaluation results;
[0060] Based on the performance scores, the feature combinations are divided into different groups, and the feature combinations with similar performance are grouped together. By analyzing the clustering results, the similarity between feature combinations is evaluated and the correlation between features is measured.
[0061] Based on the calculated similarity, a feature combination with a similarity higher than a preset threshold is selected.
[0062] As a further improvement of the present invention, the process of weighting the important features of each tree by using a voting mechanism includes:
[0063] A performance index is calculated for each decision tree based on its performance on the validation set. Different voting weights are assigned to each tree based on the performance index. A feature interaction network is established, modeling the relationships between features using a graph neural network. The importance value of each tree for a feature is used as a node connected to other feature nodes to form a feature interaction relationship graph.
[0064] The prediction results of each tree are analyzed for consistency. If the prediction of a tree in a specific field is inconsistent with that of other trees, the voting influence weight of the tree is automatically reduced. Each tree gives a prediction value and a corresponding confidence level when voting. The output is a weighted sum of the feature weight, the tree's prediction value, and the corresponding confidence level.
[0065] Features are grouped heterogeneously according to their properties. For example, meteorological factors, terrain features, hydrological conditions, etc. are separated. An independent set of decision trees is designed for each feature group and trained. When making decisions, voting is performed by feature group, and feature selection and evaluation are performed within each feature group. The voting results of each feature group are summarized.
[0066] As a further improvement of the present invention, the process of implementing batch normalization and random dropout strategy includes:
[0067] Before model training, the input data is standardized. Based on the structure of deep neural networks, a batch normalization layer is inserted after each hidden layer to ensure that the distribution of features can be adjusted to a more ideal state before the output of each layer is passed to the activation function.
[0068] After forward propagation of each hidden layer, the mean and variance of the hidden layer output are calculated and the output is normalized using statistics. The normalized output is transformed using the learned scaling and offset parameters, and the global mean and variance are dynamically adjusted using the exponential moving average method.
[0069] Set the inactivation rate. In each training iteration, randomly select several neurons and set their outputs to zero. Generate a mask matrix with the same shape as the current layer output. Control which nodes are eliminated based on the selected inactivation rate so that the model uses a different network structure in each training process.
[0070] As a further improvement of the present invention, the process of predicting the maximum height of the landslide surge includes:
[0071] Obtain the latest input feature data, including landslide volume, soil type, meteorological factors, water level changes and flow rate; integrate the collected feature data into an input vector;
[0072] Load the trained and Sparrow-optimized random forest model into memory; input the constructed input feature vector into the random forest model, perform inference on each decision tree, and execute the tree's decision process based on the input features;
[0073] Each tree will traverse downward from the root node according to the value in the feature vector, splitting according to the feature threshold until it reaches a leaf node; at each leaf node, there is a predicted value, that is, the expected maximum surge height;
[0074] Collect the prediction results of all decision trees and aggregate them into a final prediction through a weighted voting mechanism; combine all trees to obtain a prediction value, which is used to represent the expected maximum surge height under a specific input feature vector; output the maximum surge height prediction result in the format of text, chart or report.
[0075] To achieve the above object, the present invention also provides the following technical solutions:
[0076] A landslide surge height prediction system based on a sparrow-optimized random forest model is applied to the landslide surge height prediction method based on a sparrow-optimized random forest model. The landslide surge height prediction system based on a sparrow-optimized random forest model includes:
[0077] The data acquisition module is used to build a corresponding landslide surge physical model for the target landslide area, set up test groups with different landslide volumes and water depths for testing, obtain the corresponding test data set, and perform proportional scaling according to the Froude similarity criterion;
[0078] The data partitioning module is used to preprocess the test data to obtain the corresponding data set, and then divide the preprocessed data set into a training set and a test set in a ratio of 7:3, that is, the test set data accounts for 30% of the total data;
[0079] The landslide surge prediction module is used to build a sparrow optimized random forest model based on the sklearn and mealpy libraries using Python language, set the input values of various parameters and train the model using the training set, and use the labeled dataset and the trained sparrow optimized random forest model to predict the maximum height of the landslide surge.
[0080] The present invention can simulate the dynamic changes of ground water bodies when landslides occur more realistically by building a physical model of landslide surge, including the processes of landslide body entering water, water body disturbance, surge formation and propagation, etc., and set up test groups with different landslide volumes and water depths to systematically examine the influence of these key factors on the characteristics of landslide surge, providing rich test data support for subsequent model establishment and data analysis, and performing proportional scaling based on the Froude similarity criterion to ensure the comparability of test data and the wide application of the model, so that the test results can more accurately reflect the actual landslide surge situation; through preprocessing operations to remove noise and fill Missing values were filled in, which improved the quality and availability of the data. The preprocessed data set was divided into training set and test set in a ratio of 7:3, which ensured the effectiveness of model training and the independence of test, and helped to evaluate the generalization ability of the model. The sparrow optimization algorithm was used to optimize the parameters of the random forest model, which improved the prediction accuracy and stability of the model. As an emerging swarm intelligence optimization algorithm, the sparrow optimization algorithm can find the optimal solution in a complex parameter space. Based on the optimized random forest model, the training set was used for training, and the labeled data set was used for prediction, which can predict the maximum height of the landslide surge. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 This is a schematic diagram of the steps of an embodiment of a method for predicting landslide surge height based on a sparrow-optimized random forest model according to the present invention;
[0082] Figure 2 This is a schematic flow chart of the steps for building a corresponding landslide surge physical model according to an embodiment of the landslide surge height prediction method based on the sparrow optimized random forest model of the present invention;
[0083] Figure 3 This is a flowchart of the steps of dividing the preprocessed data set into a training set and a test set in one embodiment of the landslide surge height prediction method based on the sparrow optimized random forest model of the present invention;
[0084] Figure 4 This is a flow chart showing the steps of predicting the maximum height of a landslide surge according to an embodiment of the landslide surge height prediction method based on the sparrow-optimized random forest model of the present invention;
[0085] Figure 5 This is a functional module diagram of an embodiment of a landslide surge height prediction system based on a sparrow-optimized random forest model according to the present invention;
[0086] Figure 6 This is a schematic structural diagram of an embodiment of an electronic device of the present invention;
[0087] Figure 7 A schematic structural diagram of an embodiment of a storage medium of the present invention;
[0088] Figure 8 This is a schematic diagram of an embodiment of a method for predicting landslide surge height based on a sparrow-optimized random forest model according to the present invention;
[0089] Figure 9 This is a comparison diagram of prediction results before and after optimization of an embodiment of a landslide surge height prediction method based on a sparrow-optimized random forest model of the present invention;
[0090] Figure 10 A schematic diagram of a random forest learning curve of an embodiment of a method for predicting landslide surge height based on a sparrow-optimized random forest model according to the present invention;
[0091] Figure 11 A schematic diagram of a sparrow optimized random forest learning curve according to an embodiment of a landslide surge height prediction method based on a sparrow optimized random forest model of the present invention;
[0092] Figure 12 This is a comparison chart of feature importance coefficients of an embodiment of a landslide surge height prediction method based on a sparrow-optimized random forest model according to the present invention. DETAILED DESCRIPTION
[0093] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0094] The terms "first", "second" and "third" in the present invention are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first", "second" and "third" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "multiple" is at least two, for example, two, three, etc., unless otherwise clearly and specifically defined. All directional indications in the embodiments of the present invention (such as up, down, left, right, front, back...) are only used to explain the relative position relationship, movement, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices.
[0095] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0096] like Figure 1 As shown, this embodiment provides an embodiment of a method for predicting landslide surge height based on a sparrow-optimized random forest model. In this embodiment, the method for predicting landslide surge height based on a sparrow-optimized random forest model specifically includes the following steps:
[0097] Step S1: Build a corresponding landslide surge physical model for the target landslide area, set up test groups with different landslide volumes and water depths for testing, obtain the corresponding test data set, and perform proportional scaling according to the Froude similarity criterion;
[0098] Step S2: Preprocess the test data to obtain the corresponding data set, and divide the preprocessed data set into a training set and a test set in a ratio of 7:3, that is, the test set data accounts for 30% of the total data;
[0099] Step S3: Use Python language to build a sparrow optimized random forest model based on sklearn and mealpy libraries, set the input values of various parameters and use the training set to train the sparrow optimized random forest model, and use the labeled dataset and the trained sparrow optimized random forest model to predict the maximum height of the landslide surge.
[0100] Preferably, step S1 of this embodiment can simulate the dynamic changes of ground water bodies when landslides occur more realistically by building a physical model of landslide surges, including processes such as landslide body entering water, water body disturbance, surge formation and propagation, and setting test groups with different landslide volumes and water depths to systematically examine the influence of these key factors on landslide surge characteristics, providing rich test data support for subsequent model establishment and data analysis, and performing proportional scaling based on the Froude similarity criterion to ensure the comparability of test data and the wide application of the model, so that the test results can more accurately reflect the actual landslide surge situation; step S2 removes the landslide surges through preprocessing operations. Noise removal and missing value filling are performed to improve the quality and availability of the data. The preprocessed data set is divided into a training set and a test set in a ratio of 7:3, which ensures the effectiveness of model training and the independence of test, and helps to evaluate the generalization ability of the model. Step S3 uses the sparrow optimization algorithm to optimize the parameters of the random forest model, which improves the prediction accuracy and stability of the model. As an emerging swarm intelligence optimization algorithm, the sparrow optimization algorithm can find the optimal solution in a complex parameter space. Based on the optimized random forest model, the training set is used for training, and the labeled data set is used for prediction, which can predict the maximum height of the landslide surge.
[0101] In summary, this embodiment provides a reliable experimental basis for the prediction and evaluation of landslide surges, helps to deeply understand the physical mechanism of landslide surges, and lays a solid foundation for subsequent data analysis and model training through the accumulation of experimental data; data preprocessing is an important link in the machine learning process, which directly affects the training effect and prediction accuracy of the model. Through reasonable data division, the performance of the model can be evaluated more objectively, ensuring the reliability of the model in practical applications; the construction of the sparrow optimized random forest model achieves the efficiency and accuracy of landslide surge prediction, and provides strong technical support for the early warning and air defense of landslide disasters. Through the application of the model, the degree of damage caused by landslide surges can be predicted in advance, providing a scientific basis for relevant departments to formulate emergency plans and take preventive measures, which helps to reduce the threat of landslide disasters to the safety of life and property of the people.
[0102] This embodiment proposes a landslide surge height prediction method based on the sparrow optimized random forest model, and establishes a landslide surge physical model test based on the Froude similarity criterion; the actual maximum landslide surge height at different points is obtained through experiments, and a sparrow optimized random forest model is established through the random forest algorithm and the sparrow optimization algorithm, which can accurately predict the maximum height of landslide surge at different points in the study area; this embodiment introduces the sparrow optimization algorithm to optimize the random forest model and applies it to actual engineering cases; the research results show that the sparrow optimized random forest model has better performance when processing data with a small deviation rate and can significantly improve the accuracy of the model; at the same time, the sparrow optimized random forest model can reduce the standard deviation of the predicted value and improve the density of the predicted value; compared with the random forest model before optimization, the performance evaluation indicators of the optimized model are better than those of the pre-optimization model, and it shows better adaptability in predicting the maximum height of landslide surge and can predict the maximum height of landslide surge more accurately.
[0103] Further, if Figure 2 As shown, the process of building the corresponding landslide surge physical model in step S1 specifically includes the following steps:
[0104] Step S11: For the target landslide area, test data such as landslide volume and water depth are collected, and fluid-solid coupling simulation is performed on the collected data. Different models of the landslide in the water entry stage and when the landslide is completely immersed in water are calculated based on the fluid-solid coupling mathematical model;
[0105] Step S12: According to the characteristics of landslide surge, a three-dimensional landslide surge control equation is established, and the two-dimensional unsteady separated implicit PISO algorithm is used to solve the equation;
[0106] Step S13: according to the scale of the landslide, water entry velocity, water depth, river width, etc., the initial conditions and boundary conditions of the fluid-solid coupling mathematical model are set, and the fluid-solid coupling mathematical model is verified and calibrated through existing test results.
[0107] In this embodiment, step S11 fluid-solid coupling simulation:
[0108] Fluid dynamics equations:
[0109]
[0110] Where u represents the fluid velocity vector, p represents the pressure, ρ represents the fluid density, v represents the dynamic viscosity, and g represents the gravitational acceleration vector;
[0111] Solid mechanics equations (dynamic equilibrium equations):
[0112]
[0113] Where: σ represents the stress tensor, b represents the body force (such as gravity), ρ p represents the density of solid material, u p Represents the displacement of solid materials;
[0114] Step S12: Landslide surge control equation:
[0115] Surge governing equations:
[0116]
[0117] Where: η represents the water surface height
[0118] Two-dimensional unsteady separated implicit PISO algorithm:
[0119] Continuity equation for water flow:
[0120]
[0121] The momentum equation is solved using a step-by-step solution consisting of a predictor-corrector step:
[0122] Prediction step: Calculate the intermediate speed u * ;
[0123] Correction steps: Correct the velocity field by pressure;
[0124] Step S13: Model initial and boundary conditions:
[0125] Initial conditions:
[0126] u(x,y,0)=u0(x,y),η(x,y,0)=eta0(x,y)
[0127] Where u0 and η0 represent the initial velocity and water surface height distribution;
[0128] Boundary conditions:
[0129] Inlet condition (inflow): set flow rate u = u in ;
[0130] Outlet conditions (outflow): set free flow conditions
[0131] Preferably, step S11 of this embodiment involves data collection and fluid-solid coupling simulation. Key physical parameters are obtained by collecting experimental data such as landslide volume and water depth in the target landslide area, providing the necessary input information for the model. A fluid-solid coupling mathematical model is constructed to simulate the complex process of landslide movement in water, including the different dynamic effects of the water entry and full immersion stages. This model takes into account the interaction between liquid and solid materials, providing a physical basis for further analysis. Significance: The collected landslide volume and water depth data provide a reference for subsequent model construction, ensuring the model's scientific and applicable nature. The fluid-solid coupling simulation provides a deep understanding of the landslide's entry and immersion processes, helping to identify the characteristics of landslide movement under different water depths and flow conditions, thereby forming a more realistic physical model. Step S12 establishes and solves the governing equations for landslide surges. Based on the dynamic characteristics of landslide surges, a governing equation for landslide surges is established for three-dimensional flow, accurately describing the movement and propagation of landslide surges. A two-dimensional unsteady separated implicit PISO algorithm is used to solve the governing equations, providing an efficient and accurate computational method for handling complex boundaries and flows. Significance: By establishing professional control equations, the landslide phenomenon is successfully transformed into a mathematical problem, which is convenient for in-depth analysis; using numerical calculation methods, accurate numerical solutions can be obtained in a relatively short time, which is suitable for rapid simulation and prediction under different working conditions. Step S13 sets the initial and boundary conditions and verifies the model. The initial and boundary conditions of the model are set according to physical parameters such as the scale of the landslide body, water entry speed, water depth, and river width to ensure the accuracy and rationality of the simulation process; by comparing with existing test results, the simulation results are combined with the actual situation, and the model is verified and calibrated to improve the credibility of the model. Significance: Ensuring the setting of reasonable initial and boundary conditions makes the model closer to the actual situation, thereby enhancing its adaptability under various conditions; through comparison and verification, it is confirmed that the model can accurately reflect the characteristics of landslide surges, increasing its reliability in scientific research and engineering applications.
[0132] In summary, the construction of the landslide surge physical model in this example completes the entire process, from data collection, mathematical model development, and model validation. This not only provides a rich physical background and data support for the model, but also improves its computational efficiency and predictive capabilities. Ultimately, this provides an important scientific basis for in-depth research on landslide surge, disaster assessment, and the development of preventive measures.
[0133] Furthermore, the specific steps of the process of performing proportional scaling according to the Froude similarity criterion in step S1 include:
[0134] Step S14: Calculate the Froude number of the target landslide area and the model; introduce the characteristic flow frequency F c and the flow resistance factor D:
[0135]
[0136] The frequency characteristics of the flow and the influence of flow resistance are taken into account, making the similarity of the model more accurate;
[0137] Matching of the Froude number of the model and the target landslide area: The Froude number of the test model is equal to that of the target landslide area, that is:
[0138]
[0139] Among them, V m and V p are the flow velocities in the model and target areas respectively; L m and L p are the characteristic lengths of the model and target regions, respectively;
[0140] Step S15: determining the set scaling ratio λ of the model and the target landslide area;
[0141] Scaling ratio determination: According to the aforementioned Froude number formula, the scaling ratio is defined as:
[0142]
[0143] Update flow rate: Update the flow rate of the model to meet the proportional scaling condition and introduce an adjustment coefficient C v Consider the differences between actual flows:
[0144]
[0145] Adjustment coefficient C v Adjusted based on previous trial data and similarity;
[0146] Step S16: The test flow rate, pressure, gravitational acceleration and other related physical quantities are adjusted proportionally;
[0147] The influence of the density λ′ and dynamic viscosity ρ of the introduced fluid on the pressure:
[0148]
[0149] Among them C p is the pressure adjustment coefficient, which can be calibrated by test data;
[0150] After adjusting all relevant physical quantities, experiments are conducted on the model, and data are collected and recorded accordingly; the experimental data are scaled to the target landslide area in each experiment:
[0151] Prediction result = C data Actual test results
[0152] Among them C datais the conversion factor used to fit the model to the experimental data.
[0153] Preferably, step S14 of this embodiment calculates the Froude numbers for the target landslide region and the model. By introducing characteristic flow frequencies and flow resistance factors, the traditional Froude number calculation is modified, making it more adaptable and capable of better reflecting the characteristics under different flow conditions. By ensuring that the Froude numbers of the test model and the target landslide region are equal, the consistency of the flow pattern is tested. Significance: This improves the similarity of the model, enabling it to accurately simulate real-world conditions under various flow conditions. This is crucial for predicting landslide surge behavior and helps avoid experimental errors caused by insufficient similarity. Step S15 determines the collective scaling ratio λ of the model and the target landslide region. By calculating the scaling ratio, the proportional relationship between the model and the target region is clarified, and the consistency of flow velocity and geometric features is tested. By introducing an adjustment coefficient, the flow velocity can be adjusted in real time based on previous test data to reflect the complexity of actual conditions. Significance: By using an appropriate scaling ratio, the flow conditions of the test model are matched to those of the actual landslide region, thereby enhancing the practical application value of the test data and enabling effective prediction and analysis in different scenarios. In step S16, relevant physical quantities such as test flow rate, pressure, and gravitational acceleration are adjusted proportionally. Each relevant physical quantity (flow rate, pressure, gravity, etc.) is adjusted in the same proportion, ensuring that the test conditions on the test model reflect the actual conditions to the greatest extent possible. The effects of fluid density and dynamic viscosity are introduced to increase the accuracy of the description of flow phenomena, thereby enhancing the effectiveness of the model. Significance achieved: By precisely adjusting these physical quantities, the test model can truly reflect the physical characteristics of actual landslide surges, resulting in more credible test results. More accurate dynamic characteristics of landslide surges can be obtained, assisting in effective engineering design and safety assessments. Test data collection and result scaling: By collecting and recording model test data, a connection can be established between the model test results and the actual landslide area. The use of adjustment factors enhances the reliability and effectiveness of test results in practical applications. Significance achieved: Test data can be effectively converted into prediction results for the actual landslide area, improving the accuracy and usability of model predictions. This increases the guiding significance of model research for practical engineering applications, enabling it to play an important role in areas such as landslide early warning and post-disaster assessment.
[0154] In summary, each step in this embodiment is designed to enhance the similarity, accuracy, and credibility of the model, thereby enabling the experiment to obtain meaningful results in landslide surge prediction, helping to better understand the complexity of landslide surges, and providing reliable data support for addressing related engineering problems.
[0155] Further, if Figure 3As shown, the process of dividing the preprocessed data set into a training set and a test set in step S2 specifically includes the following steps:
[0156] Step S21: Represent the data set as a set, including multiple samples, each sample consists of features and a corresponding target variable (surge height); preprocess the data to generate a feature matrix and target variables; calculate a weight for each sample, the weight reflecting the importance of the sample;
[0157] Step S22: Calculate the cumulative distribution function of the samples based on the weights, use the cumulative distribution function to perform weighted random sampling, randomly select samples with a ratio of 7:3 as the training set, and use the remaining samples as the test set;
[0158] Step S23: The obtained training set and test set meet the required ratio in terms of sample number.
[0159] Preferably, in step S21 of this embodiment, data representation and preprocessing are performed by representing the dataset as a set, forming a feature matrix and a target variable. This ensures that the data has a clear structure before analysis, making subsequent operations more systematic. Through data preprocessing, the feature matrix is ensured to contain all necessary information for modeling, and the target variable clearly indicates the output that the model needs to predict. A weight is calculated for each sample. This weight reflects the difference in importance of the sample in the entire dataset, allowing the model to focus more on samples that have a greater impact on the prediction results during training. Significance: The clear data structure provides a foundation for subsequent data processing and modeling, ensuring the accuracy of analysis. By assigning weights to samples, the model's ability to learn important samples is improved, which is particularly beneficial when the sample distribution is uneven, and can significantly improve model performance. Step S22 is the cumulative distribution function of the sample and weighted random sampling. The cumulative distribution function (CDF) of each sample is calculated based on the weights. This method provides a basis for subsequent sampling and ensures that the importance of each sample is considered during the sampling process. Using the cumulative distribution for random sampling, the training and test set samples are selected while maintaining a 7:3 ratio. Compared with simple random sampling, this is more scientific and improves the representativeness of the sampled data. Significance: Weighted sampling takes into account the importance of samples, so that the training set and test set contain more samples that have a significant impact on the model effect, thereby improving the effectiveness of the model; by introducing weights in sample selection, it avoids the sample selection bias that may be caused by ordinary random sampling, and ensures that the model performs well in both the training and testing stages. Step S23 summarizes the sample size ratio to ensure that the final training set and test set precisely meet the set 7:3 ratio in terms of sample size, providing a good foundation for model training and evaluation. Significance: By strictly adhering to the proportional division, the sample size of the training and test sets is balanced, which is conducive to the stability of model training and the generalization ability of unknown data; reasonable proportional distribution helps to provide a more objective and comprehensive performance evaluation in the model evaluation stage, making subsequent model adjustment and development more scientific.
[0160] In this embodiment, the data set is represented as follows: let the original data set be D, which contains N samples, each sample is a vector (d i ), (i=1,2,…,N). Assume D=d1,d2,…,d N Represents all the data of the original dataset;
[0161] Before partitioning, the data set is preprocessed to obtain the feature matrix F and target variable T:
[0162] Feature Matrix Where M represents the feature dimension;
[0163] Target variable Where T[i] represents d ithe corresponding landslide surge height;
[0164] Introduce a method based on weighted random sampling to divide the data set without affecting the validity of the data; assume that a weight parameter w is used i Indicates the reliability or importance of each sample. The sample selection is carried out according to the following steps:
[0165] Calculate weight: Calculate the weight w for each sample i , which can be determined by the physical parameters influencing factors of the landslide surge model, for example, w i =f(d i ), where f is a nonlinear function determined by the degree of feature influence, that is:
[0166] w i =α·Feature1 i +β·Feature2 i +…+γ·FeatureM i
[0167] α, β, γ are feature weights;
[0168] Total weight calculation: Define the total weight W as the sum of all sample weights:
[0169]
[0170] Calculate the sampling ratio: Based on the ratio of 7:3, the number of samples required for the training set is N train =0.7×N; select N by random sampling train samples, and perform weighted random sampling considering the influence of weights;
[0171] The steps to use the sampling method are as follows: Generate a cumulative distribution function based on the weights:
[0172]
[0173] Sampling process: for j = 1 to N train , generate a uniformly distributed random number r such that r∈[0,1]; determine the selected sample based on the cumulative distribution function P(i):
[0174] d train,j =d k where k = argmin i (|P(i)-r|)
[0175] Determine the test set: Samples that are not selected as the training set naturally become the test set:
[0176] D test =D\D train
[0177] The final generated training set D train and the test set D test The sample size satisfies |D traun |≈0.7×N and |D test |≈0.3×N, and the random sampling of weights ensures that the importance of samples is reflected in the training and testing stages.
[0178] In summary, through the implementation of these three steps, this embodiment can ensure that the division of the data set is not only scientific but also effective. Each step provides a guarantee for the performance improvement and accurate prediction of the model, thereby playing an important role in the prediction of landslide surge height. Through the standardization of sample structure, the introduction of weights and the normalization of proportional division, the overall process lays a solid foundation for subsequent modeling and analysis. Through the above-mentioned weighted random sampling method, not only the training set and test set are obtained, but also the importance of the samples is fully considered, avoiding the imbalance caused by simple random division; it provides a new idea for the data division process, thereby enhancing the effect of the subsequent landslide surge height prediction model.
[0179] Further, such as Figure 4 As shown, the process of predicting the maximum height of the landslide surge in step S3 specifically includes the following steps:
[0180] Step S31: A performance monitoring layer is introduced into the Sparrow Optimization Algorithm to track the performance of the slope surge height prediction model on the validation set in real time during the optimization process; during the search process, the number of global searches is dynamically increased. Whenever the Sparrow Optimization Algorithm detects a local optimal bottleneck, it enters the global reinforcement stage to increase the mutation rate and crossover rate;
[0181] Step S32: A feature sharing network is established between individual sparrows, allowing them to share and receive feature information through a weighted adjacency graph. Each sparrow receives the excellent features found by its neighbors through a graph propagation mechanism, converging a set of features that are effective for swell height prediction. After completing their respective tasks, each sparrow uses a feature evaluation and voting algorithm to decide whether to adopt a feature, thereby increasing the frequency of use of the optimal feature and achieving efficient integration of information sharing within the group.
[0182] Step S33: In the feature integration stage, different sparrow individuals are allowed to introduce heterogeneous features of their own fields (for example, soil type, meteorological factors, water level changes, etc.) to form diversified feature clusters; the effects of different feature combinations are clustered, and the optimal feature combination is selected for training through similarity measurement;
[0183] Step S34: During the operation of the random forest model, the decision tree monitors the performance of split nodes through weighted hysteresis tracking, and determines in real time whether nodes under certain conditions need to be reconstructed. Each node in the decision tree has a lifespan, and a new tree structure is dynamically created. When constructing each tree, a feature subset is selected based on feature importance, and a voting mechanism is used to comprehensively weight the important features of each tree. During feature analysis, the introduction and fusion of new features are achieved through sequential dependency selection.
[0184] Step S35: Generate new samples for the training dataset using a generative adversarial network, and construct different versions of the training set by combining the generated samples with the actual samples; set the parameters of the standard distribution selected in the pre-training phase of the generator; implement batch normalization and random dropout strategies in the training of the sparrow optimized random forest model;
[0185] Step S36: Obtain data from each experiment, including landslide volume and water depth, to form an input vector. The corresponding maximum surge height is the label. The input vector and the label are combined to form a labeled dataset; the labeled dataset is represented by a one-dimensional array, in which each element corresponds to the maximum surge height of each set of feature data; the labeled dataset and the trained sparrow-optimized random forest model are used to predict the maximum height of landslide surges.
[0186] Preferably, step S31 of this embodiment introduces a performance monitoring layer to track the performance of the model on the validation set in real time, allowing dynamic adjustment of the optimization strategy; dynamically increase the number of global searches, and effectively alleviate the problem of local optimality. Significance: Improve the adaptive ability of the model so that it can make timely adjustments when performance degrades to find a better solution; enhance the model's ability to track and predict performance in complex environments, especially when facing changes in landslide characteristics, it can adapt and optimize the search process more quickly. Step S32 establishes a feature sharing network so that each sparrow can efficiently communicate and integrate excellent features; through the graph propagation mechanism, it allows individuals in the group to share successful features, thereby improving feature utilization efficiency. Significance: Promote group intelligence, improve the group's overall feature detection ability and learning accuracy; enhance feature diversity, discover more potential important features through collective evaluation and voting mechanisms, and improve the model's generalization ability. Step S33 introduces heterogeneous features to form diverse feature clusters, combining knowledge from different fields; select the optimal feature combination through clustering and similarity metrics. Significance: Enhances the representativeness and completeness of features, fully utilizing effective information from various fields to adapt to different prediction scenarios; improves the model's feature learning capabilities, ensuring it can integrate diverse information to produce more accurate swell height predictions. Step S34 monitors split node performance in real time and performs dynamic tree reconstruction to ensure continuous model optimization; setting a node lifetime enhances the flexibility of the tree structure. Significance: Improves the adaptability of the decision tree, enabling it to self-adjust based on specific data changes and maintain efficient prediction capabilities; through dynamic structural changes, it adapts to data changes in real applications and improves the long-term effectiveness of the model. New samples are generated by combining with a generative adversarial network to expand the diversity of the training set; batch normalization and random dropout are implemented during training to enhance the robustness of the model. Significance: By generating samples to expand the dataset, the model captures more potential patterns and improves generalization ability; batch normalization and random dropout further reduce the risk of overfitting during training, ensuring the model's effectiveness in real-world environments. Step S35 selects a feature subset based on feature importance, utilizing a voting mechanism to enhance the scientific nature of feature selection; and new features are selected through sequential dependency, enabling flexible feature introduction and fusion. Significance: Enhances the model's sensitivity and responsiveness to features, ensuring optimal utilization of important features. Provides a flexible feature engineering mechanism to help the model more effectively adapt to new input features, facilitating accurate predictions in scenarios with changing data. Step S36 forms an input vector and label dataset for model training and testing; maximum surge height is used as a label to define the prediction target. Significance: Clarifies the relationship between the model's learning objectives and input features, providing a clear data structure and direction for subsequent model training. Promotes systematization during the model development phase, enabling data to fully leverage its potential, helping to establish effective prediction models and ultimately improving early warning capabilities and the scientific nature of decision-making.
[0187] In summary, the prediction process of this embodiment forms a coordinated, efficient, and dynamically adjusted prediction system through various steps, aiming to more accurately predict the maximum height of landslide surges and maximize the adaptability and accuracy of the model.
[0188] Furthermore, the process of achieving efficient integration of information sharing in the group in step S32 specifically includes the following steps:
[0189] Step S321: Assume that there are multiple sparrow individuals. Each sparrow individual is connected to its neighbor set, and an adjacency matrix is defined. At the same time, the feature vector of the sparrow individual is defined. After one propagation, the feature of the individual is obtained by weighted average of the features of its neighbors.
[0190] There are N′ sparrow individuals, each sparrow individual a and its neighbor set N a ′ connection. We define the adjacency matrix in:
[0191] A ab = {1, if sparrow, b is a sparrow, a's neighbor 0, otherwise
[0192] Define the eigenvector of sparrow individual a as Where M′ is the feature dimension. After one propagation, the feature of individual a can be obtained by weighted average of the features of its neighbors:
[0193] f a″ =σ(∑b∈N′ a w ab f b )
[0194] Among them, w ab is the influence weight of sparrow b on sparrow a, σ(·) is the activation function used to introduce nonlinearity;
[0195] Step S322: For the newly generated feature, each sparrow calculates its prediction ability and the prediction error of the sparrow for the new feature. All sparrows vote based on the error evaluation results to select the feature;
[0196] Among them, for the newly generated feature f a″ , each sparrow calculates its prediction ability, let ε a is the prediction error of sparrow a for the new feature, and the formula is as follows:
[0197]
[0198] Where K is the number of validation set samples, is the output of predicting sample k using the new feature, y k is the actual maximum height of the swell;
[0199] All sparrows vote based on the error evaluation results to select features. After the features pass the evaluation, the number of votes for each feature is V. c , then the probability P of a feature c being adopted b for:
[0200]
[0201] Step S323: Sparrow passes the adoption probability P b The features that are greater than the set threshold θ form a new feature set for subsequent training.
[0202] Preferably, the feature propagation mechanism in step S321 of this embodiment establishes connections between individual sparrows by defining an adjacency matrix, enabling feature sharing within a clear network structure. Each individual sparrow's feature vector is a weighted sum of its neighbors' features, enabling efficient information dissemination within the group. The aggregation feature enables each individual to absorb important feature information from diverse sources. The use of activation functions introduces nonlinearity, enhancing the model's expressiveness and making feature transformations more flexible. Significance: Through the feature propagation mechanism, individual sparrows can share information within the group, improving overall intelligence. This allows the group to demonstrate more precise aggregation capabilities when faced with complex problems. Because the network connection considers the features of multiple neighbors, it reduces the risk of bias from a single sparrow, while simultaneously collecting information from different features and promoting diversity. In step S322, the feature evaluation and voting mechanism involves each sparrow calculating a prediction error based on a new feature, providing a specific quantitative indicator of feature effectiveness. By voting on the feature evaluation results, each sparrow makes a collective decision on feature selection, making feature selection more scientific. Significance: Error-based evaluation provides an objective reference, enabling the group to more accurately select effective features, thereby avoiding the risk of personal bias affecting decision-making; through continuous evaluation and voting, the effectiveness of features is constantly tested and adjusted, which improves the model's adaptability to useful features and helps improve the final model performance. Step S323 Feature selection and update, Sparrow selects features through the set adoption probability to ensure that only features that exceed the threshold in effectiveness can be adopted for subsequent training; a new feature set is formed to facilitate adaptation to different data scenarios and trends, ensuring feature updates and iterations of the model at different stages. Significance: By dynamically adopting effective features, the final feature set is more in line with the actual situation, and the model's ability to cope with changing environments is improved; the selection of effective features reduces unnecessary calculations during training, improves training efficiency, makes the model more efficient, and reduces the risk of overfitting.
[0203] In summary, the above steps in this embodiment constitute a collaborative and dynamic feature sharing and selection mechanism, which promotes the effectiveness of the sparrow optimization algorithm in complex scenarios. Through feature propagation, error assessment, and dynamic feature adoption, the group can efficiently integrate information, improving the model's ability to predict the maximum height of landslide surges. This demonstrates significant advantages in information sharing, effective decision-making, and model optimization, providing a new perspective and approach for solving complex problems.
[0204] Furthermore, the process of selecting the optimal feature combination for training by similarity measurement in step S33 specifically includes the following steps:
[0205] Step S331: Generate various combinations from all available features, each feature combination represents a specific selection of a set of features, and uses a binary method to indicate whether each feature is selected; use a predefined random forest to train each feature combination, and evaluate the performance on the validation set, and assign a performance score to each feature combination based on the evaluation results;
[0206] Step S332: Based on the performance scores, the feature combinations are divided into different groups, and the feature combinations with similar performance are grouped together. By analyzing the clustering results, the similarity between the feature combinations is evaluated and the correlation between the features is measured;
[0207]
[0208] Where, Sim(C x ,C y ) represents the feature combination C x and C y The higher the value, the more similar the two feature combinations are; C x ·C y Represents the feature combination C x and C y The dot product of (i.e., the intersection between them) is actually the sum of the products of the corresponding eigenvalues in the eigenvectors. x ||C y | respectively represent feature combinations C x and C y The number of features contained in, that is, the norm of the feature vector (the size of the vector), K is the number of validation set samples, indicating the total number of features, that is, the number of all possible features, I xk and I yk is an indicator variable, representing the feature combination C x and C y Whether the kth feature is included in the combination C x is included in, then I xk =1, otherwise 0. Similarly, for combination C ySo too; Represents combination C x and C y The number of features contained in them is calculated in this part, which is their intersection. and Calculate the combination C of the two items separately x and C y The norm of the feature;
[0209] Step S333: Based on the calculated similarity, select a feature combination whose similarity is higher than a preset threshold.
[0210] Preferably, step S331 of this embodiment involves generating all possible combinations from all available features, using a binary representation of feature selection. This allows for systematic exploration of the feature space, taking into account the interactions between different features. Each feature combination is trained using a predefined random forest model and the model performance is evaluated on a validation set. The model evaluation results are quantified as performance scores, providing a data foundation for subsequent analysis. Significance: By generating multiple feature combinations, it is possible to ensure that different feature combinations are considered, which is crucial in machine learning, as interactions between features may affect the model's final performance. By training a model for each combination and evaluating its performance, an objective basis is provided for subsequent feature selection, ensuring that the selected combinations are validated. Step S332 involves grouping and similarity assessment of feature combinations. Based on performance scores, feature combinations are divided into different groups, with combinations with similar performance grouped together, helping to identify which combinations perform similarly on certain performance metrics. The similarity between feature combinations is calculated, using a custom similarity metric to quantify the correlation between the feature combinations. The higher the similarity, the more likely the two combinations contain similar features. Significance: Cluster analysis can help identify feature combinations with similar performance, simplify the subsequent decision-making process, and may reveal potential redundant combinations; similarity assessment provides a tool to understand which feature combinations are functionally similar, which is very important for feature selection and model interpretation; it can reduce model complexity and improve efficiency. Step S333 selects the optimal feature combination, and selects those feature combinations with similarity higher than the preset threshold based on the calculated similarity. This step is the key to achieving feature simplification and optimization, making the model more efficient. Significance: Selecting feature combinations with high similarity can avoid excessive feature redundancy, thereby improving the training and inference efficiency of the model. At the same time, this selection ensures the generalization ability of the model, because the selected combinations may perform well on their own similar feature sets; by limiting the number of feature combinations and the redundancy of features, the risk of model overfitting is reduced, thereby improving performance on unseen data.
[0211] In summary, this example systematically explores and evaluates feature combinations, ensuring that the selected feature combinations both enhance model performance and optimize model complexity. This process not only achieves excellent model performance but also prioritizes feature relevance and model understandability, supporting the effectiveness of subsequent decision-making and application.
[0212] Furthermore, the process of weighting the important features of each tree using the voting mechanism in step S34 specifically includes the following steps:
[0213] Step S341: Calculate a performance index for each decision tree based on the tree's performance on the validation set; assign different voting weights to each tree based on the performance index; establish a feature interaction network, modeling the relationship between each feature using a graph neural network, with the importance value of each tree for the feature as a node connected to other feature nodes to form a feature interaction relationship graph;
[0214] Step S342: The prediction results of each tree are analyzed for consistency. If the prediction of a tree in a specific area (such as a certain type of soil condition) is inconsistent with that of other trees, the voting influence weight of the tree is automatically reduced. Each tree provides a prediction value and a corresponding confidence level when voting. The output is a weighted sum of the feature weight, the tree's prediction value, and the corresponding confidence level.
[0215] Step S343: Group the features heterogeneously according to their properties, such as separating meteorological factors, terrain features, hydrological conditions, etc., and design an independent decision tree set for each feature group for training; when making decisions, vote by feature group, perform feature selection and evaluation within each feature group; and summarize the voting results of each feature group.
[0216] Preferably, step S341 of this embodiment calculates a performance index and establishes a feature interaction network. A performance index is calculated for each tree based on its performance on the validation set. This index, which comprehensively evaluates metrics such as precision, recall, and F1 score, provides a quantitative measure of each tree's voting weight in the decision-making process. Based on the calculated performance index, different decision trees are assigned different voting weights, with superior trees receiving higher influence and weaker trees receiving lower weights, thereby ensuring that the final decision better reflects the strengths of the overall model. A relationship model between features is established, using a graph neural network to capture the interactions and influences between features. The feature importance values of each decision tree serve as network nodes, and by connecting different feature nodes, a feature interaction graph is formed, allowing the model to understand the complex relationships between features. The significance achieved: By dynamically adjusting tree weights, the model's performance advantages can be better utilized, improving overall prediction performance. This ensures the model's adaptability to new data and enhances its intelligence and pertinence. The feature interaction network enables the model to learn and utilize the cooperative and competitive relationships between features, providing insight into which features have the greatest impact on decision-making in specific contexts, thereby improving the relevance and accuracy of prediction results. Step S342 is consistency analysis and confidence output. By performing consistency analysis on the prediction results of each tree, the voting weight of trees with large performance differences under specific conditions (such as soil type) is automatically reduced, abnormal trees are clearly identified, and the impact of noise is reduced. In addition to giving the prediction value, each tree also outputs the corresponding confidence level, making the final decision more transparent. When weighting the vote, the feature weight, the tree's prediction value and confidence level are combined to obtain the final result through weighted summation. The significance achieved: consistency analysis ensures that the voting of the decision tree is carried out on a more rational basis, thereby improving the stability of the prediction; the weight adjustment of the abnormally performing tree avoids the noise component in the decision; the introduction of confidence not only improves the interpretability of the prediction, but also provides users with a basis for trust in the model decision-making process. For users or domain experts, this transparency helps to make reasonable decisions based on data and facilitates the improvement of subsequent features and models. Step S343: Heterogeneous feature grouping and voting. This mechanism groups features according to their classification (e.g., meteorological factors, topographical features, hydrological conditions, etc.). This allows for the training of independent decision tree sets at a more refined level, reflecting the influence of features in specific domains or conditions. When making decisions, voting is performed by feature group. Features within each feature group are independently evaluated, and their respective voting results are synthesized based on their contribution to their specific domain. The voting results of all groups are then aggregated.Significance: Heterogeneous grouping ensures that the importance of features is fully accounted for, particularly the influence of features under different conditions. Decision trees that perform well on specific subsets of features can make more targeted predictions for specific situations, improving the model's adaptability in various scenarios. Cross-group aggregate voting fully utilizes the diversity of various features, enabling a more comprehensive capture of the potential information in the data. This enhances the overall model's generalization capabilities, enabling better predictions of unseen samples and improving practical application effectiveness.
[0217] In summary, the implementation of steps S341, S342, and S343 of this embodiment not only enhances different aspects of the intelligent decision-making process, but also forms a systematic and well-interpretable model integration strategy, thereby promoting the effective prediction of the maximum height of landslide surges.
[0218] Furthermore, the process of implementing batch normalization and random dropout strategy in step S35 specifically includes the following steps:
[0219] Step S351: Before model training, the input data is normalized. Based on the structure of the deep neural network, a batch normalization layer is inserted after each hidden layer to ensure that the distribution of features can be adjusted to a more ideal state before the output of each layer is passed to the activation function.
[0220] Step S352: After forward propagation in each hidden layer, the mean and variance of the hidden layer output are calculated and the output is normalized using statistics; the normalized output is transformed using the learned scaling and offset parameters, and the global mean and variance are dynamically adjusted using the exponential moving average method;
[0221] Step S353: Set the deactivation rate. In each training iteration, randomly select several neurons and set their outputs to zero to generate a mask matrix with the same shape as the current layer output. According to the selected deactivation rate, control which nodes are eliminated so that the model uses a different network structure in each training process.
[0222] Preferably, step S351 of this embodiment performs input data standardization and insertion of a batch normalization layer to standardize the input features (for example, normalize to a similar numerical range) so that different features have similar scales and distributions; this will facilitate the subsequent training process, as widely different features may lead to instability in the training process; a batch normalization layer is inserted after each hidden layer to ensure that the features are adjusted to an ideal state, i.e., a mean of zero and a variance of one, before the output is passed to the activation function. Significance achieved: By standardizing the input data and adjusting the feature distribution in a timely manner after the output of each layer, the training fluctuations caused by changes in data distribution can be effectively reduced, thereby improving the stability of model training; batch normalization can accelerate the convergence speed of the network, reduce training time, and enable the deep learning network to learn effective features faster; by ensuring that the activation value is continuously within an appropriate range, the gradient vanishing problem that may occur in the deep network during training is alleviated. Step S352: Post-forward propagation normalization and application of learned parameters. During the forward propagation process, the mean and variance of each hidden layer's output are calculated and used to normalize the output. Learnable scaling and offset parameters are used to transform the normalized output to a state suitable for model learning. An exponential moving average is used to dynamically adjust the global mean and variance. Significance: By introducing learnable parameters, the model can not only adapt to different input features but also self-adjust its feature distribution, enhancing the flexibility and adaptability of the model structure. The standardized features are more concentrated and stable, helping the model better capture complex patterns and improving predictive performance. Step S353: Implement random dropout. Set a dropout rate. In each training iteration, randomly select some neurons and set their outputs to zero. Generate a mask matrix to control which nodes are activated or removed based on the set dropout rate. Significance: Random dropout breaks the interdependence between neurons, encouraging the model to learn more robust and diverse feature combinations, effectively reducing overfitting. Using a different network structure in each training round allows the model to explore a richer range of feature representations during learning, thereby enhancing the model's generalization ability.
[0223] In summary, this approach optimizes the deep learning model training process. By standardizing input data, implementing batch normalization, and introducing a random dropout mechanism, we create a more stable, flexible, and robust learning environment. This allows the model to more effectively adapt to complex characteristics and make accurate judgments when predicting the maximum height of landslide surges. This improves the model's performance and generalization capabilities across a variety of environments, providing a solid foundation for practical applications.
[0224] Furthermore, the process of predicting the maximum height of the landslide surge in step S36 specifically includes the following steps:
[0225] Step S361: Acquire the latest input feature data, including landslide volume, soil type, meteorological factors (such as rainfall, temperature), water level changes and flow rate, etc.; integrate the collected feature data into an input vector;
[0226] Step S362: Load the trained and optimized random forest model into memory; input the constructed input feature vector into the random forest model, perform inference on each decision tree, and execute the decision process of the tree based on the input feature;
[0227] Each tree will traverse downward from the root node according to the value in the feature vector, splitting according to the feature threshold until it reaches a leaf node; at each leaf node, there is a predicted value, that is, the expected maximum surge height;
[0228] Step S363: Collect the prediction results of all decision trees and aggregate the results into a final prediction through a weighted voting mechanism; synthesize all trees to obtain a prediction value, which is used to represent the expected maximum surge height under a specific input feature vector; output the maximum surge height prediction result in the format of text, chart or report.
[0229] Preferably, step S361 of this embodiment obtains input feature data and constructs an input vector. By monitoring and collecting the latest feature data (such as landslide volume, soil type, meteorological factors, water level changes and flow rate), it ensures that the model uses current and relevant input information; the collected multi-dimensional features are organized into a unified input vector, making it suitable for subsequent model predictions. The structured approach enables the model to more easily parse and use data. The significance achieved: The latest feature data ensures that the model prediction is based on the current real environment, greatly improving the accuracy and reliability of the prediction; integrating multiple features into a vector facilitates model recognition and processing, promotes the simplification of data preprocessing, and ensures that the input meets the model requirements. Step S362 loads the model and performs decision tree inference, ensuring that the sparrow-optimized random forest model is accurately loaded and ready for inference; after inputting the feature vector, each decision tree in the random forest will make a path decision based on the feature value, traverse from the root node downward, split according to the feature threshold, and finally output the prediction result at the leaf node. Significance achieved: By loading the optimized random forest model, the model can fully apply the previous training results and utilize the learned feature importance and decision paths; the decision tree inference process allows the model to estimate the output based on detailed feature splitting, so that each tree can independently evaluate and predict the target variable, enhancing the dispersion and robustness of the predictions. Step S363: Result aggregation and output. The prediction values of all decision trees are aggregated together through a weighted voting mechanism to obtain the final overall prediction result; the advantages of each tree can be effectively superimposed, thereby improving the reliability of the final output; the prediction results are intuitively displayed in the form of text, charts or reports for application in actual decision-making. Significance achieved: By aggregating the results of multiple decision trees, the characteristics of ensemble learning significantly enhance the generalization ability and prediction accuracy of the model and reduce the limitations of a single tree; the results are output in a variety of forms, allowing users from different backgrounds (such as decision makers, engineers, researchers, etc.) to easily understand the results and formulate corresponding countermeasures and decisions based on them, promoting the social application value of the results.
[0230] In summary, this embodiment effectively systematizes the process of predicting the maximum height of landslide surges through the above steps, maximizing the use of existing knowledge and data while prioritizing the real-time and structured nature of the data. This entire process not only improves the scientific nature and accuracy of predictions but also provides a more reliable tool for environmental monitoring, risk management, and emergency response, with far-reaching practical significance.
[0231] like Figure 5As shown, this embodiment also provides an embodiment of a landslide surge height prediction system based on a sparrow optimized random forest model. In this embodiment, the landslide surge height prediction system based on a sparrow optimized random forest model is applied to the landslide surge height prediction method based on a sparrow optimized random forest model in the above embodiment. The landslide surge height prediction system based on a sparrow optimized random forest model includes a data acquisition module 1, a data division module 2, and a landslide surge prediction module 3 electrically connected in sequence;
[0232] Among them, the data acquisition module 1 is used to build a corresponding landslide surge physical model for the target landslide area, set up test groups with different landslide volumes and water depths for testing, obtain the corresponding test data set, and scale it proportionally according to the Froude similarity criterion; the data partitioning module 2 is used to preprocess the test data to obtain the corresponding data set, and divide the preprocessed data set into a training set and a test set in a ratio of 7:3, that is, the test set data accounts for 30% of the total data; the landslide surge prediction module 3 is used to use Python language to build a sparrow optimized random forest model based on the sklearn and mealpy libraries, set the input values of various parameters and use the training set to train the model, and use the labeled data set and the trained sparrow optimized random forest model to predict the maximum height of the landslide surge.
[0233] Preferably, the data acquisition module of this embodiment establishes a corresponding landslide surge physical model for the target landslide area, accurately simulating the wave characteristics of the surge during the landslide. Comparative experiments are conducted under different landslide volume, water depth, and other conditions to obtain a rich test data set. The design of the test group ensures that a variety of possible landslide and hydrological conditions are covered, providing a foundation for subsequent modeling. The test data is scaled according to the Froude similarity criterion, allowing the experimental results to be extrapolated across different scales, making them more applicable. Significance: Real data obtained through physical experiments ensures the foundation and reliability of subsequent predictions and can effectively reflect the real-world landslide surge behavior. The scaling technique using the similarity criterion allows the obtained data to be applied to landslide scenarios of multiple scales, improving the versatility of the model. The data partitioning module cleans and preprocesses the obtained test data, removing outliers, filling missing values, and standardizing them, laying the foundation for effective data analysis and modeling. The dataset is divided into a training set and a test set in a 7:3 ratio, ensuring that the model can both learn features and effectively evaluate the model's generalization ability through the test set. Significance Achieved: A reasonable partitioning approach not only enables the model to learn, but also allows for independent test sets to validate its performance, avoiding overfitting and enhancing its credibility. Independent evaluation of the test set yields realistic model performance metrics, making it more valuable for practical applications. The landslide surge prediction module uses Python combined with the sklearn and mealpy libraries to build a sparrow-optimized random forest model. This model enhances performance by setting various parameter input values to ensure adaptation to data characteristics. The trained model is applied to the labeled dataset to provide an estimated maximum landslide surge height, effectively predicting the impact based on a new input feature vector. Significance Achieved: By leveraging machine learning algorithms, particularly the ensemble learning properties of the random forest model, precise modeling and processing of complex problems are achieved, providing fast and accurate prediction results. The process of building the prediction module demonstrates the application of modern data science and artificial intelligence in complex systems, improving the automation of landslide surge prediction and, in turn, promoting intelligent landslide disaster management and risk assessment.
[0234] In summary, the various modules within the landslide surge prediction system of this embodiment work together to create a valuable tool for landslide surge safety assessment. The data acquisition module ensures the quality and applicability of basic data, the data partitioning module ensures the effectiveness of model training and evaluation, and the prediction module enables efficient landslide surge height prediction. The effective operation of these modules not only improves the systematicity and accuracy of research but also provides a scientific basis and decision-making support for future disaster response.
[0235] like Figure 6As shown, this embodiment provides an embodiment of an electronic device. In this embodiment, the electronic device 4 includes a processor 41 and a memory 42 coupled to the processor 41.
[0236] The memory 42 stores program instructions for implementing the landslide surge height prediction method based on the sparrow optimized random forest model according to any of the above embodiments.
[0237] The processor 41 is used to execute program instructions stored in the memory 42 to predict the landslide surge height based on the sparrow optimized random forest model.
[0238] The processor 41 may also be referred to as a CPU (Central Processing Unit). The processor 41 may be an integrated circuit chip having signal processing capabilities. The processor 41 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The general-purpose processor may be a microprocessor or any conventional processor.
[0239] Furthermore, Figure 7 This is a schematic diagram of the structure of a storage medium in an embodiment of the present application. The storage medium 5 in the embodiment of the present application stores program instructions 51 that can implement all the above methods, wherein the program instructions 51 can be stored in the above storage medium in the form of a software product, including a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or terminal devices such as a computer, server, mobile phone, and tablet.
[0240] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0241] In addition, the functional units in the various embodiments of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units. The above is only an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
[0242] like Figure 8 As shown, this embodiment also provides an embodiment of a method for predicting landslide surge height based on a sparrow-optimized random forest model. In this embodiment, for a target landslide area, the following steps are performed to obtain a sparrow-optimized random forest model for predicting the maximum landslide surge height, and to predict the maximum landslide surge height:
[0243] Step 1: Build a corresponding landslide surge physical model for the target landslide area. In this embodiment, a landslide in the reservoir area of a hydropower station on the Jinsha River is selected as the research object. The landslide is 1.5 km away from the study area and about 90 km away from the hydropower station. A corresponding landslide surge physical model is built at a scale of 1:150 and a landslide surge test is carried out. The landslide is simulated by a free slide of a stone pile down the slope. Eight wave height meters are arranged at different locations near the study area. The wave height acquisition system is DYS50-3000. The wave height meters are numbered as 8, 9, 14, 15, 16, 17, 18, and 19. Test groups with different landslide volumes and water depths are set up for testing. The test grouping is shown in Table 1. The corresponding test data sets are obtained. There are 112 groups of test data in total. According to the Froude similarity criterion, the test data are proportionally amplified using a length scale and a velocity scale. The data after proportional amplification are shown in Table 2.
[0244] Table 1 Test conditions
[0245] Test number <![CDATA[Volume (m 3 )]]> Water depth (m) 1 181 800 2 181 815 3 181 825 4 123 800 5 123 815 6 123 825
[0246] Table 2 Test data after proportional amplification
[0247]
[0248] Step 2: Preprocess the test data, call the LabelEncoder() function to create a "LabelEncoder" object, use the "fit_transform" method to encode the "measurement point label" column, and use the standardization tool "scaler" to scale the features to a mean of 0 and a variance of 1; use the "maximum landslide surge height" column as the label data set, and the remaining columns of data as the feature data set, call the train_test_split() function to divide the data set into training set and test set in a ratio of 7:3, that is, the number of test sets accounts for 30% of the total data set. When dividing the data set, call the parameter shuffle to shuffle the data and randomly group them.
[0249] Step 3: Use Python language to build a sparrow optimized random forest model based on sklearn and mealpy libraries, set the input values of various parameters and use the training set to train the model.
[0250] In step 3, the following steps are performed to obtain the sparrow-optimized random forest model for landslide surge height prediction:
[0251] Step 1: Import the sklearn and mealpy libraries, and call the def() function to define the objective function "objective_function". The input value of the objective function is "params". Construct a tuple "params" containing the number of decision trees (n_estimators) and the depth of the decision tree (max_depth), call the RandomForestRegressor() function to build a random forest model containing the number of decision trees and the depth of the decision tree, set the input value of the parameter "random_state" to 1234, and the input value of the parameter "criterion" to "absolute_error", then call the fit() function to train the model, and call the predict() function to predict the result; call the mean_squared_error() function and the sqrt() function to calculate the root mean square error (RMSE) between the predicted result and the true value, call the mean_absolute_error() function to calculate the mean absolute error (MAE), calculate the sum of the two and use it as the return value of the objective function;
[0252] Step 2: Define the problem dictionary "problem_dict", which has three entries: "bounds", "minmax", and "obj_func". "bounds" specifies the range of the number of decision trees (n_estimators) and the depth of the decision tree (max_depth) of the random forest model parameters. The range of the number of decision trees (n_estimators) is 10 to 1000, and the range of the depth of the decision tree (max_depth) is 3 to 10. "minmax" sets the optimization goal of the objective function to the minimum of the sum of the root mean square error (RMSE) and the mean absolute error (MAE). "obj_func" specifies the objective function "objective_function" as the optimization goal.
[0253] Step 3: Call the SSA.DevSSA() function to build the Sparrow Optimization Algorithm (SSA) optimization model, set the parameter training rounds (epoch) to 5, population size (pop_size) to 100, step size factor (ST) to 0.8, exploration factor (PD) to 0.5, and tolerance factor (SD) to 0.5, and call the solve() method to optimize the model "promblem_dict" and obtain its optimal parameter combination "best_params";
[0254] Step 4: Get the optimal number of decision trees "best_n_estimators" and the optimal decision tree depth "best_max_depth" from the optimal parameter combination "best_params", call the RandomForestRegressor() function according to the optimal parameter combination to recreate a random forest model, call the fit() function to train the model, and obtain the optimized random forest model.
[0255] Step 4: Use the test set and the trained sparrow optimized random forest model to call the predict() function to predict the maximum height of the landslide surge. The comparison between the predicted value and the actual value of the model before and after optimization is as follows: Figure 9 As shown, the random forest model learning curve is as follows Figure 10 As shown, the learning curve of the sparrow optimized random forest model is as follows Figure 11 As shown in the figure, the feature importance coefficient comparison is as follows Figure 12 The performance evaluation indicators of the model before and after optimization are shown in Table 3.
[0256] Table 3 SSA-RF model performance evaluation indicators
[0257] Model Type RMSE MAE <![CDATA[R 2 ]]> MAPE RF 1.55075 1.14929 0.56867 35.58368 SSA-RF 1.54120 1.10510 0.57397 31.62375
[0258] This embodiment proposes a landslide surge height prediction method based on a sparrow optimized random forest model, establishes a landslide surge physical model test based on the Froude similarity criterion, obtains the actual maximum landslide surge height at different points through the test, and establishes a sparrow optimized random forest model through a random forest algorithm and a sparrow optimization algorithm, which can accurately predict the maximum height of landslide surges at different points in the study area; the present invention introduces a sparrow optimization algorithm to optimize the random forest model and applies it to actual engineering cases. The research results show that the sparrow optimized random forest model has better performance when processing data with a smaller deviation rate and can significantly improve the accuracy of the model. At the same time, the sparrow optimized random forest model can reduce the standard deviation of the predicted value and increase the density of the predicted value. Compared with the random forest model before optimization, the performance evaluation indicators of the optimized model are all better than those of the model before optimization, and it shows better adaptability in predicting the maximum height of landslide surges, and can more accurately predict the maximum height of landslide surges.
[0259] The above detailed description of the specific embodiments of the invention is intended to be illustrative only, and the present invention is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications or substitutions to the invention are also within the scope of the present invention. Therefore, equivalent changes, modifications, and improvements made without departing from the spirit and scope of the present invention are also encompassed within the scope of the present invention.
Claims
1. A landslide surge height prediction method based on sparrow optimized random forest model, characterized in that: The landslide surge height prediction method based on the sparrow optimized random forest model includes: For the target landslide area, a corresponding landslide surge physical model was built, and test groups with different landslide volumes and water depths were set up for testing to obtain the corresponding test data set, which was then scaled according to the Froude similarity criterion. The experimental data were preprocessed to obtain the corresponding data set, which was then divided into a training set and a test set in a ratio of 7:3, i.e. the test set data accounted for 30% of the total data; Using Python language, a sparrow optimized random forest model was built based on sklearn and mealpy libraries. The input values of various parameters were set and the training set was used to train the sparrow optimized random forest model. The labeled dataset and the trained sparrow optimized random forest model were used to predict the maximum height of the landslide surge.
2. The landslide surge height prediction method based on the sparrow optimized random forest model according to claim 1 is characterized in that: The process of scaling according to the Froude similarity criterion includes: Calculate the Froude number of the target landslide area and the model; introduce the characteristic flow frequency F c and the flow resistance factor D: Matching of the Froude number of the model and the target landslide area: The Froude number of the test model is equal to that of the target landslide area, that is: Among them, V m and V p are the flow velocities in the model and target areas respectively; L m and L p are the characteristic lengths of the model and target regions, respectively; Determine the ensemble scaling ratio λ of the model and the target landslide area; According to the Froude number formula, the scaling factor is defined as: Update the flow rate of the model to meet the proportional scaling condition and introduce an adjustment coefficient C v Consider the differences between actual flows: Adjustment coefficient C v Adjusted based on previous trial data and similarity; The test flow rate, pressure, gravitational acceleration and other related physical quantities are adjusted proportionally; The influence of the density λ′ and dynamic viscosity of the introduced fluid on the pressure is: Among them C p is the pressure adjustment coefficient, which is calibrated by test data; After adjusting all relevant physical quantities, experiments are conducted on the model, and data are collected and recorded accordingly; the experimental data are scaled to the target landslide area in each experiment: Prediction result = C data Actual test results Among them C data is the conversion factor used to fit the model to the experimental data.
3. The landslide surge height prediction method based on the sparrow optimized random forest model according to claim 1 is characterized in that: The preprocessed dataset is divided into training and testing sets, including: Represent the dataset as a set containing multiple samples, each of which consists of features and corresponding target variables; preprocess the data to generate a feature matrix and target variables; calculate a weight for each sample, which reflects the importance of the sample; The cumulative distribution function of the samples is calculated based on the weights, and the cumulative distribution function is used for weighted random sampling. Samples are randomly selected in a 7:3 ratio as the training set, and the remaining samples are used as the test set; The obtained training set and test set meet the required ratio in terms of sample number.
4. The landslide surge height prediction method based on the sparrow optimized random forest model according to claim 1 is characterized in that: The process of predicting the maximum height of landslide surge includes: A performance monitoring layer is introduced into the Sparrow Optimization Algorithm to track the performance of the slope surge height prediction model on the validation set in real time during the optimization process. The number of global searches is dynamically increased during the search process. Whenever the Sparrow Optimization Algorithm detects a local optimal bottleneck, it enters a global reinforcement phase to increase the mutation rate and crossover rate. A feature sharing network is established between individual sparrows, allowing them to share and receive feature information through a weighted adjacency graph. Each sparrow receives excellent features found by its neighbors through a graph propagation mechanism, converging on a set of features that are effective for predicting swell height. After completing their respective tasks, sparrows use a feature evaluation and voting algorithm to decide whether to adopt a feature, increasing the frequency of use of the optimal feature and achieving efficient information sharing and integration within the group. In the feature integration stage, different sparrow individuals are allowed to introduce heterogeneous features from their own fields to form diversified feature clusters. They are grouped by clustering the effects of different feature combinations, and the optimal feature combination is selected for training through similarity metrics. During the operation of the random forest model, the decision tree monitors the performance of split nodes through weighted hysteresis tracking, and determines in real time whether nodes under certain conditions need to be reconstructed. Each node in the decision tree has a lifespan, and a new tree structure is dynamically created. When constructing each tree, a feature subset is selected based on feature importance, and a voting mechanism is used to weight the important features of each tree. During feature analysis, new features are introduced and integrated through sequential dependency selection. Generate new samples for the training dataset using a generative adversarial network, and construct different versions of the training set by combining the generated samples with the actual samples; set the parameters of the standard distribution selected in the generator pre-training phase; and implement batch normalization and random dropout strategies in the training of the Sparrow Optimized Random Forest model. The landslide volume and water depth data for each experiment are obtained to form an input vector. The corresponding maximum surge height is the label. The input vector and the label are combined to form a labeled dataset. The labeled dataset is represented by a one-dimensional array, in which each element corresponds to the maximum surge height of each set of feature data. The maximum height of landslide surge is predicted using the labeled dataset and the trained sparrow-optimized random forest model.
5. The landslide surge height prediction method based on the sparrow optimized random forest model according to claim 4 is characterized in that: Achieving efficient and integrated information sharing within a group, including: Suppose there are multiple sparrow individuals, each sparrow individual is connected to its neighbor set, an adjacency matrix is defined, and the feature vector of the sparrow individual is also defined. After one propagation, the individual's features are obtained by weighted average of the features of its neighbors; There are N′ sparrow individuals, each sparrow individual a and its neighbor set N a 'Connect, define the adjacency matrix in: A ab = {1, if sparrow, b is a sparrow, a's neighbor 0, otherwise Define the eigenvector of sparrow individual a as Where M′ is the feature dimension. After one propagation, the feature of individual a is obtained by weighted average of the features of its neighbors: f a″ =σ(∑b∈N′ a w ab f b ) Among them, w ab is the influence weight of sparrow b on sparrow a, σ(·) is the activation function used to introduce nonlinearity; For the newly generated features, each sparrow calculates its prediction ability and the prediction error of the sparrow for the new features. All sparrows vote based on the error evaluation results to select the features. Among them, for the newly generated feature f a″ , each sparrow calculates its prediction ability, let ε a is the prediction error of sparrow a for the new feature, and the formula is as follows: Where K is the number of validation set samples, is the output of predicting sample k using the new feature, y k is the actual maximum height of the swell; All sparrows vote based on the error evaluation results to select features. After the features pass the evaluation, the number of votes for each feature is V. c , then the probability P of a feature c being adopted b for: The sparrow adopts the probability P b The features that are greater than the set threshold θ form a new feature set for subsequent training.
6. The landslide surge height prediction method based on the sparrow optimized random forest model according to claim 4 is characterized in that: The process of selecting the optimal feature combination for training through similarity measurement includes: Generate various combinations from all available features, each feature combination represents a specific selection of a set of features, using a binary method to indicate whether each feature is selected; use a predefined random forest to train each feature combination and evaluate the performance on the validation set, assigning a performance score to each feature combination based on the evaluation results; Based on the performance scores, the feature combinations are divided into different groups, and the feature combinations with similar performance are grouped together. By analyzing the clustering results, the similarity between feature combinations is evaluated and the correlation between features is measured. Based on the calculated similarity, a feature combination with a similarity higher than a preset threshold is selected.
7. The method for predicting landslide surge height based on the sparrow optimized random forest model according to claim 4 is characterized in that: The process of using a voting mechanism to comprehensively consider the important features of each tree for weighting includes: A performance index is calculated for each decision tree based on its performance on the validation set. Different voting weights are assigned to each tree based on the performance index. A feature interaction network is established, modeling the relationships between features using a graph neural network. The importance value of each tree for a feature is used as a node connected to other feature nodes to form a feature interaction relationship graph. The prediction results of each tree are analyzed for consistency. If the prediction of a tree in a specific field is inconsistent with that of other trees, the voting influence weight of the tree is automatically reduced. Each tree gives a prediction value and a corresponding confidence level when voting. The output is a weighted sum of the feature weight, the tree's prediction value, and the corresponding confidence level. The features are heterogeneously grouped according to their properties, and meteorological factors, terrain characteristics, and hydrological conditions are separated. An independent set of decision trees is designed for each feature group and trained. When making decisions, voting is performed by feature group, and feature selection and evaluation are performed within each feature group. The voting results of each feature group are summarized.
8. The method for predicting landslide surge height based on the sparrow optimized random forest model according to claim 4 is characterized in that: The process of implementing batch normalization and random dropout strategies includes: Before model training, the input data is standardized. Based on the structure of deep neural networks, a batch normalization layer is inserted after each hidden layer to ensure that the distribution of features can be adjusted to a more ideal state before the output of each layer is passed to the activation function. After forward propagation of each hidden layer, the mean and variance of the hidden layer output are calculated and the output is normalized using statistics. The normalized output is transformed using the learned scaling and offset parameters, and the global mean and variance are dynamically adjusted using the exponential moving average method. Set the inactivation rate. In each training iteration, randomly select several neurons and set their outputs to zero. Generate a mask matrix with the same shape as the current layer output. Control which nodes are eliminated based on the selected inactivation rate so that the model uses a different network structure in each training process.
9. The method for predicting landslide surge height based on the sparrow optimized random forest model according to claim 4, characterized in that: The process of predicting the maximum height of landslide surge includes: Obtain the latest input feature data, including landslide volume, soil type, meteorological factors, water level changes and flow rate; integrate the collected feature data into an input vector; Load the trained and Sparrow-optimized random forest model into memory; input the constructed input feature vector into the random forest model, perform inference on each decision tree, and execute the tree's decision process based on the input features; Each tree will traverse downward from the root node according to the value in the feature vector, splitting according to the feature threshold until it reaches a leaf node; at each leaf node, there is a predicted value, that is, the expected maximum surge height; Collect the prediction results of all decision trees and aggregate them into a final prediction through a weighted voting mechanism; combine all trees to obtain a prediction value, which is used to represent the expected maximum surge height under a specific input feature vector; output the maximum surge height prediction result in the format of text, chart or report.
10. A landslide surge height prediction system based on a sparrow-optimized random forest model, which is applied to the landslide surge height prediction method based on a sparrow-optimized random forest model as claimed in any one of claims 1 to 9, characterized in that: The landslide surge height prediction system based on the sparrow optimized random forest model includes: The data acquisition module is used to build a corresponding landslide surge physical model for the target landslide area, set up test groups with different landslide volumes and water depths for testing, obtain the corresponding test data set, and perform proportional scaling according to the Froude similarity criterion; The data partitioning module is used to preprocess the test data to obtain the corresponding data set, and then divide the preprocessed data set into a training set and a test set in a ratio of 7:3, that is, the test set data accounts for 30% of the total data; The landslide surge prediction module is used to build a sparrow optimized random forest model based on the sklearn and mealpy libraries using Python language, set the input values of various parameters and train the model using the training set, and use the labeled dataset and the trained sparrow optimized random forest model to predict the maximum height of the landslide surge.
Citation Information
Patent Citations
Model test device and model test method for researching landslide surge energy dissipation effect and wave height prediction
CN110188443A
Landslide-surge maximum height exceedance probability prediction method, device and storage device
CN117172368B
Method for judging type of surge induced by landslide in water area and predicting height of surge
CN118133380A
Simulation computing method and device of landslide surge disaster
CN108073767A
Landslide-surge climbing disaster exceeding probability evaluation method and device and storage device
CN117195766A